Network data flow surveying and mapping method, device, computer equipment and storage medium

By acquiring and decrypting network data packets, and combining DAVA algorithm and data asset lists to generate dynamic data flow diagrams, the problems of inefficiency and insufficient accuracy of network data flow surveying and mapping in the prior art are solved, and real-time and accurate data flow path monitoring and security warning are achieved.

CN119051976BActive Publication Date: 2025-08-08HANGZHOU MEICHUANG DIGITAL INTELLIGENCE TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411481512.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-23
Publication Date
2025-08-08
Estimated Expiration
2044-10-23

AI Technical Summary

Technical Problem

The existing network data flow surveying and mapping technology is inefficient and insufficiently accurate, and it is impossible to monitor the data flow status in real time, resulting in the inability to detect potential security risks in time and increase data security risks.

Method used

By acquiring network data packets and in-app data activities, decrypting and aggregating, forming a structural data set, analyzing node data using DAVA algorithm, generating dynamic data flow diagrams with data asset lists, performing similarity analysis, and tracking weak links and risk points in real time.

Benefits of technology

It realizes real-time and accurate tracking of weak links and risk points in the data flow path, timely warning, and ensure data security.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119051976B_ABST
    Figure CN119051976B_ABST
Patent Text Reader

Abstract

The embodiment of the present invention discloses a network data flow mapping method, apparatus, computer equipment and storage medium. The method includes: obtaining network data packets and data activity within an application to obtain a first definition data set and a second definition data set; decrypting specific traffic and performing aggregation processing to obtain a structured data set; analyzing and judging the node data in the first definition data set to obtain a data processing point set; obtaining data tables and field information of a preset data service server to obtain a data asset list; performing data flow analysis on the data processing point set, and performing similarity analysis and processing of data records to obtain a fourth definition data set; generating a dynamic data flow diagram; and outputting a dynamic data flow diagram. By implementing the method of the embodiment of the present invention, it is possible to achieve real-time and accurate tracking of weak links and risk points in the data flow path, thereby achieving timely warnings and ensuring data security.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data security technology, and more specifically to a network data flow surveying and mapping method, apparatus, computer equipment, and storage medium. Background Art

[0002] In recent years, the rapid development of technologies such as the internet, big data, cloud computing, the Internet of Things, and artificial intelligence has dramatically changed how data is generated, collected, stored, processed, and analyzed. These technological advances have led to an exponential increase in the speed and frequency of data generation and circulation, imbuing data with greater value and complexity.

[0003] However, current network data flow mapping technology primarily relies on manual correlation analysis of various logs and the creation of flow maps based on these analysis results. This approach suffers from significant inefficiencies, lacks integrity, and low accuracy, preventing users from monitoring the status of data flows in real time. This delay often prevents the timely detection of potential security risks, increasing data security risks and potentially causing economic losses and adverse social impacts.

[0004] Therefore, it is necessary to design a new method to achieve real-time and accurate tracking of weak links and risk points in the data flow path, so as to achieve timely warning and ensure data security. Summary of the Invention

[0005] The purpose of the present invention is to overcome the defects of the prior art and provide a network data flow surveying and mapping method, device, computer equipment and storage medium.

[0006] To achieve the above-mentioned purpose, the present invention adopts the following technical solution: a network data flow surveying and mapping method, comprising:

[0007] Acquire network data packets and data activity within the application to obtain a first defined data set and a second defined data set;

[0008] Decrypting the specific traffic in the first defined data set and the second defined data set, and performing aggregation processing to obtain a structured data set;

[0009] Analyzing and judging the node data in the first defined data set to obtain a data processing point set;

[0010] Obtain data table and field information from the preset data service server to obtain a data asset list;

[0011] Performing data flow analysis on the data processing point set to obtain a third defined data set, and performing similarity analysis and processing on the data records to obtain a fourth defined data set;

[0012] Generate a dynamic data flow diagram based on the fourth definition data set, combined with the structure data set and the data asset list;

[0013] Output the dynamic data flow graph.

[0014] A further technical solution is that the acquisition of network data packets and data activity within the application to obtain the first defined data set and the second defined data set includes:

[0015] Scan the preset network address segment and mirror the switch traffic to obtain network data packets;

[0016] Parsing the network data packet to obtain a source IP address, a source port, a destination IP address, a destination port, data packets, and device data of a network node, and merging the data packets to obtain a first defined data set;

[0017] The data activity in the application is acquired by embedding the generated specific software package into the business application and collected within a preset time period to obtain a second defined data set.

[0018] A further technical solution is: decrypting the specific traffic in the first defined data set and the second defined data set and performing aggregation processing to obtain a structured data set, including:

[0019] decrypting specific traffic in the first defined data set and the second defined data set;

[0020] The decrypted first definition data set and the second definition data set are merged to remove duplicate data to obtain a structured data set, and the structured data set is used to draw a first-layer static network node and relationship diagram.

[0021] A further technical solution is: analyzing and judging the node data in the first defined data set to obtain a data processing point set, including:

[0022] The DAVA algorithm is used to analyze the node data in the first defined data set to determine all processing points that the data passes through in the network system to form a data processing point set.

[0023] A further technical solution is: obtaining the data table and field information of the preset data service server to obtain the data asset list includes:

[0024] Scan the preset data service servers within the organization to identify data tables and field information, and form a data asset list. Label the data assets in the data asset list and create a data fingerprint information table.

[0025] A further technical solution is: performing data flow analysis on the data processing point set to obtain a third defined data set, and performing similarity analysis and processing on the data records to obtain a fourth defined data set, including:

[0026] Analyzing the flow carrier, characteristics, and route direction of data at each data processing point in the set of data processing points to record relevant time series information, thereby obtaining data records at each processing point to form a third defined data set;

[0027] A cosine similarity algorithm is used to perform similarity analysis on the data records in the third definition data set one by one, and identical and similar data are eliminated to obtain a fourth definition data set.

[0028] Its further technical solution is: the directed lines of the data flow directed graph express the attribute information of the volume and sensitivity of the data flow through the shape of the lines, wherein the thickness of the lines represents the size of the data flow, and the color of the lines represents the sensitivity of the data.

[0029] The present invention also provides a network data flow surveying and mapping device, comprising:

[0030] A data packet acquisition unit, configured to acquire network data packets and data activity within an application to obtain a first defined data set and a second defined data set;

[0031] a processing unit, configured to decrypt the specific traffic in the first defined data set and the second defined data set, and perform aggregation processing to obtain a structured data set;

[0032] An analysis and judgment unit, configured to analyze and judge the node data in the first defined data set to obtain a data processing point set;

[0033] An information acquisition unit is used to acquire data table and field information of a preset data service server to obtain a data asset list;

[0034] a data analysis unit, configured to perform data flow analysis on the data processing point set to obtain a third defined data set, and perform similarity analysis and processing on the data records to obtain a fourth defined data set;

[0035] a flow diagram generating unit, configured to generate a dynamic data flow diagram based on the fourth definition data set, in combination with the structure data set and the data asset list;

[0036] An output unit is used to output the dynamic data flow graph.

[0037] The present invention further provides a computer device, comprising a memory and a processor, wherein a computer program is stored in the memory, and the processor implements the above method when executing the computer program.

[0038] The present invention also provides a storage medium, wherein the storage medium stores a computer program, and the computer program implements the above method when executed by a processor.

[0039] The beneficial effects of the present invention compared with the existing technology are as follows: the present invention obtains and decrypts network data packets and data activities within applications, aggregates them into structured data sets, forms a basic data set and a set of data processing points, analyzes the set of data processing points, and combines it with the data asset list to generate a data flow diagram; performs similarity analysis on the data flow path to determine potential weak links and risk points in the data flow; based on the dynamic data flow diagram, provides real-time warnings and corrections to problems in the data flow to ensure data security; and achieves real-time and accurate tracking of weak links and risk points in the data flow path, thereby achieving timely warnings and ensuring data security.

[0040] The present invention will be further described below with reference to the accompanying drawings and specific embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0042] Figure 1 A schematic diagram of an application scenario of the network data flow surveying and mapping method provided by an embodiment of the present invention;

[0043] Figure 2 A schematic diagram of a flow chart of a network data flow surveying and mapping method provided by an embodiment of the present invention;

[0044] Figure 3 A schematic block diagram of a network data flow surveying and mapping device provided by an embodiment of the present invention;

[0045] Figure 4 A schematic block diagram of a computer device provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0046] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0047] It will be understood that when used in this specification and the appended claims, the terms “comprises” and “comprising” indicate the presence of described features, integers, steps, operations, elements and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof.

[0048] It should also be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the present invention. As used in the specification and appended claims, the singular forms "a," "an," and "the" are intended to include the plural forms unless the context clearly indicates otherwise.

[0049] It should be further understood that the term "and / or" used in the present description and appended claims refers to and includes any and all possible combinations of one or more of the associated listed items.

[0050] See also Figure 1 and Figure 2 , Figure 1 Schematic diagram of an application scenario of the network data flow surveying and mapping method provided by an embodiment of the present invention. Figure 2 A schematic flow chart of a network data flow mapping method provided in an embodiment of the present invention. The network data flow mapping method is applied in a server. The server interacts with the terminal to perform data exchange, collects network data packets and data activities within the application, and generates a first and a second definition data set; decrypts specific traffic and aggregates the processing data set to obtain a structured data set; analyzes the node data in the first definition data set to identify a set of data processing points; obtains data tables and field information from the data service server to generate a data asset list; analyzes the set of data processing points to generate a third definition data set, and performs similarity analysis to obtain a fourth definition data set; combines the structured data set and the data asset list to generate and output a dynamic data flow diagram. This enables real-time and accurate tracking of weak links and risk points in the data flow path, thereby providing timely warnings and ensuring data security.

[0051] Figure 2 FIG. 1 is a flow chart of a network data flow surveying and mapping method provided by an embodiment of the present invention. Figure 2 As shown, the method includes the following steps S110 to S170.

[0052] S110 , obtaining network data packets and data activity status within the application to obtain a first defined data set and a second defined data set.

[0053] In this embodiment, the first defined data set refers to the source IP address, source port, destination IP address, destination port, data packet, and device data of the network node. Device data refers to information related to the network device, such as the device model, operating system version, hardware specifications, device location, and network interface status. This information helps understand and analyze the source and destination of data packets.

[0054] The second defined data set refers to the data activities within the application that are embedded into the business application by generating a specific software package and collected within a preset time period.

[0055] In one embodiment, the aforementioned step S110 may include steps S111 to S113.

[0056] S111 , scanning a preset network address segment and mirroring switch traffic to obtain network data packets.

[0057] In this example, a network scanning tool is used to scan a specified IP address range. This involves probing all devices on the network and identifying their IP addresses, MAC addresses, and open ports. This helps determine active devices on the network and their communication ports.

[0058] Mirroring the switch's traffic to a monitoring port allows you to capture all network packets transmitted through the switch. This process typically involves configuring the switch to copy network traffic to a monitoring port for further analysis.

[0059] It can identify all devices and their open ports in the network, providing basic data for network topology and security analysis.

[0060] The captured network packets can be used for subsequent data analysis to help understand the patterns and sources of network traffic.

[0061] S112: Parse the network data packet to obtain the source IP address, source port, destination IP address, destination port, data packet and device data of the network node, and merge them to obtain a first defined data set.

[0062] In this embodiment, the captured network data packets are decoded to extract the source IP address, source port, destination IP address, destination port, data packet content, and device data. This requires the use of a suitable protocol parsing tool, such as Wireshark or tcpdump.

[0063] The extracted data are integrated to form a unified dataset (the first defined dataset), which contains the communication information of nodes in the network and detailed information of devices.

[0064] It provides detailed information about network traffic, including detailed data on communication sources and destinations, which helps with network analysis and troubleshooting. By integrating data from different sources, you can gain a more comprehensive understanding of network behavior and performance.

[0065] S113: Obtaining data activity within the application that is generated by embedding the specific software package into the business application and collected within a preset time period to obtain a second defined data set.

[0066] In this embodiment, specific software packages are embedded in the business applications, which collect data activities when the applications are running. These software packages can be specially designed agents or plug-ins for collecting real-time data within the applications.

[0067] Specifically, this system generates a specific software package (e.g., a JAR file) and embeds it into business applications to achieve automatic data collection. The application collection subunit is built using a specific programming language (e.g., Java) and uses bytecode injection technology to seamlessly integrate the collection logic into the application, enabling non-invasive collection of data activities within the application. The collection unit can collect data using protocols such as TCP and UDP, depending on the network environment.

[0068] Data activities within the application are collected within a preset time period, such as application requests, responses, user interactions, etc. These data are recorded and organized into a second defined data set.

[0069] Obtaining activity data within an application can help understand application performance, user behavior, and potential problems. By embedding monitoring tools in the application, data can be obtained in real time, helping to quickly respond and adjust application performance.

[0070] S120: Decrypt the specific traffic in the first defined data set and the second defined data set, and perform aggregation processing to obtain a structured data set.

[0071] In this embodiment, the structured data set refers to a data set representing nodes and the relationships between nodes.

[0072] In one embodiment, the aforementioned step S120 may include steps S121 - S122 .

[0073] S121. Decrypt specific traffic in the first defined data set and the second defined data set.

[0074] In this embodiment, in some cases, network traffic may be encrypted. In order to identify such encrypted traffic, the corresponding certificate is used to decrypt the specific traffic to ensure that all traffic data can be correctly parsed and identified.

[0075] Specifically, network traffic is sometimes encrypted, such as with SSL / TLS. In this case, the packet contents are encrypted, and reading them directly results in garbled text. It's necessary to identify which packets are encrypted and which are unencrypted. For encrypted network traffic, the corresponding decryption certificate (such as an SSL certificate or private key) is used to decrypt the data. This typically requires configuring a decryption proxy or a man-in-the-middle (MITM) tool, which can be inserted into the communication link to decrypt the transmitted data stream.

[0076] After decryption, the packet contents become readable, allowing for further parsing and analysis of detailed information within network traffic, such as the specific content of requests and responses.

[0077] Decrypted data makes all network traffic content readable, ensuring the comprehensiveness and accuracy of the data and avoiding situations where it cannot be identified or parsed due to encryption. Through decryption, traffic content can be analyzed more deeply, potential security threats or performance bottlenecks can be identified, and more accurate information can be provided for decision-making and optimization.

[0078] S122: Merge the decrypted first definition data set and the second definition data set, remove duplicate data, obtain a structured data set, and use the structured data set to draw a first-layer static network node and relationship diagram.

[0079] In this embodiment, the two definition data sets are merged and deduplicated to form a structured data set. These data are used to draw a graph of static nodes in the network and their relationships, laying the foundation for further analysis.

[0080] Specifically, the decrypted first-defined dataset (network packet data) and the second-defined dataset (in-application data activity) are integrated. This step involves aligning and merging data from different sources according to the same standards to form a comprehensive dataset. During the merging process, duplicate data records may be encountered. Duplicate data items need to be eliminated to ensure the accuracy and cleanliness of the dataset. The merged and deduplicated data is organized into a structured dataset, which typically includes nodes (such as devices and applications) and the relationships between them. Using the structured dataset, a first-level static network node and relationship diagram is drawn. This is part of network visualization, used to display the various nodes in the network and their relationships, such as the connections between devices and the flow of data.

[0081] Through integration and deduplication, a comprehensive and accurate view of network data is obtained, providing a reliable data foundation for network management and analysis; the drawn network node and relationship diagram can intuitively display the network topology, helping to identify key nodes and potential network path problems in the network; structured data sets and visual charts provide strong support for further network performance analysis, troubleshooting and security assessment.

[0082] S130: Analyze and judge the node data in the first defined data set to obtain a data processing point set.

[0083] In this embodiment, the data processing point set refers to a set consisting of all processing points that the data passes through in the network system.

[0084] Specifically, the DAVA (Data Activities Variety Analysis) algorithm is used to analyze the node data in the first defined dataset to identify all processing points that the data passes through in the network system, forming a data processing point set. All processing points identified, such as applications, servers, and databases, are aggregated and further data stream classification and feature analysis algorithms are used to calibrate and enhance the data set of data points and data streams.

[0085] The DAVA algorithm (Data Activities Variety Analysis) is a technique for analyzing and calculating the diversity of collected sample data. The algorithm is implemented as follows:

[0086] Data Record Set (DRS):

[0087] ;

[0088] Among them, DRS represents data record set, which is a set of data consisting of network data processing points and data flow feature records.

[0089] Data element (t): A specific data element used to describe network traffic, including the five-tuple information of the traffic, such as source IP address, destination IP address, protocol, source port, and destination port.

[0090] Set members (f):

[0091] dataAssest (data assets): represents various types or resources of data.

[0092] flowNodes: refers to the nodes in the data flow or the nodes in the data flow path.

[0093] nodeTimes (node time): time information related to data flow.

[0094] flowVarieties (data flow diversity attributes): Represents different types or categories of data flows or activity types.

[0095] The DAVA algorithm uses these parameters to perform detailed technical analysis on the collected sample data to calculate and understand the diversity of data activities.

[0096] S140: Obtain data tables and field information of a preset data service server to obtain a data asset list.

[0097] In this example, a data asset inventory is a systematic record that lists all data tables and their fields within an organization. It typically includes each table's name, field name, field data type, field description, and other relevant information. The purpose of a data asset inventory is to provide a comprehensive view to help manage, protect, and utilize data assets.

[0098] Specifically, the preset data service server within the organization is scanned to identify data tables and field information, and a data asset list is formed. The data assets in the data asset list are labeled to create a data fingerprint information table.

[0099] In this embodiment, a scan is performed on a pre-defined data service server within an organization. The purpose of the scan is to identify and collect information about all data tables and fields on the server. This includes listing detailed information such as the name of each data table, field name, field type, and field length.

[0100] Based on the scanned information, a list of all data tables and fields is generated. This list is called the data asset list. It systematically records all data resources on the data service server, forming a comprehensive view of data assets.

[0101] Label the data items in your data asset inventory. This involves assigning labels to each data table and field to describe its content, purpose, or classification. These labels can be predefined categories or customized based on business needs.

[0102] Through labeling, a data fingerprint table is created. This table records the tag information of each data asset, allowing for rapid identification and management of data assets. This fingerprint information helps effectively classify and manage data during data processing, analysis, and protection.

[0103] S150: Perform data flow analysis on the data processing point set to obtain a third definition data set, and perform similarity analysis and processing on the data records to obtain a fourth definition data set.

[0104] In this embodiment, the third defined data set refers to a data set formed by recording the time series information and data flow characteristics of each data processing point, including detailed records of each data at the data processing point and data flow trajectory information.

[0105] The fourth defined dataset is a dataset obtained by performing similarity analysis and processing on the data records in the third defined dataset. The cosine similarity algorithm is used to eliminate duplicate or similar data to ensure that the final dataset contains more unique records.

[0106] In one embodiment, the aforementioned step S150 may include steps S151 - S152 .

[0107] S151. Analyze the flow carrier, characteristics, and route direction of data at each data processing point in the data processing point set to record relevant time series information, thereby obtaining data records at each processing point to form a third defined data set.

[0108] In this embodiment, the data record includes various activity record data.

[0109] Specifically, data flow analysis is performed on the data processing point set to track the flow of data at each processing point. By analyzing the data flow path, characteristics, and timing information, a third defined data set can be formed.

[0110] The third definition data set provides detailed data flow information, which helps to accurately manage and monitor data flow.

[0111] S152: Use a cosine similarity algorithm to perform similarity analysis on the data records in the third definition data set one by one, and eliminate identical and similar data to obtain a fourth definition data set.

[0112] The cosine similarity algorithm is used to analyze the data records in the third defined data set. By eliminating the identical and similar data, a fourth defined data set is obtained, which contains a more accurate and unique set of data records.

[0113] Through similarity analysis, the fourth defined dataset reduces data duplication, improving data quality and accuracy. Eliminating duplicate data makes data analysis more efficient and avoids invalid data interference. Reducing redundant data helps optimize storage space and resource usage.

[0114] Specifically, the cosine similarity algorithm has the following calculation steps: First, each subset of the DRS is vectorized, for example, the PCA principal component analysis method is used to convert the data subset into a principal component space vector of a lower dimension:

[0115] ;

[0116] in, is the principal component space vector of the data subset, W is the weight matrix of the principal component, is the original data subset. Next, the principal component space vectors of each data subset are matched one by one for cosine similarity. The calculation formula is as follows: ;

[0117] in: and are the principal component space vectors of the two data subsets; Represents a vector and The dot product (inner product) of and Represents vectors and The Euclidean norm of (i.e., the length of the vector).

[0118] The cosine similarity value ranges from -1 to 1. When the angle between two vectors is 0 degrees, the cosine similarity is 1, indicating that they are exactly the same, that is, the two nodes or data streams have the same content, which is redundant data; when the angle is 180 degrees, the cosine similarity is -1, indicating that they are completely opposite. If the data set represents the content of the data stream, it represents the two directions of data flow. If the data set represents the data node, it means that there is no connection between the two and they should be represented as two different nodes in different time dimensions; when the angle is 90 degrees, the cosine similarity is 0, indicating that they are orthogonal, that is, there is no correlation, and they represent two different data processing nodes or data stream contents.

[0119] Each subset of the data set with network data processing points and data flow feature records is vectorized, and the principal component space vectors of each data subset are matched one by one with cosine similarity calculations to obtain accurate data processing points and data flow results.

[0120] S160: Generate a dynamic data flow diagram based on the fourth definition data set, combined with the structure data set and the data asset list.

[0121] In this embodiment, the directed lines of the data flow directed graph express attribute information of the volume and sensitivity of the data flow through the shape of the lines, wherein the thickness of the lines represents the size of the data flow, and the color of the lines represents the sensitivity of the data.

[0122] Specifically, a directed graph is created using the fourth definition data set, the structure data set, and the data asset list, wherein nodes represent data processing points or data assets.

[0123] Directed lines: Represent the path of data flow. Line thickness and color are used to express data flow and sensitivity. Line thickness: The thickness of the line indicates the amount of data flow. Thick lines indicate high flow, thin lines indicate low flow. Line chroma: The color of the line indicates the sensitivity of the data. Dark colors indicate sensitive data, while light colors indicate non-sensitive data.

[0124] Specifically, graphical modeling includes:

[0125] Node definition: Data processing points or data assets are used as nodes in the graph.

[0126] Edge definition: A directed edge represents a data flow path. Edge attributes include: Thickness: Indicates the size of the data flow. Color: Indicates the sensitivity of the data.

[0127] Graph drawing involves using graph drawing tools or visualization libraries (such as D3.js or Graphviz) to generate dynamic data flow graphs. Setting edge thickness and color attributes can be adjusted based on data flow and sensitivity.

[0128] Dynamic updating includes: implementing a dynamic updating mechanism for graphics so as to automatically adjust the graphics display as the data flow changes.

[0129] By displaying data flow and sensitivity, you can optimize data management and protection strategies and ensure that highly sensitive data is properly protected. Dynamically updated data flow diagrams help monitor data flow and changes in real time, improving response speed and decision-making efficiency. They help decision-makers identify bottlenecks and risk points in data flow, supporting optimized data processing and resource allocation.

[0130] S170: Output the dynamic data flow graph.

[0131] The dynamic data flow diagram is output to the terminal for display, which can be used for early warning and correction of problems in data flow.

[0132] The method of this embodiment can use network traffic analysis and active sniffing to detect the data flow situation in the network, including the source of the data, the processing node of the data flow, and the data transmission destination. It can then further use analysis technologies such as DAVA to automatically analyze the path and flow direction of the data in the network, combine with time series data, automatically draw data flow diagrams at each time point, and support dynamic display, thereby achieving the goal of allowing users to understand the data flow security status in a timely and intuitive manner.

[0133] The above-mentioned network data flow mapping method obtains and decrypts network data packets and data activities within applications, aggregates them into structured data sets, forms basic data sets and data processing point sets, analyzes the data processing point sets, and combines them with the data asset list to generate a data flow diagram; performs similarity analysis on data flow paths to determine potential weak links and risk points in data flow; based on the dynamic data flow diagram, provides real-time warnings and corrections to problems in data flow to ensure data security; and achieves real-time and accurate tracking of weak links and risk points in data flow paths, thereby achieving timely warnings and ensuring data security.

[0134] Figure 3 FIG. 3 is a schematic block diagram of a network data flow mapping device 300 provided by an embodiment of the present invention. Figure 3 As shown, corresponding to the above network data flow mapping method, the present invention also provides a network data flow mapping device 300. The network data flow mapping device 300 includes a unit for executing the above network data flow mapping method, and the device can be configured in a server. Figure 3 The network data flow mapping device 300 includes a data packet acquisition unit 301, a processing unit 302, a judgment unit 303, an information acquisition unit 304, a data analysis unit 305, a flow map generation unit 306 and an output unit 307.

[0135] The data packet acquisition unit 301 is used to acquire network data packets and data activities within the application to obtain a first definition data set and a second definition data set; the processing unit 302 is used to decrypt specific traffic in the first definition data set and the second definition data set, and perform aggregation processing to obtain a structured data set; the analysis unit 303 is used to analyze the node data in the first definition data set to obtain a data processing point set; the information acquisition unit 304 is used to obtain data tables and field information of a preset data service server to obtain a data asset list; the data analysis unit 305 is used to perform data flow analysis on the data processing point set to obtain a third definition data set, and perform similarity analysis and processing of data records to obtain a fourth definition data set; the flow graph generation unit 306 is used to generate a dynamic data flow graph based on the fourth definition data set, combined with the structured data set and the data asset list; the output unit 307 is used to output the dynamic data flow graph.

[0136] In one embodiment, the data packet acquiring unit 301 includes:

[0137] The network data packet acquisition subunit is used to scan the preset network address segment and mirror the switch traffic to obtain the network data packet; the parsing subunit is used to parse the network data packet to obtain the source IP address, source port, destination IP address, destination port, data packet and device data of the network node, and merge them to obtain the first defined data set; the data acquisition subunit is used to obtain the data activity within the application collected within a preset time period by generating a specific software package embedded in the business application to obtain the second defined data set.

[0138] In one embodiment, the processing unit 302 includes:

[0139] The decryption subunit is used to decrypt the specific traffic in the first definition data set and the second definition data set; the merging and elimination subunit is used to merge the decrypted first definition data set and the second definition data set, eliminate duplicate data, obtain a structured data set, and use the structured data set to draw the first-layer static network node and relationship diagram.

[0140] In one embodiment, the analysis and judgment unit 303 is configured to analyze the node data in the first defined data set using a DAVA algorithm to determine all processing points that the data passes through in the network system to form a data processing point set.

[0141] In one embodiment, the information acquisition unit 304 is used to scan the preset data service server within the organization to identify data tables and field information, form a data asset list, and label the data assets in the data asset list to create a data fingerprint information table.

[0142] In one embodiment, the data analysis unit 305 includes:

[0143] an analysis subunit, configured to analyze the flow carrier, characteristics, and route direction of the data at each data processing point in the set of data processing points, to record relevant timing information, to obtain data records at each processing point, and to form a third defined data set; and a similarity processing subunit, configured to perform similarity analysis on the data records in the third defined data set one by one using a cosine similarity algorithm, to eliminate identical and similar data, and to obtain a fourth defined data set.

[0144] It should be noted that those skilled in the art can clearly understand that the specific implementation process of the above-mentioned network data flow surveying and mapping device 300 and each unit can refer to the corresponding description in the aforementioned method embodiment. For the convenience and brevity of description, it will not be repeated here.

[0145] The network data flow surveying and mapping device 300 can be implemented in the form of a computer program. The computer program can be used in Figure 4 Runs on the computer equipment shown.

[0146] See also Figure 4 , Figure 4 1 is a schematic block diagram of a computer device provided in an embodiment of the present application. The computer device 500 may be a server, wherein the server may be an independent server or a server cluster composed of multiple servers.

[0147] See Figure 4 The computer device 500 includes a processor 502 , a memory, and a network interface 505 connected via a system bus 501 , wherein the memory may include a non-volatile storage medium 503 and an internal memory 504 .

[0148] The non-volatile storage medium 503 can store an operating system 5031 and a computer program 5032. The computer program 5032 includes program instructions, which, when executed, can enable the processor 502 to execute a network data flow mapping method.

[0149] The processor 502 is used to provide computing and control capabilities to support the operation of the entire computer device 500.

[0150] The internal memory 504 provides an environment for the operation of the computer program 5032 in the non-volatile storage medium 503. When the computer program 5032 is executed by the processor 502, the processor 502 can execute a network data flow surveying and mapping method.

[0151] The network interface 505 is used to communicate with other devices through the network. Figure 4 The structure shown in the figure is merely a block diagram of a portion of the structure related to the solution of the present application, and does not constitute a limitation on the computer device 500 to which the solution of the present application is applied. The specific computer device 500 may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.

[0152] The processor 502 is configured to execute a computer program 5032 stored in the memory to implement the following steps:

[0153] Obtain network data packets and data activities within the application to obtain a first definition data set and a second definition data set; decrypt specific traffic in the first definition data set and the second definition data set, and perform aggregation processing to obtain a structured data set; analyze and judge the node data in the first definition data set to obtain a data processing point set; obtain data tables and field information of a preset data service server to obtain a data asset list; perform data flow analysis on the data processing point set to obtain a third definition data set, and perform similarity analysis and processing of data records to obtain a fourth definition data set; generate a dynamic data flow graph based on the fourth definition data set, combined with the structured data set and the data asset list; and output the dynamic data flow graph.

[0154] The directed lines of the data flow directed graph express attribute information of the volume and sensitivity of the data flow through the shape of the lines, wherein the thickness of the lines represents the size of the data flow, and the color of the lines represents the sensitivity of the data.

[0155] In one embodiment, when the processor 502 implements the step of obtaining network data packets and data activity within the application to obtain the first definition data set and the second definition data set, the processor 502 specifically implements the following steps:

[0156] Scanning a preset network address segment and mirroring switch traffic to obtain network data packets; parsing the network data packets to obtain the source IP address, source port, destination IP address, destination port, data packet and device data of the network node, and merging them to obtain a first defined data set; obtaining the data activity within the application that is embedded in the business application by generating a specific software package and collected within a preset time period to obtain a second defined data set.

[0157] In one embodiment, when the processor 502 implements the step of decrypting the specific traffic in the first defined data set and the second defined data set and performing aggregation processing to obtain the structured data set, the processor 502 specifically implements the following steps:

[0158] Decrypt the specific traffic in the first definition data set and the second definition data set; merge the decrypted first definition data set and the second definition data set, eliminate duplicate data, obtain a structured data set, and use the structured data set to draw a first-layer static network node and relationship diagram.

[0159] In one embodiment, when the processor 502 implements the step of analyzing the node data in the first definition data set to obtain a data processing point set, the processor 502 specifically implements the following steps:

[0160] The DAVA algorithm is used to analyze the node data in the first defined data set to determine all processing points that the data passes through in the network system to form a data processing point set.

[0161] In one embodiment, when the processor 502 implements the step of obtaining the data table and field information of the preset data service server to obtain the data asset list, it specifically implements the following steps:

[0162] Scan the preset data service servers within the organization to identify data tables and field information, and form a data asset list. Label the data assets in the data asset list and create a data fingerprint information table.

[0163] In one embodiment, when the processor 502 performs the steps of performing data flow analysis on the data processing point set to obtain a third defined data set, and performing similarity analysis and processing on the data records to obtain a fourth defined data set, the processor 502 specifically performs the following steps:

[0164] Analyze the flow carrier, characteristics, and route direction of the data at each data processing point in the data processing point set to record relevant timing information, so as to obtain the data record at each processing point and form a third defined data set; use the cosine similarity algorithm to perform similarity analysis on the data records in the third defined data set one by one, eliminate the same and similar data, and obtain a fourth defined data set.

[0165] It should be understood that in the embodiment of the present application, the processor 502 may be a central processing unit (CPU) 302. The processor 502 may also be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor.

[0166] Those skilled in the art will appreciate that all or part of the steps in the method of the above-described embodiment can be implemented by instructing the relevant hardware through a computer program. The computer program includes program instructions, which can be stored in a storage medium that is computer-readable. The program instructions are executed by at least one processor in the computer system to implement the steps in the method of the above-described embodiment.

[0167] Therefore, the present invention also provides a storage medium. The storage medium may be a computer-readable storage medium. The storage medium stores a computer program, wherein when the computer program is executed by a processor, the processor performs the following steps:

[0168] Obtain network data packets and data activities within the application to obtain a first definition data set and a second definition data set; decrypt specific traffic in the first definition data set and the second definition data set, and perform aggregation processing to obtain a structured data set; analyze and judge the node data in the first definition data set to obtain a data processing point set; obtain data tables and field information of a preset data service server to obtain a data asset list; perform data flow analysis on the data processing point set to obtain a third definition data set, and perform similarity analysis and processing of data records to obtain a fourth definition data set; generate a dynamic data flow graph based on the fourth definition data set, combined with the structured data set and the data asset list; and output the dynamic data flow graph.

[0169] The directed lines of the data flow directed graph express attribute information of the volume and sensitivity of the data flow through the shape of the lines, wherein the thickness of the lines represents the size of the data flow, and the color of the lines represents the sensitivity of the data.

[0170] In one embodiment, when the processor executes the computer program to implement the step of obtaining network data packets and data activity within the application to obtain the first defined data set and the second defined data set, the processor specifically implements the following steps:

[0171] Scanning a preset network address segment and mirroring switch traffic to obtain network data packets; parsing the network data packets to obtain the source IP address, source port, destination IP address, destination port, data packet and device data of the network node, and merging them to obtain a first defined data set; obtaining the data activity within the application that is embedded in the business application by generating a specific software package and collected within a preset time period to obtain a second defined data set.

[0172] In one embodiment, when the processor executes the computer program to implement the step of decrypting the specific traffic in the first defined data set and the second defined data set and performing aggregation processing to obtain a structured data set, the processor specifically implements the following steps:

[0173] Decrypt the specific traffic in the first definition data set and the second definition data set; merge the decrypted first definition data set and the second definition data set, eliminate duplicate data, obtain a structured data set, and use the structured data set to draw a first-layer static network node and relationship diagram.

[0174] In one embodiment, when the processor executes the computer program to implement the step of analyzing the node data in the first definition data set to obtain a data processing point set, the processor specifically implements the following steps:

[0175] The DAVA algorithm is used to analyze the node data in the first defined data set to determine all processing points that the data passes through in the network system to form a data processing point set.

[0176] In one embodiment, when the processor executes the computer program to implement the step of obtaining data table and field information of a preset data service server to obtain a data asset list, the processor specifically implements the following steps:

[0177] Scan the preset data service servers within the organization to identify data tables and field information, and form a data asset list. Label the data assets in the data asset list and create a data fingerprint information table.

[0178] In one embodiment, when the processor executes the computer program to implement the steps of performing data flow analysis on the data processing point set to obtain a third defined data set, and performing similarity analysis and processing on the data records to obtain a fourth defined data set, the processor specifically implements the following steps:

[0179] Analyze the flow carrier, characteristics, and route direction of the data at each data processing point in the data processing point set to record relevant timing information, so as to obtain the data record at each processing point and form a third defined data set; use the cosine similarity algorithm to perform similarity analysis on the data records in the third defined data set one by one, eliminate the same and similar data, and obtain a fourth defined data set.

[0180] The storage medium may be any computer-readable storage medium that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a magnetic disk, or an optical disk.

[0181] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the composition and steps of each example according to function. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.

[0182] In the several embodiments provided herein, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the various units is merely a logical functional division, and actual implementation may employ other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be omitted or not implemented.

[0183] The steps in the method of the embodiment of the present invention may be adjusted in order, combined, or deleted as needed. The units in the apparatus of the embodiment of the present invention may be combined, divided, or deleted as needed. In addition, the functional units in the various embodiments of the present invention may be integrated into a single processing unit 302, each unit may exist physically separately, or two or more units may be integrated into a single unit.

[0184] If this integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product, stored in a storage medium, includes instructions for enabling a computer device (such as a personal computer, terminal, or network device) to execute all or part of the steps of the method described in various embodiments of the present invention.

[0185] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and such modifications or substitutions are intended to be within the scope of protection of the present invention. Therefore, the scope of protection of the present invention shall be subject to the scope of protection of the claims.

Claims

1. A network data flow surveying and mapping method, characterized in that: include: Acquire network data packets and data activity within the application to obtain a first defined data set and a second defined data set; Decrypting the specific traffic in the first defined data set and the second defined data set, and performing aggregation processing to obtain a structured data set; Analyzing and judging the node data in the first defined data set to obtain a data processing point set; Obtain data table and field information from the preset data service server to obtain a data asset list; Performing data flow analysis on the data processing point set to obtain a third defined data set, and performing similarity analysis and processing on the data records to obtain a fourth defined data set; Generate a dynamic data flow diagram based on the fourth definition data set, combined with the structure data set and the data asset list; Outputting the dynamic data flow graph; The acquiring of network data packets and data activity within the application to obtain the first defined data set and the second defined data set includes: Scan the preset network address segment and mirror the switch traffic to obtain network data packets; Parsing the network data packet to obtain a source IP address, a source port, a destination IP address, a destination port, data packets, and device data of a network node, and merging the data packets to obtain a first defined data set; Obtaining data activity within the application, which is embedded in the business application by generating a specific software package and collected within a preset time period, to obtain a second defined data set; The performing data flow analysis on the data processing point set to obtain a third defined data set, and performing similarity analysis and processing on the data records to obtain a fourth defined data set, includes: Analyzing the flow carrier, characteristics, and route direction of data at each data processing point in the set of data processing points to record relevant time series information, thereby obtaining data records at each processing point to form a third defined data set; Performing similarity analysis on each of the data records in the third defined data set using a cosine similarity algorithm, and eliminating identical and similar data to obtain a fourth defined data set; The directed lines of the data flow graph express the attribute information of the volume and sensitivity of the data flow through the shape of the lines, wherein the thickness of the lines represents the size of the data flow, and the color of the lines represents the sensitivity of the data; The analyzing and judging the node data in the first defined data set to obtain a data processing point set includes: Analyzing the node data in the first defined data set using the DAVA algorithm to determine all processing points that the data passes through in the network system to form a data processing point set; DRS = {t│f=〈dataAssest, flowNodes, nodeTimes, flowVarieties〉} Among them, DRS stands for Data Record Set, which is a set of data consisting of network data processing points and data flow feature records; Data element t identifies a specific data element used to describe network traffic, including the five-tuple information of the traffic: source IP address, destination IP address, protocol, source port, and destination port; The set members f include: dataAssest represents various types or resources of data; flowNodes refers to nodes in a data flow or nodes in a data flow path; nodeTimes involves the time information of the data stream; flowVarieties represent different types or categories of data flows or activity types.

2. The network data flow surveying and mapping method according to claim 1, characterized in that: The decrypting and aggregating the specific traffic in the first defined data set and the second defined data set to obtain a structured data set includes: decrypting specific traffic in the first defined data set and the second defined data set; The decrypted first definition data set and the second definition data set are merged to remove duplicate data to obtain a structured data set, and the structured data set is used to draw a first-layer static network node and relationship diagram.

3. The network data flow surveying and mapping method according to claim 1, characterized in that: The step of obtaining data tables and field information of a preset data service server to obtain a data asset list includes: Scan the preset data service servers within the organization to identify data tables and field information, and form a data asset list. Label the data assets in the data asset list and create a data fingerprint information table.

4. A network data flow surveying and mapping device, wherein the device uses the network data flow surveying and mapping method according to any one of claims 1 to 3, characterized in that: include: A data packet acquisition unit, configured to acquire network data packets and data activity within an application to obtain a first defined data set and a second defined data set; a processing unit, configured to decrypt the specific traffic in the first defined data set and the second defined data set, and perform aggregation processing to obtain a structured data set; An analysis and judgment unit, configured to analyze and judge the node data in the first defined data set to obtain a data processing point set; An information acquisition unit is used to acquire data table and field information of a preset data service server to obtain a data asset list; a data analysis unit, configured to perform data flow analysis on the data processing point set to obtain a third defined data set, and perform similarity analysis and processing on the data records to obtain a fourth defined data set; a flow diagram generating unit, configured to generate a dynamic data flow diagram based on the fourth definition data set, in combination with the structure data set and the data asset list; An output unit, configured to output the dynamic data flow graph; The data packet acquisition unit includes: The network data packet acquisition subunit is used to scan the preset network address segment and mirror the switch traffic to obtain the network data packet; the parsing subunit is used to parse the network data packet to obtain the source IP address, source port, destination IP address, destination port, data packet and device data of the network node, and merge them to obtain a first defined data set; the data acquisition subunit is used to obtain the data activity status within the application collected within a preset time period by generating a specific software package embedded in the business application to obtain a second defined data set; The data analysis unit includes: an analysis subunit, configured to analyze the flow carrier, characteristics, and route direction of the data at each data processing point in the set of data processing points, to record relevant time series information, to obtain data records at each processing point, and to form a third defined data set; a similarity processing subunit, configured to perform similarity analysis on each data record in the third defined data set using a cosine similarity algorithm, to eliminate identical and similar data, and to obtain a fourth defined data set; The directed lines of the data flow diagram express attribute information of the volume and sensitivity of data flow through the shape of the lines, wherein the thickness of the lines represents the size of the data flow, and the color of the lines represents the sensitivity of the data.

5. A computer device, characterized in that: The computer device includes a memory and a processor, the memory stores a computer program, and the processor implements the method according to any one of claims 1 to 3 when executing the computer program.

6. A storage medium, characterized in that The storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 3 is implemented.

Citation Information

Patent Citations

  • Data flow path drawing method and device, storage medium and electronic equipment

    CN113691423A

  • Data asset surveying and mapping method and system based on mysql protocol analysis

    CN116489044A