An intelligent data acquisition and management system and method for big data.
By combining data from information big data warehouses and switching equipment, an intelligent network security threat collection model is constructed, which solves the problems of insufficient adaptability and response speed of existing network security systems in dynamic environments. It realizes intelligent management and real-time monitoring of network threats and adapts to emerging technology environments such as cloud computing, big data and the Internet of Things.
Patent Information
- Application Number
- CN202510175220.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-18
- Publication Date
- 2025-11-14
- Estimated Expiration
- 2045-02-18
AI Technical Summary
Existing network security systems struggle to monitor real-time topology changes in dynamic network environments and lack the ability to intelligently analyze diverse security events. This results in insufficient adaptability and response speed of protection systems, making them unable to effectively address the challenges posed by emerging technology environments such as cloud computing, big data, and the Internet of Things.
By acquiring information from a big data warehouse, extracting and maintaining data dictionary features, and combining this with data from switch devices to perform network topology analysis and traffic statistics, an intelligent network security threat acquisition model is constructed to achieve dynamic monitoring and intelligent management of network security threats.
It enhances the flexibility and real-time performance of network security protection systems, enabling timely identification and response to security threats in complex network environments. It also strengthens the system's adaptability and scalability, adapting to the challenges of emerging technology environments.
Smart Images

Figure CN120034375B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of network security technology, and in particular to an intelligent information big data collection and management system and method. Background Technology
[0002] Existing methods largely rely on static network topology diagrams, making it difficult to monitor real-time topology changes in dynamic network environments. Their ability to detect network transmission anomalies is limited, easily overlooking dynamic network topology changes and potential security threats, thus reducing the accuracy of network security protection. Traditional methods largely rely on manual rules for network security threat detection, resulting in relatively simple response mechanisms that struggle to respond promptly and accurately to diverse security events. Furthermore, they lack intelligent analysis capabilities for different threat types, leading to poor adaptability of the protection system. Existing information collection and management systems are typically based on fixed frameworks and processes, lacking high scalability and adaptability, and unable to effectively address the challenges of emerging technology environments (such as cloud computing, big data, and the Internet of Things). Insufficient system flexibility prevents rapid adjustments to response strategies in the face of constantly changing network environments, reducing the system's real-time performance and adaptability. Summary of the Invention
[0003] Therefore, it is necessary for the present invention to provide an intelligent information big data collection and management system and method to solve at least one of the above-mentioned technical problems.
[0004] To achieve the above objectives, an intelligent data collection and management method for big data includes the following steps:
[0005] Step S1: Obtain the information big data warehouse, and extract data dictionary features based on the information big data warehouse to obtain data dictionary data; maintain the data dictionary data to obtain data dictionary maintenance data.
[0006] Step S2: Obtain switch device data and perform network topology analysis based on the switch device data to obtain switch device network topology data; perform transmission anomaly analysis based on the switch device network topology data to obtain network transmission anomaly topology data.
[0007] Step S3: Perform traffic statistics based on the switch device data to obtain switch traffic data; perform network security threat analysis based on the switch traffic data to obtain network security threat data.
[0008] Step S4: Based on the network security threat data, classify the abnormal network transmission topology to obtain network security threat data topology data;
[0009] Step S5: Construct a cybersecurity threat intelligent collection model based on the data dictionary maintenance data and the cybersecurity threat data topology structure data to obtain the cybersecurity threat intelligent collection model; manage cybersecurity threats in the information big data warehouse based on the cybersecurity threat intelligent collection model to obtain cybersecurity threat management data.
[0010] This invention effectively manages data structure and features by acquiring and maintaining a big data warehouse and extracting and maintaining data dictionary features, thereby improving data understandability and consistency. Based on this, network topology analysis and transmission anomaly analysis using the big data warehouse allow for dynamic monitoring of topology changes and abnormal transmissions in the network environment, effectively identifying potential network problems and overcoming the limitations of traditional static topology maps in adapting to dynamic environments. By performing traffic statistics and network security threat analysis on switch device data, anomalies in network traffic can be comprehensively identified and assessed, and security threats can be detected promptly, providing a strong basis for implementing protective measures. Furthermore, by classifying network security threat data into topological structures, different types of security threats and their impact range can be accurately located, ensuring more targeted security protection. The intelligent acquisition model constructed by combining data dictionary maintenance data and network security threat topology data enables intelligent security threat identification and response, thereby improving the system's adaptability and real-time response capabilities. Finally, managing network security threats in the big data warehouse based on the intelligent acquisition model allows for continuous optimization of protection strategies, ensuring the system maintains high-efficiency protection capabilities in the face of complex network environments and emerging threats. This method has extremely high scalability and adaptability, and can effectively cope with the challenges of emerging technology environments such as cloud computing, big data, and the Internet of Things. It improves the flexibility and real-time performance of network security protection systems and solves many limitations of traditional methods in dynamic environments.
[0011] Optionally, step S1 specifically includes:
[0012] Step S11: Obtain the information big data warehouse, and extract data dictionary features based on the information big data warehouse to obtain data dictionary data;
[0013] Step S12: Perform field standardization on the data dictionary data to obtain field-standardized data;
[0014] Step S13: Perform encoding mapping based on the data dictionary data to obtain the standard mapping code value for the digital warehouse;
[0015] Step S14: Perform dictionary maintenance based on the standardized field data and the standard mapping code value of the digital warehouse to obtain data dictionary maintenance data.
[0016] This invention provides a unified structural framework for data by acquiring information from a big data warehouse and extracting features from a data dictionary, ensuring that data from different sources and types can be managed and analyzed uniformly. By standardizing data dictionary fields, the problem of inconsistent data formats and standards is solved, enabling more efficient unified calculation and analysis in subsequent processing, thereby enhancing the overall system synergy and data consistency. Encoding and mapping the data dictionary effectively achieves standardized conversion of various types of data, optimizing the accuracy and efficiency of data processing and ensuring that data is mapped more accurately to the digital warehouse, laying the foundation for subsequent data processing and analysis. Further dictionary maintenance of standardized data and mapping code values ensures continuous updating and improvement of the data dictionary, enabling the entire system to quickly adapt to and handle new data problems in the face of constantly changing network environments, improving the flexibility of data processing and the adaptability of the system. The implementation of this series of steps breaks through the limitations of traditional static topology diagrams and manual rules, better supporting real-time data monitoring and security threat detection in dynamic environments, improving the system's real-time performance, adaptability, and response speed, and providing strong technical support for achieving more intelligent and flexible network security protection.
[0017] Optionally, step S2 specifically includes:
[0018] Step S21: Obtain switch device data and perform node identification based on the switch device data to obtain switch node data;
[0019] Step S22: Extract port connection features based on switch device data to obtain port connection data;
[0020] Step S23: Establish a primary topology based on switch node data and port connection data to obtain primary topology data;
[0021] Step S24: Redundant path identification is performed on the primary topology data to obtain redundant path data;
[0022] Step S25: Loop path identification is performed on the primary topology data to obtain loop path data;
[0023] Step S26: Construct the network topology of the switch device based on the loop path data and redundant path data, thereby obtaining the network topology data of the switch device.
[0024] Step S27: Perform transmission anomaly analysis based on the network topology data of the switch device to obtain network transmission anomaly topology data.
[0025] This invention, through the acquisition and node identification of switch device data, can efficiently identify the location and connection relationships of each switch node in the network, laying the foundation for subsequent network topology construction. Furthermore, by extracting port connection features from the switch device data, it can accurately capture the connection status between switches in the network, providing strong data support for constructing an accurate network topology map. Based on node and port connection data, a preliminary topology structure can be constructed, providing basic information for further identification of redundant and loop paths, enhancing the accuracy and scalability of the topology structure. The identification of redundant and loop paths not only ensures high availability and stability in the network but also provides more detailed and accurate topology information for subsequent network transmission anomaly analysis. Combining loop and redundant path data allows for a more accurate construction of the complete network topology structure, enabling early identification and avoidance of potential anomalies and bottlenecks, thereby improving the system's ability to detect network anomalies. By performing transmission anomaly analysis based on a complete topology structure, abnormal phenomena in network transmission can be detected in a timely manner, providing real-time data support for security protection in dynamic network environments and effectively avoiding the shortcomings of traditional static topology maps that cannot reflect network changes in real time, thus enhancing the accuracy and flexibility of network security protection. Furthermore, this series of steps, through efficient topology construction and anomaly analysis, can provide strong support for addressing the challenges of emerging technology environments (such as cloud computing, big data, and the Internet of Things), improve the system's real-time performance, adaptability, and intelligent processing capabilities, and ensure that it can adjust and respond to diverse security threats in a timely manner when facing a constantly changing network environment.
[0026] Optionally, step S27 specifically includes:
[0027] Step S271: Identify switch port faults based on the network topology data of the switch device to obtain switch port fault data;
[0028] Step S272: Calculate the bandwidth utilization of the switch port fault data to obtain the bandwidth utilization rate;
[0029] Step S273: Perform numerical statistics based on bandwidth utilization to obtain high bandwidth utilization;
[0030] Step S274: Analyze network performance anomalies based on the network topology data of the switch device according to the high bandwidth utilization, thereby obtaining network performance anomaly data;
[0031] Step S275: Perform link health anomaly detection on the network topology data of the switch device to obtain link health anomaly data;
[0032] Step S276: Based on the abnormal link health data and abnormal network performance data, perform network transmission abnormal topology analysis on the network topology of the switch device to obtain abnormal network transmission topology data.
[0033] This invention identifies switch port faults by analyzing network topology data from switch devices, enabling timely detection of port failures and providing first-hand data support for subsequent network performance optimization. Bandwidth utilization calculation comprehensively assesses network resource usage, providing effective early warnings of potential bandwidth bottlenecks and offering more accurate data for network performance analysis. Numerical statistics based on bandwidth utilization quickly identify nodes with high bandwidth utilization, further uncovering potential issues of excessive network load and providing a reference for anomaly detection. Utilizing high bandwidth utilization data for network performance anomaly analysis helps identify performance bottlenecks caused by high loads and provides effective directions for system optimization. Link health anomaly detection monitors the health of network links in real time, detecting link anomalies and providing early warnings, preventing network interruptions or performance degradation due to link failures. Furthermore, combining link health anomaly data with network performance anomaly data for network transmission anomaly topology analysis allows for in-depth analysis and revelation of network transmission faults, thus providing strong support for stable network operation. Through the effective implementation of the above steps, the system can not only monitor network performance and link health status in real time, but also identify and analyze network transmission anomalies in a timely manner, thereby improving the accuracy and response speed of network security protection. It avoids the limitations of traditional static methods in the face of dynamic network environments and enhances the system's flexibility, adaptability and ability to cope with emerging technology challenges.
[0034] Optionally, step S3 specifically includes:
[0035] Step S31: Perform traffic statistics based on the switch device data to obtain switch traffic data;
[0036] Step S32: Detect phishing attacks based on switch traffic data to obtain phishing attack data;
[0037] Step S33: Perform traffic anomaly monitoring based on switch traffic data to obtain traffic anomaly data;
[0038] Step S34: Integrate cybersecurity threat data based on phishing attack data and abnormal traffic data to obtain cybersecurity threat data.
[0039] This invention, by performing traffic statistics on switch device data, can comprehensively understand the distribution and usage of network traffic, providing crucial data support for subsequent security analysis. Based on this, phishing attack detection using traffic data can promptly identify abnormal traffic patterns and potential phishing attack threats, thus providing early warnings for preventing such attacks. Further monitoring of switch traffic anomalies allows for dynamic monitoring of traffic fluctuations, timely detection of traffic anomalies such as bandwidth abuse or malicious traffic, thereby reducing the impact of potential risks on network security. By integrating phishing attack data with traffic anomaly data, more accurate network security threat identification can be achieved, improving the detection accuracy and response speed of the network protection system. Overall, the system can efficiently monitor network traffic and detect anomalies. Combining phishing attack detection and traffic anomaly monitoring, it captures various potential threats in the network in real time and provides comprehensive and dynamic network security threat identification and protection through data integration. This overcomes the limitations of traditional methods where static network topology maps are difficult to adapt to dynamic network environments, enhancing the intelligence, flexibility, and adaptability of the network security protection system.
[0040] Optionally, step S32 specifically includes:
[0041] Step S321: Perform DNS request analysis based on switch traffic data to obtain DNS request data;
[0042] Step S322: Perform request volume statistics on the Domain Name System (DNS) request data to obtain high request volume data for the DNS;
[0043] Step S323: Identify suspicious external IP addresses based on high request volume data from the Domain Name System to obtain suspicious external IP address data;
[0044] Step S324: Perform short link detection based on switch traffic data to obtain short link data;
[0045] Step S325: Identify abnormal Uniform Resource Locators (URLs) in the short link data to obtain abnormal URL data;
[0046] Step S326: Based on the abnormal Uniform Resource Locator (URI) data and short link data, determine the phishing attack data of the suspicious external IP address data.
[0047] This invention analyzes Domain Name System (DNS) requests from switch traffic data to gain insights into domain name resolution behavior patterns within the network, providing foundational data for identifying potential anomaly requests. Furthermore, statistical analysis of DNS request volume helps identify peak periods of abnormal traffic, revealing potential malicious activities such as DDoS attacks or domain name abuse. This analysis process can accurately pinpoint suspicious external IP addresses with high request volumes, providing crucial clues for subsequent security protection and identifying attack sources or malicious manipulation. Additionally, short link detection of switch traffic data effectively identifies malicious short links associated with phishing attacks or other network attacks. Further analysis of short link data using anomalous Uniform Resource Locators (URLs) reveals abnormal URL patterns, providing clues for tracking malicious activities. Combining anomalous URLs with short link data allows for more accurate confirmation of the association between suspicious external IP addresses and phishing attacks, ultimately determining whether phishing attacks exist within the network. Overall, this process enhances the in-depth analysis of network traffic, helps identify complex security threats, especially in dynamically changing network environments, enables real-time responses to emerging security threats, and improves the system's intelligence, flexibility, and ability to protect against diverse attacks.
[0048] Optionally, step S33 specifically includes:
[0049] Step S331: Preprocess the switch traffic data to obtain preprocessed switch traffic data;
[0050] Step S332: Establish a traffic pattern baseline on the preprocessed switch traffic data to obtain switch traffic pattern baseline data;
[0051] Step S333: Perform real-time anomaly monitoring on switch traffic data based on the switch traffic pattern baseline data to obtain real-time anomaly switch traffic data.
[0052] Step S334: Classify the anomaly types based on the real-time abnormal switch traffic data to obtain bandwidth abuse anomaly data and network worm anomaly data;
[0053] Step S335: Perform correlation analysis based on bandwidth abuse anomaly data and network worm anomaly data to obtain potential security threat data;
[0054] Step S336: Perform traffic anomaly detection on the switch traffic data based on potential security threat data to obtain traffic anomaly data.
[0055] This invention preprocesses switch traffic data to remove noise and standardize the data, providing a more accurate and clear data foundation for subsequent analysis. Furthermore, establishing a traffic pattern baseline on the preprocessed traffic data helps identify normal network traffic patterns, laying the groundwork for anomaly detection. Real-time anomaly monitoring based on the traffic pattern baseline can promptly capture abnormal changes in traffic, identify network attacks or performance issues, and generate real-time anomaly traffic data. By classifying real-time anomaly traffic data into different types, different types of anomalies, such as bandwidth abuse or network worm attacks, can be clearly distinguished, providing more detailed information for subsequent defense measures. Correlation analysis combining bandwidth abuse anomaly data and network worm anomaly data can reveal potential security threats, helping security teams to promptly detect attacks that have not yet manifested in the network. Finally, traffic anomaly detection based on potential security threat data can further confirm the presence of abnormal traffic in the network, ensuring that the security protection system can more comprehensively monitor and respond to threats in the network. Overall, this process enables more efficient and accurate real-time monitoring and anomaly identification, improves the intelligence, response speed, and adaptability of network protection, adapts to dynamically changing network environments, and prevents potential threats from harming network security.
[0056] Optionally, step S332 specifically includes:
[0057] Traffic rate features and data packet features are extracted from the preprocessed switch traffic data to obtain traffic rate data and data packet data.
[0058] Obtain historical traffic data from the switch;
[0059] Traffic fluctuation range data is obtained by statistically analyzing the historical traffic data of the switch.
[0060] A traffic fluctuation baseline is constructed based on the traffic fluctuation range data to obtain the traffic fluctuation baseline data;
[0061] A traffic rate baseline is constructed from the traffic rate data to obtain the traffic rate baseline data;
[0062] Based on the data packet data, traffic source optimization is performed on the traffic fluctuation baseline data to obtain optimized traffic fluctuation baseline data;
[0063] Based on the traffic fluctuation baseline optimization data and the traffic rate baseline data, the switch traffic pattern baseline is integrated to obtain the switch traffic pattern baseline data.
[0064] This invention preprocesses switch traffic data and extracts traffic rate and packet characteristics to more accurately capture key patterns and behaviors of network traffic, thereby helping to identify potential abnormal traffic. By acquiring historical switch traffic data and statistically analyzing its traffic fluctuation range, the normal fluctuation range of network traffic can be revealed, providing a benchmark for further traffic analysis. After establishing a traffic fluctuation baseline, the normal range of traffic fluctuations under different network conditions can be clearly defined, helping to identify abnormal fluctuations that exceed the normal range. By constructing a traffic rate baseline from traffic rate data, the normal range of traffic rate values can be clearly defined, further improving the sensitivity to abnormal traffic rates. Optimizing the traffic fluctuation baseline by combining packet data allows for more precise baseline adjustments, reducing misjudgments caused by changes in the network environment. Finally, by integrating optimized traffic fluctuation baseline data and traffic rate baseline data, a comprehensive switch traffic pattern baseline is constructed, enabling more accurate identification of abnormal traffic patterns in the network. These steps improve the system's adaptability to dynamic network environments, enhance its ability to monitor and respond to complex and changing network conditions, effectively improve the system's speed of identifying and responding to potential security threats, reduce misjudgments or omissions caused by static methods, and enhance the intelligence and real-time performance of network security protection.
[0065] Optionally, step S4 specifically includes:
[0066] Step S41: Extract link disconnection features and malware detection features based on network security threat data to obtain link disconnection data and malware detection data;
[0067] Step S42: Identify security threats and risks in the abnormal network transmission topology based on the link disconnection data, thereby obtaining link disconnection security threat and risk data;
[0068] Step S43: Identify security threat risks of abnormal network transmission topology based on malware detection data, thereby obtaining malware detection security threat risk data;
[0069] Step S44: Integrate security threat risk data based on link disconnection security threat risk data and malware detection security threat risk data to obtain security threat risk data;
[0070] Step S45: Classify network security threats based on the abnormal network transmission topology of security threat risk data, thereby obtaining network security threat data topology data.
[0071] This invention, through link disconnection feature extraction and malware detection feature extraction from network security threat data, can accurately identify potential threats in the network, such as link interruptions and malware activity. This helps the system more quickly capture abnormal events affecting network transmission, preventing network performance degradation or data leakage caused by network interruptions or malware intrusion. Security threat risk identification based on link disconnection data can identify security vulnerabilities caused by link disconnections, ensuring the stability and security of network communication. Risk identification of abnormal network transmission topologies using malware detection data can promptly detect potential malware attacks or infections, thereby enabling effective defensive measures to prevent further spread of malware or damage to the network. Integrating link disconnection security threat risk data and malware detection security threat risk data helps to comprehensively assess various security threats in the network, optimize risk identification strategies, and improve network protection capabilities. Finally, by analyzing security threat risk data and classifying security threats based on abnormal network transmission topologies, the distribution of different types of security threats in the network can be clearly identified, providing a precise basis for subsequent security response and protection strategy formulation. This strengthens the dynamic adaptability and protection capabilities of the network system, improves the overall security protection level, and reduces the difficulty of adapting to constantly changing network environments.
[0072] Optionally, this specification also provides an intelligent information big data acquisition and management system for executing the intelligent information big data acquisition and management method described above. This intelligent information big data acquisition and management system includes:
[0073] The data dictionary maintenance module is used to acquire information from a big data warehouse, extract data dictionary features based on the big data warehouse to obtain data dictionary data, and maintain the data dictionary data to obtain data dictionary maintenance data.
[0074] The transmission anomaly analysis module is used to acquire switch device data, perform network topology analysis based on the switch device data to obtain network topology data of the switch device; and perform transmission anomaly analysis based on the network topology data of the switch device to obtain network transmission anomaly topology data.
[0075] The network security threat analysis module is used to perform traffic statistics based on switch device data to obtain switch traffic data; and to perform network security threat analysis based on switch traffic data to obtain network security threat data.
[0076] The network security threat classification module is used to classify network security threats based on the abnormal network transmission topology structure according to network security threat data, thereby obtaining network security threat data topology structure data;
[0077] The intelligent network security threat acquisition model construction module is used to construct an intelligent network security threat acquisition model based on data dictionary maintenance data and network security threat data topology structure data, thereby obtaining an intelligent network security threat acquisition model; and to manage network security threats in the information big data warehouse based on the intelligent network security threat acquisition model, thereby obtaining network security threat management data.
[0078] This invention discloses an intelligent information big data acquisition and management system. This system can implement any of the intelligent information big data acquisition and management methods of this invention. It is used to combine the operation and signal transmission media between various modules to complete the intelligent information big data acquisition and management method. The internal modules of the system cooperate with each other, thereby improving the efficiency of identifying network security threats. Attached Figure Description
[0079] Other features, objects, and advantages of the invention will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings:
[0080] Figure 1 This is a flowchart illustrating the steps of the intelligent data collection and management method for big data of the present invention.
[0081] Figure 2 This is a detailed flowchart of step S1 in the present invention;
[0082] Figure 3 This is a detailed flowchart of step S2 in this invention.
[0083] The realization of the objective, functional features and advantages of the present invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0084] The technical method of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. Based on the embodiments of this invention, all other embodiments obtained by those skilled in the art without inventive effort are within the scope of protection of this invention.
[0085] Furthermore, the accompanying drawings are merely illustrative of the invention and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and therefore repeated descriptions of them will be omitted. Some block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities can be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor methods and / or microcontroller methods.
[0086] It should be understood that although the terms "first," "second," etc., may be used herein to describe various units, these units should not be limited by these terms. These terms are used merely to distinguish one unit from another. For example, without departing from the scope of the exemplary embodiments, a first unit may be referred to as a second unit, and similarly, a second unit may be referred to as a first unit. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.
[0087] To achieve the above objectives, please refer to Figures 1 to 3 This invention provides an intelligent data collection and management method for big data, the method comprising the following steps:
[0088] Step S1: Obtain the information big data warehouse, and extract data dictionary features based on the information big data warehouse to obtain data dictionary data; maintain the data dictionary data to obtain data dictionary maintenance data.
[0089] In this embodiment, raw information containing metadata and data table structures is extracted from the big data warehouse by connecting to its API interface. This process typically involves SQL queries and ETL (Extract, Transform, Load) techniques. A preliminary data dictionary is constructed by specifying the table names, field names, and data types of the data warehouse. Specific operations for data dictionary feature extraction include analyzing field names, data types (such as integer, string, date, etc.), and field constraints (such as primary keys, foreign keys, unique constraints, etc.). To obtain an effective data dictionary, the extraction process must adhere to data dictionary specifications, such as field names not exceeding 50 characters in length and field types conforming to standard database types (e.g., VARCHAR(255) or INT). Next, based on these data dictionary features, data dictionary maintenance is performed. Regular expressions are used to standardize the field names, ensuring consistency. For example, if fields "User_Id" and "user_id" have the same meaning, they are unified as "user_id". Then, using data dictionary maintenance rules, the fields are standardized and updated in the data dictionary maintenance table, recording the standard name, mapping relationship, and other auxiliary information for each field.
[0090] Step S2: Obtain switch device data and perform network topology analysis based on the switch device data to obtain switch device network topology data; perform transmission anomaly analysis based on the switch device network topology data to obtain network transmission anomaly topology data.
[0091] In this embodiment, real-time operational data of the switch devices is obtained via SNMP (Simple Network Management Protocol) or CLI (Command Line Interface), including port status, link status, MAC address table, VLAN information, etc. Based on this data, the switch network topology is first analyzed. Topology analysis mainly relies on the connection information between switch ports and the physical links between switches. Specifically, a preliminary topology map is constructed using the switch port connection data. For example, the network topology is generated by analyzing the MAC address table of the switch ports, the port connection type (e.g., Ethernet or fiber optic connection), and the status of each port (e.g., "up" or "down"). After the topology is constructed, transmission anomalies of the switch devices are analyzed. Anomaly detection is performed based on the bandwidth utilization and transmission latency of the switch ports. The bandwidth utilization threshold can be set to 90%. If the bandwidth utilization of a port continuously exceeds 90% for more than 5 minutes, the port is considered to have a transmission anomaly. Combined with the link's RTT (Round-Trip Time) value, if the link's RTT is greater than 50ms, it is further marked as network transmission anomaly topology data and stored.
[0092] Step S3: Perform traffic statistics based on the switch device data to obtain switch traffic data; perform network security threat analysis based on the switch traffic data to obtain network security threat data.
[0093] In this embodiment, traffic data for each port is obtained through the switch's traffic statistics interface (such as NetFlow or sFlow protocol), including the number of packets, bytes, packet loss rate, and traffic direction (inbound or outbound) for each port. Traffic statistics need to be grouped by time window, typically 5 minutes per window, to statistically analyze the traffic of each port, recording the maximum, minimum, average, and burst traffic within each time period. After completing the traffic statistics, further network security threat analysis is performed. The core of this part is to identify potential security threats through traffic pattern recognition, such as phishing attacks and DDoS attacks. Taking phishing attacks as an example, firstly, DNS query logs are analyzed to identify the number of abnormal domain name requests. A threshold (e.g., 100 requests) is set; if the number of requests for a domain name exceeds this threshold within a short period, the domain name is associated with a phishing attack. Furthermore, abnormal fluctuations in traffic data are analyzed; for example, a traffic increase exceeding 50% within a short period is considered abnormal and associated with a DDoS attack. These analysis results are then summarized to form network security threat data, recording the attack type and its occurrence time and location.
[0094] Step S4: Based on the network security threat data, classify the abnormal network transmission topology to obtain network security threat data topology data;
[0095] In this embodiment, potential attack sources or security risk nodes are identified by analyzing the network security threat data generated in the previous stage. Specifically, network security threat data (such as phishing attacks, DDoS attacks, etc.) is matched with the network topology data of the switching devices. Threat nodes are determined by analyzing the IP addresses of the attack sources and the flow of abnormal traffic. Then, these nodes are segmented based on the data of abnormal transmission topology. Specifically, abnormal links in the network (such as abnormal bandwidth, high packet loss rate, etc.) are associated with the identified threat data. For example, if a link has excessively high bandwidth utilization and is associated with a known malicious IP address, that link is marked as a security threat link. Ultimately, this data forms a network security threat data topology, helping to locate potential attack sources and providing a basis for further defensive measures.
[0096] Step S5: Construct a cybersecurity threat intelligent collection model based on the data dictionary maintenance data and the cybersecurity threat data topology structure data to obtain the cybersecurity threat intelligent collection model; manage cybersecurity threats in the information big data warehouse based on the cybersecurity threat intelligent collection model to obtain cybersecurity threat management data.
[0097] In this embodiment, an intelligent data acquisition model is constructed using data dictionary maintenance data and network security threat data topology structure data. This model is based on machine learning algorithms (such as decision trees, random forests, or neural networks), and its input features include node information in the network topology structure (such as switch status, port status, bandwidth utilization, etc.) and indicators of network security threats (such as traffic volatility, DNS request anomalies, etc.). By training the model, pattern learning can be performed based on historical security threat data to generate a model capable of identifying and classifying security threats in real time. To ensure the effectiveness of the model, hyperparameter tuning is required, such as setting the depth of the decision tree and the number of trees in the random forest. After the model is built, it is used to manage network security threats in a large data warehouse. The model acquires new network traffic data from the data warehouse in real time and analyzes it using a classifier to identify new security threats. If a security threat is detected, the model will trigger an alarm and update the network security threat management data, including information such as affected devices, attack type, and time of occurrence. This information will help administrators take necessary protective measures, such as isolating infected nodes or strengthening firewall rules.
[0098] Optionally, step S1 specifically includes:
[0099] Step S11: Obtain the information big data warehouse, and extract data dictionary features based on the information big data warehouse to obtain data dictionary data;
[0100] In this embodiment, structured data needs to be retrieved from a big data warehouse. This process is typically accomplished by connecting to a database management system (such as MySQL, Oracle, or Hadoop), relying on SQL queries or directly calling APIs to obtain metadata within the warehouse (e.g., table names, field names, data types, table relationships, etc.). Assuming the data warehouse uses MySQL, the SQL query `SHOW TABLES` is first used to retrieve all table names, and then `DESCRIBE` is used...<table_name> Obtain the field characteristics of each table, extracting information such as field name, data type, whether NULL is allowed, and field length. For field data types, special attention should be paid to data standardization, such as unifying "VARCHAR(100)" and "TEXT" to string type (STRING). In addition, it is necessary to extract field index information, primary and foreign key relationships, and other structured data, which are important components of building the data dictionary. During the data dictionary feature extraction process, all field names must be recorded in a uniform format, such as using lowercase letters and underscores as separators (e.g., user_id instead of UserId), to ensure data standardization.
[0101] Step S12: Perform field standardization on the data dictionary data to obtain field-standardized data;
[0102] In this embodiment, fields extracted from the data dictionary are standardized. The purpose of field standardization is to ensure consistency in field naming and data format, facilitating subsequent data processing and analysis. Specifically, this involves uniformly converting field names to a standardized format. For example, if a field name uses camelCase (e.g., UserName), it is converted to underscore naming (e.g., user_name). For data type standardization, different types of fields need to be standardized. For example, all date fields (e.g., DATE, DATETIME, TIMESTAMP) are standardized to the "DATE" type, and all numeric fields (e.g., INT, BIGINT, FLOAT) are standardized to "INTEGER" or "FLOAT". Furthermore, field constraints (e.g., "NOTNULL" or "UNIQUE") should be preserved during standardization to ensure data integrity. The standardization process also checks whether field lengths conform to database design specifications and makes appropriate adjustments. For example, if a field length exceeds the recommended value, it is corrected to ensure efficient table storage and retrieval. After standardization, all fields should conform to a unified naming convention, and the attributes (type, length, constraints) of each field should be in a standardized format.
[0103] Step S13: Perform encoding mapping based on the data dictionary data to obtain the standard mapping code value for the digital warehouse;
[0104] In this embodiment, the fields in the data dictionary need to be encoded and mapped. The goal of encoding and mapping is to convert non-numeric data (such as strings, dates, etc.) in the fields into a processable numeric form. First, for text fields (such as user_name, product_type), the unique values in the field (such as user name, product type, etc.) are mapped to unique numeric IDs. Specifically, for each text value (such as "admin" or "user"), a unique number is assigned to it, for example, "admin" corresponds to 1, and "user" corresponds to 2. This mapping typically uses hash algorithms, dictionary mapping, or category encoding methods. During the encoding process, for the data in each field, all its unique values are first counted, and these values are assigned a unique numeric code in alphabetical or numeric order. In addition, for date type fields (such as created_at), the date can be converted into a custom numeric format, such as converting the date field into a timestamp (the number of seconds from January 1, 1970 to the current date). This method ensures that the data is stored in numeric format in subsequent processing, facilitating calculation and analysis. Finally, a mapping dictionary or mapping table is generated through encoding mapping to record the relationship between the original value of each field and the numeric code, which facilitates subsequent querying and data restoration.
[0105] Step S14: Perform dictionary maintenance based on the standardized field data and the standard mapping code value of the digital warehouse to obtain data dictionary maintenance data.
[0106] In this embodiment, the key to dictionary maintenance is combining the standardized field data with the results of the encoding mapping to ensure continuous updating and maintenance of the data dictionary. Specifically, based on the data dictionary and mapping relationships generated in steps S12 and S13, the existing data dictionary needs to be supplemented and updated. The update process includes merging the standardized field names, type information, and corresponding encoding mapping values into the dictionary. For example, if the field `user_name` is renamed to `user_name` during standardization and its text value (e.g., "admin") is converted to a numeric ID (e.g., 1) through encoding mapping, it is added as a new dictionary item to the data dictionary. Furthermore, the dictionary needs to be checked and maintained regularly to ensure its integrity and consistency. For example, when a new field is added to the data table, it should first be standardized, and its encoding mapping added to the dictionary. For deleted fields, the relevant entries need to be removed from the dictionary. Data dictionary maintenance also needs to include field description information (e.g., the field's purpose, scope, limitations, etc.) to ensure that each field has a detailed description for easy subsequent data management and use. Through this process, the final data dictionary maintenance data includes not only field names, data types, and encoding mappings, but also field descriptions, constraints, and other auxiliary information, ensuring the standardization and efficiency of data management.
[0107] Optionally, step S2 specifically includes:
[0108] Step S21: Obtain switch device data and perform node identification based on the switch device data to obtain switch node data;
[0109] In this embodiment, network topology information needs to be obtained from the switch devices. This data can be collected through management protocols (such as SNMP, LLDP) or the switch's syslog logs. Switch device data includes the switch's MAC address, IP address, port number, and connected device information. After data acquisition, this information is used for node identification. A node is a device in the network capable of receiving and sending data (such as a switch port, server, or terminal). Based on the switch device data, the device connected to each port is identified by matching the MAC address and IP address of each switch port. Accurate mapping algorithms are necessary for accurate node identification. Node data includes a unique identifier for each port (such as switch1_port2) and information about the connected devices, which can be other switches, servers, or network terminals. During node identification, network topology discovery protocols (such as CDP, LLDP, etc.) are also needed to confirm the relationship between the switch and other network devices. The identified node data will be stored in a standardized format, such as JSON or CSV files, to ensure consistent subsequent processing of the topology data.
[0110] Step S22: Extract port connection features based on switch device data to obtain port connection data;
[0111] In this embodiment, information related to each port in the switch data is parsed, such as port number, port type (e.g., Ethernet, SFP), connection status (enabled or disabled), and bandwidth. The connection characteristics of each port involve several key parameters: physical connection status (UP / DOWN), data transmission rate, port type (e.g., 10 / 100 / 1000Mbps, fiber optic port), device bound to the MAC address, and historical port performance (e.g., bandwidth utilization, packet loss rate). This information can be captured by capturing the output of the switch's SNMP interface or CLI commands. The feature extraction process is performed by parsing the switch device configuration file (e.g., the output of the `show interfaces` command). For each port, special attention is paid to bandwidth, link status, packet loss rate, etc. This data is recorded and stored as port connection characteristics and output in JSON format, such as {"port_id":"port1","status":"UP","bandwidth":"1000Mbps","connected_device":"device123"}. This port connection data provides a foundation for subsequent topology establishment.
[0112] Step S23: Establish a primary topology based on switch node data and port connection data to obtain primary topology data;
[0113] In this embodiment, the node data of the switch, namely the switch's port identifiers and the information of the connected devices, is used to form a preliminary topology. For each pair of devices connected via a port, a virtual connection is created, recording the connected port number, the connected device information, and the connection status. A connection matrix is built using this information, where each item in the matrix represents a connected port pair. The topology data model should include node identifiers (such as switch port IDs), connection status (UP / DOWN), bandwidth, and link quality. During the initial construction of the topology, a graphical representation method is used, such as using an adjacency matrix or adjacency list to represent the connections between devices. This data structure ensures that the connection relationships between each device (node) and other devices are clearly represented, for example, {"device_id":"switch1","connected_ports":["port1","port2"],"device_status":"active"}. At this stage, the topology can be checked and adjusted using network topology visualization tools to ensure the accuracy and completeness of the data structure.
[0114] Step S24: Redundant path identification is performed on the primary topology data to obtain redundant path data;
[0115] In this embodiment, multiple paths connecting the same node exist, typically used to improve network reliability. Based on the initial topology data, graph theory algorithms (such as Dijkstra's algorithm or Floyd-Warshall algorithm) are applied to identify the shortest path between all devices. If multiple equally weighted paths connecting the same node are found, they are considered redundant paths. Redundant path identification focuses on bandwidth, link status, and path reliability. In the path selection for each node, if a backup path exists (e.g., two different links connecting the same node), that path is considered redundant. Redundant path identification must ensure no impact on data transmission efficiency; redundant paths are marked and their information recorded in the data structure. For example, a redundant path can be represented using the following format: {"path_id":"path1","redundant":true,"nodes":["switch1","switch2","switch3"],"status":"active"}. In this way, which paths in the network are redundant can be identified, allowing for further optimization of network performance.
[0116] Step S25: Loop path identification is performed on the primary topology data to obtain loop path data;
[0117] In this embodiment, a loop refers to a network topology where data packets can be transmitted cyclically between multiple paths, leading to network congestion or data packet loss. Loop path identification techniques include graph search algorithms such as Depth-First Search (DFS) and Breadth-First Search (BFS). In practice, each node and all its connected ports are first checked, traversing each path to find loops. If a path ultimately returns to the source node, it is determined to be a loop path. After identifying a loop path, the loop's start point, end point, traversed nodes, and bandwidth utilization should be recorded. Loop path data should be represented in a structured manner, such as {"loop_id":"loop1","nodes":["switch1","switch2","switch3"],"path":["port1","port2","port3"],"status":"detected"}. After loop path identification, the network design can be further adjusted to avoid network performance problems caused by loops.
[0118] Step S26: Construct the network topology of the switch device based on the loop path data and redundant path data, thereby obtaining the network topology data of the switch device.
[0119] In this embodiment, loop paths and redundant paths are removed to ensure the uniqueness and efficiency of the network topology. For redundant paths, the need to enable link aggregation or load balancing techniques should be considered. For loop paths, loop connections need to be removed using a spanning tree algorithm (such as Spanning Tree Protocol, STP) to retain the optimal working path. The constructed network topology should include a unique identifier for each device (node), port connection status, link status, and bandwidth information for each path. In the topology data model, the active connection status and redundant connections of each node should be clearly marked to ensure smooth data transmission. The topology data can be stored in a graph database or matrix representation, such as {"network_id":"network1","nodes":[{"id":"switch1","ports":["port1","port2"],"status":"active"}],"links":[{"from":"switch1","to":"switch2","status":"active"}]}, which facilitates network optimization and maintenance.
[0120] Step S27: Perform transmission anomaly analysis based on the network topology data of the switch device to obtain network transmission anomaly topology data.
[0121] In this embodiment, transmission anomaly analysis technology is used to examine the constructed network topology to identify network performance problems or faults. During anomaly analysis, metrics such as bandwidth utilization, packet loss rate, and latency of switch ports are monitored first. This data is collected periodically via protocols such as SNMP or NetFlow. A threshold (e.g., bandwidth utilization exceeding 85% or packet loss rate exceeding 5%) is set as the standard for anomaly detection in transmission anomaly analysis. If the performance metrics of a link or port exceed the threshold, the path is marked as an abnormal path. Furthermore, trend analysis is performed using historical performance data to identify potential performance bottlenecks or fault points. Network transmission anomaly topology data should include node information of the abnormal link, the anomaly type (bandwidth problem, packet loss, link disconnection, etc.), and the time of occurrence. Anomaly data should be represented as follows: {"link_id":"link1","status":"error","error_type":"high_bandwidth_utilization","value":"90%"}, providing a basis for subsequent fault diagnosis.
[0122] Optionally, step S27 specifically includes:
[0123] Step S271: Identify switch port faults based on the network topology data of the switch device to obtain switch port fault data;
[0124] In this embodiment, port faults are identified by extracting port status information from the network topology data of the switch device. First, the connection status information of each switch port needs to be obtained from the network topology data, such as whether the port is in an "UP" or "DOWN" state. If a port is in a "DOWN" state and the connected device cannot be accessed, the port is identified as faulty. In addition to port status, port error counters should also be monitored, including indicators such as packet loss rate, error frames, CRC errors, and operational errors. By analyzing the switch's CLI command output or SNMP monitoring data (such as showinterfaces), the error counter value of each port is checked. If the error count exceeds a preset threshold (e.g., error frame count greater than 1000), the port is considered faulty. Fault identification data should include the port ID, fault type (e.g., physical disconnection, link error, etc.), and the time of fault occurrence, in the format {"port_id":"port1","status":"DOWN","error_type":"CRC_error","error_count":1200}.
[0125] Step S272: Calculate the bandwidth utilization of the switch port fault data to obtain the bandwidth utilization rate;
[0126] In this embodiment, real-time traffic data of the switch ports is collected via SNMP, NetFlow, or sFlow protocols. This data includes the number of input and output bytes of the switch ports. Port traffic refers to the amount of data transmitted through a port per unit time. Then, the maximum bandwidth of each switch port is determined. This value is set by the switch configuration and is typically the port speed, such as 1000Mbps or 10000Mbps. Based on the collected traffic data and the maximum bandwidth of the port, the bandwidth utilization rate of that port is calculated. During the calculation process, to ensure data accuracy, the bandwidth utilization rate monitoring cycle is typically set to once every 5 minutes, continuously monitored for 30 minutes, to capture changes in port bandwidth utilization. During monitoring, dynamic changes in the network topology also need to be considered, such as changes in the devices connected to the ports or adjustments to the bandwidth configuration; therefore, the data needs to be updated in real time to ensure its timeliness. After the bandwidth utilization rate calculation is completed, the relevant data will be stored in a specific format. For example, the storage results can record the bandwidth utilization and related time points for each port, with the data format {"port_id":"port1","bandwidth_utilization":"20%","time_period":"2025-01-07 15:00:00"}, for subsequent analysis and processing.
[0127] Step S273: Perform numerical statistics based on bandwidth utilization to obtain high bandwidth utilization;
[0128] In this embodiment, a threshold for high bandwidth utilization is defined; for example, bandwidth utilization exceeding 80% is considered high bandwidth utilization. Next, the bandwidth utilization data of all ports is traversed, and ports with bandwidth utilization exceeding 80% are extracted. During the statistical process, statistical measures such as standard deviation and mean are used to analyze the bandwidth utilization of each port. For example, if the port list of a switch is [10%, 20%, 90%, 85%, 70%], then based on the threshold, the ports with high bandwidth utilization are identified as 90% and 85%. The statistical results will be output in list form, including port ID, bandwidth utilization, and anomaly type (such as "high bandwidth utilization"). For example, the statistical result is {"high_utilization_ports":[{"port_id":"port3","bandwidth_utilization":90%},{"port_id":"port4","bandwidth_utilization":85%}]}.
[0129] Step S274: Analyze network performance anomalies based on the network topology data of the switch device according to the high bandwidth utilization, thereby obtaining network performance anomaly data;
[0130] In this embodiment, based on high bandwidth utilization data, the system analyzes whether these ports pose a risk of bandwidth bottlenecks or network performance degradation. During the analysis, bandwidth utilization is combined with performance indicators such as packet loss rate and latency. If high bandwidth utilization is accompanied by an increasing packet loss rate (e.g., packet loss rate exceeding 2%), it is considered a performance anomaly. Furthermore, by comparing historical performance data with real-time traffic data, anomaly trend analysis is performed to identify links or nodes causing performance degradation. The anomaly analysis considers peak and average network traffic to promptly identify temporary or long-term performance issues. For example, if port3 has a bandwidth utilization of 90% and a packet loss rate of 3%, then the network performance anomaly data for this port is {"port_id":"port3","bandwidth_utilization":90%,"packet_loss":3%,"status":"performance_issue"}. This data can be used to further optimize network performance.
[0131] Step S275: Perform link health anomaly detection on the network topology data of the switch device to obtain link health anomaly data;
[0132] In this embodiment, link health detection analyzes metrics such as packet loss rate, latency, error frames, and link status. First, standard thresholds are set for the health of each link. For example, a link is considered unhealthy if the packet loss rate exceeds 2%, the link status is not "UP", or the number of error frames exceeds 1000. Second, by periodically monitoring the error count and link status of switch ports, if a link's error frames, packet loss rate, or latency exceeds the threshold, it is recorded as unhealthy. Link health data should include link ID, health status, number of error frames, and packet loss rate. The format of unhealthy link data is as follows: {"link_id":"link1","status":"unhealthy","error_count":1500,"packet_loss":3%,"latency":"200ms"}. By detecting the health status of links, potential link faults or performance bottlenecks can be identified in advance, providing data support for subsequent network optimization.
[0133] Step S276: Based on the abnormal link health data and abnormal network performance data, perform network transmission abnormal topology analysis on the network topology of the switch device to obtain abnormal network transmission topology data.
[0134] In this embodiment, the network performance anomaly ports discovered in step S274 are associated with the link health anomalies discovered in step S275. If a port or link exhibits performance anomalies, and this port or link is part of a critical path in the network, then this path is determined to be a network transmission anomaly path. Using the shortest path algorithm in graph theory, the affected paths in the network topology are analyzed, and the anomaly links and related nodes are identified. Anomaly data should include multi-dimensional information such as path, node, link status, bandwidth, packet loss rate, and latency. For example, if a path from switch1 to switch2 passes through a link with high bandwidth utilization and poor link health, then the network transmission anomaly data for this path is {"path_id":"path1","status":"abnormal","nodes":["switch1","switch2"],"link_status":"unhealthy","packet_loss":5%,"bandwidth_utilization":90%}. This data can be further used to analyze network bottlenecks and fault points, aiding in network optimization.
[0135] Optionally, step S3 specifically includes:
[0136] Step S31: Perform traffic statistics based on the switch device data to obtain switch traffic data;
[0137] In this embodiment, traffic data from the switch devices is collected. Traffic statistics rely on protocols such as SNMP (Simple Network Management Protocol), NetFlow, or sFlow, which can provide real-time input and output data byte counts for ports. By collecting this traffic data, the data transmission volume of each switch port within a certain time period is statistically analyzed. These statistical values include, but are not limited to, input and output byte counts, packet counts, and traffic rates. Based on this real-time data, the overall traffic situation of the switch can be obtained. The data collection cycle is set to once every 5 minutes to ensure that traffic fluctuations in different time periods are captured.
[0138] Step S32: Detect phishing attacks based on switch traffic data to obtain phishing attack data;
[0139] In this embodiment, phishing attacks are identified by analyzing the characteristics of switch traffic, such as DNS request frequency and abnormal volume of external IP requests. DNS query logs and external IP address access patterns are used as the basis for judgment. For example, if the number of domain name requests from an IP address exceeds a normal threshold (e.g., more than 10 requests / minute), the IP address is considered suspicious and a source of a phishing attack. Simultaneously, the suspiciousness of external URLs is also monitored, analyzing for the presence of a large number of shortened links or malicious URLs. This data is further filtered through correlation analysis, ultimately identifying phishing attack data and generating phishing attack alerts.
[0140] Step S33: Perform traffic anomaly monitoring based on switch traffic data to obtain traffic anomaly data;
[0141] In this embodiment, a baseline is established for switch traffic data to calculate the normal traffic fluctuation range. A traffic baseline, such as the average traffic value and standard deviation, is established using historical traffic data. Real-time traffic data is compared to this baseline, and a traffic anomaly alarm is triggered when the traffic deviates from the normal baseline by more than a set threshold (e.g., more than twice the standard deviation). Furthermore, different types of traffic anomaly monitoring indicators can be set, such as bandwidth abuse, DDoS attack traffic, or malicious traffic patterns. The specific threshold for abnormal traffic is dynamically adjusted based on historical data and network load. The monitoring system performs a data check every 5 minutes to capture and record potential traffic anomalies.
[0142] Step S34: Integrate cybersecurity threat data based on phishing attack data and abnormal traffic data to obtain cybersecurity threat data.
[0143] In this embodiment, correlation analysis is performed by combining suspicious IPs, domains, and types of traffic anomalies associated with phishing attacks. For example, when the access volume of a suspicious IP address suddenly increases and abnormal traffic characteristics are detected, the system will automatically mark the IP as high-risk and prioritize its handling. Furthermore, contextual information of other security threats, such as device location and timestamps, needs to be considered. This information is combined with identified phishing attacks and traffic anomaly data to generate a comprehensive cybersecurity threat report. This data is stored in a structured format, such as {"threat_type":"phishing","ip_address":"192.168.1.100","anomaly_type":"high_traffic","timestamp":"2025-01-0715:00:00"}, for subsequent security analysis and response.
[0144] Optionally, step S32 specifically includes:
[0145] Step S321: Perform DNS request analysis based on switch traffic data to obtain DNS request data;
[0146] In this embodiment, Domain Name System (DNS) request data is extracted from switch traffic data. This request data is obtained through switch traffic monitoring tools (such as NetFlow, sFlow, or using the SNMP protocol). In switch traffic data, DNS requests typically appear as request packets sent to DNS servers. Each request contains information such as the source IP address, the target DNS server's IP address, the requested domain name, and a timestamp. By extracting these requests, a request dataset is established. Detailed information for each request (such as the requested domain name, time, and requested IP address) is recorded and stored in chronological order for subsequent analysis and processing. The monitoring data collection frequency is set to once every 5 minutes to ensure timely capture of dynamic changes in DNS requests.
[0147] Step S322: Perform request volume statistics on the Domain Name System (DNS) request data to obtain high request volume data for the DNS;
[0148] In this embodiment, the number of requests for each DNS domain name is statistically analyzed at regular time windows (e.g., hourly or daily). If the request volume for a domain name fluctuates drastically over a period of time (e.g., the request volume exceeds twice the normal range), the request volume for that domain name is considered abnormal. During the statistical analysis, independent statistical analysis is performed for each domain name, recording the number of requests within each time period. When the request volume reaches a specific threshold (e.g., exceeding 50 requests per minute), it is considered that the request volume for that domain name is abnormal. Based on this, high request volume data of the Domain Name System is generated for subsequent analysis.
[0149] Step S323: Identify suspicious external IP addresses based on high request volume data from the Domain Name System to obtain suspicious external IP address data;
[0150] In this embodiment, suspicious external IP addresses are identified based on high request volume data. For each DNS request, its source IP address is checked to see if it is an external IP address, and whether the request frequency of that IP address within the statistical period exceeds a preset threshold. For example, if an external IP address makes more than 100 requests within 10 minutes, it is considered a suspicious IP address. This threshold is dynamically adjusted based on historical data and the average level of network traffic patterns. If the request volume of an IP address is higher than normal, and the domain name it requests is suspicious (such as shortened links, malicious websites, etc.), then the IP address is marked as a potential attack source. Data on all suspicious external IP addresses (such as IP address, number of requests, requested domain name, etc.) is recorded as input data for the next step of analysis.
[0151] Step S324: Perform short link detection based on switch traffic data to obtain short link data;
[0152] In this embodiment, short link detection relies on the analysis of the requested domain name. When a DNS request contains a domain name generated by a short link service (such as bit.ly, goo.gl, etc.), the system marks that domain name as a short link domain. By analyzing the domain name of the DNS request and comparing it with the short link service database, requests belonging to short link services are identified as short link data. All detected short links are recorded, including the target URL of the short link, the source IP address of the request, and the timestamp. During the short link domain detection process, the system needs to continuously update the domain name database of short link service providers to maintain the accuracy of short link identification.
[0153] Step S325: Identify abnormal Uniform Resource Locators (URLs) in the short link data to obtain abnormal URL data;
[0154] In this embodiment, the target URL pointed to by the shortened link is obtained. A shortened link parsing tool is used to parse the actual target URL corresponding to the shortened link. These target URLs are then subjected to feature analysis to check for typical malicious URL characteristics, such as: URLs containing malicious scripts, domains with phishing website characteristics (e.g., URLs impersonating legitimate websites), and unknown or uncommon top-level domains (e.g., .xyz). If a URL matches a known malicious URL database or contains suspicious characteristics (e.g., excessively long paths, random characters), it is considered an abnormal Uniform Resource Locator (URL). Detailed information for each abnormal URL is recorded, such as the target URL, the source IP address, and the access time, to facilitate subsequent analysis and response.
[0155] Step S326: Based on the abnormal Uniform Resource Locator (URI) data and short link data, determine the phishing attack data of the suspicious external IP address data.
[0156] In this embodiment, when a suspicious IP address is associated with an abnormal URL, and the resource pointed to by that URL exhibits typical phishing attack characteristics, the system will associate the suspicious IP address with a phishing attack. For example, when an external IP address frequently requests shortened links, and these shortened links point to URLs containing phishing characteristics, it can be determined that the IP address has launched a phishing attack. All information related to phishing attacks, including IP addresses, requested domain names, and URLs, will be aggregated and recorded as phishing attack data. The generated phishing attack data includes the attack source IP, the attack target URL, the time of the attack, and related network activity data. This data will serve as early warning information for the network security protection system, used for further protection responses.
[0157] Optionally, step S33 specifically includes:
[0158] Step S331: Preprocess the switch traffic data to obtain preprocessed switch traffic data;
[0159] In this embodiment, the main purpose of preprocessing is to remove redundant data, fill in missing values, and standardize the data format to ensure the accuracy of subsequent analysis. Using real-time switch traffic data collected through network monitoring tools (such as SNMP, sFlow, or NetFlow), the data is first integrated in a time series to ensure consistency in traffic data across each collection period. For each switch traffic record, including fields such as source IP, destination IP, number of bytes transmitted, and transmission duration, unit conversion is performed, for example, converting the number of bytes to megabytes for subsequent analysis. Then, missing or outlier values (such as negative or excessively large values) are checked in the data and filled or removed according to predefined rules. Finally, all cleaned data is stored in chronological order, and the data is output in a uniform format (such as CSV or JSON) for easy subsequent processing.
[0160] Step S332: Establish a traffic pattern baseline on the preprocessed switch traffic data to obtain switch traffic pattern baseline data;
[0161] In this embodiment, by aggregating and analyzing large-scale historical data, typical traffic characteristics of each switch port or device in different time periods (e.g., hours, days, weeks) are extracted. Clustering analysis methods, such as the K-means algorithm, can then be used to identify different types of traffic patterns, such as normal traffic, peak traffic, and off-peak traffic. The average bandwidth utilization, peak traffic volume, and fluctuation range of each port are calculated. By classifying and summarizing these patterns, a baseline traffic model is established. During this process, some key thresholds are defined, such as: under normal traffic conditions, bandwidth utilization does not exceed 60%, and traffic fluctuation range is within 20%. When certain traffic exceeds this standard range, it needs to be marked as abnormal data.
[0162] Step S333: Perform real-time anomaly monitoring on switch traffic data based on the switch traffic pattern baseline data to obtain real-time anomaly switch traffic data.
[0163] In this embodiment, the system compares real-time traffic data to determine whether the current traffic deviates from the established traffic pattern baseline. Whenever a new record of switch traffic data is collected, the system compares the real-time data with historical traffic patterns according to a set monitoring period (e.g., once per minute). If the real-time traffic data (such as traffic rate, bandwidth utilization) exceeds the baseline threshold (e.g., bandwidth utilization exceeds 80% of the baseline), the traffic is recorded as abnormal. During this process, a sliding window technique is used to monitor each traffic sampling point in real time and record traffic changes within each monitoring period. All real-time abnormal switch traffic data will be collected and stored for further analysis.
[0164] Step S334: Classify the anomaly types based on the real-time abnormal switch traffic data to obtain bandwidth abuse anomaly data and network worm anomaly data;
[0165] In this embodiment, abnormal traffic can be categorized into different types by analyzing its anomaly characteristics. For example, if the bandwidth utilization of abnormal traffic is significantly higher than the normal value (exceeding 90% of the baseline), it is marked as bandwidth abuse anomaly data. Bandwidth abuse data can be identified by recognizing large-volume transmissions (such as large file downloads and video transmissions). Another common type of anomaly is network worm attack traffic. Network worm attacks typically manifest as rapidly spreading abnormal traffic, usually characterized by a large number of outgoing connection requests within a short period. Therefore, if a rapid increase in the number of connections and data packets occurs in real-time traffic (e.g., the number of connection requests increases by more than 200% within 5 minutes), it can be marked as network worm anomaly data. Detailed characteristics of each type of abnormal traffic (such as source IP, destination port, traffic size, etc.) will be categorized and stored for subsequent correlation analysis.
[0166] Step S335: Perform correlation analysis based on bandwidth abuse anomaly data and network worm anomaly data to obtain potential security threat data;
[0167] In this embodiment, a multi-dimensional data correlation model is established to cross-analyze bandwidth abuse anomaly traffic and network worm anomaly traffic. By combining traffic data with network topology data, it is examined whether these anomaly traffic events are concentrated in specific switches or subnets. In particular, it checks whether any source IP addresses appear simultaneously in both bandwidth abuse traffic and worm attack traffic, or whether the traffic patterns of certain IP addresses change drastically within a short period of time. For anomaly traffic that meets the criteria, correlation analysis is performed, for example, using the Pearson correlation coefficient to determine the relationship between different anomaly traffic events. If a high correlation is found between anomaly traffic patterns, it indicates that these traffic events are likely part of the same attack or security threat. Based on this, potential security threat data is generated, recording relevant source IPs, target ports, attack patterns, and other information.
[0168] Step S336: Perform traffic anomaly detection on the switch traffic data based on potential security threat data to obtain traffic anomaly data.
[0169] In this embodiment, the system will monitor real-time traffic again based on the characteristics in the potential security threat data and in conjunction with historical traffic baselines. Traffic anomaly detection is performed by comparing potential threat characteristics (such as abnormal traffic on specific ports, abnormal behavior of source IPs, etc.). If real-time traffic exhibits characteristics related to known potential security threats during the monitoring period (such as a sudden increase in traffic from suspicious IPs or a large number of connection requests), the traffic will be marked as abnormal. This detection process is implemented using real-time traffic analysis and historical traffic data comparison techniques to ensure timely detection and response to potential security threats. All detected abnormal traffic data will be recorded in detail and further analyzed or alerted.
[0170] Optionally, step S332 specifically includes:
[0171] Traffic rate features and data packet features are extracted from the preprocessed switch traffic data to obtain traffic rate data and data packet data.
[0172] In this embodiment, network monitoring tools (such as SNMP, NetFlow, or sFlow) are used to collect traffic data from switch ports, including the flow rate per second and the size of each data packet. The flow rate can be calculated by the number of bytes transmitted in each time interval (e.g., bytes transmitted per second), while data packet characteristics can be extracted by capturing the size and type of each data packet (e.g., TCP, UDP, ICMP). For flow rate data, the flow rate of each port is recorded according to a set time window (e.g., per second, per minute), and statistical data such as average, maximum, and minimum values are calculated based on the actual flow rate. During data packet characteristic extraction, the size and frequency of each data packet are analyzed, as well as whether there are abnormal data packet sizes (e.g., excessively large or small packets). The extracted data is stored in a database in the format: {"port_id":"port1","flow_rate":500Mbps,"packet_size":1500,"protocol":"TCP","timestamp":"2025-01-07 12:00:00"}.
[0173] Obtain historical traffic data from the switch;
[0174] In this embodiment, historical traffic data is acquired through continuous network traffic collection tools, typically by periodically capturing traffic statistics from switch devices. The collected historical data includes information such as real-time traffic, number of bytes transmitted, and number of packets for each port. Historical traffic data is usually stored in a centralized data warehouse in the format: {"port_id":"port1","timestamp":"2025-01-06 00:00:00","total_bytes":1000MB,"total_packets":5000}. During this process, the collected data is organized by time period to construct a historical traffic dataset, providing a data foundation for subsequent traffic fluctuation statistics.
[0175] Traffic fluctuation range data is obtained by statistically analyzing the historical traffic data of the switch.
[0176] In this embodiment, the traffic fluctuation range of each port is calculated based on historical traffic data. For example, the traffic fluctuation range of a port can be obtained by comparing the maximum and minimum traffic within each time period. If the traffic value within a certain time period exceeds the normal fluctuation range of the port (e.g., a fluctuation exceeding 30%), it is considered a potential anomaly. During this process, the traffic fluctuation amplitude of each port is statistically analyzed, and the normal fluctuation range of the port is defined based on this. This data can be obtained by calculating methods such as standard deviation and variance to quantify the degree of traffic fluctuation. For example, the standard deviation can be used to calculate the dispersion of traffic fluctuations, and the fluctuation range is represented by the ratio of the standard deviation to the average traffic value.
[0177] A traffic fluctuation baseline is constructed based on the traffic fluctuation range data to obtain the traffic fluctuation baseline data;
[0178] In this embodiment, the traffic fluctuation baseline for the entire network or a specific switch is calculated by aggregating the traffic fluctuation range data for each port. For example, assuming the traffic fluctuation range for multiple ports of a switch is [10%-50%], then the traffic fluctuation baseline for that switch is 50%. Constructing the traffic fluctuation baseline requires estimation based on fluctuation amplitudes in historical data, setting an acceptable standard as the normal fluctuation range for network traffic. This baseline will serve as a reference value for subsequent detection of abnormal traffic fluctuations. By analyzing traffic fluctuations over different time periods, stable traffic patterns are identified and categorized. For example, traffic fluctuations are larger on weekdays and smaller on holidays.
[0179] A traffic rate baseline is constructed from the traffic rate data to obtain the traffic rate baseline data;
[0180] In this embodiment, the normal speed range for each port or switch is calculated based on traffic rate data. For example, for a specific port on a switch, the average traffic rate over the past week is calculated, typically 30 Mbps, while the maximum rate is 100 Mbps. The traffic rate baseline is determined by analyzing the highest and lowest traffic rates in historical data to establish the normal operating range for each port. A corresponding rate baseline is constructed based on set threshold standards (e.g., traffic rates exceeding 80% of the baseline are considered abnormal). The construction of the traffic rate baseline relies not only on historical data but also on the performance parameters of network devices, such as the maximum bandwidth of switch ports, to ensure that the traffic rate remains within a reasonable range.
[0181] Based on the data packet data, traffic source optimization is performed on the traffic fluctuation baseline data to obtain optimized traffic fluctuation baseline data;
[0182] In this embodiment, the traffic fluctuation baseline is optimized by analyzing information such as source address, destination address, and packet size in historical traffic data. This step pays particular attention to traffic with abnormal fluctuations and unknown sources. If traffic fluctuations from certain source IPs are found to significantly exceed the predetermined baseline, they are marked as potential sources of abnormal traffic. Traffic source optimization not only helps improve the accuracy of the traffic fluctuation baseline but also identifies sources of security threats or abnormal traffic patterns. By analyzing the source and type of each packet and combining it with traffic fluctuation data, the traffic fluctuation baseline is updated to adapt to the constantly changing network environment.
[0183] Based on the traffic fluctuation baseline optimization data and the traffic rate baseline data, the switch traffic pattern baseline is integrated to obtain the switch traffic pattern baseline data.
[0184] In this embodiment, a comprehensive switch traffic pattern baseline is established by combining the optimized traffic fluctuation baseline and traffic rate baseline. The integration process requires analyzing traffic data from all ports to identify different traffic patterns and comparing them with historical baselines to further optimize the overall network traffic pattern. After integration, the traffic pattern baseline will help the real-time monitoring system more accurately identify traffic fluctuations and rate anomalies, facilitating subsequent abnormal traffic detection and security threat identification. Finally, the integrated traffic pattern baseline is stored as a standard dataset for subsequent real-time monitoring and security analysis.
[0185] Optionally, step S4 specifically includes:
[0186] Step S41: Extract link disconnection features and malware detection features based on network security threat data to obtain link disconnection data and malware detection data;
[0187] In this embodiment, link disconnection feature extraction and malware detection feature extraction are performed to obtain link disconnection data and malware detection data. Link disconnection feature extraction involves real-time collection of link status monitoring information from network devices, such as determining link normality through switch port status. Link disconnection events are typically reported by switches via the SNMP protocol; when a link disconnection is detected, information such as the port where the disconnection occurred, the time of disconnection, and the duration is recorded. Malware detection feature extraction detects malware activity by analyzing abnormal behavior in traffic, particularly by using Deep Packet Inspection (DPI) technology to identify the fingerprints of viruses, Trojans, or other malware. Malware characteristics include specific IP communication patterns, large amounts of abnormal traffic, and malicious requests. Malware detection feature extraction relies on a predefined malware fingerprint database and uses traffic analysis tools, such as IDS / IPS systems, for real-time traffic detection and analysis. These methods are used to extract link disconnection data and malware detection data, in the following formats: {"timestamp":"2025-01-07 13:00:00","event":"link_disruption","port_id":"port1","duration":"15seconds"} and {"timestamp":"2025-01-07 13:00:00","ip":"192.168.1.100","protocol":"TCP","malware_type":"worm"}.
[0188] Step S42: Identify security threats and risks in the abnormal network transmission topology based on the link disconnection data, thereby obtaining link disconnection security threat and risk data;
[0189] In this embodiment, based on port information, disconnection time, and duration from link disconnection data, combined with network topology information, it is determined which links in the network are abnormal. By analyzing the frequency and location of link disconnection events, the impact of the event on network transmission is assessed. Furthermore, by calculating the impact of link disconnections on network topology, potential security threats in the network are identified. For example, if certain links frequently disconnect and occur near important network nodes, it indicates the presence of external attacks or internal equipment failures. In this case, the security risk of link disconnection events can be quantified by setting a threshold, such as setting a threshold (e.g., ports with more than 3 disconnection events / hour are high-risk ports) for risk assessment, and marking these ports and links as potential sources of security threats.
[0190] Step S43: Identify security threat risks of abnormal network transmission topology based on malware detection data, thereby obtaining malware detection security threat risk data;
[0191] In this embodiment, potential security threats are identified by analyzing abnormal behavior in network traffic, especially traffic characteristics related to malware. If the traffic pattern of a certain IP or port is detected to be similar to known malware behavior (such as a large number of TCP connection requests, malicious DNS queries, etc.), then that node can be considered risky. Furthermore, by combining the network topology, the path of malware propagation and key nodes in the network are identified. For example, if an external IP is identified as a malicious source and communicates with multiple internal nodes in the network, then that external IP can be considered a potential source of security threats. Based on thresholds, such as traffic exceeding three times the normal range or request frequency exceeding ten times, it can be marked as malicious traffic and a security risk assessment can be performed.
[0192] Step S44: Integrate security threat risk data based on link disconnection security threat risk data and malware detection security threat risk data to obtain security threat risk data;
[0193] In this embodiment, link disconnection data and malware detection data are processed separately, and risk information is extracted from each type of data. Then, the two are integrated, and the location, time, and type of the risk are compared. If a link disconnection event and malware detection data occur within the same time period and are located on the same or nearby nodes in the network, these two risk events can be considered as related security threats and comprehensively assessed. During integration, a weighted method or priority ranking method can be used to comprehensively consider the duration and frequency of the link disconnection event and the abnormality of malware traffic, ultimately forming a comprehensive security threat risk score. For example, the risk score for a link disconnection event is 3 (in a risk scoring system from 1 to 5), while the risk score for malware detection is 4, resulting in a final integrated risk score of 7.
[0194] Step S45: Classify network security threats based on the abnormal network transmission topology of security threat risk data, thereby obtaining network security threat data topology data.
[0195] In this embodiment, network topology information is used to divide each node in the network. Based on security threat risk scores from link disconnection and malware detection, nodes with high security risks are identified. During this process, nodes exhibiting both link disconnection events and malware activity are given special attention and categorized as high-risk areas. Multiple risk levels, such as low, medium, and high, can be set during the segmentation process to allow for different levels of security protection measures. The output of the security threat data topology structure is typically presented as a network graph, where high-risk nodes and links are marked in red or other warning colors. Ultimately, this data will be used for subsequent security response and monitoring decisions. For example, a formatted output like {"node_id":"node1","risk_level":"high","associated_threats":["link_disruption","malware"]} indicates that node 1 presents a security threat related to link disconnection and malware.
[0196] Optionally, this specification also provides an intelligent information big data acquisition and management system for executing the intelligent information big data acquisition and management method described above. This intelligent information big data acquisition and management system includes:
[0197] The data dictionary maintenance module is used to acquire information from a big data warehouse, extract data dictionary features based on the big data warehouse to obtain data dictionary data, and maintain the data dictionary data to obtain data dictionary maintenance data.
[0198] The transmission anomaly analysis module is used to acquire switch device data, perform network topology analysis based on the switch device data to obtain network topology data of the switch device; and perform transmission anomaly analysis based on the network topology data of the switch device to obtain network transmission anomaly topology data.
[0199] The network security threat analysis module is used to perform traffic statistics based on switch device data to obtain switch traffic data; and to perform network security threat analysis based on switch traffic data to obtain network security threat data.
[0200] The network security threat classification module is used to classify network security threats based on the abnormal network transmission topology structure according to network security threat data, thereby obtaining network security threat data topology structure data;
[0201] The intelligent network security threat acquisition model construction module is used to construct an intelligent network security threat acquisition model based on data dictionary maintenance data and network security threat data topology structure data, thereby obtaining an intelligent network security threat acquisition model; and to manage network security threats in the information big data warehouse based on the intelligent network security threat acquisition model, thereby obtaining network security threat management data.
[0202] This invention discloses an intelligent information big data acquisition and management system. This system can implement any of the intelligent information big data acquisition and management methods of this invention. It is used to combine the operation and signal transmission media between various modules to complete the intelligent information big data acquisition and management method. The internal modules of the system cooperate with each other, thereby improving the efficiency of identifying network security threats.
[0203] Therefore, the embodiments should be considered as exemplary and non-limiting in all respects, and the scope of the invention is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of the equivalents of the application are intended to be included within the invention.
[0204] The above description is merely a specific embodiment of the present invention, enabling those skilled in the art to understand or implement the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the present invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features of the invention herein.
Claims
1. A method for intelligent collection and management of big data information, characterized in that, Includes the following steps: Step S1: Obtain the information big data warehouse, and extract data dictionary features based on the information big data warehouse to obtain data dictionary data; maintain the data dictionary data to obtain data dictionary maintenance data. Step S2: Obtain switch device data and perform network topology analysis based on the switch device data to obtain switch device network topology data; perform transmission anomaly analysis based on the switch device network topology data to obtain network transmission anomaly topology data. Step S3: Perform traffic statistics based on the switch device data to obtain switch traffic data; perform network security threat analysis based on the switch traffic data to obtain network security threat data. Step S3 specifically involves: Step S31: Perform traffic statistics based on the switch device data to obtain switch traffic data; Step S32: Detect phishing attacks based on switch traffic data to obtain phishing attack data; Step S33: Perform traffic anomaly monitoring based on switch traffic data to obtain traffic anomaly data. Step S33 specifically involves: Step S331: Preprocess the switch traffic data to obtain preprocessed switch traffic data; Step S332: Establish a traffic pattern baseline on the preprocessed switch traffic data to obtain switch traffic pattern baseline data. Step S332 specifically involves: Traffic rate features and data packet features are extracted from the preprocessed switch traffic data to obtain traffic rate data and data packet data. Obtain historical traffic data from the switch; Traffic fluctuation range data is obtained by statistically analyzing the historical traffic data of the switch. A traffic fluctuation baseline is constructed based on the traffic fluctuation range data to obtain the traffic fluctuation baseline data; A traffic rate baseline is constructed from the traffic rate data to obtain the traffic rate baseline data; Based on the data packet data, traffic source optimization is performed on the traffic fluctuation baseline data to obtain optimized traffic fluctuation baseline data; Based on the traffic fluctuation baseline optimization data and the traffic rate baseline data, the switch traffic pattern baseline is integrated to obtain the switch traffic pattern baseline data. Step S333: Perform real-time anomaly monitoring on switch traffic data based on the switch traffic pattern baseline data to obtain real-time anomaly switch traffic data. Step S334: Classify the anomaly types based on the real-time abnormal switch traffic data to obtain bandwidth abuse anomaly data and network worm anomaly data; Step S335: Perform correlation analysis based on bandwidth abuse anomaly data and network worm anomaly data to obtain potential security threat data; Step S336: Perform traffic anomaly detection on the switch traffic data based on potential security threat data to obtain traffic anomaly data; Step S34: Integrate network security threat data based on phishing attack data and abnormal traffic data to obtain network security threat data; Step S4: Based on the network security threat data, classify the abnormal network transmission topology to obtain network security threat data topology data; Step S5: Construct a network security threat intelligent collection model based on the data dictionary maintenance data and the network security threat data topology structure data, thereby obtaining the network security threat intelligent collection model; The network security threat management data is obtained by using an intelligent network security threat collection model to manage network security threats in a big data information warehouse.
2. The intelligent data acquisition and management method for big data according to claim 1, characterized in that, Step S1 is as follows: Step S11: Obtain the information big data warehouse, and extract data dictionary features based on the information big data warehouse to obtain data dictionary data; Step S12: Perform field standardization on the data dictionary data to obtain field-standardized data; Step S13: Perform encoding mapping based on the data dictionary data to obtain the standard mapping code value for the digital warehouse; Step S14: Perform dictionary maintenance based on the standardized field data and the standard mapping code value of the digital warehouse to obtain data dictionary maintenance data.
3. The intelligent data acquisition and management method for information big data according to claim 1, characterized in that, Step S2 is as follows: Step S21: Obtain switch device data and perform node identification based on the switch device data to obtain switch node data; Step S22: Extract port connection features based on switch device data to obtain port connection data; Step S23: Establish a primary topology based on switch node data and port connection data to obtain primary topology data; Step S24: Redundant path identification is performed on the primary topology data to obtain redundant path data; Step S25: Loop path identification is performed on the primary topology data to obtain loop path data; Step S26: Construct the network topology of the switch device based on the loop path data and redundant path data, thereby obtaining the network topology data of the switch device. Step S27: Perform transmission anomaly analysis based on the network topology data of the switch device to obtain network transmission anomaly topology data.
4. The intelligent data acquisition and management method for information big data according to claim 3, characterized in that, Step S27 is as follows: Step S271: Identify switch port faults based on the network topology data of the switch device to obtain switch port fault data; Step S272: Calculate the bandwidth utilization of the switch port fault data to obtain the bandwidth utilization rate; Step S273: Perform numerical statistics based on bandwidth utilization to obtain high bandwidth utilization; Step S274: Analyze network performance anomalies based on the network topology data of the switch device according to the high bandwidth utilization, thereby obtaining network performance anomaly data; Step S275: Perform link health anomaly detection on the network topology data of the switch device to obtain link health anomaly data; Step S276: Based on the abnormal link health data and abnormal network performance data, perform network transmission abnormal topology analysis on the network topology of the switch device to obtain abnormal network transmission topology data.
5. The intelligent data acquisition and management method for big data according to claim 1, characterized in that, Step S32 is as follows: Step S321: Perform DNS request analysis based on switch traffic data to obtain DNS request data; Step S322: Perform request volume statistics on the Domain Name System (DNS) request data to obtain high request volume data for the DNS; Step S323: Identify suspicious external IP addresses based on high request volume data from the Domain Name System to obtain suspicious external IP address data; Step S324: Perform short link detection based on switch traffic data to obtain short link data; Step S325: Identify abnormal Uniform Resource Locators (URLs) in the short link data to obtain abnormal URL data; Step S326: Based on the abnormal Uniform Resource Locator (URI) data and short link data, determine the phishing attack data of the suspicious external IP address data.
6. The intelligent data acquisition and management method for information big data according to claim 1, characterized in that, Step S4 is as follows: Step S41: Extract link disconnection features and malware detection features based on network security threat data to obtain link disconnection data and malware detection data; Step S42: Identify security threats and risks in the abnormal network transmission topology based on the link disconnection data, thereby obtaining link disconnection security threat and risk data; Step S43: Identify security threat risks of abnormal network transmission topology based on malware detection data, thereby obtaining malware detection security threat risk data; Step S44: Integrate security threat risk data based on link disconnection security threat risk data and malware detection security threat risk data to obtain security threat risk data; Step S45: Classify network security threats based on the abnormal network transmission topology of security threat risk data, thereby obtaining network security threat data topology data.
7. An intelligent data acquisition and management system for information big data, characterized in that, For executing the intelligent data acquisition and management method for information big data as described in claim 1, the intelligent data acquisition and management system for information big data includes: The data dictionary maintenance module is used to acquire information from a big data warehouse, extract data dictionary features based on the big data warehouse to obtain data dictionary data, and maintain the data dictionary data to obtain data dictionary maintenance data. The transmission anomaly analysis module is used to acquire switch device data, perform network topology analysis based on the switch device data to obtain network topology data of the switch device; and perform transmission anomaly analysis based on the network topology data of the switch device to obtain network transmission anomaly topology data. The network security threat analysis module is used to perform traffic statistics based on switch device data to obtain switch traffic data; and to perform network security threat analysis based on switch traffic data to obtain network security threat data. The network security threat classification module is used to classify network security threats based on the abnormal network transmission topology structure according to network security threat data, thereby obtaining network security threat data topology structure data; The intelligent network security threat acquisition model construction module is used to construct an intelligent network security threat acquisition model based on data dictionary maintenance data and network security threat data topology structure data, thereby obtaining an intelligent network security threat acquisition model; and to manage network security threats in the information big data warehouse based on the intelligent network security threat acquisition model, thereby obtaining network security threat management data.
Citation Information
Patent Citations
Computer network anomaly detection method
CN118784364A