Information big data intelligent acquisition management system and method

By performing data dictionary feature extraction and network topology analysis on information big data warehouse and switch equipment data, combined with intelligent acquisition model, the problem of difficulty in real-time monitoring of topology changes and identifying security threats in dynamic network environments is solved, and efficient and accurate network security protection is achieved.

CN120034375AActive Publication Date: 2025-05-23SHENZHEN IDEAL POWER INFORMATION TECHNOLOGY CO LTD

Patent Information

Application Number
CN202510175220.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-18
Publication Date
2025-05-23
Estimated Expiration
2045-02-18

AI Technical Summary

Technical Problem

Existing network security technologies are difficult to monitor topological changes and identify potential security threats in real time in dynamic network environments, resulting in insufficient adaptability and response speed of protection systems.

Method used

By obtaining information big data warehouses, data dictionary feature extraction and maintenance, network topology analysis and transmission abnormality analysis are carried out in combination with switch equipment data, and an intelligent acquisition model is built to identify and manage network security threats.

Benefits of technology

It realizes accurate management of data structures and features, dynamically monitors network topology changes and abnormal transmission, improves the accuracy and response speed of network security protection, and enhances the system's adaptability and real-time performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120034375A_ABST
    Figure CN120034375A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of network security, in particular to an information big data intelligent acquisition management system and method. The method comprises the following steps of obtaining an information big data warehouse, and performing data dictionary feature extraction according to the information big data warehouse so as to obtain data dictionary data; carrying out maintenance according to the data dictionary data so as to obtain data dictionary maintenance data; obtaining switch equipment data, and carrying out network topology structure analysis according to the switch equipment data so as to obtain switch equipment network topology structure data; performing transmission anomaly analysis according to the network topology structure data of the switch equipment so as to obtain network transmission anomaly topology structure data; carrying out traffic statistics according to the switch equipment data so as to obtain switch traffic data; and carrying out network security threat analysis according to the switch flow data so as to obtain network security threat data. According to the invention, the network security threat identification efficiency is improved based on the network security technology.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of network security technology, and in particular to an information big data intelligent collection and management system and method. Background Art

[0002] Existing methods mostly rely on static network topology maps, which makes it difficult to achieve real-time topology change monitoring in dynamic network environments. The ability to detect network transmission anomalies is limited, and it is easy to ignore dynamic changes in network topology structures and potential security threats, reducing the accuracy of network security protection. In traditional methods, most network security threat detection relies on manual rules, and the response mechanism is relatively simple. It is difficult to respond to a variety of security events in a timely and accurate manner, and lacks intelligent analysis capabilities for different types of threats, resulting in poor adaptability of the protection system. Existing information collection and management systems are usually based on fixed frameworks and processes, lack high scalability and adaptability, and cannot effectively respond to the challenges of emerging technology environments (such as cloud computing, big data, the Internet of Things, etc.). The lack of flexibility of the system makes it impossible to quickly adjust the response strategy in the face of a changing network environment, reducing the real-time and adaptability of the system. Summary of the invention

[0003] Based on this, it is necessary for the present invention to provide an information big data intelligent collection and management system and method to solve at least one of the above technical problems.

[0004] To achieve the above purpose, a method for intelligent collection and management of information big data includes the following steps:

[0005] Step S1: obtaining an information big data warehouse, and performing data dictionary feature extraction based on the information big data warehouse, thereby obtaining data dictionary data; performing maintenance based on the data dictionary data, thereby obtaining data dictionary maintenance data;

[0006] Step S2: acquiring switch device data, and performing network topology analysis based on the switch device data, thereby obtaining switch device network topology data; performing transmission anomaly analysis based on the switch device network topology data, thereby obtaining network transmission anomaly topology data;

[0007] Step S3: Performing traffic statistics according to the switch device data to obtain switch traffic data; performing network security threat analysis according to the switch traffic data to obtain network security threat data;

[0008] Step S4: performing network security threat classification on the abnormal network transmission topology structure according to the network security threat data, thereby obtaining network security threat data topology structure data;

[0009] Step S5: construct a network security threat intelligent collection model based on the data dictionary maintenance data and the network security threat data topology structure data, thereby obtaining the network security threat intelligent collection model; perform network security threat management on the information big data warehouse based on the network security threat intelligent collection model, thereby obtaining network security threat management data.

[0010] The present invention effectively realizes the precise management of data structure and features by acquiring the information big data warehouse and performing data dictionary feature extraction and maintenance, and improves the comprehensibility and consistency of data. On this basis, network topology structure analysis and transmission anomaly analysis are performed according to the information big data warehouse, and the topology changes and abnormal transmission conditions in the network environment can be dynamically monitored, and potential network problems can be effectively identified, thereby overcoming the defect that the traditional static topology map cannot adapt to the dynamic environment. By performing traffic statistics and network security threat analysis on the switch equipment data, the abnormal conditions in the network traffic can be fully identified and evaluated, and security threats can be discovered in time, thereby providing a strong basis for the implementation of protective measures. Further, by performing topological structure division on the network security threat data, different types of security threats and their impact ranges can be accurately located to ensure that security protection is more targeted. In combination with the data dictionary maintenance data and the network security threat topology structure data, the constructed intelligent acquisition model can realize intelligent security threat identification and response, thereby improving the system's adaptive ability and real-time response ability. Finally, the network security threats of the information big data warehouse are managed based on the intelligent acquisition model, and the protection strategy can be continuously optimized to ensure that the system still maintains efficient protection capabilities when facing complex network environments and emerging threats. This method has extremely high scalability and adaptability, and can effectively respond to challenges in emerging technology environments such as cloud computing, big data, and the Internet of Things. It improves the flexibility and real-time performance of network security protection systems and solves the various limitations of traditional methods in dynamic environments.

[0011] Optionally, step S1 specifically includes:

[0012] Step S11: obtaining an information big data warehouse, and performing data dictionary feature extraction based on the information big data warehouse, thereby obtaining data dictionary data;

[0013] Step S12: standardizing the fields of the data dictionary data to obtain field standardized data;

[0014] Step S13: Perform code mapping according to the data dictionary data, thereby mapping the code value to the digital warehouse standard;

[0015] Step S14: Perform dictionary maintenance based on the field standardization data and the digital warehouse standard mapping code value to obtain data dictionary maintenance data.

[0016] The present invention can provide a unified structural framework for data by acquiring information big data warehouse and extracting data dictionary features, ensuring that data of different sources and types can be uniformly managed and analyzed. Through the field standardization of data dictionary, the problem of inconsistency between different data formats and standards is solved, so that data can be more efficiently and uniformly calculated and analyzed in the subsequent processing process, thereby enhancing the overall coordination of the system and the consistency of data. By encoding and mapping the data dictionary, the standardized conversion of various types of data is effectively realized, thereby optimizing the accuracy and efficiency in the data processing process, ensuring that the data can be more accurately mapped to the digital warehouse, and laying the foundation for subsequent data processing and analysis. Further dictionary maintenance of standardized data and mapping code values ​​ensures the continuous updating and improvement of the data dictionary, so that the entire system can quickly adapt to and handle new data problems in the face of a constantly changing network environment, and improves the flexibility of data processing and the adaptability of the system. The implementation of this series of steps breaks through the limitations of traditional static topology maps and artificial rules, can better support real-time data monitoring and security threat detection in a dynamic environment, improves the real-time performance, adaptability and response speed of the system, and provides a strong technical support for realizing more intelligent and flexible network security protection.

[0017] Optionally, step S2 specifically includes:

[0018] Step S21: acquiring switch device data, and performing node identification according to the switch device data, thereby obtaining switch node data;

[0019] Step S22: extracting port connection features according to the switch device data, thereby obtaining port connection data;

[0020] Step S23: Establishing a primary topology structure according to the switch node data and the port connection data, thereby obtaining primary topology structure data;

[0021] Step S24: performing redundant path identification on the primary topology structure data, thereby obtaining redundant path data;

[0022] Step S25: performing loop path identification on the primary topology structure data, thereby obtaining loop path data;

[0023] Step S26: constructing a switch device network topology structure according to the loop path data and the redundant path data, thereby obtaining switch device network topology structure data;

[0024] Step S27: Perform transmission anomaly analysis based on the network topology data of the switch device, thereby obtaining network transmission anomaly topology data.

[0025] Through the acquisition and node identification of the switch device data, the present invention can efficiently identify the positions and connection relationships of each switch node in the network, laying a foundation for the subsequent construction of the network topology. Further, by extracting the port connection characteristics from the switch device data, the connection situation between switches in the network can be accurately captured, providing strong data support for constructing an accurate network topology map. Based on the node data and port connection data, a primary topology structure can be constructed, and basic information for further identifying redundant paths and loop paths is provided, enhancing the accuracy and scalability of the topology structure. The identification of redundant paths and loop paths can not only ensure high availability and stability in the network but also provide more detailed and accurate topology information for subsequent network transmission anomaly analysis. Combining the data of loop paths and redundant paths can more accurately construct the complete topology structure of the network, enabling potential anomalies and bottlenecks in the network to be identified and avoided in advance, thereby improving the system's detection ability for network anomalies. Through transmission anomaly analysis based on a complete topology structure, abnormal phenomena occurring in network transmission can be detected in a timely manner, providing real-time data support for security protection in a dynamic network environment and effectively avoiding the defect that traditional static topology maps cannot reflect network changes in real time, thereby enhancing the accuracy and flexibility of network security protection. In addition, through efficient topology construction and anomaly analysis, this series of steps can provide strong support for coping with the challenges of emerging technology environments (such as cloud computing, big data, Internet of Things, etc.), improving the system's real-time performance, adaptability, and intelligent processing ability, and ensuring that it can adjust and respond to diverse security threats in a timely manner when facing a constantly changing network environment.

[0026] Optionally, step S27 is specifically as follows:

[0027] Step S271: Identify switch port failures based on the switch device network topology structure data, so as to obtain switch port failure data;

[0028] Step S272: Calculate the bandwidth utilization rate for the switch port failure data, so as to obtain the bandwidth utilization rate;

[0029] Step S273: Conduct numerical statistics based on the bandwidth utilization rate, so as to obtain a high bandwidth utilization rate;

[0030] Step S274: Conduct network performance anomaly analysis on the switch device network topology structure data based on the high bandwidth utilization rate, so as to obtain network performance anomaly data;

[0031] Step S275: Detect link health anomalies for the switch device network topology structure data, so as to obtain link health anomaly data;

[0032] Step S276: Performing a network transmission abnormal topology structure analysis on the switch device network topology structure according to the link health abnormality data and the network performance abnormality data, thereby obtaining network transmission abnormal topology structure data.

[0033] The present invention can timely discover port faults by identifying switch port faults on the network topology data of the switch device, and provide first-hand data support for subsequent network performance optimization. The calculation of bandwidth utilization can comprehensively evaluate the use of network resources, thereby providing effective early warning for potential bandwidth bottlenecks and providing a more accurate basis for network performance analysis. Based on the numerical statistics of bandwidth utilization, nodes with high bandwidth utilization can be quickly identified, potential problems of excessive network load can be further explored, and reference can be provided for anomaly detection. Using data with high bandwidth utilization to perform network performance anomaly analysis helps to identify performance bottlenecks caused by high load in the network and provide effective directions for system optimization. Link health anomaly detection can monitor the health status of network links in real time, discover link anomalies and issue early warnings, avoiding network interruption or performance degradation caused by link failures. In addition, the network transmission abnormal topology structure analysis based on the integrated link health anomaly data and network performance anomaly data can deeply explore and reveal network transmission failures, thereby providing a strong guarantee for the stable operation of the network. Through the effective implementation of the above steps, the system can not only monitor network performance and link health status in real time, but also promptly identify and analyze network transmission anomalies, thereby improving the accuracy and response speed of network security protection, avoiding the limitations of traditional static methods in the face of dynamic network environments, and enhancing the system's flexibility, adaptability and ability to cope with emerging technology challenges.

[0034] Optionally, step S3 specifically includes:

[0035] Step S31: Perform traffic statistics according to the switch device data to obtain switch traffic data;

[0036] Step S32: Perform phishing attack detection according to the switch traffic data, thereby obtaining phishing attack data;

[0037] Step S33: monitoring traffic anomalies according to the switch traffic data, thereby obtaining traffic anomaly data;

[0038] Step S34: Integrate network security threats based on phishing attack data and traffic anomaly data to obtain network security threat data.

[0039] The present invention can comprehensively grasp the distribution and usage of network traffic by performing traffic statistics on switch equipment data, and provide important data support for subsequent security analysis. On this basis, phishing attack detection based on traffic data can timely discover abnormal traffic patterns and identify potential phishing attack threats, thereby providing early warning for preventing such attacks. Further monitoring of traffic anomalies on switch traffic data can dynamically monitor traffic fluctuations and timely capture traffic anomalies, such as bandwidth abuse or malicious traffic, thereby reducing the impact of potential risks on network security. By integrating phishing attack data with traffic anomaly data, a more accurate network security threat identification can be formed, and the detection accuracy and response speed of the network protection system can be improved. Overall, the system can efficiently perform network traffic monitoring and anomaly detection, combine phishing attack detection and traffic anomaly monitoring, capture various potential threats in the network in real time, and provide comprehensive and dynamic network security threat identification and protection through data integration, making up for the limitation that static network topology maps in traditional methods are difficult to adapt to dynamic network environments, and enhancing the intelligence, flexibility and adaptability of network security protection systems.

[0040] Optionally, step S32 is specifically:

[0041] Step S321: performing domain name system request analysis according to the switch traffic data, thereby obtaining domain name system request data;

[0042] Step S322: performing request volume statistics on the domain name system request data, thereby obtaining high request volume data of the domain name system;

[0043] Step S323: Identify suspicious external IP addresses based on the high request volume data of the domain name system, thereby obtaining suspicious external IP address data;

[0044] Step S324: Perform short link detection according to the switch traffic data to obtain short link data;

[0045] Step S325: performing abnormal uniform resource locator identification on the short link data, thereby obtaining abnormal uniform resource locator data;

[0046] Step S326: determine whether the suspicious external IP address data is a phishing attack based on the abnormal uniform resource locator data and the short link data, thereby obtaining phishing attack data.

[0047] The present invention can deeply understand the behavior pattern of domain name resolution in the network by performing domain name system (DNS) request analysis on switch flow data, and provide basic data for identifying potential abnormal requests. On this basis, the request volume statistics of DNS request data are helpful to find the peak period of abnormal traffic, thereby revealing potential malicious activities, such as DDoS attacks or domain name abuse. This analysis process can accurately locate suspicious external IP addresses with high request volumes, provide important clues for subsequent security protection, and identify attack sources or malicious manipulation behaviors. In addition, by performing short link detection on switch flow data, malicious short links can be effectively identified, which are related to phishing attacks or other network attacks. Further, by performing abnormal uniform resource locator (URL) identification on short link data, abnormal URL patterns can be revealed, thereby providing clues for tracking malicious activities. When the abnormal uniform resource locator is combined with the short link data, the association between the suspicious external IP address and the phishing attack can be more accurately confirmed, and finally it is determined whether there is a phishing attack in the network. Overall, this process strengthens the in-depth analysis of network traffic and helps identify complex security threats. Especially in a dynamically changing network environment, it can respond in real time and quickly respond to emerging security threats, thereby improving the system's intelligence, flexibility and ability to protect against diverse attacks.

[0048] Optionally, step S33 is specifically:

[0049] Step S331: pre-processing the switch flow data to obtain pre-processed switch flow data;

[0050] Step S332: Establishing a traffic pattern baseline for the pre-processed switch traffic data, thereby obtaining switch traffic pattern baseline data;

[0051] Step S333: performing real-time abnormal monitoring on the switch traffic data according to the switch traffic pattern baseline data, thereby obtaining real-time abnormal switch traffic data;

[0052] Step S334: classifying abnormal types according to the real-time abnormal switch traffic data, thereby obtaining bandwidth abuse abnormal data and network worm abnormal data;

[0053] Step S335: performing correlation analysis based on the bandwidth abuse abnormal data and the network worm abnormal data, thereby obtaining potential security threat data;

[0054] Step S336: Perform flow anomaly detection on the switch flow data according to the potential security threat data, thereby obtaining flow anomaly data.

[0055] The present invention can remove noise and normalize data by preprocessing the switch flow data, thereby providing a more accurate and clear data basis for subsequent analysis. Further establishing a flow pattern baseline for the preprocessed flow data can help identify normal network flow patterns and lay the foundation for the detection of abnormal flow. Real-time abnormal monitoring based on the flow pattern baseline can timely capture abnormal changes in flow, identify network attacks or performance problems, and generate real-time abnormal flow data. By classifying the real-time abnormal flow data into abnormal types, different types of abnormalities, such as bandwidth abuse or network worm attacks, can be clearly distinguished, thereby providing more detailed information for subsequent defense measures. Combining bandwidth abuse abnormal data with network worm abnormal data for correlation analysis can reveal potential security threats and help the security team to promptly discover attack behaviors that have not yet appeared in the network. Finally, flow anomaly detection based on potential security threat data can further confirm the abnormal flow in the network and ensure that the security protection system can more comprehensively monitor and respond to threats in the network. Overall, this process can achieve more efficient and accurate real-time monitoring and anomaly identification, improve the intelligence, response speed and adaptability of network protection, adapt to the dynamically changing network environment, and prevent potential threats from causing harm to network security.

[0056] Optionally, step S332 is specifically:

[0057] Performing flow rate feature extraction and data packet feature extraction on the pre-processed switch flow data, thereby obtaining flow rate data and data packet data;

[0058] Get the historical traffic data of the switch;

[0059] The flow fluctuation range is counted based on the historical flow data of the switch, thereby obtaining the flow fluctuation range data;

[0060] A flow fluctuation baseline is constructed according to the flow fluctuation range data, thereby obtaining flow fluctuation baseline data;

[0061] constructing a flow rate baseline for the flow rate data, thereby obtaining flow rate baseline data;

[0062] Optimizing the traffic source of the traffic fluctuation baseline data according to the data packet data, thereby obtaining the traffic fluctuation baseline optimization data;

[0063] The switch traffic pattern baseline data is integrated according to the traffic fluctuation baseline optimization data and the traffic rate baseline data, so as to obtain the switch traffic pattern baseline data.

[0064] The present invention can more accurately capture the key patterns and behaviors of network traffic by preprocessing the switch traffic data and extracting traffic rate and data packet features, thereby helping to identify potential abnormal traffic. By obtaining the historical traffic data of the switch and performing traffic fluctuation range statistics, the normal fluctuation range of the network traffic can be revealed, providing a benchmark for further traffic analysis. After establishing a traffic fluctuation baseline, the normal range of traffic fluctuations under different network states can be clearly defined to help discover abnormal fluctuations that exceed the normal range. By constructing a traffic rate baseline for traffic rate data, the normal value range of the traffic rate can be clarified, further improving the sensitivity to abnormal traffic rates. By optimizing the traffic fluctuation baseline in combination with data packet data, the baseline can be adjusted more accurately to reduce misjudgments caused by changes in the network environment. Finally, by integrating the traffic fluctuation baseline optimization data and the traffic rate baseline data, a comprehensive switch traffic pattern baseline is constructed, which can more accurately identify abnormal traffic patterns in the network. This series of steps improves the system's adaptability to dynamic network environments, enhances its monitoring and response capabilities to complex and changeable network conditions, effectively improves the system's identification and response speed to potential security threats, while reducing misjudgments or missed judgments caused by static methods, and improving the intelligence and real-time nature of network security protection.

[0065] Optionally, step S4 is specifically:

[0066] Step S41: extracting link disconnection features and malware detection features according to the network security threat data, thereby obtaining link disconnection data and malware detection data;

[0067] Step S42: Identify the security threat risk of abnormal network transmission topology according to the link disconnection data, thereby obtaining link disconnection security threat risk data;

[0068] Step S43: identifying security threat risks of abnormal network transmission topology according to the malware detection data, thereby obtaining malware detection security threat risk data;

[0069] Step S44: performing security threat risk integration according to the link disconnection security threat risk data and the malware detection security threat risk data, thereby obtaining security threat risk data;

[0070] Step S45: network security threats are divided according to the abnormal topological structure of network transmission of security threat risk data, thereby obtaining network security threat data topological structure data.

[0071] The present invention can accurately identify potential threats in the network, such as link interruption and malware activities, by extracting link disconnection features and malware detection features from network security threat data. This helps the system to more quickly capture abnormal events that affect network transmission and prevent network performance degradation or data leakage caused by network interruption or malware intrusion. Security threat risk identification based on link disconnection data can identify security risks caused by link disconnection and ensure the stability and security of network communication. Risk identification of abnormal network transmission topology structure through malware detection data can timely discover potential malware attacks or infections, thereby taking effective defense measures to prevent malware from further spreading or damaging the network. Integrating link disconnection security threat risk data and malware detection security threat risk data can help to comprehensively evaluate various security threats in the network, optimize risk identification strategies, and enhance network protection capabilities. Finally, by analyzing security threat risk data and combining network transmission abnormal topology structure for security threat classification, the distribution of different types of security threats in the network can be clarified, providing accurate basis for subsequent security response and protection strategy formulation, thereby enhancing the dynamic adaptability and protection capabilities of the network system, improving the overall security protection level, and reducing the difficulty of adapting to the ever-changing network environment.

[0072] Optionally, this specification also provides an information big data intelligent collection and management system, which is used to execute the information big data intelligent collection and management method as described above, and the game information big data intelligent collection and management system includes:

[0073] The data dictionary maintenance module is used to obtain the information big data warehouse, and perform data dictionary feature extraction based on the information big data warehouse to obtain data dictionary data; perform maintenance based on the data dictionary data to obtain data dictionary maintenance data;

[0074] The transmission anomaly analysis module is used to obtain switch device data, and perform network topology analysis based on the switch device data, thereby obtaining switch device network topology data; perform transmission anomaly analysis based on the switch device network topology data, thereby obtaining network transmission anomaly topology data;

[0075] The network security threat analysis module is used to perform traffic statistics based on the switch device data, thereby obtaining the switch traffic data; perform network security threat analysis based on the switch traffic data, thereby obtaining the network security threat data;

[0076] A network security threat classification module is used to classify network security threats based on network security threat data to obtain network security threat data topology data;

[0077] The network security threat intelligent collection model construction module is used to construct the network security threat intelligent collection model based on the data dictionary maintenance data and the network security threat data topology structure data, so as to obtain the network security threat intelligent collection model; according to the network security threat intelligent collection model, the network security threat management of the information big data warehouse is performed, so as to obtain the network security threat management data.

[0078] The present invention provides an information big data intelligent collection and management system, which can implement any information big data intelligent collection and management method of the present invention, and is used to combine the operation and signal transmission medium between various modules to complete the information big data intelligent collection and management method. The internal modules of the system cooperate with each other, thereby improving the efficiency of identifying network security threats. BRIEF DESCRIPTION OF THE DRAWINGS

[0079] Other features, objects and advantages of the present invention will become more apparent from the detailed description of non-limiting embodiments thereof made with reference to the following drawings:

[0080] Figure 1 This is a schematic diagram of the steps of the information big data intelligent collection and management method of the present invention;

[0081] Figure 2 Detailed step flow diagram of step S1 in the present invention;

[0082] Figure 3 Detailed step flow diagram of step S2 in the present invention.

[0083] The realization of the purpose, functional features and advantages of the present invention will be further explained in conjunction with embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION

[0084] The following is a clear and complete description of the technical method of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by technicians in this field without creative work are within the scope of protection of the present invention.

[0085] In addition, the accompanying drawings are only schematic illustrations of the present invention and are not necessarily drawn to scale. The same reference numerals in the figures represent the same or similar parts, and their repeated description will be omitted. Some of the block diagrams shown in the accompanying drawings are functional entities and do not necessarily correspond to physically or logically independent entities. The functional entities can be implemented in software form, or implemented in one or more hardware modules or integrated circuits, or implemented in different networks and / or processor methods and / or microcontroller methods.

[0086] It should be understood that, although the terms "first", "second", etc. may be used herein to describe various units, these units should not be limited by these terms. These terms are used only to distinguish one unit from another unit. For example, without departing from the scope of the exemplary embodiments, the first unit may be referred to as the second unit, and similarly the second unit may be referred to as the first unit. The term "and / or" used herein includes any and all combinations of one or more of the listed associated items.

[0087] To achieve this, please refer to Figures 1 to 3 The present invention provides a method for intelligent collection and management of information big data, the method comprising the following steps:

[0088] Step S1: obtaining an information big data warehouse, and performing data dictionary feature extraction based on the information big data warehouse, thereby obtaining data dictionary data; performing maintenance based on the data dictionary data, thereby obtaining data dictionary maintenance data;

[0089] In this embodiment, the original information including metadata and data table structure is extracted from the big data warehouse by connecting to the API interface of the information big data warehouse. This process usually involves SQL query and ETL (extraction, transformation, loading) technology, and a preliminary data dictionary is constructed by specifying the table name, field name and data type information of the data warehouse. The specific operation of data dictionary feature extraction includes analyzing the name of the field, the data type (such as integer, string, date type, etc.), and the constraints of the field (such as primary key, foreign key, unique constraint, etc.). In order to obtain a valid data dictionary, the extraction process needs to follow the data dictionary specification, such as the field name length does not exceed 50 characters, and the field type conforms to the standard database type (such as VARCHAR (255) or INT). Next, based on these data dictionary features, data dictionary maintenance is performed, and regular expressions are used to normalize the format of the field name to ensure the uniformity of the field name. For example, if there are fields "User_Id" and "user_id" with the same meaning, they are unified as "user_id". At this time, the data dictionary maintenance rules are used to standardize the fields and update them to the data dictionary maintenance table, recording the standard name, mapping relationship and other auxiliary information of each field.

[0090] Step S2: acquiring switch device data, and performing network topology analysis based on the switch device data, thereby obtaining switch device network topology data; performing transmission anomaly analysis based on the switch device network topology data, thereby obtaining network transmission anomaly topology data;

[0091] In this embodiment, the real-time operation data of the switch device is obtained through SNMP (Simple Network Management Protocol) or CLI (Command Line Interface), including port status, link status, MAC address table, VLAN information, etc. Through these data, the switch network topology structure is first analyzed. The topology structure analysis mainly depends on the connection information between the switch ports and the physical links between the switches. The specific method is to build a preliminary topology map through the port connection data of the switch, for example, by analyzing the MAC address table of the switch port, the port connection type (for example, using Ethernet or optical fiber connection) and the status of each port (such as "up" or "down") to generate the network topology structure. After the topology structure is built, the transmission anomaly of the switch device is analyzed. At this time, it is necessary to perform anomaly detection based on the bandwidth utilization and transmission delay of the switch port. The bandwidth utilization threshold can be set to 90%. If the bandwidth utilization of a port exceeds 90% continuously and exceeds 5 minutes, it is considered that the port has a transmission anomaly. Combined with the RTT (round trip time) value of the link, if the RTT of the link is greater than 50ms, it is further marked as network transmission abnormality topology structure data and stored.

[0092] Step S3: Performing traffic statistics according to the switch device data to obtain switch traffic data; performing network security threat analysis according to the switch traffic data to obtain network security threat data;

[0093] In this embodiment, the flow data of each port is obtained through the flow statistics interface of the switch (such as NetFlow or sFlow protocol), including the number of packets, number of bytes, packet loss rate, flow direction (inbound or outbound), etc. of each port. The flow statistics data need to be grouped by time window, usually 5 minutes as a time window, and the flow of each port is counted, and the maximum, minimum, average flow and burst flow in each time period are recorded. After the flow statistics data is completed, the network security threat analysis is further performed. The core of this part is to identify potential security threats through traffic pattern recognition, such as phishing attacks, DDoS attacks, etc. Taking phishing attacks as an example, first identify the number of abnormal domain name requests through DNS query log analysis. Set a threshold (for example, 100 requests). If the number of requests for a domain name in a short period of time exceeds the threshold, the domain name is related to the phishing attack. In addition, analyze the abnormal fluctuations of the flow data, such as the increase in flow in a short period of time exceeds 50%, it is determined to be abnormal flow, which is related to the DDoS attack. After these analysis results are summarized, network security threat data is formed, recording the type of attack and the time and place of its occurrence.

[0094] Step S4: performing network security threat classification on the abnormal network transmission topology structure according to the network security threat data, thereby obtaining network security threat data topology structure data;

[0095] In this embodiment, by analyzing the network security threat data generated in the previous stage, potential attack sources or security risk nodes are identified. The specific method is to match the network security threat data (such as phishing attacks, DDoS attacks, etc.) with the network topology data of the switch device, and determine the threat node by analyzing the attack source IP address and the direction of abnormal traffic. Then, these nodes are divided according to the data of the abnormal topology structure. The specific division method is to associate the abnormal links in the network (such as bandwidth abnormalities, high packet loss rate, etc.) with the identified threat data. For example, when the bandwidth utilization of a link is too high and there is an association with a known malicious IP address, the link is marked as a security threat link. Ultimately, these data will form a network security threat data topology structure, which will help locate potential attack sources and provide a basis for further defense measures.

[0096] Step S5: construct a network security threat intelligent collection model based on the data dictionary maintenance data and the network security threat data topology structure data, thereby obtaining the network security threat intelligent collection model; perform network security threat management on the information big data warehouse based on the network security threat intelligent collection model, thereby obtaining network security threat management data.

[0097] In this embodiment, the data dictionary maintenance data and the network security threat data topology data are used to build an intelligent collection model. This model is based on a machine learning algorithm (such as a decision tree, a random forest, or a neural network), and its input features include node information in the network topology (such as switches, port status, bandwidth utilization, etc.), and indicators of network security threats (such as traffic volatility, DNS request anomalies, etc.). By training the model, pattern learning can be performed based on historical security threat data to generate a model that can identify and classify security threats in real time. In order to ensure the effectiveness of the model, hyperparameter tuning is required, such as setting the depth of the decision tree, the number of trees in the random forest, etc. After the model is built, the model is used to manage network security threats in the information big data warehouse. The model obtains new network traffic data from the data warehouse in real time, and analyzes it through a classifier to identify new security threats. If a security threat is detected, the model will trigger an alarm and update the network security threat management data, including information such as the affected device, the type of attack, and the time of occurrence. This information will help managers take necessary protective measures, such as isolating infected nodes or enhancing firewall rules.

[0098] Optionally, step S1 specifically includes:

[0099] Step S11: obtaining an information big data warehouse, and performing data dictionary feature extraction based on the information big data warehouse, thereby obtaining data dictionary data;

[0100] In this embodiment, it is necessary to obtain structured data from the information big data warehouse. This process is usually completed by connecting to a database management system (such as MySQL, Oracle, or Hadoop), relying on SQL queries or directly calling APIs to obtain metadata in the warehouse (for example, table names, field names, data types, relationships between tables, etc.). Assuming that the data warehouse uses MySQL, first use the SQL query statement SHOW TABLES to obtain all table names, and then use DESCRIBE<table_name> Get the field features of each table, extract the field name, data type, whether NULL is allowed, field length and other information. For field data types, special attention should be paid to the standardization of data types, such as unifying "VARCHAR(100)" and "TEXT" into string types (STRING). In addition, it is necessary to extract structural data such as field index information, primary and foreign key relationships, etc., which are important components of building a data dictionary. During the data dictionary feature extraction process, all field names must be recorded in a unified format, such as using lowercase letters and underscores to separate (such as user_id instead of UserId) to ensure data normalization.

[0101] Step S12: standardizing the fields of the data dictionary data to obtain field standardized data;

[0102] In this embodiment, the fields extracted from the data dictionary are standardized. The purpose of field standardization is to ensure the consistency of field naming and data format, which is convenient for subsequent data processing and analysis. The specific operation is to convert the field name to a unified format. For example, if the field name has a camel case naming (such as UserName), it is converted to an underscore nomenclature (such as user_name). For the standardization of data types, it is necessary to unify different types of fields. For example, all date fields (such as DATE, DATETIME, TIMESTAMP) are standardized to the "DATE" type, and all numeric fields (such as INT, BIGINT, FLOAT) are standardized to the "INTEGER" or "FLOAT" type. In addition, the constraints of the field (such as "NOTNULL" or "UNIQUE") should be retained during the standardization process to ensure the integrity of the data. The standardization operation also needs to check whether the field length meets the database design specifications and make appropriate adjustments. For example, if the field length exceeds the recommended value, it is corrected to ensure efficient storage and query of the table. After standardization, all fields should meet the unified naming specifications, and the attributes (type, length, constraints) of each field are in a standardized format.

[0103] Step S13: Perform code mapping according to the data dictionary data, thereby mapping the code value to the digital warehouse standard;

[0104] In this embodiment, it is necessary to encode and map the fields in the data dictionary. The goal of encoding and mapping is to convert non-digital data (such as strings, dates, etc.) in the field into a processable digital form. First, for text fields (such as user_name, product_type), the unique values ​​in the field (such as user name, product type, etc.) are mapped to unique digital IDs. Specifically, for each text value (such as "admin" or "user"), a unique number is assigned to it, for example, "admin" corresponds to 1, and "user" corresponds to 2. This mapping usually adopts a hash algorithm, a dictionary mapping or a category coding method. In the encoding process, for the data of each field, all its unique values ​​are first counted, and a unique digital code is assigned to these values ​​in alphabetical or numerical order. In addition, for date type fields (such as created_at), the date can be converted into a custom digital format, such as converting the date field into a timestamp (the number of seconds from January 1, 1970 to the current date). This method can ensure that the data is stored in a digital format in the subsequent processing process, which is convenient for calculation and analysis. Finally, a mapping dictionary or mapping table is generated through code mapping to record the relationship between the original value and digital code of each field, which is convenient for subsequent query and data restoration.

[0105] Step S14: Perform dictionary maintenance based on the field standardization data and the digital warehouse standard mapping code value to obtain data dictionary maintenance data.

[0106] In this embodiment, the key to dictionary maintenance is to combine the standardized field data with the results of the code mapping to ensure that the data dictionary is continuously updated and maintained. In the specific implementation, it is first necessary to supplement and update the original data dictionary based on the data dictionary and mapping relationship generated in steps S12 and S13. The update process includes merging the field name, type information and corresponding code mapping value after field standardization into the dictionary. For example, if the field user_name is renamed to user_name during the standardization process, and its text value (such as "admin") is converted to a digital ID (such as 1) through code mapping, it is added to the data dictionary as a new dictionary item. On this basis, it is also necessary to regularly check and maintain the dictionary to ensure the integrity and consistency of the dictionary. For example, when a new field is added to a data table, it should first be standardized and its code mapping should be added to the dictionary. For deleted fields, the relevant entries need to be cleared from the dictionary. Data dictionary maintenance also needs to include field description information (such as the purpose, scope, restrictions, etc. of the field) to ensure that each field has a detailed description for subsequent data management and use. Through this process, the data dictionary maintenance data finally obtained not only contains field names, data types and coding mappings, but also field descriptions, constraints and other auxiliary information, ensuring the standardization and efficiency of data management.

[0107] Optionally, step S2 specifically includes:

[0108] Step S21: acquiring switch device data, and performing node identification according to the switch device data, thereby obtaining switch node data;

[0109] In this embodiment, it is necessary to obtain network topology information from the switch device, and this data can be collected through management protocols (such as SNMP, LLDP) or the syslog log of the switch. The switch device data includes the MAC address, IP address, port number, connected device information, etc. of the switch. After the data is acquired, this information is used to identify the node. A node is a device in the network that can receive and send data (such as a switch port, server, or terminal). According to the switch device data, the device connected to each port is identified by matching the MAC address and IP address of each switch port. In order to accurately identify the node, an accurate mapping algorithm must be used. The node data includes a unique identifier for each port (such as switch1_port2) and the device information connected to it. The device is another switch, server, or network terminal. In the process of node identification, it is also necessary to rely on the network topology discovery protocol (such as CDP, LLDP, etc.) to confirm the relationship between the switch and other network devices. The data after node identification will be stored in a standardized format, such as JSON or CSV files, to ensure that the subsequent processing of topology data can be unified.

[0110] Step S22: extracting port connection features according to the switch device data, thereby obtaining port connection data;

[0111] In this embodiment, the relevant information of each port in the switch data is parsed, such as the port number, port type (such as Ethernet, SFP), connection status (whether enabled) and bandwidth. The connection characteristics of each port involve the following key parameters: the physical connection status of the port (UP / DOWN), data transmission rate, port type (such as 10 / 100 / 1000Mbps, fiber port), the device bound to the MAC address, and the historical performance of the port (such as bandwidth utilization, packet loss rate). This information can be captured by the output of the SNMP interface of the switch or the CLI command of the switch. The feature extraction process is performed by parsing the switch device configuration file (such as the show interfaces command output). For each port, special attention is paid to data such as bandwidth, link status, packet loss rate, etc., and these data are recorded and stored as port connection characteristics, and the data is output in JSON format, such as {"port_id":"port1","status":"UP","bandwidth":"1000Mbps","connected_device":"device123"}. Through these port connection data, it is possible to provide a basis for the subsequent topology structure to be established.

[0112] Step S23: Establishing a primary topology structure according to the switch node data and the port connection data, thereby obtaining primary topology structure data;

[0113] In this embodiment, the node data of the switch, that is, the port identifier of the switch and the device information connected to it, is used to form a preliminary topology. For each pair of devices connected through the port, a virtual connection is created to record the connected port number, the information of the connected device, and the connection status. A connection matrix is ​​established through this information, in which each item of the matrix represents a connected port pair. The data model of the topology structure should include a node identifier (such as a switch port ID), a connection status (UP / DOWN), bandwidth, and link quality. When the topology structure is initially constructed, a graphical representation method is used, such as using an adjacency matrix or an adjacency list to represent the connection between devices. This data structure ensures that the connection relationship between each device (node) and other devices is clearly represented, for example, {"device_id":"switch1","connected_ports":["port1","port2"],"device_status":"active"}. The topology structure at this time can be checked and adjusted by a network topology visualization tool to ensure the accuracy and integrity of the data structure.

[0114] Step S24: performing redundant path identification on the primary topology structure data, thereby obtaining redundant path data;

[0115] In this embodiment, there are multiple paths connecting the same node, which are usually used to improve the reliability of the network. According to the primary topology data, a graph theory algorithm (such as the Dijkstra algorithm or the Floyd-Warshall algorithm) is applied to identify the shortest path between all devices. If it is found that there are multiple equal-weight paths connecting the same node, it can be determined as a redundant path. When identifying redundant paths, focus on bandwidth, link status, and path reliability. In the path selection of each node, if there is a backup path (for example, two different links connecting the same node), the path is a redundant path. The identification of redundant paths needs to ensure that the data transmission efficiency is not affected, and the redundant paths are marked in the data structure and their information is recorded. For example, redundant paths can be represented in the following format: {"path_id":"path1","redundant":true,"nodes":["switch1","switch2","switch3"],"status":"active"}. In this way, it is possible to identify which paths in the network are redundant in order to further optimize network performance.

[0116] Step S25: performing loop path identification on the primary topology structure data, thereby obtaining loop path data;

[0117] In this embodiment, a loop refers to a network topology in which a data packet can be transmitted cyclically between multiple paths, resulting in network congestion or data packet loss. The techniques used to identify the loop path include graph search algorithms such as depth-first search (DFS) and breadth-first search (BFS). In the specific implementation, each node and all ports connected to it are first checked, and each path is traversed to find the loop. If a path eventually returns to the source node, it is determined to be a loop path. After identifying the loop path, the starting point, end point, nodes passed through, and bandwidth utilization of the loop should be recorded. The data of the loop path should be represented in a structured manner, such as {"loop_id":"loop1","nodes":["switch1","switch2","switch3"],"path":["port1","port2","port3"],"status":"detected"}. After the loop path is identified, the network design can be further adjusted to avoid network performance problems caused by the loop.

[0118] Step S26: constructing a switch device network topology structure according to the loop path data and the redundant path data, thereby obtaining switch device network topology structure data;

[0119] In this embodiment, loop paths and redundant paths are removed to ensure the uniqueness and efficiency of the network topology. For redundant paths, it should be considered whether link aggregation or load balancing technology needs to be enabled. For loop paths, it is necessary to remove loop connections through a spanning tree algorithm (such as Spanning Tree Protocol, STP) to retain the optimal working path. The constructed network topology should contain the unique identifier of each device (node), the port connection status, the link status, and the bandwidth information of each path. In the topology data model, the active connection status and redundant connections of each node should be clearly marked to ensure smooth data transmission. The data of the topology structure can be stored in a graph database or matrix representation, such as {"network_id":"network1","nodes":[{"id":"switch1","ports":["port1","port2"],"status":"active"}],"links":[{"from":"switch1","to":"switch2","status":"active"}]}, and it is convenient for network optimization and maintenance.

[0120] Step S27: Perform transmission anomaly analysis based on the network topology data of the switch device, thereby obtaining network transmission anomaly topology data.

[0121] In this embodiment, the constructed network topology is checked using transmission anomaly analysis technology to identify network performance problems or faults. During the anomaly analysis process, the bandwidth utilization, packet loss rate, and delay of the switch port are first monitored. These data are collected regularly through protocols such as SNMP or NetFlow. In the transmission anomaly analysis, a threshold (such as bandwidth utilization exceeding 85% or packet loss rate exceeding 5%) is set as the standard for anomaly detection. If the performance indicator of a link or port exceeds the threshold, the path is marked as an abnormal path. In addition, trend analysis is required in combination with historical performance data to find potential performance bottlenecks or failure points. The network transmission abnormal topology data should include the node information of the abnormal link, the abnormal type (bandwidth problem, packet loss, link disconnection, etc.) and the time of occurrence. The abnormal data should be represented as follows: {"link_id":"link1","status":"error","error_type":"high_bandwidth_utilization","value":"90%"}, and provide a basis for subsequent fault diagnosis.

[0122] Optionally, step S27 is specifically:

[0123] Step S271: performing switch port fault identification according to the switch device network topology data, thereby obtaining switch port fault data;

[0124] In this embodiment, the port fault is identified by extracting the port status information from the network topology data of the switch device. First, the connection status information of each switch port needs to be obtained from the network topology data, such as whether the port is in the "UP" or "DOWN" state. If the status of a port is "DOWN" and the connected device is inaccessible, it is indicated that the port has a fault. In addition to the port status, the error counter of the port should also be monitored, including indicators such as packet loss rate, error frame, CRC error, and operation error. By analyzing the CLI command output or SNMP monitoring data (such as showinterfaces) of the switch, the error counter value of each port is checked. If the error count exceeds the preset threshold (such as the number of error frames is greater than 1000), it is determined that the port has a fault. The fault identification data should include the port ID, the fault type (such as physical disconnection, link error, etc.) and the time when the fault occurs, and the format is {"port_id":"port1","status":"DOWN","error_type":"CRC_error","error_count":1200}.

[0125] Step S272: Calculate the bandwidth utilization of the switch port fault data to obtain the bandwidth utilization;

[0126] In this embodiment, real-time traffic data of the switch port is collected through SNMP, NetFlow or sFlow protocol, and these data include the number of input and output bytes of the switch port. Port traffic refers to the amount of data transmitted through the port per unit time. Then, the maximum bandwidth of each switch port is determined. This value is set by the switch configuration, usually the rate of the port, such as 1000Mbps, 10000Mbps, etc. Based on the collected traffic data and the maximum bandwidth of the port, the bandwidth utilization of the port is calculated. During the calculation process, in order to ensure the accuracy of the data, the monitoring period of bandwidth utilization is usually set to once every 5 minutes, and continuous monitoring is carried out for 30 minutes to capture changes in port bandwidth utilization. During the monitoring process, it is also necessary to pay attention to the dynamic changes of the network topology, such as changes in the device connected to the port or adjustments to the bandwidth configuration, so the data needs to be updated in real time to ensure its timeliness. After the bandwidth utilization calculation is completed, the relevant data will be stored in a certain format. For example, the storage result can record the bandwidth utilization and related time points for each port, and the data format is {"port_id":"port1","bandwidth_utilization":"20%","time_period":"2025-01-07 15:00:00"}, for subsequent analysis and processing.

[0127] Step S273: Perform numerical statistics according to the bandwidth utilization, so as to obtain high bandwidth utilization;

[0128] In this embodiment, a threshold for high bandwidth utilization is defined, such as a bandwidth utilization exceeding 80% is considered to be high bandwidth utilization. Next, the bandwidth utilization data of all ports are traversed to extract those ports whose bandwidth utilization exceeds 80%. During the statistical process, statistics such as standard deviation and average value are used to analyze the bandwidth utilization of each port. For example, if the port list of a switch is [10%, 20%, 90%, 85%, 70%], the ports with high bandwidth utilization are identified as 90% and 85% according to the threshold. The statistical results will be output in the form of a list and include port ID, bandwidth utilization and abnormal type (such as "high bandwidth utilization"). For example, the statistical results are {"high_utilization_ports":[{"port_id":"port3","bandwidth_utilization":90%},{"port_id":"port4","bandwidth_utilization":85%}]}.

[0129] Step S274: performing network performance anomaly analysis on the network topology data of the switch device according to the high bandwidth utilization, thereby obtaining network performance anomaly data;

[0130] In this embodiment, based on the data of high bandwidth utilization, it is analyzed whether these ports have the risk of bandwidth bottleneck or network performance degradation. During the analysis, the bandwidth utilization is combined with performance indicators such as packet loss rate and delay. If the bandwidth utilization is high and accompanied by an increase in packet loss rate (such as a packet loss rate of more than 2%), it is considered to be a performance abnormality. Further, by comparing historical performance data with real-time traffic data, an abnormal trend analysis is performed to identify the links or nodes that cause performance degradation. In the abnormal analysis, the peak and average values ​​of network traffic are considered to timely discover temporary or long-term performance problems. For example, if the bandwidth utilization of port3 is 90% and the packet loss rate is 3%, the network performance abnormality data of the port is {"port_id":"port3","bandwidth_utilization":90%,"packet_loss":3%,"status":"performance_issue"}. This data can be used to further optimize network performance.

[0131] Step S275: performing link health anomaly detection on the network topology data of the switch device, thereby obtaining link health anomaly data;

[0132] In this embodiment, link health detection is performed by collecting and analyzing indicators such as the packet loss rate, delay, error frame, and link status of the link. First, a standard threshold for the health of each link is set. For example, when the packet loss rate is greater than 2%, the link status is not "UP", or the number of error frames exceeds 1000, the link is judged to be abnormally healthy. Secondly, by regularly monitoring the error count and link status of the switch port, if the error frame, packet loss rate, or delay of a link exceeds the threshold, it is recorded as a link health abnormality. Link health data should include information such as link ID, health status, number of error frames, and packet loss rate. The format of link health abnormality data is as follows: {"link_id":"link1","status":"unhealthy","error_count":1500,"packet_loss":3%,"latency":"200ms"}. By detecting the health status of the link, potential link failures or performance bottlenecks can be identified in advance, providing data support for subsequent network optimization.

[0133] Step S276: Performing a network transmission abnormal topology structure analysis on the switch device network topology structure according to the link health abnormality data and the network performance abnormality data, thereby obtaining network transmission abnormal topology structure data.

[0134] In this embodiment, the network performance abnormal port found in step S274 is associated with the link health abnormal link in step S275. If a port or link has a performance abnormality, and the port or link is part of a critical path in the network, the path is determined to be a network transmission abnormal path. Through the shortest path algorithm in graph theory, the affected paths in the network topology are analyzed to identify abnormal links and related nodes. Abnormal data should include multi-dimensional information such as path, node, link status, bandwidth, packet loss rate, and delay. For example, if a path from switch1 to switch2 passes through a link with high bandwidth utilization and poor link health, the network transmission abnormal data of the path is {"path_id":"path1","status":"abnormal","nodes":["switch1","switch2"],"link_status":"unhealthy","packet_loss":5%,"bandwidth_utilization":90%}. This data can be further used to analyze network bottlenecks and fault points to help network tuning.

[0135] Optionally, step S3 specifically includes:

[0136] Step S31: Perform traffic statistics according to the switch device data to obtain switch traffic data;

[0137] In this embodiment, the flow data of the switch device is collected. The flow statistics rely on protocols such as SNMP (Simple Network Management Protocol), NetFlow or sFlow, which can provide the real-time input and output data bytes of the port. By collecting these flow data, the data transmission volume of each switch port in a certain period of time is counted. These statistical values ​​include but are not limited to the number of input and output bytes, the number of data packets, the flow rate, etc. Based on these real-time data, the overall situation of the switch flow can be obtained. The data collection period is set to once every 5 minutes to ensure that the flow fluctuations in different time periods are captured.

[0138] Step S32: Perform phishing attack detection according to the switch traffic data, thereby obtaining phishing attack data;

[0139] In this embodiment, phishing attacks are identified by analyzing the characteristics of switch traffic, such as DNS request frequency, abnormal amount of external IP requests, etc. DNS query logs and external IP address access patterns are used as the basis for judgment. For example, if the number of domain names requested by a certain IP address exceeds the conventional threshold (such as more than 10 requests / minute), the IP address is considered suspicious and is the source of a phishing attack. At the same time, it is also necessary to monitor the suspiciousness of external URLs and analyze whether there are a large number of short links or malicious URLs. These data are further screened through correlation analysis, and finally phishing attack data is identified, and a phishing attack alert is generated.

[0140] Step S33: monitoring traffic anomalies according to the switch traffic data, thereby obtaining traffic anomaly data;

[0141] In this embodiment, a baseline is established for the switch flow data, and the normal flow fluctuation range is calculated. A flow baseline is established through historical flow data, such as the mean and standard deviation of the flow. Real-time flow data is compared with the baseline, and when the flow deviates from the normal baseline by more than a set threshold (such as more than 2 times the standard deviation), a flow anomaly alarm is triggered. In addition, different types of flow anomaly monitoring indicators can be set, such as bandwidth abuse, DDoS attack traffic, or malicious traffic patterns. The specific threshold of abnormal traffic is dynamically adjusted based on historical data and network load conditions. The monitoring system performs a data check every 5 minutes to capture and record potential flow anomalies.

[0142] Step S34: Integrate network security threats based on phishing attack data and traffic anomaly data to obtain network security threat data.

[0143] In this embodiment, association analysis is performed in combination with the suspicious IP, domain name, and type of traffic anomaly of the phishing attack. For example, when the access volume of a suspicious IP address increases suddenly and abnormal traffic characteristics are detected, the system will automatically mark the IP as high risk and give priority to it. In addition, it is also necessary to consider the contextual information of other security threats, such as device location, timestamp, etc., and combine this information with the identified phishing attacks and traffic anomaly data to generate a comprehensive network security threat report. These data are stored in a structured format, such as {"threat_type":"phishing","ip_address":"192.168.1.100","anomaly_type":"high_traffic","timestamp":"2025-01-0715:00:00"}, for subsequent security analysis and response.

[0144] Optionally, step S32 is specifically:

[0145] Step S321: performing domain name system request analysis according to the switch traffic data, thereby obtaining domain name system request data;

[0146] In this embodiment, Domain Name System (DNS) request data is extracted from the switch traffic data. These request data are obtained through switch traffic monitoring tools (such as NetFlow, sFlow or using SNMP protocol). In the switch traffic data, DNS requests are usually expressed as request packets sent to the DNS server. Each request contains information such as the source IP address, the IP address of the target DNS server, the requested domain name, and a timestamp. By extracting these requests, a request data set is established. The detailed information of each request (such as the requested domain name, time, and requested IP address) is recorded and can be stored in chronological order for subsequent analysis and processing. The collection frequency of monitoring data is set to once every 5 minutes to ensure that the dynamic changes of DNS requests are captured in a timely manner.

[0147] Step S322: performing request volume statistics on the domain name system request data, thereby obtaining high request volume data of the domain name system;

[0148] In this embodiment, the number of requests for each DNS domain name is counted every certain time window (such as every hour or every day). If the request volume of a domain name fluctuates drastically over a period of time (for example, the request volume exceeds twice the normal range), the request volume of the domain name is considered abnormal. When counting, an independent statistical analysis is performed for each domain name, and the number of requests in each time period is recorded. When the request volume reaches a specific threshold (such as more than 50 requests per minute), it is considered that the amount of requests for the domain name is abnormal. On this basis, high request volume data of the domain name system is generated for subsequent analysis.

[0149] Step S323: Identify suspicious external IP addresses based on the high request volume data of the domain name system, thereby obtaining suspicious external IP address data;

[0150] In this embodiment, suspicious external IP addresses are identified based on high request volume data. For each DNS request, check whether its source IP address is an external IP address, and whether the request frequency of the IP address during the statistical period exceeds the preset threshold. For example, if an external IP address sends more than 100 requests within 10 minutes, the IP address is considered to be a suspicious IP. The threshold is dynamically adjusted based on historical data and the average level of network traffic patterns. If the request volume of the IP address is higher than the normal range, and the domain name it requests is suspicious (such as short links, malicious websites, etc.), the IP address is marked as a potential source of attack. All data of suspicious external IP addresses (such as IP addresses, number of requests, requested domain names, etc.) will be recorded as input data for the next step of analysis.

[0151] Step S324: Perform short link detection according to the switch traffic data to obtain short link data;

[0152] In this embodiment, short link detection relies on the analysis of the requested domain name. When the domain name contained in a DNS request is a link generated by a short link service (such as bit.ly, goo.gl, etc.), the system will mark the domain name as a short link domain name. By analyzing the domain name of the DNS request and comparing it with the short link service database, the request belonging to the short link service is detected as short link data. All detected short links are recorded, including the target URL of the short link, the source IP address of the request, and the timestamp. During the short link domain name detection process, the system needs to continuously update the domain name database of the short link service provider to maintain the accuracy of short link identification.

[0153] Step S325: performing abnormal uniform resource locator identification on the short link data, thereby obtaining abnormal uniform resource locator data;

[0154] In this embodiment, the target URL pointed to by the short link is obtained. The actual target URL corresponding to the short link is parsed through the short link parsing tool. These target URLs are feature analyzed to check whether there are typical malicious URL features, such as: malicious scripts contained in the URL, domain names with phishing website features (such as URLs that imitate regular websites), and unknown or uncommon top-level domain names (such as .xyz, etc.). If the URL matches a known malicious URL library, or contains suspicious features (such as an overly long path, random characters, etc.), the URL is considered to be an abnormal uniform resource locator. Record detailed information for each abnormal URL, such as the target URL, access source IP, access time, etc., to facilitate subsequent analysis and response.

[0155] Step S326: determine whether the suspicious external IP address data is a phishing attack based on the abnormal uniform resource locator data and the short link data, thereby obtaining phishing attack data.

[0156] In this embodiment, when a suspicious IP address is associated with an abnormal URL, and the resource pointed to by the URL has typical phishing attack characteristics, the system will associate the suspicious IP address with the phishing attack. For example, when an external IP address frequently requests short links, and these short links point to URLs containing phishing characteristics, it can be determined that the IP address has launched a phishing attack. All information such as IP addresses, requested domain names, and URLs involved in phishing attacks will be summarized and recorded as phishing attack data. The generated phishing attack data includes the attack source IP, attack target URL, the time when the attack occurred, and related network activity data. These data will be used as early warning information for the network security protection system for further protection response.

[0157] Optionally, step S33 is specifically:

[0158] Step S331: pre-processing the switch flow data to obtain pre-processed switch flow data;

[0159] In this embodiment, the main purpose of preprocessing is to remove redundant data, fill missing values ​​and standardize the data format to ensure the accuracy of subsequent analysis. By using the switch flow data collected in real time by network monitoring tools (such as SNMP, sFlow or NetFlow), the data is first time-series integrated to ensure the uniformity of flow data within each collection cycle. For each switch flow record, including fields such as source IP, target IP, number of bytes transmitted, and transmission duration, unit conversion is performed, such as converting the number of bytes to megabytes for subsequent analysis. Then, check the missing values ​​or abnormal values ​​in the data (such as negative values ​​or excessive values), and fill or remove them according to the set rules. Finally, all cleaned data is stored in chronological order, and ensure that the data is output in a unified format (such as CSV, JSON) for subsequent processing.

[0160] Step S332: Establishing a traffic pattern baseline for the pre-processed switch traffic data, thereby obtaining switch traffic pattern baseline data;

[0161] In this embodiment, by performing aggregation analysis on large-scale historical data, typical traffic characteristics of each switch port or device in different time periods (for example, hours, days, and weeks) are extracted. At this time, cluster analysis methods, such as the K-means algorithm, can be used to identify different types of traffic patterns, such as normal traffic, peak traffic, and low peak traffic. The average bandwidth utilization, traffic peak, and fluctuation range of each port are calculated. By classifying and summarizing these patterns, a baseline traffic model is established. In this process, some key thresholds are defined, such as: in normal traffic mode, the bandwidth utilization does not exceed 60%, and the traffic fluctuation range is within 20%. When some traffic exceeds this standard range, it needs to be marked as abnormal data.

[0162] Step S333: performing real-time abnormal monitoring on the switch traffic data according to the switch traffic pattern baseline data, thereby obtaining real-time abnormal switch traffic data;

[0163] In this embodiment, the system will determine whether the current traffic deviates from the established traffic pattern baseline by comparing the real-time traffic data. Whenever a new record of switch traffic data is collected, the system will compare the difference between the real-time data and the historical traffic pattern according to the set monitoring period (for example, once a minute). If it is found that the real-time traffic data (such as traffic rate, bandwidth utilization) exceeds the threshold set by the baseline (for example, the bandwidth utilization exceeds 80% of the baseline), the traffic is recorded as abnormal traffic. In this process, the sliding window technology is used to monitor each traffic sampling point in real time, and the traffic change data within each monitoring period is recorded. All real-time abnormal switch traffic data will be collected and stored for further analysis.

[0164] Step S334: classifying abnormal types according to the real-time abnormal switch traffic data, thereby obtaining bandwidth abuse abnormal data and network worm abnormal data;

[0165] In this embodiment, by analyzing the abnormal characteristics of real-time traffic, abnormal traffic can be divided into different types. For example, if the bandwidth utilization of abnormal traffic is much higher than the normal value (more than 90% of the baseline), it is marked as bandwidth abuse abnormal data. Bandwidth abuse data can be determined by identifying large traffic transmissions (such as large file downloads, video transmissions). Another common type of abnormality is network worm attack traffic. Network worm attacks are usually manifested as abnormal traffic that spreads rapidly, usually with a large number of connection requests initiated outward in a short period of time. Therefore, if there is a rapid increase in the number of connections and data packets in real-time traffic (for example, the number of connection requests increases by more than 200% in 5 minutes), it can be marked as network worm abnormal data. The detailed characteristics of each type of abnormal traffic (such as source IP, target port, traffic size, etc.) will be stored in categories to facilitate subsequent correlation analysis.

[0166] Step S335: performing correlation analysis based on the bandwidth abuse abnormal data and the network worm abnormal data, thereby obtaining potential security threat data;

[0167] In this embodiment, a multi-dimensional data association model is established to cross-analyze bandwidth abuse abnormal traffic and network worm abnormal traffic. By combining traffic data with network topology data, it is checked whether these abnormal traffics are concentrated in specific switches or subnets. In particular, it is checked whether there are source IP addresses that appear in both bandwidth abuse traffic and worm attack traffic, or whether the traffic patterns of certain IP addresses change dramatically in a short period of time. For abnormal traffic that meets the conditions, a correlation analysis is performed, for example, the Pearson correlation coefficient is used to determine the relationship between different abnormal traffics. If it is found that the abnormal traffic patterns are highly correlated, it indicates that these traffics are very likely to be components of the same attack behavior or security threat. On this basis, potential security threat data is generated, and relevant source IP, target port, attack pattern and other information are recorded.

[0168] Step S336: Perform flow anomaly detection on the switch flow data according to the potential security threat data, thereby obtaining flow anomaly data.

[0169] In this embodiment, the system will monitor the real-time traffic again based on the features in the potential security threat data and in combination with the historical traffic baseline. Traffic anomaly detection is performed by comparing the potential threat features (such as abnormal traffic on a specific port, abnormal behavior of the source IP, etc.). If the real-time traffic shows features related to known potential security threats during the monitoring period (such as a sudden increase in traffic from a suspicious IP or a large number of connection requests), the traffic will be marked as abnormal traffic. This detection process is achieved by using real-time traffic analysis and historical traffic data comparison technology to ensure timely detection and response to potential security threats. All detected traffic anomaly data will be recorded in detail and further analyzed or alarmed.

[0170] Optionally, step S332 is specifically:

[0171] Performing flow rate feature extraction and data packet feature extraction on the pre-processed switch flow data, thereby obtaining flow rate data and data packet data;

[0172] In this embodiment, the flow data of the switch port is collected by a network monitoring tool (such as SNMP, NetFlow or sFlow), including the flow rate per second and the size of each data packet. The flow rate can be calculated by the number of bytes transmitted in each time interval (such as the number of bytes transmitted per second), and the data packet features can be extracted by capturing the size and type of each data packet (such as TCP, UDP, ICMP). For the flow rate data, the flow rate of each port is recorded according to the set time window (such as every second, every minute), and its average value, maximum value, minimum value and other statistical data are calculated according to the actual flow rate. When extracting the data packet features, the size and frequency of each data packet and whether there is an abnormal data packet size (such as an oversized or undersized packet) are analyzed. The extracted data will be saved in the database in the format of: {"port_id":"port1","flow_rate":500Mbps,"packet_size":1500,"protocol":"TCP","timestamp":"2025-01-07 12:00:00"}.

[0173] Get the historical traffic data of the switch;

[0174] In this embodiment, historical traffic data is obtained through a continuous network traffic collection tool, which is usually achieved by regularly capturing traffic statistics of switch devices. The collected historical data will involve information such as the real-time traffic of each port, the number of bytes and packets transmitted, etc. Historical traffic data is usually stored in a centralized data warehouse in the format of: {"port_id":"port1","timestamp":"2025-01-06 00:00:00","total_bytes":1000MB,"total_packets":5000}. In this process, by arranging the collected data by time period, a historical traffic data set is constructed to provide a data basis for subsequent traffic fluctuation range statistics.

[0175] The flow fluctuation range is counted based on the historical flow data of the switch, thereby obtaining the flow fluctuation range data;

[0176] In this embodiment, the flow fluctuation range of each port is calculated based on the historical flow data of each port. For example, by comparing the maximum flow and the minimum flow in each time period, the flow fluctuation range of the port can be obtained. If the flow value in a certain time period exceeds the normal fluctuation range of the port (for example, a fluctuation of more than 30%), it is regarded as a potential anomaly. In this process, the flow fluctuation amplitude of each port is statistically calculated, and the normal fluctuation range of the port is defined based on this. The data can be obtained by calculating the standard deviation, variance, etc. to quantify the degree of flow fluctuation. For example, the standard deviation can calculate the discreteness of the flow fluctuation, and the fluctuation range is represented by the ratio of the standard deviation to the flow average value.

[0177] A flow fluctuation baseline is constructed according to the flow fluctuation range data, thereby obtaining flow fluctuation baseline data;

[0178] In this embodiment, the traffic fluctuation baseline of the entire network or a specific switch device is calculated by summarizing the traffic fluctuation range data of each port. For example, assuming that the traffic fluctuation range of multiple ports of a switch is [10%-50%], the traffic fluctuation baseline of the switch is 50%. The construction of the traffic fluctuation baseline needs to be estimated by the fluctuation amplitude in historical data, and an acceptable standard is set as the normal fluctuation range of network traffic. This baseline will serve as a reference value for subsequent detection of abnormal traffic fluctuations. By analyzing the traffic fluctuations in different time periods, stable traffic patterns are identified and these patterns are classified. For example, the traffic fluctuation range on weekdays is larger, while the fluctuation on holidays is smaller.

[0179] constructing a flow rate baseline for the flow rate data, thereby obtaining flow rate baseline data;

[0180] In this embodiment, the normal rate range of each port or switch is calculated based on the traffic rate data. For example, for a port of a switch, the average traffic rate of the port in the past week is calculated, which is usually 30Mbps, and the maximum rate is 100Mbps. The traffic rate baseline determines the normal working range of each port by analyzing the highest and lowest traffic rates of historical data. According to the set threshold standard (such as abnormal when the traffic rate exceeds 80% of the baseline), the corresponding rate baseline is constructed. The construction of the traffic rate baseline not only depends on historical data, but also needs to consider the performance parameters of the network equipment, such as the maximum bandwidth of the switch port, to ensure that the traffic rate is within a reasonable range.

[0181] Optimizing the traffic source of the traffic fluctuation baseline data according to the data packet data, thereby obtaining the traffic fluctuation baseline optimization data;

[0182] In this embodiment, the traffic fluctuation baseline is optimized by analyzing the source address, destination address, size of the data packet and other information in the historical traffic data. In this step, special attention is paid to the traffic with abnormal traffic fluctuation and unknown source. If the traffic fluctuation from certain source IPs is found to be far beyond the predetermined baseline, it is marked as a potential source of traffic anomaly. Traffic source optimization not only helps improve the accuracy of the traffic fluctuation baseline, but also identifies the source of security threats or abnormal traffic patterns. By analyzing the source and type of each data packet and combining the traffic fluctuation data, the traffic fluctuation baseline is updated to adapt to the ever-changing network environment.

[0183] The switch traffic pattern baseline data is integrated according to the traffic fluctuation baseline optimization data and the traffic rate baseline data, so as to obtain the switch traffic pattern baseline data.

[0184] In this embodiment, a comprehensive switch traffic pattern baseline is established by combining the optimized traffic fluctuation baseline and traffic rate baseline. The integration process requires analyzing the traffic data of all ports, identifying different traffic patterns, and comparing them with historical baselines to further optimize the overall network traffic pattern. After the traffic pattern baseline is integrated, it will help the real-time monitoring system to more accurately identify traffic fluctuations and rate anomalies, facilitating subsequent abnormal traffic detection and security threat identification. Finally, the integrated traffic pattern baseline is stored as a standard data set for subsequent real-time monitoring and security analysis.

[0185] Optionally, step S4 is specifically:

[0186] Step S41: extracting link disconnection features and malware detection features according to the network security threat data, thereby obtaining link disconnection data and malware detection data;

[0187] In this embodiment, link disconnection feature extraction and malware detection feature extraction are performed to obtain link disconnection data and malware detection data. Link disconnection feature extraction collects link status monitoring information of network devices in real time, such as judging whether the link is normal by the port status of the switch. Link disconnection events are usually reported by the switch through the SNMP protocol. When a link disconnection is detected, information such as the port where the disconnection occurred, the time of disconnection, and the duration of the disconnection are recorded. For malware detection feature extraction, malware activities are detected by analyzing abnormal behaviors in traffic, especially by identifying fingerprints of viruses, Trojans, or other malware through deep packet inspection (DPI) technology. Malware features include specific IP communication modes, large amounts of abnormal traffic, malicious requests, etc. Malware detection feature extraction needs to rely on a pre-defined malware fingerprint library and use traffic analysis tools, such as IDS / IPS systems, to perform real-time traffic detection and analysis. Through these means, link disconnection data and malware detection data are extracted in the following formats: {"timestamp":"2025-01-0713:00:00","event":"link_disruption","port_id":"port1","duration":"15seconds"} and {"timestamp":"2025-01-07 13:00:00","ip":"192.168.1.100","protocol":"TCP","malware_type":"worm"}.

[0188] Step S42: Identify the security threat risk of abnormal network transmission topology according to the link disconnection data, thereby obtaining link disconnection security threat risk data;

[0189] In this embodiment, based on the port information, disconnection time and duration in the link disconnection data, combined with the network topology information, it is determined which links in the network have abnormalities. By analyzing the frequency and location of link disconnection events, the impact of the event on network transmission is evaluated. Further, by calculating the impact of link disconnection on the network topology, potential security threats in the network are identified. For example, if certain links are frequently disconnected and occur near important network nodes, it indicates that there is an external attack or internal equipment failure. At this point, the security risk of the link disconnection event can be quantified by threshold setting, such as setting a threshold (for example, ports with more than 3 disconnection events / hour are high-risk ports) to perform risk assessment, and mark these ports and links as potential sources of security threats.

[0190] Step S43: identifying security threat risks of abnormal network transmission topology according to the malware detection data, thereby obtaining malware detection security threat risk data;

[0191] In this embodiment, potential security threats are identified by analyzing abnormal behaviors in network traffic, especially traffic characteristics related to malware. If the traffic pattern of a certain IP or port is detected to be similar to known malware behavior (such as a large number of TCP connection requests, malicious DNS queries, etc.), it can be considered that the node is at risk. Further, by combining the network topology, the path of malware propagation and the key nodes in the network are identified. For example, if an external IP is identified as a malicious source and communicates with multiple internal nodes in the network, the external IP can be considered a potential source of security threats. According to the threshold, if the traffic exceeds the normal range by 3 times or the request frequency exceeds 10 times, it can be marked as malicious traffic and a security risk assessment can be performed.

[0192] Step S44: performing security threat risk integration according to the link disconnection security threat risk data and the malware detection security threat risk data, thereby obtaining security threat risk data;

[0193] In this embodiment, the link disconnection data and malware detection data are processed separately, and the risk information in each data is extracted. Then, the two are integrated to compare the location, time and type of risk. If the link disconnection event and the malware detection data occur in the same time period and are located at the same node or a close node in the network, the two risk events can be regarded as related security threats and comprehensively evaluated. When integrating, the duration, frequency and abnormality of the link disconnection event and the malware traffic can be comprehensively considered through weighting or priority sorting methods, and finally a comprehensive security threat risk score is formed. For example, the risk score of the link disconnection event is 3 (risk scoring system from 1 to 5), and the risk score of the malware detection is 4, and the final integrated risk score is 7.

[0194] Step S45: network security threats are divided according to the abnormal topological structure of network transmission of security threat risk data, thereby obtaining network security threat data topological structure data.

[0195] In this embodiment, each node in the network is divided using network topology information, and the security threat risk scores based on link disconnection and malware detection are used to determine which nodes have higher security risks. In this process, focus is placed on nodes that have both link disconnection events and malware activities, and these nodes are classified as high-risk areas. During the division process, multiple levels of risk levels can be set, such as low risk, medium risk, and high risk, so as to implement different levels of security protection measures. The output of security threat data topology structure data is usually presented in the form of a network diagram, in which high-risk nodes and links are marked in red or other warning colors. Ultimately, these data will be used for subsequent security response and monitoring decisions. For example, the formatted output is: {"node_id":"node1","risk_level":"high","associated_threats":["link_disruption","malware"]}, indicating that node 1 has security threats related to link disconnection and malware.

[0196] Optionally, this specification also provides an information big data intelligent collection and management system, which is used to execute the information big data intelligent collection and management method as described above, and the game information big data intelligent collection and management system includes:

[0197] The data dictionary maintenance module is used to obtain the information big data warehouse, and perform data dictionary feature extraction based on the information big data warehouse to obtain data dictionary data; perform maintenance based on the data dictionary data to obtain data dictionary maintenance data;

[0198] The transmission anomaly analysis module is used to obtain switch device data, and perform network topology analysis based on the switch device data, thereby obtaining switch device network topology data; perform transmission anomaly analysis based on the switch device network topology data, thereby obtaining network transmission anomaly topology data;

[0199] The network security threat analysis module is used to perform traffic statistics based on the switch device data, thereby obtaining the switch traffic data; perform network security threat analysis based on the switch traffic data, thereby obtaining the network security threat data;

[0200] A network security threat classification module is used to classify network security threats based on network security threat data on abnormal network transmission topology structures, thereby obtaining network security threat data topology structure data;

[0201] The network security threat intelligent collection model construction module is used to construct the network security threat intelligent collection model based on the data dictionary maintenance data and the network security threat data topology structure data, so as to obtain the network security threat intelligent collection model; according to the network security threat intelligent collection model, the network security threat management of the information big data warehouse is performed, so as to obtain the network security threat management data.

[0202] The present invention provides an information big data intelligent collection and management system, which can implement any information big data intelligent collection and management method of the present invention, and is used to combine the operation and signal transmission medium between various modules to complete the information big data intelligent collection and management method. The internal modules of the system cooperate with each other, thereby improving the efficiency of identifying network security threats.

[0203] Therefore, the embodiments should be regarded as illustrative and non-restrictive from all points, and the scope of the present invention is limited by the appended claims rather than the above description, and it is therefore intended that all changes falling within the meaning and range of equivalent elements of the application documents are included in the present invention.

[0204] The above description is only a specific embodiment of the present invention, so that those skilled in the art can understand or implement the present invention. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but should conform to the widest scope consistent with the principles and novel features invented herein.

Claims

1. A method for intelligent collection and management of information big data, characterized in that: The following steps are involved: Step S1: obtaining an information big data warehouse, and performing data dictionary feature extraction based on the information big data warehouse, thereby obtaining data dictionary data; performing maintenance based on the data dictionary data, thereby obtaining data dictionary maintenance data; Step S2: acquiring switch device data, and performing network topology analysis based on the switch device data, thereby obtaining switch device network topology data; performing transmission anomaly analysis based on the switch device network topology data, thereby obtaining network transmission anomaly topology data; Step S3: Perform traffic statistics according to the switch device data to obtain switch traffic data; Perform network security threat analysis based on switch traffic data to obtain network security threat data; Step S4: performing network security threat classification on the abnormal network transmission topology structure according to the network security threat data, thereby obtaining network security threat data topology structure data; Step S5: constructing a network security threat intelligent collection model according to the data dictionary maintenance data and the network security threat data topology structure data, thereby obtaining a network security threat intelligent collection model; Network security threat management is performed on the information big data warehouse based on the network security threat intelligent collection model to obtain network security threat management data.

2. The method for intelligent collection and management of information big data according to claim 1 is characterized in that: Step S1 is specifically as follows: Step S11: obtaining an information big data warehouse, and performing data dictionary feature extraction based on the information big data warehouse, thereby obtaining data dictionary data; Step S12: standardizing the fields of the data dictionary data to obtain field standardized data; Step S13: Perform code mapping according to the data dictionary data, thereby mapping the code value to the digital warehouse standard; Step S14: Perform dictionary maintenance based on the field standardization data and the digital warehouse standard mapping code value to obtain data dictionary maintenance data.

3. The method for intelligent collection and management of information big data according to claim 1 is characterized in that: Step S2 is specifically as follows: Step S21: acquiring switch device data, and performing node identification according to the switch device data, thereby obtaining switch node data; Step S22: extracting port connection features according to the switch device data, thereby obtaining port connection data; Step S23: Establishing a primary topology structure according to the switch node data and the port connection data, thereby obtaining primary topology structure data; Step S24: performing redundant path identification on the primary topology structure data, thereby obtaining redundant path data; Step S25: performing loop path identification on the primary topology structure data, thereby obtaining loop path data; Step S26: constructing a switch device network topology structure according to the loop path data and the redundant path data, thereby obtaining switch device network topology structure data; Step S27: Perform transmission anomaly analysis based on the network topology data of the switch device, thereby obtaining network transmission anomaly topology data.

4. The method for intelligent collection and management of information big data according to claim 3 is characterized in that: Step S27 is specifically as follows: Step S271: performing switch port fault identification according to the switch device network topology data, thereby obtaining switch port fault data; Step S272: Calculate the bandwidth utilization of the switch port fault data to obtain the bandwidth utilization; Step S273: Perform numerical statistics according to the bandwidth utilization, so as to obtain high bandwidth utilization; Step S274: performing network performance anomaly analysis on the network topology data of the switch device according to the high bandwidth utilization, thereby obtaining network performance anomaly data; Step S275: performing link health anomaly detection on the network topology data of the switch device, thereby obtaining link health anomaly data; Step S276: Performing a network transmission abnormal topology structure analysis on the switch device network topology structure according to the link health abnormality data and the network performance abnormality data, thereby obtaining network transmission abnormal topology structure data.

5. The method for intelligent collection and management of information big data according to claim 1 is characterized in that: Step S3 is specifically as follows: Step S31: Perform traffic statistics according to the switch device data to obtain switch traffic data; Step S32: Perform phishing attack detection according to the switch traffic data, thereby obtaining phishing attack data; Step S33: monitoring traffic anomalies according to the switch traffic data, thereby obtaining traffic anomaly data; Step S34: Integrate network security threats based on phishing attack data and traffic anomaly data to obtain network security threat data.

6. The method for intelligent collection and management of information big data according to claim 5 is characterized in that: Step S32 is specifically as follows: Step S321: performing domain name system request analysis according to the switch traffic data, thereby obtaining domain name system request data; Step S322: performing request volume statistics on the domain name system request data, thereby obtaining high request volume data of the domain name system; Step S323: Identify suspicious external IP addresses based on the high request volume data of the domain name system, thereby obtaining suspicious external IP address data; Step S324: Perform short link detection according to the switch traffic data to obtain short link data; Step S325: performing abnormal uniform resource locator identification on the short link data, thereby obtaining abnormal uniform resource locator data; Step S326: determine whether the suspicious external IP address data is a phishing attack based on the abnormal uniform resource locator data and the short link data, thereby obtaining phishing attack data.

7. The method for intelligent collection and management of information big data according to claim 5 is characterized in that: Step S33 is specifically as follows: Step S331: pre-processing the switch flow data to obtain pre-processed switch flow data; Step S332: Establishing a traffic pattern baseline for the pre-processed switch traffic data, thereby obtaining switch traffic pattern baseline data; Step S333: performing real-time abnormal monitoring on the switch traffic data according to the switch traffic pattern baseline data, thereby obtaining real-time abnormal switch traffic data; Step S334: classifying abnormal types according to the real-time abnormal switch traffic data, thereby obtaining bandwidth abuse abnormal data and network worm abnormal data; Step S335: performing correlation analysis based on the bandwidth abuse abnormal data and the network worm abnormal data, thereby obtaining potential security threat data; Step S336: Perform flow anomaly detection on the switch flow data according to the potential security threat data, thereby obtaining flow anomaly data.

8. The method for intelligent collection and management of information big data according to claim 7 is characterized in that: Step S332 is specifically as follows: Performing flow rate feature extraction and data packet feature extraction on the pre-processed switch flow data, thereby obtaining flow rate data and data packet data; Get the historical traffic data of the switch; The flow fluctuation range is counted based on the historical flow data of the switch, thereby obtaining the flow fluctuation range data; A flow fluctuation baseline is constructed according to the flow fluctuation range data, thereby obtaining flow fluctuation baseline data; constructing a flow rate baseline for the flow rate data, thereby obtaining flow rate baseline data; Optimizing the traffic source of the traffic fluctuation baseline data according to the data packet data, thereby obtaining the traffic fluctuation baseline optimization data; The switch traffic pattern baseline data is integrated according to the traffic fluctuation baseline optimization data and the traffic rate baseline data, so as to obtain the switch traffic pattern baseline data.

9. The method for intelligent collection and management of information big data according to claim 1 is characterized in that: Step S4 is specifically as follows: Step S41: extracting link disconnection features and malware detection features according to the network security threat data, thereby obtaining link disconnection data and malware detection data; Step S42: Identify the security threat risk of abnormal network transmission topology according to the link disconnection data, thereby obtaining link disconnection security threat risk data; Step S43: identifying security threat risks of abnormal network transmission topology according to the malware detection data, thereby obtaining malware detection security threat risk data; Step S44: performing security threat risk integration according to the link disconnection security threat risk data and the malware detection security threat risk data, thereby obtaining security threat risk data; Step S45: network security threats are divided according to the abnormal topological structure of network transmission of security threat risk data, thereby obtaining network security threat data topological structure data.

10. An intelligent information big data collection and management system, characterized in that: Used to execute the information big data intelligent collection and management method as claimed in claim 1, the information big data intelligent collection and management system comprises: The data dictionary maintenance module is used to obtain the information big data warehouse, and perform data dictionary feature extraction based on the information big data warehouse to obtain data dictionary data; perform maintenance based on the data dictionary data to obtain data dictionary maintenance data; The transmission anomaly analysis module is used to obtain switch device data, and perform network topology analysis based on the switch device data, thereby obtaining switch device network topology data; perform transmission anomaly analysis based on the switch device network topology data, thereby obtaining network transmission anomaly topology data; The network security threat analysis module is used to perform traffic statistics based on the switch device data, thereby obtaining the switch traffic data; perform network security threat analysis based on the switch traffic data, thereby obtaining the network security threat data; A network security threat classification module is used to classify network security threats based on network security threat data on abnormal network transmission topology structures, thereby obtaining network security threat data topology structure data; The network security threat intelligent collection model construction module is used to construct the network security threat intelligent collection model based on the data dictionary maintenance data and the network security threat data topology structure data, so as to obtain the network security threat intelligent collection model; according to the network security threat intelligent collection model, the network security threat management of the information big data warehouse is performed, so as to obtain the network security threat management data.

Citation Information

Patent Citations

  • Software defined networking (SDN) network topology flow visual monitoring method and control terminal

    CN106130796A

  • Fault diagnosis method and system based on network traffic data

    CN109150619A

  • Network security protection method and system

    CN117879970A

  • Computer network security intelligent analysis system and method based on big data

    CN117896137A

  • Computer network anomaly detection method

    CN118784364A

Cited By

  • Intelligent detection method and system for security vulnerabilities of terminal layer of power internet of things

    CN120455112A