Network data security positioning method and system based on deep learning algorithm
Through the network data security positioning method based on deep learning algorithms, the traditional method's shortcomings in identifying security threats when processing massive network traffic data is solved, and efficient and accurate network data security positioning is achieved, which improves the recognition rate and reduces the false alarm rate.
Patent Information
- Application Number
- CN202510243880.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-03
- Publication Date
- 2025-06-10
AI Technical Summary
Traditional network data positioning methods cannot identify potential security threats in a timely and accurate manner when processing massive network traffic data, and lack effective feature extraction and classification capabilities in the face of unknown attacks and complex network traffic data behaviors, resulting in low recognition rate and high false alarm rate, affecting the overall security positioning effect of network data.
The network data security positioning method based on deep learning algorithm is adopted. By obtaining network traffic data, extracting multiple network features, building multiple subsets of network features, calculating the correlation index and security positioning index of each subset, selecting the largest correlation index threshold and security positioning index as the target feature subset, and using deep learning algorithms to safely locate network data.
When processing massive network traffic data, it can timely and accurately identify potential security threats, improve the network's defense capabilities when encountering attacks, provide effective feature extraction and classification capabilities, improve the overall security positioning effect of network data, reduce false alarm rates and improve identification rates.
Smart Images

Figure CN120128367A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of network security technology, and particularly to a network data security positioning method and system based on deep learning algorithms. Background Art
[0002] With the rapid development of network technology and the increasingly complex environment, network security faces unprecedented challenges. The means of network attacks are constantly evolving, and traditional protection measures are increasingly difficult to effectively cope with new threats. In the context of big data and cloud computing, network traffic is huge and complex, making security positioning face a more severe test. Therefore, achieving efficient and accurate network data security positioning has become a key issue that urgently needs to be solved in the current network security field.
[0003] In existing technical solutions, traditional network data positioning methods are unable to timely and accurately identify potential security threats when dealing with massive network traffic data, resulting in insufficient defense capabilities of the network when under attack.
[0004] In addition, in the face of unknown attacks and complex network traffic data behaviors, existing technologies often lack effective feature extraction and classification capabilities, leading to low recognition rates and high false alarm rates, which in turn affect the overall network data security positioning effect. Summary of the Invention
[0005] In order to solve the technical problems that traditional network data positioning methods are unable to timely and accurately identify potential security threats when dealing with massive network traffic data, resulting in insufficient defense capabilities of the network when under attack, and lack effective feature extraction and classification capabilities in the face of unknown attacks and complex network traffic data behaviors, leading to low recognition rates and high false alarm rates, which in turn affect the overall network data security positioning effect, the present invention provides a network data security positioning method and system based on deep learning algorithms.
[0006] The technical solutions provided by the embodiments of the present invention are as follows:
[0007] First aspect:
[0008] A network data security positioning method based on deep learning algorithms provided by an embodiment of the present invention includes:
[0009] S1: Obtain network traffic data;
[0010] S2: Extract multiple network features of the network traffic data;
[0011] S3: Randomly combine the multiple extracted network features to construct multiple network feature subsets;
[0012] S4: Calculate the correlation index of each network feature subset;
[0013] S5: Select a network feature subset whose correlation index is greater than or equal to the correlation index threshold;
[0014] S6: Calculate the security positioning index of each selected network feature subset;
[0015] S7: Select the network feature subset with the largest security positioning index as the target network feature subset;
[0016] S8: Based on the target network feature subset, perform network data security positioning through a deep learning algorithm.
[0017] Second aspect:
[0018] A network data security positioning system based on a deep learning algorithm provided by an embodiment of the present invention includes: a memory and one or more processors;
[0019] One or more application programs are stored in the memory, and the one or more application programs are adapted to be executed by the one or more processors to implement the above-mentioned network data security positioning method based on a deep learning algorithm.
[0020] The beneficial effects brought by the technical solution provided by the embodiment of the present invention at least include:
[0021] In the present invention, by calculating the correlation index of each network feature subset, selecting a network feature subset whose correlation index is greater than or equal to the correlation index threshold, calculating the security positioning index of each selected network feature subset, and selecting the network feature subset with the largest security positioning index as the target network feature subset, when processing a large amount of network traffic data, potential security threats can be identified in a timely and accurate manner, enabling the network to have sufficient defense capabilities when encountering attacks. Based on the target network feature subset, network data security positioning is performed through a deep learning algorithm, providing effective feature extraction and classification capabilities in the face of unknown attacks and complex network traffic data behaviors, with a high recognition rate and a low false alarm rate, thereby improving the overall network data security positioning effect. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings without creative efforts based on these drawings.
[0023] Figure 1 It is a schematic flowchart of a network data security positioning method based on a deep learning algorithm provided by an embodiment of the present invention;
[0024] Figure 2 This is a schematic structural diagram of a network data security positioning system based on a deep learning algorithm provided by an embodiment of the present invention. Detailed implementation manners
[0025] The following describes the technical solutions in the present invention with reference to the accompanying drawings.
[0026] In the embodiments of the present invention, words such as "exemplarily" and "for example" are used to represent examples, illustrations or explanations. Any embodiment or design solution described as an "example" in the present invention should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of the word "example" is intended to present concepts in a specific manner. In addition, in the embodiments of the present invention, the meaning expressed by "and / or" can be both, or either one of the two.
[0027] To make the technical problems, technical solutions and advantages to be solved by the present invention clearer, the following will be described in detail with reference to the accompanying drawings and specific embodiments.
[0028] Refer to the accompanying drawings of the specification Figure 1 which shows a schematic flowchart of a network data security positioning method based on a deep learning algorithm provided by an embodiment of the present invention.
[0029] The embodiments of the present invention provide a network data security positioning method based on a deep learning algorithm. This method can be implemented by a network data security positioning device based on a deep learning algorithm, and the network data security positioning device based on a deep learning algorithm can be a terminal or a server. The processing flow of the network data security positioning method based on a deep learning algorithm may include the following steps:
[0030] S1: Obtain network traffic data.
[0031] Specifically, by using a network packet capture tool (such as Wireshark or tcpdump), network traffic data transmitted in the network is collected in real time.
[0032] Among them, Wireshark is an open-source network protocol analysis tool widely used for network troubleshooting, analysis and development. It can capture and display detailed information of network traffic in real time, including the source and destination addresses, protocol types, data contents, etc. of each data packet. Users can conveniently browse, filter and analyze the captured data through its graphical user interface to help identify problems in the network, monitor network security and conduct protocol development. Wireshark supports the parsing of multiple protocols, making it an important tool for network engineers and security experts.
[0033] Among them, tcpdump is a command-line network packet capture tool used to capture and analyze the data packets transmitted through the computer network interface. It can run on multiple operating systems and is mainly used for network troubleshooting, security monitoring, and traffic analysis. Users can use tcpdump to specify capture filtering conditions to capture data packets of specific protocols, ports, or addresses. The captured data can be displayed in real-time or saved to a file for subsequent analysis. Due to its lightweight and powerful functions, tcpdump is widely used in the fields of network management and security.
[0034] In the present invention, by capturing network traffic data in real-time, it is possible to instantaneously monitor and analyze the network status, quickly identify and respond to potential network problems and security threats, thereby reducing the impact of network failures and security incidents.
[0035] S2: Extract multiple network features of the network traffic data.
[0036] Specifically, by parsing the obtained network traffic data, multiple network features are extracted.
[0037] It should be noted that network features refer to various attributes or metrics used to describe and analyze network traffic data, and these features can help identify the behavior patterns of traffic and their potential security threats.
[0038] Optionally, the network features specifically include: source IP address, destination IP address, source port, destination port, transport protocol (such as TCP, UDP), packet size, traffic duration, traffic direction (inbound or outbound), number of sessions, flag bits (such as TCP flag bits), and byte and packet counts in the traffic.
[0039] In the present invention, by carefully analyzing key network features such as source IP address, destination IP address, port, and transport protocol, network attacks such as botnet activities, hacker intrusions, and denial-of-service attacks (DDoS) can be effectively identified and responded to. These features help the security team quickly locate the source and nature of abnormal behaviors, thereby taking measures to defend against or mitigate these attacks.
[0040] S3: Randomly combine the multiple extracted network features to construct multiple network feature subsets.
[0041] Specifically, based on the multiple extracted network features, multiple network feature subsets are constructed in a random combination manner, and each network feature subset contains different numbers of network features.
[0042] In the present invention, by generating multiple feature subsets containing different numbers and types of network features, diversity can be introduced during the data analysis and model training processes. This helps to explore the utility of different feature combinations, find the most effective feature combination, and improve the robustness and generalization ability of the model.
[0043] S4: Calculate the correlation index of each network feature subset.
[0044] It should be noted that the correlation index is an indicator used to measure the relationship between a specific network feature subset and the target category, and it reflects the effectiveness of this feature subset in distinguishing different categories. Specifically, it considers the correlation between each feature in the feature subset and the target category, as well as the mutual correlation between features, so as to evaluate the contribution of features to the classification task.
[0045] In the present invention, by evaluating the correlation between the feature subset and the target category, the correlation index helps to identify which features are most effective in distinguishing different categories. Doing so can improve the model's depth of understanding of the data, thereby achieving higher accuracy in the classification task. The correlation index provides a quantitative basis for feature selection, making the feature selection process more scientific and systematic. Selecting a feature subset highly correlated with the target category can reduce the model complexity and improve the training efficiency.
[0046] Optionally, the correlation index is specifically:
[0047]
[0048] where r cj represents the correlation index between the c-th network feature subset and the j-th category, k represents the number of network features in the network feature subset, represents the average correlation index between all network features in the network feature subset and the j-th category, represents the average mutual correlation index between the individual network features in the network feature subset.
[0049] In the present invention, by quantifying the degree of association between the feature subset and the classification target, this formula helps to determine which feature subsets are most critical for distinguishing specific categories. This precise feature selection can directly affect the performance of the classification model and improve the classification accuracy. By effectively identifying and utilizing features highly correlated with the target category, the use of computational resources of the model can be optimized, because the model can rely only on the most influential features for training and prediction, thereby reducing unnecessary computational overhead.
[0050] S5: Select the network feature subsets whose correlation index is greater than or equal to the correlation index threshold.
[0051] It should be noted that those skilled in the art can set the size of the correlation index threshold according to actual needs, and the present invention does not make any limitations in this regard.
[0052] In the present invention, by setting a threshold to screen out highly correlated feature subsets, it can be ensured that the selected features are more effective in distinguishing different categories, thereby improving the accuracy and efficiency of the model in classifying data.
[0053] S6: Calculate the security positioning index of each selected network feature subset.
[0054] It should be noted that the security positioning index is an effectiveness indicator used to evaluate the network feature subset in the security data positioning task. It comprehensively considers the positioning accuracy rate and information gain of the feature subset to measure the contribution of this feature combination to improving the security positioning performance.
[0055] Among them, the positioning accuracy rate refers to the effectiveness of the network feature subset in correctly classifying network traffic data, which is specifically calculated through the number of true positives and true negatives, and is used to evaluate the accuracy of this feature subset in identifying potential security threats.
[0056] Among them, the information gain is an indicator used to measure the importance of the feature subset for the classification decision. It reflects the contribution of this feature subset to improving the classification performance by calculating the degree of reduction in the entropy value of the target variable (category) when using this feature subset.
[0057] In the present invention, through the security positioning index, it is possible to accurately identify which feature subsets are most critical for detecting potential threats in the network. This helps to specifically strengthen the monitoring of those features that have a significant impact on the security situation, thereby enhancing the sensitivity and response speed of the security monitoring system. Using the security positioning index allows model developers to adjust the feature subset according to the actual network environment and threat landscape, enabling the security model to flexibly adapt to different security requirements and environmental changes, and enhancing the effectiveness of the model in the face of new or changing threats.
[0058] Optionally, the security positioning index is specifically:
[0059] S i = ω 1 A i + ω 2 IG i
[0060] Among them, S i represents the security positioning index of the i-th network feature subset, A i represents the positioning accuracy rate of the i-th network feature subset, IG i represents the information gain of the i-th network feature subset, ω 1The weight coefficient representing the positioning accuracy rate, ω 2 The weight coefficient representing the information gain.
[0061] In the present invention, by accurately calculating and utilizing the security positioning index, it can be ensured that the selected feature subset can effectively improve the system's ability to identify and respond to network threats. This method helps to construct a more accurate security model, thereby reducing security vulnerabilities and risks. In the case of resource constraints, optimizing the selection of the feature subset can ensure that computing resources are concentrated on the processing of the most critical features, thus improving the overall efficiency and performance of the system.
[0062] Among them, the positioning accuracy rate is specifically:
[0063]
[0064] Among them, A i represents the positioning accuracy rate of the i-th network feature subset, TP i represents the number of true positive examples of the i-th network feature subset, TN i represents the number of true negative examples of the i-th network feature subset, FP i represents the number of false positive examples of the i-th network feature subset, FN i represents the number of false negative examples of the i-th network feature subset.
[0065] In the present invention, by calculating the positioning accuracy rate, the effectiveness of each feature subset in distinguishing positive and negative categories can be intuitively evaluated. This helps to select the feature combination that can most correctly identify positive and negative samples, thereby improving the performance of the model in practical applications. The positioning accuracy rate provides a direct evaluation criterion for feature selection, helping developers to select the best-performing combination from numerous feature subsets. This selection is data-driven and can significantly improve the scientific nature and effectiveness of decision-making.
[0066] Among them, the information gain is specifically:
[0067] IG i =H(Y)-H(Y|M i )
[0068] Among them, IG i represents the information gain of the i-th network feature subset, H(Y) represents the entropy value of the category, Y represents the category, H(Y|M i ) represents the conditional entropy of the category under the network feature subset, M i represents the selected i-th network feature subset.
[0069] In the present invention, enhancing the effectiveness of feature selection: Information gain is a quantitative metric that measures the increase in the amount of information in the classification result by a feature subset. Through this metric, it is possible to identify which feature subsets are most effective in reducing the uncertainty in the classification process, thus helping to select the features that can most improve the classification accuracy. Information gain directly reflects the contribution of a feature subset to the final classification decision. Feature subsets with high information gain are of great value for understanding the classification mechanism in the data and enhancing the transparency of the model.
[0070] S7: Select the network feature subset with the largest security positioning index as the target network feature subset.
[0071] In the present invention, maximizing performance optimization: By selecting the feature subset with the highest security positioning index, it is ensured that the selected feature subset can provide optimal performance in the security positioning task. This is because the security positioning index comprehensively considers the positioning accuracy rate and information gain, and both jointly reflect the ability of the feature subset in terms of correct classification and increasing decision-making information. Feature subsets with high security positioning index can more effectively help the model distinguish different network traffic categories, especially when distinguishing normal traffic from potential threat traffic. This improvement in efficiency and precision is crucial for quickly responding to network security threats. By focusing on the feature subsets that can most enhance the model's performance, unnecessary consumption of computing resources and processing time can be reduced, thus making the entire security system more economical and efficient.
[0072] In a possible implementation manner, S7 is specifically:
[0073] Select the network feature subset with the largest security positioning index as the target network feature subset. When there are multiple network feature subsets with the same maximum security positioning index, select the network feature subset with the smallest number of network features as the target network feature subset.
[0074] In the present invention, selecting the feature subset with the highest security positioning index ensures the effectiveness of the features, and further selecting the subset with the smallest number of features among these effective feature subsets helps to simplify the model. This strategy not only maintains high performance but also reduces the complexity of the model and improves the computing efficiency. A smaller number of features means less computing resources are required in the data processing and model training processes. This is especially important for resource-constrained environments, as it can reduce hardware resource consumption and save energy. On the premise of maintaining a high security positioning index, reducing the number of features helps to reduce the risk of overfitting. This is because the model no longer relies on too many potentially irrelevant features to make decisions, but only relies on the most critical information, thereby enhancing the generalization ability of the model on unseen data.
[0075] S8: Based on the target network feature subset, network data security positioning is performed through deep learning algorithms.
[0076] It should be noted that deep learning algorithm is a subfield of machine learning. It builds multiple layers of neural network models to automatically learn and extract complex features from data. It can learn end-to-end through a large amount of training data and perform feature abstraction layer by layer, thereby achieving efficient classification, regression and generation tasks. Deep learning has achieved remarkable success in fields such as computer vision, natural language processing and speech recognition because it can handle nonlinear relationships and high-dimensional data, giving the model powerful expressiveness and generalization capabilities.
[0077] In the present invention, the deep learning algorithm can effectively learn and extract complex data features through its multi-layer structure. This is particularly important in the field of network security, because network traffic data often contains complex and implicit behavior patterns, and deep learning can automatically identify these patterns without manually setting specific rules. Network traffic data is usually high-dimensional and contains many types of features. Deep learning algorithms are naturally suitable for processing such data, and can mine and utilize the complex relationships in these high-dimensional features to enhance the accuracy of detection algorithms.
[0078] In a possible implementation manner, S8 specifically includes:
[0079] Based on the target network feature subset, network data security positioning is performed through convolutional neural network.
[0080] It should be noted that convolutional neural network (CNN) is a deep learning algorithm that is particularly suitable for processing data with a grid structure. It effectively extracts spatial features from input data by using multiple hierarchical structures such as convolutional layers, pooling layers, and fully connected layers. The convolutional layer extracts features from the input through convolution operations, while the pooling layer reduces the computational complexity and improves the robustness of features through dimensionality reduction. This hierarchical feature learning process enables CNN to perform well in tasks such as image classification, object detection, and image segmentation. Due to its powerful ability to automatically learn features, convolutional neural networks have been widely used in computer vision and related fields.
[0081] Furthermore, the convolutional neural network (CNN) has been improved by introducing feedback layers and emphasis layers to enhance its feature extraction and classification performance. The feedback layer enables high-level prediction information to be passed back to the lower layers, thereby guiding the learning of underlying features and helping the model to better identify difficult-to-distinguish objects. At the same time, the emphasis layer dynamically weights the feature map based on the feedback information, enhancing the expression of important features and suppressing interfering features.
[0082] In the present invention, through the feedback layer, the model not only makes responses based on the results of feedforward feature extraction, but also can reversely adjust the feature extraction process according to the prediction results. Such dynamic adjustment makes the CNN more accurate in network data security positioning, especially when differentiating similar or subtle network behaviors. The introduction of the emphasis layer enables the network to dynamically adjust the weights of each feature according to the current task requirements and feedback information. This not only enhances the sensitivity of the model to key features, but also suppresses those irrelevant features that may cause misjudgment, thereby improving the overall decision-making quality of the model. The combination of the feedback layer and the emphasis layer not only improves the adaptability of the model to the current data, but also enhances the learning and adaptation ability of the model in the face of new threats. This enables the security system to continuously evolve to cope with the changing network threat environment.
[0083] In a possible implementation manner, the convolutional neural network includes an input layer, a plurality of convolutional layers, a plurality of pooling layers, a fully connected layer, and an output layer. The convolutional layers include a first convolutional layer, a second convolutional layer, and a third convolutional layer. The pooling layers include a first pooling layer, a second pooling layer, and a third pooling layer. S8 specifically includes sub-steps S801 to S810:
[0084] S801: In the input layer, input the target network feature subset;
[0085] S802: In the first convolutional layer, perform feature extraction on the target network feature subset to generate a first local feature map;
[0086] It should be noted that the first convolutional layer is a convolutional layer including 8 7×7 convolutional kernels.
[0087] S803: In the first pooling layer, perform a max-pooling operation on the first local feature map to generate a first pooled feature map;
[0088] It should be noted that the first pooling layer is a pooling layer including a 2×2 max-pooling window with a stride of 2.
[0089] S804: In the second convolutional layer, perform feature extraction on the first pooled feature map to generate a second local feature map;
[0090] It should be noted that the second convolutional layer is a convolutional layer including 12 5×5 convolutional kernels.
[0091] S805: In the second pooling layer, perform a max-pooling operation on the second local feature map to generate a second pooled feature map;
[0092] It should be noted that the second pooling layer is a pooling layer including a 2×2 max-pooling window with a stride of 2.
[0093] S806: In the third convolutional layer, perform feature extraction on the second pooled feature map to generate a third local feature map;
[0094] It should be noted that the third convolutional layer is a convolutional layer containing 16 3×3 convolutional kernels.
[0095] S807: In the third pooling layer, perform a max pooling operation on the third local feature map to generate a third pooled feature map;
[0096] It should be noted that the third pooling layer is a pooling layer containing a 2×2 max pooling window with a stride of 2.
[0097] S808: Perform weighted fusion on the first pooled feature map, the second pooled feature map, and the third pooled feature map to generate a target feature map;
[0098] In a possible implementation manner, the target feature map is specifically:
[0099]
[0100] where H represents the target feature map, W 1 represents the weight coefficient of the first pooled feature map, A represents the activation function, F 1 represents the first pooled feature map, represents the concatenation operation, W 2 represents the weight coefficient of the second pooled feature map, F 2 represents the second pooled feature map, W 3 represents the weight coefficient of the third pooled feature map, F 3 represents the third pooled feature map.
[0101] S809: In the fully connected layer, perform network data security positioning according to the target feature map;
[0102] S810: In the output layer, output the network data security positioning result.
[0103] In the present invention, through convolutional kernels of different sizes and multi-layer pooling operations, the network can extract multi-level features from different scales, ensuring that the network can capture various information of the input data. Weighted fusion of multi-layer feature maps can combine features of different scales, provide a more accurate and comprehensive feature representation, and enhance the classification ability of the network. The pooling operation reduces the size of the feature map, reduces the complexity of the model, and helps improve the generalization ability of the model. By comprehensively integrating multi-layer features, the network can better identify potential security threats in network data and complete efficient network data security positioning.
[0104] Optionally, the training method of the convolutional neural network is specifically:
[0105] With the goal of minimizing the mean squared error loss function, the convolutional neural network is trained through the cuckoo search algorithm to determine the optimal network parameters of the convolutional neural network.
[0106] It should be noted that the mean squared error (MSE) loss function is a commonly used loss function for measuring the difference between predicted values and actual values. It evaluates the prediction performance of the model by calculating the average of the squared differences between the predicted values and the actual values, and is usually used in regression tasks. The smaller the value of MSE, the closer the predicted result of the model is to the true result, that is, the better the performance of the model. In a neural network, using the mean squared error loss function can help optimize the model parameters and minimize the prediction error.
[0107] It should be noted that the cuckoo search algorithm is a population-based optimization technique inspired by the parasitic breeding behavior of cuckoos. The algorithm simulates the behavior of cuckoos laying eggs in the nests of other birds, using slightly mutated eggs to replace the eggs of other birds. In the algorithm, each solution is regarded as an egg in a bird's nest, and new solutions are generated through random walks and Lévy flights, and a probability parameter is used to decide whether to accept the new solution, so as to explore the solution space and find the global optimal solution. The cuckoo search algorithm is widely used to solve various optimization problems because of its simplicity, efficiency and ease of implementation.
[0108] Among them, the Lévy flight operation is a random walking strategy, the characteristic of which is that the step size follows the Lévy distribution, and this distribution allows the step size to occasionally make very large jumps. This enables Lévy flights to effectively avoid local optima when exploring complex search spaces and enhances the global search ability. In optimization algorithms, especially when using Lévy flights in the cuckoo search algorithm, it can help the algorithm jump from the current solution to a potentially better solution, thus increasing the probability and speed of finding the global optimal solution. This strategy is especially suitable for solving complex optimization problems with multiple peaks.
[0109] Furthermore, the cuckoo search algorithm is improved by introducing an adaptive adjustment strategy, so that the step size factor, discovery probability and scaling factor can be dynamically adjusted according to the number of iterations, improving the convergence speed and solution quality of the algorithm.
[0110] In the present invention, the mean squared error loss function is used as the optimization objective to help ensure that the error between the neural network output and the actual value is minimized, thereby improving the prediction accuracy of the model. This is particularly important for applications that require high accuracy such as network data security positioning. The cuckoo search algorithm combined with the mean squared error loss function can effectively handle high-dimensional and complex-structured data processed by convolutional neural networks. The flexibility of the optimization algorithm and its powerful global search ability make it particularly suitable for such problems.
[0111] Specifically, aiming to minimize the mean squared error loss function, an objective function is constructed:
[0112]
[0113] where min represents taking the minimum value, MSE represents the objective function, θ represents the set of network parameters of the convolutional neural network, R represents the total number of network traffic data, y r represents the true detection result of the r-th network traffic data, and
[0114] represents the actual detection result of the r-th network traffic data.
[0115] Furthermore, according to the objective function, the convolutional neural network is trained by the cuckoo search algorithm to determine the optimal network parameters of the convolutional neural network.
[0116] Optionally, the cuckoo search algorithm specifically includes:
[0117] Initializing the parameters and setting the maximum number of iterations of the cuckoo search algorithm.
[0118] Randomly generating an initial population, where the initial population includes multiple bird nests, and each bird nest represents a set of feasible network parameters of the convolutional neural network.
[0119] Using the objective function as the fitness function, calculating the fitness values of each bird nest, and taking the bird nest with the maximum fitness value as the current bird nest.
[0120]
[0121] where α t represents the step size factor at the t-th iteration, α max represents the maximum value of the step size factor, T represents the maximum number of iterations, α min represents the minimum value of the step size factor, represents the discovery probability of the i-th bird nest at the t-th iteration, P max represents the maximum value of the discovery probability, P min represents the minimum value of the discovery probability, represents the fitness value of the i-th bird nest at the t-th iteration, represents the maximum fitness value at the t-th iteration, the scaling factor of the i-th bird nest at the t-th iteration, γ max represents the maximum value of the scaling factor, γ min represents the minimum value of the scaling factor, represents the minimum fitness value at the t-th iteration.
[0122] Update the current bird nest through Levy flight operation according to the adaptively adjusted step size factor:
[0123]
[0124] where, represents the i-th bird nest at the (t + 1)-th iteration, represents the i-th bird nest at the t-th iteration, α represents the adaptively adjusted step size factor, represents the dot product operation, levy(β) represents the random step size generated by Levy flight operation.
[0125] Calculate the fitness value of the updated bird nest and compare the fitness value of the updated bird nest with that of the current bird nest. When the fitness value of the updated bird nest is greater than that of the current bird nest, replace the current bird nest with the updated bird nest. When the fitness value of the updated bird nest is less than or equal to that of the current bird nest, keep the current bird nest unchanged.
[0126] Generate a random number in the range of 0 to 1 and compare the random number with the adaptively adjusted discovery probability. When the random number is greater than the adaptively adjusted discovery probability, the current bird nest is discovered and the current bird nest is updated according to the adaptively adjusted scaling factor. When the random number is less than or equal to the adaptively adjusted discovery probability, the current bird nest is not discovered and the current bird nest remains unchanged.
[0127] where, when the random number is greater than the adaptively adjusted discovery probability, the current bird nest is discovered and the current bird nest is updated according to the adaptively adjusted scaling factor, specifically:
[0128]
[0129] where, represents the i-th bird nest at the (t + 1)-th iteration after the current bird nest is discovered, represents the i-th bird nest at the t-th iteration after the current bird nest is discovered, γ represents the adaptively adjusted scaling factor, and represent two individuals randomly selected from the population at the t-th iteration after the current bird nest is discovered, l 1 and l 2 .
[0130] Repeat the above steps until the maximum number of iterations is reached.
[0131] In the present invention, the cuckoo search algorithm adopts a random search strategy combined with Lévy flight, enabling the algorithm to jump out of local optimal solutions and explore a wider solution space. This method helps to find the global optimal solution. Meanwhile, by adaptively adjusting the step size factor and discovery probability, the algorithm can quickly explore in the initial stage and focus on optimization in the later stage, effectively accelerating the convergence speed. Through the cuckoo search algorithm, the parameters of the convolutional neural network, including weights and biases, can be finely adjusted to ensure that the network structure achieves the best performance when processing specific tasks. This is particularly important for network data security positioning because the network environment is complex and changeable, and the need for parameter adjustment is more precise. Optimizing network parameters through a global search algorithm can more effectively avoid overfitting problems compared with traditional local optimization methods (such as gradient descent), especially in network security data analysis with a large amount of data or complex features.
[0132] The beneficial effects brought by the technical solution provided in the embodiment of the present invention at least include:
[0133] In the present invention, by calculating the correlation index of each network feature subset, selecting the network feature subsets with a correlation index greater than or equal to the correlation index threshold, calculating the security positioning index of the selected network feature subsets, and selecting the network feature subset with the largest security positioning index as the target network feature subset, when processing a large amount of network traffic data, potential security threats can be identified in a timely and accurate manner, enabling the network to have sufficient defense capabilities when encountering attacks. According to the target network feature subset, through a deep learning algorithm, network data security positioning is carried out, providing effective feature extraction and classification capabilities in the face of unknown attacks and complex network traffic data behaviors, with a high recognition rate and a low false alarm rate, thereby improving the overall network data security positioning effect.
[0134] Refer to the attached Figure 2 illustrates the structural schematic diagram of a network data security positioning system provided by the present invention based on a deep learning algorithm.
[0135] The present invention also provides a network data security positioning system 30 based on a deep learning algorithm, including: a memory 303 and one or more processors 301.
[0136] One or more application programs are stored in the memory 303, and the one or more application programs are adapted to be executed by the one or more processors 301 to implement the network data security positioning method based on the deep learning algorithm described in the method embodiment.
[0137] The network data security positioning system 30 based on a deep learning algorithm includes: a processor 301 and a memory 303. Among them, the processor 301 and the memory 303 are connected, such as connected through a bus 302.
[0138] The structure of the network data security positioning system 30 based on the deep learning algorithm does not constitute a limitation to the embodiments of the present invention.
[0139] The processor 301 can be a CPU, a general-purpose processor, a DSP, an ASIC, an FPGA, or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute various exemplary logical blocks, modules, and circuits described in connection with the disclosure of the present invention. The processor 301 can also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, etc.
[0140] The bus 302 can include a path for transmitting information between the above components. The bus 302 can be a PCI bus or an EISA bus, etc. The bus 302 can be divided into an address bus, a data bus, a control bus, etc. For the sake of simplicity of representation, only a thick line is shown in the figure, but it does not mean that there is only one bus or one type of bus.
[0141] The memory 303 can be a ROM or other types of static storage devices that can store static information and instructions, a RAM, or other types of dynamic storage devices that can store information and instructions. It can also be an EEPROM, a CD-ROM, or other optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media, or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto.
[0142] It should be noted that the network data security positioning system 30 based on the deep learning algorithm can implement the above-mentioned network data security positioning method based on the deep learning algorithm and can achieve the same or similar technical effects. To avoid repetition, the present invention will not be elaborated herein.
[0143] The beneficial effects brought by the technical solution provided by the embodiments of the present invention at least include:
[0144] In the present invention, by calculating the correlation indices of each network feature subset, selecting the network feature subsets whose correlation indices are greater than or equal to the correlation index threshold, calculating the security positioning indices of the selected network feature subsets, and selecting the network feature subset with the largest security positioning index as the target network feature subset, when processing a large amount of network traffic data, potential security threats can be identified timely and accurately, enabling the network to have sufficient defense capabilities when encountering attacks. According to the target network feature subset, through deep learning algorithms, network data security positioning is carried out, providing effective feature extraction and classification capabilities in the face of unknown attacks and complex network traffic data behaviors, with high recognition rate and low false alarm rate, thereby improving the overall network data security positioning effect.
[0145] The above is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
[0146] The following points need to be explained:
[0147] (1) The attached drawings of the embodiments of the present invention only relate to the structures involved in the embodiments of the present invention, and other structures can refer to the general design.
[0148] (2) For clarity, in the attached drawings used to describe the embodiments of the present invention, the thickness of layers or regions is enlarged or reduced, that is, these drawings are not drawn according to the actual scale. It can be understood that when an element such as a layer, film, region or substrate is referred to as being "on" or "under" another element, the element can be "directly" on or under the other element or there can be intermediate elements.
[0149] (3) Without conflict, the embodiments of the present invention and the features in the embodiments can be combined with each other to obtain new embodiments.
[0150] The above is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. The protection scope of the present invention should be subject to the protection scope of the claims.
Claims
1. A network data security positioning method based on deep learning algorithm, characterized in that: include: S1: Obtain network traffic data; S2: extracting multiple network features of the network traffic data; S3: Randomly combine the extracted multiple network features to construct multiple network feature subsets; S4: Calculating the correlation index of each of the network feature subsets; S5: Select a subset of network features whose correlation index is greater than or equal to a correlation index threshold; S6: Calculate the security positioning index of each selected network feature subset; S7: Select the network feature subset with the largest security positioning index as the target network feature subset; S8: Based on the target network feature subset, network data security positioning is performed through a deep learning algorithm.
2. The network data security positioning method based on deep learning algorithm according to claim 1 is characterized in that: The correlation index is specifically: Among them, r cj represents the correlation index between the cth network feature subset and the jth category, k represents the number of network features in the network feature subset, represents the average correlation index between all network features in the network feature subset and the jth category, Represents the average cross-correlation index between each network feature in the network feature subset.
3. The network data security positioning method based on deep learning algorithm according to claim 1 is characterized in that: The safety positioning index is specifically: S i =ω1A i +ω2IG i Among them, S i represents the security positioning index of the ith network feature subset, A i represents the positioning accuracy of the i-th network feature subset, IG i represents the information gain of the i-th network feature subset, ω1 represents the weight coefficient of positioning accuracy, and ω2 represents the weight coefficient of information gain.
4. The network data security positioning method based on deep learning algorithm according to claim 3 is characterized in that: The positioning accuracy is specifically: Among them, A i represents the positioning accuracy of the i-th network feature subset, TP i represents the true number of examples of the i-th network feature subset, TN i Represents the number of true negative examples of the i-th network feature subset, FP i represents the number of false positives of the i-th network feature subset, FN i Represents the number of false negative examples for the i-th network feature subset.
5. The network data security positioning method based on deep learning algorithm according to claim 3 is characterized in that: The information gain is specifically: IG i =H(Y)-H(Y|M i ) Among them, IG i represents the information gain of the i-th network feature subset, H(Y) represents the entropy value of the category, Y represents the category, H(Y|M i ) represents the conditional entropy of the category under the network feature subset, M i Represents the selected i-th network feature subset.
6. The network data security positioning method based on deep learning algorithm according to claim 1 is characterized in that: The S7 is specifically: A network feature subset with the largest safety positioning index is selected as the target network feature subset. When there are multiple network feature subsets with the same maximum safety positioning index, a network feature subset with the least number of network features is selected as the target network feature subset.
7. The network data security positioning method based on deep learning algorithm according to claim 1 is characterized in that: The S8 is specifically: Based on the target network feature subset, network data security positioning is performed through a convolutional neural network.
8. The network data security positioning method based on deep learning algorithm according to claim 7 is characterized in that: The convolutional neural network includes an input layer, multiple convolutional layers, multiple pooling layers, a fully connected layer and an output layer, the convolutional layer includes a first convolutional layer, a second convolutional layer and a third convolutional layer, the pooling layer includes a first pooling layer, a second pooling layer and a third pooling layer, and S8 specifically includes: S801: In the input layer, input the target network feature subset; S802: In the first convolutional layer, extract features from the target network feature subset to generate a first local feature map; S803: In the first pooling layer, performing a maximum pooling operation on the first local feature map to generate a first pooling feature map; S804: In the second convolutional layer, extract features from the first pooled feature map to generate a second local feature map; S805: In the second pooling layer, performing a maximum pooling operation on the second local feature map to generate a second pooling feature map; S806: In the third convolutional layer, extract features from the second pooled feature map to generate a third local feature map; S807: In the third pooling layer, performing a maximum pooling operation on the third local feature map to generate a third pooling feature map; S808: performing weighted fusion on the first pooling feature map, the second pooling feature map and the third pooling feature map to generate a target feature map; S809: In the fully connected layer, performing network data security positioning according to the target feature graph; S810: In the output layer, output the network data security positioning result.
9. The network data security positioning method based on deep learning algorithm according to claim 8 is characterized in that: The training method of the convolutional neural network is specifically as follows: With the goal of minimizing the mean square error loss function, the convolutional neural network is trained through a cuckoo search algorithm to determine the optimal network parameters of the convolutional neural network.
10. A network data security positioning system based on deep learning algorithm, characterized in that: include: memory and one or more processors; One or more applications are stored in the memory, and the one or more applications are suitable for being executed by the one or more processors to implement the network data security positioning method based on deep learning algorithm as described in any one of claims 1 to 9.