Network security identification method and device based on log aggregation and electronic equipment

Through a log aggregation method, the security log data of terminal devices and servers are aggregated and input into the improved SVM model, solving the shortcomings of the existing technology in identifying network attacks and achieving more efficient and flexible network threat identification.

CN120017303APending Publication Date: 2025-05-16SHANGHAI THREE ZERO GUARD INFORMATION SECURITY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411921199.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-25
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

Existing network security protection methods are difficult to effectively identify and deal with evolving network attack methods and unknown security vulnerabilities, especially when handling large amounts of message data in network requests, the computing complexity is high and the demand for hardware resources increases.

Method used

The network security identification method based on log aggregation is adopted. By obtaining the security log data of multiple terminal devices and servers, converting and mapping and aggregating, forming a network security log data set and inputting it into the improved Support Vector Machine (SVM) network threat identification classification model optimized by the Firefly algorithm.

Benefits of technology

It improves the model's sensitivity to the network threat status caused by terminal device abnormalities, reduces the demand for hardware resources, expands the scope of application of network threat identification, and improves identification efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120017303A_ABST
    Figure CN120017303A_ABST
Patent Text Reader

Abstract

The invention provides a network security identification method and device based on log aggregation and electronic equipment, and relates to the technical field of computers, and the method comprises the steps: obtaining security log data of various terminal equipment, server security log data and a preset time window initial value; converting the security log data of the to-be-aggregated terminal equipment category one by one based on the conversion mapping relationship, and summarizing the security log data with the security log data of the aggregated target terminal equipment category to obtain security log aggregated data of various terminal equipment; processing the multiple terminal equipment security log aggregated data and the server security log data based on a preset time window initial value to obtain a network security log data set; and inputting the network security log data set into the network threat identification and classification model to obtain a network threat identification result. According to the method, aggregation of heterogeneous log data is completed, input data sources of the model are enriched, and the sensitivity of the model to a network threat state caused by terminal equipment abnormity is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular to a network security identification method, device and electronic device based on log aggregation. Background Art

[0002] With the rapid development of information technology and the widespread application of technologies such as the Internet and cloud computing, the network environment is becoming increasingly complex, and network applications are highly open, which also makes network systems face unprecedented security challenges. Network attackers can use security vulnerabilities in the system to carry out various forms of network attacks, including but not limited to website paralysis, sensitive information leakage, web page content tampering, etc. These attacks not only seriously damage the information security of enterprises, but also pose a huge threat to user data privacy.

[0003] Traditional network security protection methods, such as rule-matching-based intrusion prevention systems, can resist known network threats to a certain extent, but their effects are often unsatisfactory in the face of evolving attack methods and unknown security vulnerabilities (such as 0day vulnerabilities). The high false positive and false negative rates of such systems not only increase the burden of security operations and maintenance, but may also cause new security issues due to misoperation. In addition, the formulation and maintenance of rules require highly specialized knowledge and continuous efforts, which is a considerable cost for most companies.

[0004] In order to meet this challenge, in recent years, academia and industry have begun to explore anomaly detection technologies based on deep learning and machine learning. These technologies can effectively identify unknown attacks by modeling the multi-dimensional features of network traffic or request messages, thereby significantly improving the intelligent level of network security protection. For example, an existing patent (CN116668089A) uses deep learning technologies such as CNN network model and LSTM network model to design traffic attack identification modules for different data types by identifying the type of network message data, effectively improving the efficiency of network attack classification and identification. However, this method has high computational complexity when processing scenarios with large amounts of message data in network requests, and the demand for hardware resources also increases accordingly. Another patent (CN108400995A) proposes using a network traffic capture module to capture network traffic data of multiple network nodes in real time, and identifying network attacks by calculating the similarity of traffic data under network grouping. This method can detect network attacks through abnormal changes in traffic, but its identification effect is greatly reduced when facing attacks that do not cause significant changes in network traffic.

[0005] Therefore, a network security identification method, device and electronic device based on log aggregation are proposed. Summary of the invention

[0006] This specification provides a network security identification method, device and electronic device based on log aggregation, which completes the aggregation of heterogeneous log data, enriches the input data source of the model, and improves the model's sensitivity to network threat status caused by terminal device anomalies.

[0007] This specification provides a network security identification method based on log aggregation, including:

[0008] Obtain security log data of multiple terminal devices, server security log data, and a preset time window initial value; wherein the security log data of the multiple terminal devices includes security log data of the aggregation target terminal device category and security log data of the terminal device category to be aggregated;

[0009] Based on the conversion mapping relationship, the security log data of the terminal device category to be aggregated are converted one by one, and the security log data of the aggregation target terminal device category are aggregated to obtain security log aggregation data of multiple terminal devices;

[0010] Processing the multiple terminal device security log aggregation data and the server security log data based on the preset time window initial value to obtain a network security log data set;

[0011] The network security log data set is input into a network threat identification classification model to obtain a network threat identification result.

[0012] Optionally, the process of acquiring the security log data of the aggregation target terminal device category and the security log data of the terminal device category to be aggregated includes:

[0013] Obtain security log data from various terminal devices;

[0014] The security log data are clustered based on the categories of the terminal devices, entries of the security log data of each category of the terminal devices are determined, the terminal device category with the most entries of the security log data of the terminal devices is used as the aggregation target terminal device category, and the remaining terminal device categories are used as the terminal device categories to be aggregated.

[0015] Optionally, the process of establishing the conversion mapping relationship includes:

[0016] Numbering all entries of the security log data of each type of terminal device, extracting key log information corresponding to each entry one by one, and vectorizing the key log information;

[0017] Performing a cosine similarity operation on the vectorized log key information of the terminal device type to be aggregated and the vectorized log key information of the aggregation target terminal device of all entries one by one;

[0018] The vectorized log key information of the terminal device category to be aggregated with a cosine similarity higher than a preset value is distributedly verified with the vectorized log key information of the aggregation target terminal device to establish a conversion mapping relationship of the matching vector pair.

[0019] Optionally, performing a distribution check on the vectorized log key information of the terminal device category to be aggregated whose cosine similarity is higher than a preset value and the vectorized log key information of the aggregation target terminal device to establish a conversion mapping relationship of the matching vector pair includes:

[0020] A preset number N of samples are extracted from the security log data of the terminal device category to be aggregated and the security log data of the aggregation target terminal device. r The sample data is obtained by combining a numerical matrix based on the numerical list corresponding to the sample data, and performing numerical normalization processing on the two numerical matrices to obtain a matrix A and a matrix B;

[0021] The column vectors of the matrix A are combined with the column vectors of the matrix B in pairs to obtain column vector pairs. Any column vector pair is denoted as vector a = {a1, a2, ..., a m}, b={b1,b2,...,b n}, calculate the mean difference T mean , median difference T median , variance difference T variance ;

[0022] Combine vectors a and b into a vector c of length m+n, randomly shuffle the order of the elements of vector c, form a new vector a′ with the first m elements and a new vector b′ with the last n elements, and calculate the mean difference T′ mean , median difference T′ median , variance difference T′ variance ;

[0023] Repeat the random ordering of the elements of vector c, compose the first m elements into a new vector a′, and the last n elements into a new vector b′, and calculate the mean difference T′ mean , median difference T′ median , variance difference T′ variance , record the number of repetitions as N repeat =β*N r , β is the preset parameter, Among them, Count(T′ mean >T mean ) represents T′ in the process of repeatedly shuffling elements mean Greater than T mean The number of times; similarly, When pmean 、p median 、p variance When both are greater than 0.05, that is, p total =p mean +p median +p variance ; For all vector pairs consisting of matrix A and matrix B, select p total The largest vector pair is taken as the matching vector pair;

[0024] Based on the multiple relationship between the matching vector pairs and the numerical homogeneity of the sample data, a conversion mapping relationship of the matching vector pairs is established.

[0025] Optionally, the processing of the plurality of terminal device security log aggregation data and the server security log data based on the preset time window initial value to obtain a network security log data set includes:

[0026] The server and the terminal device are time synchronized, and the multiple terminal device security log aggregate data and the server security log data are divided into time periods according to the preset time window initial value to obtain the multiple terminal device security log aggregate data within the time period and the server security log data within the time period;

[0027] Collecting statistics on the security log aggregate data of various terminal devices in each time period and the security log data of the server in each time period one by one, and summarizing the statistical information of the security log aggregate data of various terminal devices in all time periods and the statistical information of the security log data of the server in all time periods, to obtain a first statistical information matrix and a second statistical information matrix respectively;

[0028] Determining the Spearman rank correlation coefficient based on the first statistical information matrix and the second statistical information matrix;

[0029] The time window value corresponding to the maximum Spearman rank correlation coefficient is used as the statistical window duration, and the first statistical information matrix and the second statistical information matrix corresponding to the statistical window duration are merged to obtain a network security log data set.

[0030] Optionally, the process of constructing the network threat identification and classification model includes:

[0031] Finding Hyperplanes in High-Dimensional Feature Spaces Among them, ω is the normal vector in the high-dimensional space, b is the bias term in the high-dimensional space, is a nonlinear mapping function;

[0032] The constraints are

[0033] Get the value of the bias term b in the high-dimensional space through the support vector:

[0034]

[0035] Among them, SV is the index set of support vectors, n s is the number of support vectors, y i represents the classification label of the i-th sample, K(x i ,x j ) is the polynomial kernel function K(x i ,x j )=(x i ·x j +c) d ,x i ·x j represents the dot product of two sample vectors, c represents the constant offset of the polynomial kernel function, and d represents the polynomial order;

[0036] The decision function for classification prediction is:

[0037]

[0038] Among them, x represents the security log vector of the sample to be predicted, and sign() represents the indicator function;

[0039] The output of the indicator function sign() is the result of the classification prediction, that is, x represents the classification prediction corresponding to the security log vector of the sample to be predicted.

[0040] Optionally, the parameter optimization process of the network threat identification and classification model includes:

[0041] Initialize the model parameters of the firefly model, where the model parameters include the firefly population size, the maximum number of iterations, the value range of the regularization penalty parameter, the kernel function constant offset and the value range of the polynomial order, and generate the firefly population;

[0042] Calculating the fitness value of each firefly according to the position of all individuals in the firefly population;

[0043] For any i-th firefly, according to the current position X i Calculate the fitness value. The higher the fitness value, the brighter the firefly. Randomly select N fireflies in the population. sr individuals, where N sr =N train τ , τ∈(0,1), randomly select a firefly position X with a brightness higher than its own among these individuals j As the flight target, and update the flight target position:

[0044]

[0045] Where r represents X i , X j The distance between two locations, β max and β min are preset values ​​of the attraction coefficient, γ=1 is the absorption coefficient of the medium to light, and α is the step factor of the disturbance term N cir is the current iteration number, p is a preset integer, and rand(normal(0,1)) represents a random number that follows a standard normal distribution with a mean of 0 and a standard deviation of 1;

[0046] Calculate the flight target position X′ i The fitness value and the current position X i The fitness of the flight target is compared. If the flight target position is X′ i The fitness value is higher than the current position X i The firefly will fly to the target position X′ if the fitness is i , otherwise the firefly will stay at the current position X i ;

[0047] For the firefly with the highest brightness in the population, there is no flight target position X′ i :

[0048] X′ Rbest =X Rbest +α*rand(normal(0,1))

[0049] Among them, X Rbest is the position of the firefly with the highest brightness in the population within one iteration, i.e., the position of the firefly with the largest fitness value, X′ Rbest The target position of the firefly with the highest brightness;

[0050] Repeatedly return to calculate the fitness value of each firefly according to the position of all individuals in the firefly population until the maximum number of iterations is reached, and use the searched optimal firefly position as the output, which is the optimized parameter of the network threat identification and classification model.

[0051] This specification provides a network security identification device based on log aggregation, including:

[0052] An acquisition module, used to acquire security log data of multiple terminal devices, server security log data, and a preset time window initial value; wherein the security log data of the multiple terminal devices includes security log data of the aggregation target terminal device category and security log data of the terminal device category to be aggregated;

[0053] An aggregation module, used to convert the security log data of the terminal device category to be aggregated one by one based on the conversion mapping relationship, and aggregate them with the security log data of the aggregation target terminal device category to obtain security log aggregation data of multiple terminal devices;

[0054] A processing module, configured to process the plurality of terminal device security log aggregation data and the server security log data based on the preset time window initial value to obtain a network security log data set;

[0055] The identification module is used to input the network security log data set into the network threat identification classification model to obtain a network threat identification result.

[0056] Optionally, the process of acquiring the security log data of the aggregation target terminal device category and the security log data of the terminal device category to be aggregated includes:

[0057] Obtain security log data from various terminal devices;

[0058] The security log data are clustered based on the categories of the terminal devices, entries of the security log data of each category of the terminal devices are determined, the terminal device category with the most entries of the security log data of the terminal devices is used as the aggregation target terminal device category, and the remaining terminal device categories are used as the terminal device categories to be aggregated.

[0059] Optionally, the process of establishing the conversion mapping relationship includes:

[0060] Numbering all entries of the security log data of each type of terminal device, extracting key log information corresponding to each entry one by one, and vectorizing the key log information;

[0061] Performing a cosine similarity operation on the vectorized log key information of the terminal device type to be aggregated and the vectorized log key information of the aggregation target terminal device of all entries one by one;

[0062] The vectorized log key information of the terminal device category to be aggregated with a cosine similarity higher than a preset value is distributedly verified with the vectorized log key information of the aggregation target terminal device to establish a conversion mapping relationship of the matching vector pair.

[0063] Optionally, performing a distribution check on the vectorized log key information of the terminal device category to be aggregated whose cosine similarity is higher than a preset value and the vectorized log key information of the aggregation target terminal device to establish a conversion mapping relationship of the matching vector pair includes:

[0064] A preset number N of samples are extracted from the security log data of the terminal device category to be aggregated and the security log data of the aggregation target terminal device. r The sample data is obtained by combining a numerical matrix based on the numerical list corresponding to the sample data, and performing numerical normalization processing on the two numerical matrices to obtain a matrix A and a matrix B;

[0065] The column vectors of the matrix A are combined with the column vectors of the matrix B in pairs to obtain column vector pairs. Any column vector pair is denoted as vector a = {a1, a2, ..., a m}, b={b1,b2,...,b n}, calculate the mean difference T mean , median difference T median , variance difference T variance ;

[0066] Combine vectors a and b into a vector c of length m+n, randomly shuffle the order of the elements of vector c, form a new vector a′ with the first m elements and a new vector b′ with the last n elements, and calculate the mean difference T′ mean , median difference T′ median , variance difference T′ variance ;

[0067] Repeat the random ordering of the elements of vector c, compose the first m elements into a new vector a′, and the last n elements into a new vector b′, and calculate the mean difference T′ mean , median difference T′ median , variance difference T′ variance , record the number of repetitions as N repeat =β*N r , β is the preset parameter, Among them, Count(T′ mean >T mean ) represents T′ in the process of repeatedly shuffling elements mean Greater than T mean The number of times; similarly, When p mean 、p median 、p variance When both are greater than 0.05, that is, p total =p mean +p median +p variance ; For all vector pairs consisting of matrix A and matrix B, select p total The largest vector pair is taken as the matching vector pair;

[0068] Based on the multiple relationship between the matching vector pairs and the numerical homogeneity of the sample data, a conversion mapping relationship of the matching vector pairs is established.

[0069] Optionally, the processing module includes:

[0070] The server and the terminal device are time synchronized, and the multiple terminal device security log aggregate data and the server security log data are divided into time periods according to the preset time window initial value to obtain the multiple terminal device security log aggregate data within the time period and the server security log data within the time period;

[0071] Collecting statistics on the security log aggregate data of various terminal devices in each time period and the security log data of the server in each time period one by one, and summarizing the statistical information of the security log aggregate data of various terminal devices in all time periods and the statistical information of the security log data of the server in all time periods, to obtain a first statistical information matrix and a second statistical information matrix respectively;

[0072] Determining the Spearman rank correlation coefficient based on the first statistical information matrix and the second statistical information matrix;

[0073] The time window value corresponding to the maximum Spearman rank correlation coefficient is used as the statistical window duration, and the first statistical information matrix and the second statistical information matrix corresponding to the statistical window duration are merged to obtain a network security log data set.

[0074] Optionally, the process of constructing the network threat identification and classification model includes:

[0075] Finding Hyperplanes in High-Dimensional Feature Spaces Among them, ω is the normal vector in the high-dimensional space, b is the bias term in the high-dimensional space, is a nonlinear mapping function;

[0076] The constraints are

[0077] Get the value of the bias term b in the high-dimensional space through the support vector:

[0078]

[0079] Among them, SV is the index set of support vectors, n s is the number of support vectors, y i represents the classification label of the i-th sample, K(x i ,x j ) is the polynomial kernel function K(x i ,x j )=(x i ·x j +c) d ,x i ·xj represents the dot product of two sample vectors, c represents the constant offset of the polynomial kernel function, and d represents the polynomial order;

[0080] The decision function for classification prediction is:

[0081]

[0082] Among them, x represents the security log vector of the sample to be predicted, and sign() represents the indicator function;

[0083] The output of the indicator function sign() is the result of the classification prediction, that is, x represents the classification prediction corresponding to the security log vector of the sample to be predicted.

[0084] Optionally, the parameter optimization process of the network threat identification and classification model includes:

[0085] Initialize the model parameters of the firefly model, where the model parameters include the firefly population size, the maximum number of iterations, the value range of the regularization penalty parameter, the kernel function constant offset and the value range of the polynomial order, and generate the firefly population;

[0086] Calculating the fitness value of each firefly according to the position of all individuals in the firefly population;

[0087] For any i-th firefly, according to the current position X i Calculate the fitness value. The higher the fitness value, the brighter the firefly. Randomly select N fireflies in the population. sr individuals, where N sr =N train τ , τ∈(0,1), randomly select a firefly position X with a brightness higher than its own among these individuals j As the flight target, and update the flight target position:

[0088]

[0089] Where r represents X i , X j The distance between two locations, β max and β min are preset values ​​of the attraction coefficient, γ=1 is the absorption coefficient of the medium to light, and α is the step factor of the disturbance term N cur is the current iteration number, p is a preset integer, and rand(normal(0,1)) represents a random number that follows a standard normal distribution with a mean of 0 and a standard deviation of 1;

[0090] Calculate the flight target position X′i The fitness value and the current position X i The fitness of the flight target is compared. If the flight target position is X′ i The fitness value is higher than the current position X i The firefly will fly to the target position X′ if the fitness is i , otherwise the firefly will stay at the current position X i ;

[0091] For the firefly with the highest brightness in the population, there is no flight target position X′ i :

[0092] X′ Rbest =X Rbest +α*rand(normal(0,1))

[0093] Among them, X Rbest is the position of the firefly with the highest brightness in the population within one iteration, i.e., the position of the firefly with the largest fitness value, X′ Rbest The target position of the firefly with the highest brightness;

[0094] Repeatedly return to calculate the fitness value of each firefly according to the position of all individuals in the firefly population until the maximum number of iterations is reached, and use the searched optimal firefly position as the output, which is the optimized parameter of the network threat identification and classification model.

[0095] This specification also provides an electronic device, wherein the electronic device includes:

[0096] processor; and,

[0097] A memory storing computer executable instructions, which when executed cause the processor to perform any of the above methods.

[0098] The present specification also provides a computer-readable storage medium, wherein the computer-readable storage medium stores one or more programs, and when the one or more programs are executed by a processor, any of the above methods is implemented.

[0099] In the present invention, in the current complex network security environment, in order to address the challenge of difficult integration and analysis of heterogeneous threat log data generated by devices of different manufacturers, natural language processing (NLP) technology and numerical statistical methods are used to successfully achieve the aggregation of heterogeneous log data. The terminal device security log is combined with the server security log, and the statistical results of these integrated security log data are used as input to a support vector machine (SVM) network threat identification classification prediction model optimized by an improved firefly algorithm, which not only enriches the input data source of the model, but also significantly improves the sensitivity of the model to the network threat state caused by abnormal terminal devices. The input data of the model only depends on the security behavior log of the terminal device and the security status log of the server, and is not affected by the size of the network request flow or the size of the message data, thereby greatly expanding the scope of application of network threat identification. In addition, the parameters of the SVM model are optimally solved by using the improved firefly algorithm. The algorithm shows a more reasonable search distribution in the three stages of fast spatial search, disturbance attenuation, and local search, which not only improves the search efficiency of parameter optimization, but also enhances the ability of the algorithm to jump out of the local optimal solution, ensuring the high stability and performance of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0100] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0101] Figure 1 A schematic diagram of the principle of a network security identification method based on log aggregation provided in an embodiment of this specification;

[0102] Figure 2 A schematic diagram of the structure of a network security identification device based on log aggregation provided in an embodiment of this specification;

[0103] Figure 3 A schematic diagram of the structure of an electronic device provided in an embodiment of this specification;

[0104] Figure 4 A schematic diagram of a computer-readable medium provided for an embodiment of this specification. DETAILED DESCRIPTION

[0105] The following description is used to disclose the present invention so that those skilled in the art can implement the present invention. The preferred embodiments described below are only examples, and those skilled in the art can think of other obvious variations. The basic principles of the present invention defined in the following description can be applied to other embodiments, variations, improvements, equivalents, and other technical solutions that do not deviate from the spirit and scope of the present invention.

[0106] The following is combined with Figure 1-4 The exemplary embodiments of the present invention are described more fully. However, the exemplary embodiments can be implemented in various forms, and it should not be understood that the present invention is limited to the embodiments set forth herein. On the contrary, providing these exemplary embodiments can make the present invention more comprehensive and complete, and it is more convenient to fully convey the inventive concept to those skilled in the art. The same reference numerals in the figures represent the same or similar elements, components or parts, and thus their repeated description will be omitted.

[0107] Under the premise of being consistent with the technical concept of the present invention, the features, structures, characteristics or other details described in a specific embodiment do not exclude that they can be combined in one or more other embodiments in a suitable manner.

[0108] In the description of specific embodiments, the features, structures, characteristics or other details described in the present invention are intended to enable those skilled in the art to fully understand the embodiments. However, it does not exclude that those skilled in the art can practice the technical solutions of the present invention without one or more of the specific features, structures, characteristics or other details.

[0109] The flowcharts shown in the accompanying drawings are only exemplary and do not necessarily include all the contents and operations / steps, nor must they be executed in the order described. For example, some operations / steps can be decomposed, and some operations / steps can be combined or partially combined, so the actual execution order may change according to actual conditions.

[0110] The block diagrams shown in the accompanying drawings are merely functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities may be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.

[0111] The term "and / or" or "and / or" includes all combinations of any one or more of the associated listed items.

[0112] Figure 1 A schematic diagram of a network security identification method based on log aggregation provided in an embodiment of this specification, the method may include:

[0113] S110: Acquire security log data of multiple terminal devices, server security log data, and a preset time window initial value; wherein the security log data of the multiple terminal devices includes security log data of the aggregation target terminal device category and security log data of the terminal device category to be aggregated;

[0114] Optionally, the process of acquiring the security log data of the aggregation target terminal device category and the security log data of the terminal device category to be aggregated includes:

[0115] Obtain security log data from various terminal devices;

[0116] The security log data are clustered based on the categories of the terminal devices, entries of the security log data of each category of the terminal devices are determined, the terminal device category with the most entries of the security log data of the terminal devices is used as the aggregation target terminal device category, and the remaining terminal device categories are used as the terminal device categories to be aggregated.

[0117] In the specific implementation of this specification, security log data from a variety of different terminal devices are obtained, which comprehensively record the security status and events of each terminal during operation; then, based on the specific category of the terminal device, a clustering algorithm is used to classify and aggregate the collected security log data to clearly distinguish and determine the number of security log data entries corresponding to each type of terminal device. In this process, the terminal device category with the largest number of security log data entries is identified and set as the aggregation target terminal device category, which means that the security log information generated by the terminal in this category is the richest and has a high reference value for the analysis of the overall security situation; at the same time, the remaining terminal device categories with relatively few security log data entries but also containing important security information are uniformly classified as the terminal device category to be aggregated. Although the log volume of these categories to be aggregated is not as large as that of the aggregation target category, their log data is also of great significance for a comprehensive understanding of the terminal security status and the discovery of potential security risks. In the future, further data integration or comparative analysis may be carried out with the aggregation target category according to analysis needs.

[0118] S120: converting the security log data of the terminal device category to be aggregated one by one based on the conversion mapping relationship, and aggregating them with the security log data of the aggregation target terminal device category to obtain security log aggregation data of multiple terminal devices;

[0119] In a specific implementation of the present specification, data preprocessing and conversion are performed on all successfully matched log entries under the category of aggregated terminal devices according to the conversion mapping relationship, and then the converted data are aggregated and integrated with the original data of the aggregated target terminal device category to generate security log aggregation data covering a variety of terminal devices. This aggregated security log data is rich in content, specifically including the operation type, detailed information on the target object, and the numerical vector corresponding to the operation description after text vectorization processing. In addition, the numerical lists in the log are also carefully processed, and they are split by column and recombined into new numerical vectors to facilitate subsequent data analysis and mining.

[0120] Optionally, the process of establishing the conversion mapping relationship includes:

[0121] Numbering all entries of the security log data of each type of terminal device, extracting key log information corresponding to each entry one by one, and vectorizing the key log information;

[0122] Performing a cosine similarity operation on the vectorized log key information of the terminal device type to be aggregated and the vectorized log key information of the aggregation target terminal device of all entries one by one;

[0123] The vectorized log key information of the terminal device category to be aggregated with a cosine similarity higher than a preset value is distributedly verified with the vectorized log key information of the aggregation target terminal device to establish a conversion mapping relationship of the matching vector pair.

[0124] In the specific implementation of this specification, based on a number of key technologies in natural language processing (NLP), including word segmentation, named entity recognition (NER) and dependency syntax analysis, key information in each log entry can be accurately extracted one by one. This information specifically covers the operation type, target object, operation description and numerical list. In order to further realize the quantitative processing and efficient analysis of log data, advanced vectorization methods such as TF-IDF (term frequency-inverse document frequency) or Word2Vec are used to vectorize the extracted operation type, target object and operation description. The originally unstructured log text data is effectively converted into a structured numerical vector, which lays a solid foundation for subsequent tasks such as data mining, pattern recognition and machine learning.

[0125] For all terminal device categories to be aggregated, select security log entries in these categories and perform aggregation entry matching with all entries of the aggregation target terminal device category. For successfully matched aggregation entries, set a conversion mapping for these entries under the terminal device category to be aggregated, and map them to the corresponding target entries of the aggregation target terminal device category.

[0126] Calculate the cosine similarity of the digitized vectors of the two log entries to be compared on key information (including operation type, target object, and operation description). These key information have been previously digitized using vectorization methods such as TF-IDF and Word2Vec. If and only if the cosine similarity of the two log entries on the three key information is higher than the preset threshold, the two log entries are judged to be preliminarily matched. This step ensures that log entries are considered matched only when the content is highly similar, thereby improving the accuracy and effectiveness of aggregation.

[0127] Optionally, performing a distribution check on the vectorized log key information of the terminal device category to be aggregated whose cosine similarity is higher than a preset value and the vectorized log key information of the aggregation target terminal device to establish a conversion mapping relationship of the matching vector pair includes:

[0128] A preset number N of samples are extracted from the security log data of the terminal device category to be aggregated and the security log data of the aggregation target terminal device. r The sample data is obtained by combining a numerical matrix based on the numerical list corresponding to the sample data, and performing numerical normalization processing on the two numerical matrices to obtain a matrix A and a matrix B;

[0129] The column vectors of the matrix A are combined with the column vectors of the matrix B in pairs to obtain column vector pairs. Any column vector pair is denoted as vector a = {a1, a2, ..., a m}, b={b1,b2,...,b n}, calculate the mean difference T mean , median difference T median , variance difference T variance ;

[0130] Combine vectors a and b into a vector c of length m+n, randomly shuffle the order of the elements of vector c, form a new vector a′ with the first m elements and a new vector b′ with the last n elements, and calculate the mean difference T′ mean , median difference T′ median , variance difference T′ variance ;

[0131] Repeat the random ordering of the elements of vector c, compose the first m elements into a new vector a′, and the last n elements into a new vector b′, and calculate the mean difference T′ mean , median difference T′ median , variance difference T′ variance , record the number of repetitions as N repeat =β*N r , β is the preset parameter, Among them, Count(T′ mean >T mean ) represents T′ in the process of repeatedly shuffling elements mean Greater than T mean The number of times; similarly, When p mean 、p median 、p variance When both are greater than 0.05, that is, p total =p mean +p median +p variance ; For all vector pairs consisting of matrix A and matrix B, select p total The largest vector pair is taken as the matching vector pair;

[0132] Based on the multiple relationship between the matching vector pairs and the numerical homogeneity of the sample data, a conversion mapping relationship of the matching vector pairs is established.

[0133] S130: Processing the multiple terminal device security log aggregation data and the server security log data based on the preset time window initial value to obtain a network security log data set;

[0134] Optionally, the S130 includes:

[0135] The server and the terminal device are time synchronized, and the multiple terminal device security log aggregate data and the server security log data are divided into time periods according to the preset time window initial value to obtain the multiple terminal device security log aggregate data within the time period and the server security log data within the time period;

[0136] Collecting statistics on the security log aggregate data of various terminal devices in each time period and the security log data of the server in each time period one by one, and summarizing the statistical information of the security log aggregate data of various terminal devices in all time periods and the statistical information of the security log data of the server in all time periods, to obtain a first statistical information matrix and a second statistical information matrix respectively;

[0137] Determining the Spearman rank correlation coefficient based on the first statistical information matrix and the second statistical information matrix;

[0138] The time window value corresponding to the maximum Spearman rank correlation coefficient is used as the statistical window duration, and the first statistical information matrix and the second statistical information matrix corresponding to the statistical window duration are merged to obtain a network security log data set.

[0139] In the specific implementation of this specification, the statistical information of the security log data covers multiple key aspects, including the operation type, the target object, and the numerical vector corresponding to the operation description after text vectorization. In addition, the statistical information also records the total number of log records, and conducts in-depth statistical analysis on each numerical vector. These statistical values ​​specifically include statistical indicators such as the maximum value, sum, median, and variance of the numerical vector elements, which provide strong support for the comprehensive understanding and in-depth analysis of the data. For the security log of the server, its statistical data is more detailed and targeted. These data not only include the basic performance indicators of the server, such as CPU usage, memory occupancy rate, etc., to reflect the operating efficiency and health of the server; they also include indicators of the server network, such as network traffic and bandwidth usage, number of network connections, number of abnormal connections and their proportion, etc., to reveal potential threats and abnormal behaviors at the network level. At the same time, the security log statistics of the server also involve the status information of the firewall and VPN, such as the number of illegal accesses, blacklist IP access records, etc., which are of great significance for preventing external attacks and illegal intrusions. In addition, database-related indicators are also an important part of statistical data, such as the number of data connection failures, abnormal queries and time-consuming queries, which provide important references for database security and performance optimization.

[0140] Spearman rank correlation coefficient Where n is the length of the vector, R X , R Y For the order data converted from the two vector original data, calculate Where d1 and d2 are the number of column vectors of the first matrix and the second matrix respectively, r s (R X(i) ,R X(j) ) represents the Sperman correlation coefficient between the i-th vector in the first matrix and the j-th vector in the second matrix.

[0141] In [T0,T max ] interval according to the step size T d Incrementally, obtain multiple time window values, and repeat the above steps to calculate the R corresponding to all values s value, R s The time window value corresponding to the maximum value is used as the statistical window length T. Among them, the preset time window initial value T0 and step length T d , maximum value T max .

[0142] The first matrix and the second matrix when the statistical window length is T are merged as the network security log data set composed of statistical data of all time periods. The time periods with security threat events are marked with status, and the marked data label is 1, otherwise it is marked as -1, and the above-marked data set is divided into a training data set and a test data set.

[0143] S140: Inputting the network security log data set into a network threat identification classification model to obtain a network threat identification result.

[0144] Optionally, the process of constructing the network threat identification and classification model includes:

[0145] Finding Hyperplanes in High-Dimensional Feature Spaces Among them, ω is the normal vector in the high-dimensional space, b is the bias term in the high-dimensional space, is a nonlinear mapping function;

[0146] The constraints are

[0147] Get the value of the bias term b in the high-dimensional space through the support vector:

[0148]

[0149] Among them, SV is the index set of support vectors, n s is the number of support vectors, y i represents the classification label of the i-th sample, K(x i ,x j 0 is the polynomial kernel function K(x i ,x j )=(x i ·x j +c) d ,x i ·x j represents the dot product of two sample vectors, c represents the constant offset of the polynomial kernel function, and d represents the polynomial order;

[0150] The decision function for classification prediction is:

[0151]

[0152] Among them, x represents the security log vector of the sample to be predicted, and sign() represents the indicator function;

[0153] The output of the indicator function sign() is the result of the classification prediction, that is, x represents the classification prediction corresponding to the security log vector of the sample to be predicted.

[0154] In a specific implementation of this specification, in this network threat identification classification prediction model, the goal of the soft margin optimization problem can be converted into:

[0155]

[0156] Among them, J represents the objective function of the support vector machine, α i is the Lagrange multiplier, N is the total number of samples, C is the regularization parameter, and y i represents the classification label of the i-th sample, K(x i ,x j ) is the polynomial kernel function K(x i ,x j )=(x i ·x j +c) d ,x i ·x j represents the dot product of two sample vectors, c represents the constant offset of the polynomial kernel function, and d represents the polynomial order of the polynomial kernel function.

[0157] Optionally, the parameter optimization process of the network threat identification and classification model includes:

[0158] Initialize the model parameters of the firefly model, where the model parameters include the firefly population size, the maximum number of iterations, the value range of the regularization penalty parameter, the kernel function constant offset and the value range of the polynomial order, and generate the firefly population;

[0159] Calculating the fitness value of each firefly according to the position of all individuals in the firefly population;

[0160] For any i-th firefly, according to the current position X i Calculate the fitness value. The higher the fitness value, the brighter the firefly. Randomly select N fireflies in the population. sr individuals, where N sr =N train τ , τ∈(0,1), randomly select a firefly position X with a brightness higher than its own among these individuals j As the flight target, and update the flight target position:

[0161]

[0162] Where r represents X i , X j The distance between two locations, β max and β min are preset values ​​of the attraction coefficient, γ=1 is the absorption coefficient of the medium to light, and α is the step factor of the disturbance term N cur is the current iteration number, p is a preset integer, and rand(normal(0,1)) represents a random number that follows a standard normal distribution with a mean of 0 and a standard deviation of 1; calculate the flight target position X′ i The fitness value and the current position X i The fitness of the flight target is compared. If the flight target position is X′ i The fitness value is higher than the current position X i The firefly will fly to the target position X′ if the fitness is i , otherwise the firefly will stay at the current position X i ;

[0163] For the firefly with the highest brightness in the population, there is no flight target position X′ i :

[0164] X′ Rbest =X Rbest +α*rand(normal(0,1))

[0165] Among them, X Rbest is the position of the firefly with the highest brightness in the population within one iteration, i.e., the position of the firefly with the largest fitness value, X′ Rbest The target position of the firefly with the highest brightness;

[0166] Repeatedly return to calculate the fitness value of each firefly according to the position of all individuals in the firefly population until the maximum number of iterations is reached, and use the searched optimal firefly position as the output, which is the optimized parameter of the network threat identification and classification model.

[0167] In a specific implementation of this specification, the model parameters of the firefly model are initialized, wherein the model parameters include the firefly population size N p , the maximum number of iterations N iter , the value interval of the regularization penalty parameter C, the value interval of the kernel function constant offset C and the polynomial order and D, and generate the population of fireflies:

[0168]

[0169] Where D represents the number of parameters in the parameter optimization problem, (N p ,D) represents the space where the fireflies fly, rand(normal(0,1)) represents the random number X(N) that follows a standard normal distribution with a mean of 0 and a standard deviation of 1 p ,D) represents the position of all populations in space, S U , S L Represents the upper and lower bounds of the firefly flight search space.

[0170] The fitness value of each firefly is calculated according to the position of all individuals in the firefly population. The fitness value is calculated by inputting the training set data obtained by data preprocessing into the support vector machine model and calculating the fitness value based on the following formula:

[0171]

[0172] Among them, N train represents the number of training samples, Indicates the number of samples in the training sample whose prediction results are consistent with the true value of the data, R acc Represents the model classification accuracy of the corresponding parameters of individual fireflies.

[0173] In the present invention, in the current complex network security environment, in order to address the challenge of difficult integration and analysis of heterogeneous threat log data generated by devices of different manufacturers, natural language processing (NLP) technology and numerical statistical methods are used to successfully achieve the aggregation of heterogeneous log data. The terminal device security log is combined with the server security log, and the statistical results of these integrated security log data are used as input to a support vector machine (SVM) network threat identification classification prediction model optimized by an improved firefly algorithm, which not only enriches the input data source of the model, but also significantly improves the sensitivity of the model to the network threat state caused by abnormal terminal devices. The input data of the model only depends on the security behavior log of the terminal device and the security status log of the server, and is not affected by the size of the network request flow or the size of the message data, thereby greatly expanding the scope of application of network threat identification. In addition, the parameters of the SVM model are optimally solved by using the improved firefly algorithm. The algorithm shows a more reasonable search distribution in the three stages of fast spatial search, disturbance attenuation, and local search, which not only improves the search efficiency of parameter optimization, but also enhances the ability of the algorithm to jump out of the local optimal solution, ensuring the high stability and performance of the model.

[0174] Figure 2 A schematic diagram of a network security identification device based on log aggregation provided in an embodiment of this specification may include:

[0175] An acquisition module, used to acquire security log data of multiple terminal devices, server security log data, and a preset time window initial value; wherein the security log data of the multiple terminal devices includes security log data of the aggregation target terminal device category and security log data of the terminal device category to be aggregated;

[0176] An aggregation module, used to convert the security log data of the terminal device category to be aggregated one by one based on the conversion mapping relationship, and aggregate them with the security log data of the aggregation target terminal device category to obtain security log aggregation data of multiple terminal devices;

[0177] A processing module, configured to process the plurality of terminal device security log aggregation data and the server security log data based on the preset time window initial value to obtain a network security log data set;

[0178] The identification module is used to input the network security log data set into the network threat identification classification model to obtain a network threat identification result.

[0179] Optionally, the process of acquiring the security log data of the aggregation target terminal device category and the security log data of the terminal device category to be aggregated includes:

[0180] Obtain security log data from various terminal devices;

[0181] The security log data are clustered based on the categories of the terminal devices, entries of the security log data of each category of the terminal devices are determined, the terminal device category with the most entries of the security log data of the terminal devices is used as the aggregation target terminal device category, and the remaining terminal device categories are used as the terminal device categories to be aggregated.

[0182] Optionally, the process of establishing the conversion mapping relationship includes:

[0183] Numbering all entries of the security log data of each type of terminal device, extracting key log information corresponding to each entry one by one, and vectorizing the key log information;

[0184] Performing a cosine similarity operation on the vectorized log key information of the terminal device type to be aggregated and the vectorized log key information of the aggregation target terminal device of all entries one by one;

[0185] The vectorized log key information of the terminal device category to be aggregated with a cosine similarity higher than a preset value is distributedly verified with the vectorized log key information of the aggregation target terminal device to establish a conversion mapping relationship of the matching vector pair.

[0186] Optionally, performing a distribution check on the vectorized log key information of the terminal device category to be aggregated whose cosine similarity is higher than a preset value and the vectorized log key information of the aggregation target terminal device to establish a conversion mapping relationship of the matching vector pair includes:

[0187] A preset number N of samples are extracted from the security log data of the terminal device category to be aggregated and the security log data of the aggregation target terminal device. r The sample data is obtained by combining a numerical matrix based on the numerical list corresponding to the sample data, and performing numerical normalization processing on the two numerical matrices to obtain a matrix A and a matrix B;

[0188] The column vectors of the matrix A are combined with the column vectors of the matrix B in pairs to obtain column vector pairs. Any column vector pair is denoted as vector a = {a1, a2, ..., a m}, b={b1,b2,...,b n}, calculate the mean difference T mean , median difference T median , variance difference T variance ;

[0189] Combine vectors a and b into a vector c of length m+n, randomly shuffle the order of the elements of vector c, form a new vector a′ with the first m elements and a new vector b′ with the last n elements, and calculate the mean difference T′ mean , median difference T′ median , variance difference T′ variance ;

[0190] Repeat the random ordering of the elements of vector c, compose the first m elements into a new vector a′, and the last n elements into a new vector b′, and calculate the mean difference T′ mean , median difference T′ median , variance difference T′ variance , record the number of repetitions as N repeat =β*N r , β is the preset parameter, Among them, Count(T′ mean >T mean ) represents T′ in the process of repeatedly shuffling elements mean Greater than T mean The number of times; similarly, When p mean 、p median 、p variance When both are greater than 0.05, that is, p total =p mean +p median +p variance ; For all vector pairs consisting of matrix A and matrix B, select p total The largest vector pair is taken as the matching vector pair;

[0191] Based on the multiple relationship between the matching vector pairs and the numerical homogeneity of the sample data, a conversion mapping relationship of the matching vector pairs is established.

[0192] Optionally, the processing module includes:

[0193] The server and the terminal device are time synchronized, and the multiple terminal device security log aggregate data and the server security log data are divided into time periods according to the preset time window initial value to obtain the multiple terminal device security log aggregate data within the time period and the server security log data within the time period;

[0194] Collecting statistics on the security log aggregate data of various terminal devices in each time period and the security log data of the server in each time period one by one, and summarizing the statistical information of the security log aggregate data of various terminal devices in all time periods and the statistical information of the security log data of the server in all time periods, to obtain a first statistical information matrix and a second statistical information matrix respectively;

[0195] Determining the Spearman rank correlation coefficient based on the first statistical information matrix and the second statistical information matrix;

[0196] The time window value corresponding to the maximum Spearman rank correlation coefficient is used as the statistical window duration, and the first statistical information matrix and the second statistical information matrix corresponding to the statistical window duration are merged to obtain a network security log data set.

[0197] Optionally, the process of constructing the network threat identification and classification model includes:

[0198] Find the hyperplane ω*□(x)+b=0 in the high-dimensional feature space, where ω is the normal vector in the high-dimensional space, b is the bias term in the high-dimensional space, and □(x) is a nonlinear mapping function;

[0199] The constraints are

[0200] Get the value of the bias term b in the high-dimensional space through the support vector:

[0201]

[0202] Among them, SV is the index set of support vectors, n s is the number of support vectors, y i represents the classification label of the i-th sample, K(x i ,x k ) is the polynomial kernel function K(x i ,x k )=(x i ·x k +c) d ,x i ·x jrepresents the dot product of two sample vectors, c represents the constant offset of the polynomial kernel function, and d represents the polynomial order;

[0203] The decision function for classification prediction is:

[0204]

[0205] Among them, x represents the security log vector of the sample to be predicted, and sign() represents the indicator function;

[0206] The output of the indicator function sign() is the result of the classification prediction, that is, x represents the classification prediction corresponding to the security log vector of the sample to be predicted.

[0207] Optionally, the parameter optimization process of the network threat identification and classification model includes:

[0208] Initialize the model parameters of the firefly model, where the model parameters include the firefly population size, the maximum number of iterations, the value range of the regularization penalty parameter, the kernel function constant offset and the value range of the polynomial order, and generate the firefly population;

[0209] Calculating the fitness value of each firefly according to the position of all individuals in the firefly population;

[0210] For any i-th firefly, according to the current position X i Calculate the fitness value. The higher the fitness value, the brighter the firefly. Randomly select N fireflies in the population. sr individuals, where N sr =N train τ , τ∈(0,1), randomly select a firefly position X with a brightness higher than its own among these individuals j As the flight target, and update the flight target position:

[0211]

[0212] Where r represents X i , X j The distance between two locations, β max and β min are preset values ​​of the attraction coefficient, γ=1 is the absorption coefficient of the medium to light, and α is the step factor of the disturbance term N cur is the current iteration number, p is a preset integer, and rand(normal(0,1)) represents a random number that follows a standard normal distribution with a mean of 0 and a standard deviation of 1;

[0213] Calculate the flight target position X′ iThe fitness value and the current position X i The fitness of the flight target is compared. If the flight target position is X′ i The fitness value is higher than the current position X i The firefly will fly to the target position X′ if the fitness is i , otherwise the firefly will stay at the current position X i ;

[0214] For the firefly with the highest brightness in the population, there is no flight target position X′ i :

[0215] X′ Rbest =X Rbest +α*rand(normal(0,1))

[0216] Among them, X Rbest is the position of the firefly with the highest brightness in the population within one iteration, i.e., the position of the firefly with the largest fitness value, X′ Rbest The target position of the firefly with the highest brightness;

[0217] Repeatedly return to calculate the fitness value of each firefly according to the position of all individuals in the firefly population until the maximum number of iterations is reached, and use the searched optimal firefly position as the output, which is the optimized parameter of the network threat identification and classification model.

[0218] The functions of the device in the embodiment of the present invention have been described in the above method embodiment, so for details not provided in the description of this embodiment, please refer to the relevant description in the above embodiment, and no further description will be given here.

[0219] Based on the same inventive concept, an embodiment of this specification also provides an electronic device.

[0220] The following describes an electronic device embodiment of the present invention, which can be regarded as a specific physical implementation of the method and device embodiments of the present invention. The details described in the electronic device embodiment of the present invention should be regarded as a supplement to the above method or device embodiments; details not disclosed in the electronic device embodiment of the present invention can be implemented with reference to the above method or device embodiments.

[0221] Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this specification. Figure 3 The electronic device 300 according to the embodiment of the present invention is described. Figure 3 The electronic device 300 shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present invention.

[0222] like Figure 3As shown, the electronic device 300 is in the form of a general computing device. The components of the electronic device 300 may include, but are not limited to: at least one processing unit 310, at least one storage unit 320, a bus 330 connecting different system components (including the storage unit 320 and the processing unit 310), a display unit 340, etc.

[0223] The storage unit stores program codes, which can be executed by the processing unit 310, so that the processing unit 310 performs the steps according to various exemplary embodiments of the present invention described in the above processing method section of this specification. For example, the processing unit 310 can perform the following steps: Figure 1 Steps shown.

[0224] The storage unit 320 may include a readable medium in the form of a volatile storage unit, such as a random access memory unit (RAM) 3201 and / or a cache memory unit 3202 , and may further include a read-only memory unit (ROM) 3203 .

[0225] The storage unit 320 may also include a program / utility 3204 having a set (at least one) of program modules 3205, such program modules 3205 including but not limited to: an operating system, one or more application programs, other program modules and program data, each of which or some combination may include the implementation of a network environment.

[0226] Bus 330 may represent one or more of several types of bus structures, including a memory unit bus or memory unit controller, a peripheral bus, an accelerated graphics port, a processing unit, or a local bus using any of a variety of bus architectures.

[0227] The electronic device 300 may also communicate with one or more external devices 400 (e.g., keyboards, pointing devices, Bluetooth devices, etc.), one or more devices that enable viewers to interact with the electronic device 300, and / or any device that enables the electronic device 300 to communicate with one or more other computing devices (e.g., routers, modems, etc.). Such communication may be performed through an input / output (I / O) interface 350. Furthermore, the electronic device 300 may also communicate with one or more networks (e.g., a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through a network adapter 360. The network adapter 360 may communicate with other modules of the electronic device 300 through the bus 330. It should be understood that although Figure 3Not shown, other hardware and / or software modules may be used in conjunction with the electronic device 300, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.

[0228] Through the description of the above implementation methods, it is easy for those skilled in the art to understand that the exemplary embodiments described in the present invention can be implemented by software, or by software combined with necessary hardware. Therefore, the technical solution according to the implementation method of the present invention can be embodied in the form of a software product, which can be stored in a computer-readable storage medium (which can be a CD-ROM, USB flash drive, mobile hard disk, etc.) or on a network, including a number of instructions to enable a computing device (which can be a personal computer, server, or network device, etc.) to execute the above method according to the present invention. When the computer program is executed by a data processing device, the computer-readable medium can implement the above method of the present invention, that is: Figure 1 The method shown.

[0229] Figure 4 A schematic diagram of a computer-readable medium provided for an embodiment of this specification.

[0230] accomplish Figure 1 The computer program of the method shown can be stored on one or more computer readable media. The computer readable medium can be a readable signal medium or a readable storage medium. The readable storage medium can be, for example, but not limited to, a system, device or device of electricity, magnetism, light, electromagnetic, infrared, or semiconductor, or any combination of the above. More specific examples (non-exhaustive list) of readable storage media include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0231] The computer readable storage medium may include a data signal propagated in a baseband or as part of a carrier wave, wherein a readable program code is carried. This propagated data signal may take a variety of forms, including but not limited to an electromagnetic signal, an optical signal, or any suitable combination of the above. The readable storage medium may also be any readable medium other than a readable storage medium, which may send, propagate, or transmit a program for use by an instruction execution system, an apparatus, or a device or used in combination with it. The program code contained on the readable storage medium may be transmitted with any appropriate medium, including but not limited to wireless, wired, optical cable, RF, etc., or any suitable combination of the above.

[0232] Program code for performing the operations of the present invention may be written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Java, C++, etc., and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may be executed entirely on the viewer computing device, partially on the viewer device, as a stand-alone software package, partially on the viewer computing device and partially on a remote computing device, or entirely on a remote computing device or server. In cases involving a remote computing device, the remote computing device may be connected to the viewer computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (e.g., via the Internet using an Internet service provider).

[0233] In summary, the present invention can be implemented in hardware, or in a software module running on one or more processors, or in a combination thereof. It should be understood by those skilled in the art that general data processing devices such as microprocessors or digital signal processors (DSPs) can be used in practice to implement some or all of the functions of some or all of the components in the embodiments of the present invention. The present invention can also be implemented as a device or apparatus program (e.g., a computer program and a computer program product) for executing part or all of the methods described herein. Such a program implementing the present invention can be stored on a computer-readable medium, or can have the form of one or more signals. Such a signal can be downloaded from an Internet website, or provided on a carrier signal, or provided in any other form.

[0234] The specific embodiments described above further describe the purpose, technical solutions and beneficial effects of the present invention in detail. It should be understood that the present invention is not inherently related to any specific computer, virtual device or electronic device, and various general devices can also implement the present invention. The above description is only a specific embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the protection scope of the present invention.

[0235] The various embodiments in this specification are described in a progressive manner, and the same or similar parts between the various embodiments can be referenced to each other, and each embodiment focuses on the differences from other embodiments.

[0236] The above is only an embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application should be included in the scope of the claims of the present application.

Claims

1. A network security identification method based on log aggregation, characterized in that: include: Obtain security log data of multiple terminal devices, server security log data, and a preset time window initial value; wherein the security log data of the multiple terminal devices includes security log data of the aggregation target terminal device category and security log data of the terminal device category to be aggregated; Based on the conversion mapping relationship, the security log data of the terminal device category to be aggregated are converted one by one, and the security log data of the aggregation target terminal device category are aggregated to obtain security log aggregation data of multiple terminal devices; Processing the multiple terminal device security log aggregation data and the server security log data based on the preset time window initial value to obtain a network security log data set; The network security log data set is input into a network threat identification classification model to obtain a network threat identification result.

2. The network security identification method based on log aggregation according to claim 1, characterized in that: The process of obtaining the security log data of the aggregation target terminal device category and the security log data of the terminal device category to be aggregated includes: Obtain security log data from various terminal devices; The security log data are clustered based on the categories of the terminal devices, entries of the security log data of each category of the terminal devices are determined, the terminal device category with the most entries of the security log data of the terminal devices is used as the aggregation target terminal device category, and the remaining terminal device categories are used as the terminal device categories to be aggregated.

3. The network security identification method based on log aggregation as claimed in claim 2, characterized in that: The process of establishing the conversion mapping relationship includes: Numbering all entries of the security log data of each type of terminal device, extracting key log information corresponding to each entry one by one, and vectorizing the key log information; Performing a cosine similarity operation on the vectorized log key information of the terminal device type to be aggregated and the vectorized log key information of the aggregation target terminal device of all entries one by one; The vectorized log key information of the terminal device category to be aggregated with a cosine similarity higher than a preset value is distributedly verified with the vectorized log key information of the aggregation target terminal device to establish a conversion mapping relationship of the matching vector pair.

4. The network security identification method based on log aggregation as claimed in claim 3, characterized in that: The step of performing a distribution check on the vectorized log key information of the terminal device category to be aggregated and the vectorized log key information of the aggregated target terminal device, and establishing a conversion mapping relationship between matching vector pairs, includes: A preset number N of samples are extracted from the security log data of the terminal device category to be aggregated and the security log data of the aggregation target terminal device. r The sample data is obtained by combining a numerical matrix based on the numerical list corresponding to the sample data, and performing numerical normalization processing on the two numerical matrices to obtain a matrix A and a matrix B; The column vectors of the matrix A are combined with the column vectors of the matrix B in pairs to obtain column vector pairs. Any column vector pair is denoted as vector a = {a1, a2, ..., a m }, b={b1,b2,...,b n }, calculate the mean difference T mean , median difference T median , variance difference T variance ; Combine vectors a and b into a vector c of length m+n, randomly shuffle the order of the elements of vector c, form a new vector a′ with the first m elements and a new vector b′ with the last n elements, and calculate the mean difference T′ mean , median difference T′ median , variance difference T′ variance ; Repeat the random ordering of the elements of vector c, compose the first m elements into a new vector a′, and the last n elements into a new vector b′, and calculate the mean difference T′ mean , median difference T′ median , variance difference T′ variance , record the number of repetitions as N repeat =β*N r , β is the preset parameter, Among them, Count(T′ mean >T mean ) represents T′ in the process of repeatedly shuffling elements mean Greater than T mean The number of times; similarly, When p mean 、p median 、p variance When both are greater than 0.05, that is, p total =p mean +p median +p variance ; For all vector pairs consisting of matrix A and matrix B, select p total The largest vector pair is taken as the matching vector pair; Based on the multiple relationship between the matching vector pairs and the numerical homogeneity of the sample data, a conversion mapping relationship of the matching vector pairs is established.

5. The network security identification method based on log aggregation as claimed in claim 4, characterized in that: The processing of the plurality of terminal device security log aggregation data and the server security log data based on the preset time window initial value to obtain a network security log data set includes: The server and the terminal device are time synchronized, and the multiple terminal device security log aggregate data and the server security log data are divided into time periods according to the preset time window initial value to obtain the multiple terminal device security log aggregate data within the time period and the server security log data within the time period; Collecting statistics on the security log aggregate data of various terminal devices in each time period and the security log data of the server in each time period one by one, and summarizing the statistical information of the security log aggregate data of various terminal devices in all time periods and the statistical information of the security log data of the server in all time periods, to obtain a first statistical information matrix and a second statistical information matrix respectively; Determining the Spearman rank correlation coefficient based on the first statistical information matrix and the second statistical information matrix; The time window value corresponding to the maximum Spearman rank correlation coefficient is used as the statistical window duration, and the first statistical information matrix and the second statistical information matrix corresponding to the statistical window duration are merged to obtain a network security log data set.

6. The network security identification method based on log aggregation as claimed in claim 5, characterized in that: The process of constructing the network threat identification and classification model includes: Finding Hyperplanes in High-Dimensional Feature Spaces Among them, ω is the normal vector in the high-dimensional space, b is the bias term in the high-dimensional space, is a nonlinear mapping function; The constraints are Get the value of the bias term b in the high-dimensional space through the support vector: Among them, SV is the index set of support vectors, n s is the number of support vectors, y i represents the classification label of the i-th sample, K(x i ,x j ) is the polynomial kernel function K(x i ,x j )=(x i ·x j +c) d ,x i ·x j represents the dot product of two sample vectors, c represents the constant offset of the polynomial kernel function, and d represents the polynomial order; The decision function for classification prediction is: Among them, x represents the security log vector of the sample to be predicted, and sign() represents the indicator function; The output of the indicator function sign() is the result of the classification prediction, that is, x represents the classification prediction corresponding to the security log vector of the sample to be predicted.

7. The network security identification method based on log aggregation as claimed in claim 6, characterized in that: The parameter optimization process of the network threat identification and classification model includes: Initialize the model parameters of the firefly model, where the model parameters include the firefly population size, the maximum number of iterations, the value range of the regularization penalty parameter, the kernel function constant offset and the value range of the polynomial order, and generate the firefly population; Calculating the fitness value of each firefly according to the position of all individuals in the firefly population; For any i-th firefly, according to the current position X i Calculate the fitness value. The higher the fitness value, the brighter the firefly. Randomly select N fireflies in the population. sr individuals, where N sr =N train τ , τ∈(0,1), randomly select a firefly position X with a brightness higher than its own among these individuals j As the flight target, and update the flight target position: Where r represents X i , X j The distance between two locations, β max and β min are preset values ​​of the attraction coefficient, γ=1 is the absorption coefficient of the medium to light, and α is the step factor of the disturbance term N cur is the current iteration number, p is a preset integer, and rand(normal(0,1)) represents a random number that follows a standard normal distribution with a mean of 0 and a standard deviation of 1; Calculate the flight target position X′ i The fitness value and the current position X i The fitness of the flight target is compared. If the flight target position is X′ i The fitness value is higher than the current position X i The firefly will fly to the target position X′ if the fitness is i , otherwise the firefly will stay at the current position X i ; For the firefly with the highest brightness in the population, there is no flight target position X′ i : X′ Rbest =X Rbest +a*rand(normal(0,1)) Among them, X′ Rbest is the position of the firefly with the highest brightness in the population within one iteration, i.e., the position of the firefly with the largest fitness value, X Rbest The target position of the firefly with the highest brightness; Repeatedly return to calculate the fitness value of each firefly according to the position of all individuals in the firefly population until the maximum number of iterations is reached, and use the searched optimal firefly position as the output, which is the optimized parameter of the network threat identification and classification model.

8. A network security identification device based on log aggregation, characterized in that: include: An acquisition module, used to acquire security log data of multiple terminal devices, server security log data, and a preset time window initial value; wherein the security log data of the multiple terminal devices includes security log data of the aggregation target terminal device category and security log data of the terminal device category to be aggregated; An aggregation module, used to convert the security log data of the terminal device category to be aggregated one by one based on the conversion mapping relationship, and aggregate them with the security log data of the aggregation target terminal device category to obtain security log aggregation data of multiple terminal devices; A processing module, configured to process the plurality of terminal device security log aggregation data and the server security log data based on the preset time window initial value to obtain a network security log data set; The identification module is used to input the network security log data set into the network threat identification classification model to obtain a network threat identification result.

9. An electronic device, wherein: The electronic device includes: processor; and, A memory storing computer executable instructions which, when executed, cause the processor to perform a method according to any one of claims 1-7.

10. A computer-readable storage medium, wherein: The computer-readable storage medium stores one or more programs, and when the one or more programs are executed by a processor, the method of any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Network attack identification method and identification system based on traffic mode comparison

    CN108400995A

  • Network attack detection method and system based on deep learning

    CN116668089A