Network data packet association analysis method and device of unknown transmission protocol, electronic equipment and storage medium

The random forest model generates key field types and generates association rules, which solves the problems of field extraction and association analysis in packet analysis of unknown transmission protocols, and improves the analysis accuracy.

CN120067889APending Publication Date: 2025-05-30ELECTRIC POWER RES INST OF GUANGDONG POWER GRID CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510143472.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-10
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The prior art is difficult to effectively parse and analyze network packets of unknown transmission protocols, especially in the absence of transmission protocol type information, and the inability to accurately identify and extract valid fields, resulting in failure to fully reveal the potential laws and patterns between protocol fields.

Method used

Through the random forest model decision, the field type of the key fields to be extracted is generated, the key fields in the network data packet are extracted, the key fields matrix is ​​generated, the correlation coefficients and association rules between key fields are calculated, and the correlation between fields is determined.

Benefits of technology

The accuracy of network packet association analysis of unknown transmission protocols is improved, and the gap in the prior art that fails to fully reveal potential laws and patterns between protocol fields is filled.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120067889A_ABST
    Figure CN120067889A_ABST
Patent Text Reader

Abstract

The invention discloses an association analysis method and device for network data packets of an unknown transmission protocol, electronic equipment and a storage medium. The method comprises the following steps: acquiring a plurality of to-be-analyzed network data packets of the unknown transmission protocol and field types of to-be-extracted key fields; according to the field type of the to-be-extracted key field, extracting the key field in each to-be-analyzed network data packet, and generating a key field matrix; calculating correlation coefficients among the key fields according to the key field matrix, and generating a correlation coefficient matrix; according to the key field matrix, counting the occurrence frequency of each key field combination in the key field matrix, and generating an association rule between the key fields; and according to the correlation coefficient matrix and the association rule, determining the association between the fields of the network data packet to be analyzed. According to the invention, the accuracy of association analysis of the network data packet of the unknown transmission protocol can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of network transmission protocol analysis, and particularly to a method, apparatus, electronic device, and storage medium for correlative analysis of network data packets of an unknown transmission protocol. Background Art

[0002] With the diversification of network transmission protocols and the popularization of Internet of Things (IoT) devices, the types of transmission protocols used in the network are increasing continuously. The traditional transmission protocol analysis method based on a field library can no longer effectively meet the parsing requirements of unknown transmission protocols. Conducting correlative analysis on protocol fields to explore the potential rules and patterns between protocol fields is of great significance for the applications of transmission protocol parsing and traffic analysis. However, the existing technologies have not achieved effective functions in this aspect and cannot comprehensively reveal the complex correlations between protocol fields.

[0003] Currently, automatic field extraction and correlative analysis still face many challenges. Due to the diversity and flexibility of network transmission protocols, the field structures and encoding methods of data packets are often very different and lack a unified standard. Especially in the case of no transmission protocol type information, how to accurately identify and extract effective fields remains a difficult problem. The existing technologies usually ignore the correlative analysis between protocol fields and fail to fully explore the potential patterns and rules that may exist between these fields, thus limiting the accuracy of correlative analysis of network data packets of unknown transmission protocols. Summary of the Invention

[0004] Embodiments of the present invention provide a method, apparatus, electronic device, and storage medium for correlative analysis of network data packets of an unknown transmission protocol. By implementing the present invention, the accuracy of correlative analysis of network data packets of an unknown transmission protocol can be improved.

[0005] An embodiment of the present invention provides a method for correlative analysis of network data packets of an unknown transmission protocol, including:

[0006] Obtaining a plurality of network data packets to be analyzed of an unknown transmission protocol and the field types of keyword fields to be extracted; wherein, the field types of the keyword fields to be extracted are generated by decision-making of a random forest model;

[0007] Extracting the keyword fields in each network data packet to be analyzed according to the field types of the keyword fields to be extracted, and generating a keyword field matrix; wherein, each row of the keyword field matrix corresponds to a network data packet to be analyzed, and each column corresponds to a keyword field;

[0008] Calculating the correlation coefficients between the keyword fields according to the keyword field matrix, and generating a correlation coefficient matrix;

[0009] According to the keyword field matrix, count the occurrence frequency of each combination of keyword fields in the keyword field matrix, and generate the association rules between keyword fields;

[0010] According to the correlation coefficient matrix and the association rules, determine the correlation between the fields of the network data packet to be analyzed.

[0011] Further, generate the field types of the keyword fields to be extracted in the following way:

[0012] Obtain the training network data packets for training the random forest model and the corresponding protocol labels; wherein, the protocol labels include the transmission protocol information used by the training network data packets;

[0013] Extract the first protocol fields in the training network data packets; wherein, the first protocol fields refer to the fields that can be extracted from the training network data packets without parsing according to the transmission protocol information;

[0014] Generate a field matrix according to the first protocol fields; wherein, each row of the field matrix corresponds to a training network data packet, and each column corresponds to a first protocol field;

[0015] According to the field matrix and the protocol labels, calculate the reduction in the Gini index when each first protocol field is split at a node, and construct a random forest model for classifying the transmission protocols of network data packets according to the reduction in the Gini index;

[0016] Accumulate the reduction in the Gini index when each first protocol field is used as a split node in the random forest model, and calculate the importance of each first protocol field;

[0017] Take the first protocol fields with importance greater than the preset importance threshold as the field types of the keyword fields to be extracted.

[0018] Further, before extracting the first protocol fields in the training network data packets, it further includes:

[0019] Perform data denoising on the training network data packets to generate denoised training network data packets;

[0020] Perform time alignment on the denoised training network data packets to generate reorganized training network data packets;

[0021] Update the training network data packets according to the reorganized training network data packets.

[0022] Further, the calculating the correlation coefficients between keyword fields according to the keyword field matrix and generating a correlation coefficient matrix includes:

[0023] Calculate the average value of each keyword field according to the keyword field matrix;

[0024] Calculate the correlation coefficient between keyword fields using the following formula:

[0025]

[0026] where r XY is the correlation coefficient between keyword field X and keyword field Y; n is the number of rows in the keyword field matrix; is the average value of keyword field X; is the average value of keyword field Y;

[0027] Generate a correlation coefficient matrix based on the correlation coefficient between keyword fields.

[0028] Furthermore, based on the keyword field matrix, count the occurrence frequency of each keyword field combination in the keyword field matrix to generate association rules between keyword fields, including:

[0029] Based on the keyword field matrix, take each row as a keyword field combination to construct a transaction dataset;

[0030] Based on the transaction dataset, count the occurrence frequency of each keyword field combination in the keyword field matrix and calculate the support degree of each keyword field combination;

[0031] Add the keyword field combinations in the transaction dataset whose support degree is greater than the preset support degree threshold to the frequent item set;

[0032] Based on the frequent item set, determine the association relationship between keyword fields and calculate the confidence degree of each association relationship;

[0033] Determine the association rules between keyword fields for the association relationships whose confidence degree is greater than the preset confidence degree threshold.

[0034] Furthermore, after determining the relevance between each field of the network data packet to be analyzed based on the correlation coefficient matrix and the association rules, it further includes:

[0035] Accumulate the reduction amount of the Gini index when each keyword field is used as a splitting node in the random forest model, and calculate the weight of each keyword field;

[0036] Based on the keyword field matrix and the weight of the keyword field, calculate the weighted Euclidean distance between the network data packets to be analyzed and generate a distance matrix;

[0037] Based on the distance matrix, cluster the network data packets to be analyzed and generate several clusters of network data packets to be analyzed;

[0038] Based on the clusters of network data packets to be analyzed, the correlation coefficient matrix, and the association rules, perform transmission protocol parsing and traffic analysis.

[0039] Further, the weighted Euclidean distance between network data packets is calculated by the following formula:

[0040]

[0041] where d weighted (x i , x j ) is the weighted Euclidean distance between network data packet x i and network data packet x j ; n is the number of key features; w k is the weight of the k-th key feature; f ik is the k-th key feature of network data packet x i ; f jk is the k-th key feature of network data packet x j .

[0042] Based on the above method item embodiments, the present invention correspondingly provides apparatus item embodiments.

[0043] An embodiment of the present invention provides a network data packet association analysis apparatus for an unknown transmission protocol, including: a data acquisition module, a keyword field matrix generation module, a correlation coefficient matrix generation module, an association rule generation module, and a field correlation determination module;

[0044] The data acquisition module is configured to acquire a plurality of network data packets to be analyzed for an unknown transmission protocol and the field types of the keyword fields to be extracted from the network data packets to be analyzed; wherein, the field types of the keyword fields to be extracted are generated by decision-making of a random forest model;

[0045] The keyword field matrix generation module is configured to extract the keyword fields in each network data packet to be analyzed according to the field types of the keyword fields to be extracted, and generate a keyword field matrix; wherein, each row of the keyword field matrix corresponds to a network data packet to be analyzed, and each column corresponds to a keyword field;

[0046] The correlation coefficient matrix generation module is configured to calculate the correlation coefficients between the keyword fields according to the keyword field matrix, and generate a correlation coefficient matrix;

[0047] The association rule generation module is configured to count the occurrence frequencies of each keyword field combination in the keyword field matrix according to the keyword field matrix, and generate association rules between the keyword fields;

[0048] The field correlation determination module is configured to determine the correlation between each field of the network data packet to be analyzed according to the correlation coefficient matrix and the association rules.

[0049] Based on the above method item embodiments, the present invention correspondingly provides electronic device item embodiments.

[0050] An embodiment of the present invention provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it can implement the network packet association analysis method for an unknown transmission protocol described in any one of the above method item embodiments.

[0051] Based on the above method item embodiments, the present invention correspondingly provides storage medium item embodiments.

[0052] An embodiment of the present invention provides a storage medium, on which a computer program is stored. When the computer program is executed by a processor, it can implement the network packet association analysis method for an unknown transmission protocol described in any one of the above method item embodiments.

[0053] Compared with the prior art, the present invention has the following beneficial effects:

[0054] Embodiments of the present invention provide a network packet association analysis method, device, electronic device, and storage medium for an unknown transmission protocol. The method extracts key fields of packets of an unknown transmission protocol through a random forest model, and generates a key field matrix through the key fields. Then, the correlation coefficient between the key fields is calculated to generate a correlation coefficient matrix; the occurrence frequency of field combinations is counted to generate association rules.

[0055] The present invention uses a random forest model to determine the key fields in packets of an unknown transmission protocol, and generates a key field matrix according to the key fields, solving the problem in the prior art of how to accurately extract effective fields without information on the type of transmission protocol. By calculating the correlation coefficient between fields and generating association rules, the solution fills the gap in the prior art that fails to fully reveal the potential laws and patterns between protocol fields, thereby improving the accuracy of network packet association analysis for an unknown transmission protocol. Description of the Drawings

[0056] Figure 1 is a flowchart of a network packet association analysis method for an unknown transmission protocol provided by an embodiment of the present invention.

[0057] Figure 2 is a structural diagram of a network packet association analysis device for an unknown transmission protocol provided by an embodiment of the present invention. Detailed Embodiments

[0058] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0059] As Figure 1 shown, an embodiment of the present invention provides a method for associative analysis of network data packets of an unknown transmission protocol, which at least includes the following steps:

[0060] Step S1, obtain a plurality of network data packets to be analyzed of an unknown transmission protocol and the field types of fields to be extracted with keywords.

[0061] It should be explained here that the plurality of network data packets to be analyzed of the unknown transmission protocol can be captured based on the Wireshark capture tool. The field types of the fields to be extracted with keywords are generated by decision of a random forest model;

[0062] In a preferred embodiment, the field types of the fields to be extracted with keywords are generated in the following manner:

[0063] Obtain training network data packets for training the random forest model and corresponding protocol tags; wherein, the protocol tags contain the transmission protocol information used by the training network data packets;

[0064] Extract the first protocol fields in the training network data packets; wherein, the first protocol fields refer to the fields that can be extracted from the training network data packets without parsing according to the transmission protocol information;

[0065] Generate a field matrix according to the first protocol fields; wherein, each row of the field matrix corresponds to a training network data packet, and each column corresponds to a first protocol field;

[0066] Calculate the reduction in Gini index when each first protocol field is split at a node according to the field matrix and the protocol tags, and construct a random forest model for classification decision of the transmission protocol of network data packets according to the reduction in Gini index;

[0067] Accumulate the reduction in Gini index when each first protocol field is used as a split node in the random forest model, and calculate the importance of each first protocol field;

[0068] Use the first protocol fields with importance greater than a preset importance threshold as the field types of the fields to be extracted with keywords.

[0069] Specifically, the first protocol field may include, but is not limited to, the following types: packet length, packet size, time interval between packets, packet frequency, packet direction, packet rate, timestamp, source IP address, destination IP address, source port, destination port, and packet payload, etc. These fields cover the basic characteristics and transmission attributes of network packets, providing rich information dimensions and data support for protocol parsing and subsequent correlation analysis.

[0070] Optionally, before extracting the first protocol field from the training network packets, it further includes:

[0071] Performing data denoising on the training network packets to generate denoised training network packets;

[0072] Performing time alignment on the denoised training network packets to generate reorganized training network packets;

[0073] Updating the training network packets according to the reorganized training network packets.

[0074] Specifically, preprocessing the training network packets, first clearing possible outliers, redundant data or noise information in the network packets through data denoising methods to generate denoised training network packets; then performing time alignment operations on the denoised training network packets to ensure that the time series of the packets conforms to the actual transmission logic, generating reorganized training network packets after time adjustment; finally, replacing and updating the original training network packets according to the reorganized training network packets to form more accurate and standardized training data, providing a reliable basis for subsequent model construction and analysis.

[0075] Step S2: Extract the keyword fields in each network packet to be analyzed according to the field types of the keyword fields to be extracted, and generate a keyword field matrix.

[0076] Among them, each row of the keyword field matrix corresponds to a network packet to be analyzed, and each column corresponds to a keyword field;

[0077] Here, it needs to be explained that according to the field types of the keyword fields to be extracted, the keyword fields corresponding to the field types are extracted one by one from each network packet to be analyzed, and after summarization, a keyword field matrix is generated. Among them, the organization form of the keyword field matrix is: each row represents the keyword field data contained in a network packet to be analyzed, reflecting the feature combination of the packet; each column corresponds to a specific keyword field, representing the values of a specific field type in all packets. In this way, the keyword field matrix comprehensively integrates the feature information of each network packet, providing a structured input basis for subsequent data analysis and association rule mining.

[0078] Step S3: Calculate the correlation coefficients between the keyword fields according to the keyword field matrix, and generate a correlation coefficient matrix.

[0079] In a preferred embodiment, the calculating the correlation coefficients between the keyword fields according to the keyword field matrix and generating a correlation coefficient matrix includes:

[0080] Calculate the average value of each keyword field according to the keyword field matrix;

[0081] Calculate the correlation coefficients between the keyword fields through the following formula:

[0082]

[0083] where r XY is the correlation coefficient between keyword field X and keyword field Y; n is the number of rows of the keyword field matrix; is the average value of keyword field X; is the average value of keyword field Y; specifically, the correlation coefficient here is the Pearson correlation coefficient.

[0084] Generate a correlation coefficient matrix according to the correlation coefficients between the keyword fields.

[0085] Step S4: According to the keyword field matrix, count the occurrence frequencies of each keyword field combination in the keyword field matrix, and generate association rules between the keyword fields. Specifically, the FP-growth algorithm can be used here to generate the association rules between the keyword fields.

[0086] In a preferred embodiment, the counting the occurrence frequencies of each keyword field combination in the keyword field matrix and generating association rules between the keyword fields includes:

[0087] According to the keyword field matrix, take each row as a keyword field combination and construct a transaction dataset;

[0088] According to the transaction dataset, count the occurrence frequencies of each keyword field combination in the keyword field matrix and calculate the support degrees of each keyword field combination;

[0089] Add the keyword field combinations in the transaction dataset whose support degrees are greater than the preset support degree threshold to the frequent item set;

[0090] According to the frequent item set, determine the association relationships between the keyword fields and calculate the confidence degrees of each association relationship;

[0091] Determine the association rules between the keyword fields for the association relationships whose confidence degrees are greater than the preset confidence degree threshold.

[0092] Specifically, according to the keyword field matrix, each row in the matrix is regarded as a combination of keyword fields, and a transaction dataset is constructed based on this. Each combination of keyword fields is used as an item in the transaction dataset. By statistically analyzing the transaction dataset, the occurrence frequency of each combination of keyword fields in the keyword field matrix is calculated, and then the support degree of each combination of keyword fields is determined. The combinations of keyword fields with a support degree greater than the preset support degree threshold are screened and added to the frequent item set to form a representative set of combinations of keyword fields. Subsequently, based on the combinations of keyword fields in the frequent item set, the potential correlation relationships between each field are analyzed one by one, and the confidence degree of each correlation relationship is calculated. Finally, the correlation relationships with a confidence degree greater than the preset confidence degree threshold are determined as the association rules between keyword fields, identifying the significant patterns that may exist between keyword fields and providing a basis for subsequent network packet feature analysis and mining.

[0093] Optionally, the field values in each combination of keyword fields can also be discretized based on a preset threshold to meet the analysis requirements of different scenarios. For example, according to the preset threshold of the packet length, it can be divided into "long" or "short"; according to the preset threshold of the packet frequency, it can be divided into "high" or "low". Through this division method, not only can the processing of continuous data be simplified, but also the patterns of keyword fields can be made more intuitive, facilitating subsequent feature analysis, classification, or association rule mining. This process can flexibly set the threshold or classification criteria according to specific requirements to optimize the analysis effect.

[0094] Step S5: Determine the correlation between each field of the network packet to be analyzed according to the correlation coefficient matrix and the association rule.

[0095] Specifically, according to the correlation coefficient matrix, analyze the numerical correlation between each field in the network packet to be analyzed. For example, by the magnitude and sign of the correlation coefficient, judge the strength and direction of the linear relationship between fields; combined with the association rule, further explore the logical correlation between each field of the network packet to be analyzed. For example, through the support degree and confidence degree of the rule, clarify the influence pattern between fields in a specific combination. By comprehensively considering the numerical correlation and logical correlation, comprehensively determine the association characteristics between each field in the network packet to be analyzed.

[0096] In a preferred embodiment, after determining the correlation between each field of the network packet to be analyzed according to the correlation coefficient matrix and the association rule, it further includes:

[0097] Accumulate the reduction amount of the Gini index when each keyword field is used as a splitting node in the random forest model, and calculate the weight of each keyword field;

[0098] Calculate the weighted Euclidean distance between the network data packets to be analyzed according to the keyword field matrix and the weights of the keyword fields, and generate a distance matrix;

[0099] Cluster the network data packets to be analyzed according to the distance matrix, and generate several clusters of network data packets to be analyzed;

[0100] Perform transmission protocol parsing and traffic analysis according to the clusters of network data packets to be analyzed, the correlation coefficient matrix, and the association rules.

[0101] It should be noted here that by accumulating the reduction in the Gini index when each keyword field is used as a splitting node in the random forest model, the weight of each keyword field is calculated, so as to determine the importance of each field in protocol classification and provide a basis for subsequent analysis; according to the keyword field matrix and the weights of each keyword field, the weighted Euclidean distance is used to measure the similarity between the network data packets to be analyzed, and a distance matrix is generated to provide a numerical basis for the similarity analysis and subsequent processing of the data packets; based on the distance matrix, a clustering algorithm is used to group the network data packets to be analyzed, and several data packet clusters are generated, so as to identify groups of data packets with similar characteristics and provide structured data for further analysis; according to the clusters of network data packets to be analyzed, the correlation coefficient matrix, and the association rules, the correlation and similarity between each data packet are analyzed in depth, and by comparing the characteristics of different groups of data packets, potential protocol types and communication patterns are identified; combined with the correlation coefficient matrix, the correlation between fields is analyzed to reveal the interaction relationship between each keyword field in the data packet, and then the network transmission protocol is accurately inferred; through the association rules, the potential rules between the data packet characteristics are mined, so as to support the comprehensive analysis and optimization of network traffic and provide effective data support for tasks such as network security monitoring, traffic prediction, and anomaly detection.

[0102] For example, when performing transmission protocol parsing and traffic analysis of network data packets, first, the clustering algorithm is used to cluster the network data packets to be analyzed according to their keyword field characteristics, forming several data packet clusters. Suppose we cluster the data packets according to fields such as packet size, time interval, frequency, etc., and obtain two clusters: Cluster 1 contains most of the data packets, and Cluster 2 contains data packets during a few traffic peak periods. Next, by calculating the correlation coefficient matrix, the relationships between these fields are analyzed. For example, the relationship between packet size and time interval may reflect the pattern of traffic patterns. Then, using the association rule mining method, we can extract the association relationships of keyword field combinations from the frequent item sets. For example, when the packet size is greater than a certain value, the time interval is usually shorter, or between specific source IP and destination IP, the packet frequency is higher. Finally, by combining the clustering clusters, the correlation coefficient matrix, and the association rules, the behavioral characteristics of the data packets in each cluster are analyzed, thereby realizing the parsing of the transmission protocol. For example, by analyzing the data packets in these clusters, we can identify the traffic patterns of specific protocols, such as the common packet characteristics of the HTTP protocol, as well as potential abnormal traffic behaviors, which helps further traffic monitoring and network security analysis.

[0103] In one embodiment, the weighted Euclidean distance between network data packets is calculated by the following formula:

[0104]

[0105] where d weighted (x i , x j ) is the weighted Euclidean distance between network data packet x i and network data packet x j ; n is the number of key features; w k is the weight of the kth key feature; f ik is the kth key feature of network data packet x i ; f jk is the kth key feature of network data packet x j .

[0106] Based on the above method item embodiment, the present invention correspondingly provides an apparatus item embodiment.

[0107] As Figure 2 shown, an embodiment of the present invention provides a network data packet association analysis apparatus for an unknown transmission protocol, including: a data acquisition module, a keyword field matrix generation module, a correlation coefficient matrix generation module, an association rule generation module, and a field relevance determination module;

[0108] The data acquisition module is used to acquire a number of network data packets to be analyzed with unknown transmission protocols and the field types of the fields to be extracted from the network data packets to be analyzed; wherein, the field types of the fields to be extracted are generated by decision-making of a random forest model.

[0109] The keyword field matrix generation module is used to extract the keyword fields from each network data packet to be analyzed according to the field types of the fields to be extracted, and generate a keyword field matrix; wherein, each row of the keyword field matrix corresponds to a network data packet to be analyzed, and each column corresponds to a keyword field.

[0110] The correlation coefficient matrix generation module is used to calculate the correlation coefficients between keyword fields according to the keyword field matrix, and generate a correlation coefficient matrix.

[0111] The association rule generation module is used to count the occurrence frequencies of each keyword field combination in the keyword field matrix according to the keyword field matrix, and generate the association rules between keyword fields.

[0112] The field correlation determination module is used to determine the correlations between the fields of the network data packets to be analyzed according to the correlation coefficient matrix and the association rules.

[0113] It should be noted that the embodiments of the devices described above correspond to the above embodiments of the present invention, and can implement the network data packet association analysis method for unknown transmission protocols described in any one of the above of the present invention. In addition, the embodiments of the above devices are merely illustrative. The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. In addition, in the drawings of the device embodiments provided by the present invention, the connection relationships between modules indicate that they have communication connections, which can be specifically implemented as one or more communication buses or signal lines. Those of ordinary skill in the art can understand and implement without creative efforts.

[0114] Based on the above method embodiment of the present invention, an embodiment of an electronic device is correspondingly provided.

[0115] An embodiment of the present invention provides an electronic device, including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, it implements the network data packet association analysis method for unknown transmission protocols described in any one of the present invention, or when the processor executes the computer program, it implements the functions of each module in the above device embodiments.

[0116] Exemplarily, the computer program may be divided into one or more modules, and the one or more modules are stored in the memory and executed by the processor to implement the present invention. The one or more modules may be a series of computer program instruction segments capable of completing specific functions, and these instruction segments are used to describe the execution process of the computer program in the terminal device.

[0117] The terminal device may be a computing device such as a desktop computer, a notebook, a palm computer, and a cloud server. The terminal device may include, but is not limited to, a processor and a memory.

[0118] The so-called processor may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The processor is the control center of the terminal device, and connects various parts of the entire terminal device through various interfaces and lines.

[0119] The memory may be used to store the computer program and / or module. The processor realizes various functions of the terminal device by running or executing the computer program and / or module stored in the memory, and by calling the data stored in the memory. The memory may mainly include a program storage area and a data storage area. Among them, the program storage area may store an operating system, application programs required for at least one function, etc.; the data storage area may store data created according to the use of the mobile phone, etc. In addition, the memory may include high-speed random access memory, and may also include non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, at least one magnetic disk storage device, a flash memory device, or other volatile solid-state storage devices.

[0120] Based on the above method item embodiments, the present invention correspondingly provides storage medium item embodiments;

[0121] Another embodiment of the present invention provides a storage medium, which includes a stored computer program. When the computer program runs, it controls the device where the storage medium is located to execute the network data packet correlation analysis method of any one of the above unknown transmission protocols of the present invention.

[0122] Among them, the above storage medium is a computer-readable storage medium. The computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.

[0123] In the description of this specification, the description with reference to terms such as "one embodiment", "some embodiments", "example", "specific example" or "some examples" means that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. Moreover, the specific features, structures, materials or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.

[0124] The above is the preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art in this technical field, without departing from the principle of the present invention, several improvements and refinements can still be made, and these improvements and refinements are also regarded as the protection scope of the present invention.

Claims

1. A method for analyzing network data packets with unknown transmission protocols, characterized in that: include: Acquire a number of network data packets to be analyzed of an unknown transmission protocol and field types of key fields to be extracted; wherein the field types of the key fields to be extracted are generated by random forest model decision; According to the field type of the key field to be extracted, the key field in each network data packet to be analyzed is extracted to generate a key field matrix; wherein each row of the key field matrix corresponds to a network data packet to be analyzed, and each column corresponds to a key field; According to the key field matrix, calculate the correlation coefficients between the key fields and generate a correlation coefficient matrix; According to the key field matrix, the occurrence frequency of each key field combination in the key field matrix is ​​counted to generate association rules between the key fields; The correlation between the fields of the network data packet to be analyzed is determined according to the correlation coefficient matrix and the association rule.

2. The method for analyzing network data packets of unknown transmission protocol according to claim 1, characterized in that: Generate the field type of the key field to be extracted in the following way: Obtain a training network data packet and a corresponding protocol label for training a random forest model; wherein the protocol label includes transmission protocol information used by the training network data packet; Extracting a first protocol field from a training network data packet; wherein the first protocol field refers to a field that can be extracted from the training network data packet without being parsed according to the transmission protocol information; Generate a field matrix according to the first protocol field; wherein each row of the field matrix corresponds to a training network data packet, and each column corresponds to a first protocol field; According to the field matrix and the protocol label, the Gini index reduction of each first protocol field when the node is split is calculated, and a random forest model for network data packet transmission protocol classification decision is constructed according to the Gini index reduction; Accumulate the reduction of the Gini index of each first protocol field when it is used as a split node in the random forest model, and calculate the importance of each first protocol field; The first protocol field whose importance is greater than a preset importance threshold is used as the field type of the key field to be extracted.

3. The method for analyzing network data packets of unknown transmission protocol according to claim 2, characterized in that: Before extracting the first protocol field in the training network data packet, the method further includes: Perform data denoising on the training network data packets to generate denoised training network data packets; Time-align the denoised training network data packets to generate reconstructed training network data packets; The training network data packets are updated according to the reorganized training network data packets.

4. The method for analyzing network data packets of unknown transmission protocol according to claim 3, characterized in that: The step of calculating the correlation coefficients between the key fields according to the key field matrix and generating the correlation coefficient matrix includes: According to the key field matrix, calculate the average value of each key field; The correlation coefficient between key fields is calculated using the following formula: Among them, r XY is the correlation coefficient between key field X and key field Y; n is the number of rows in the key field matrix; is the average value of the key field X; is the average value of the key field Y; Generate a correlation coefficient matrix based on the correlation coefficients between key fields.

5. The method for analyzing network data packets of unknown transmission protocol according to claim 4, characterized in that: The method of counting the occurrence frequency of each key field combination in the key field matrix and generating association rules between the key fields includes: According to the key field matrix, each row is taken as a key field combination to construct the transaction data set; According to the transaction data set, the occurrence frequency of each key field combination in the key field matrix is ​​counted, and the support of each key field combination is calculated; The key field combinations whose support in the transaction data set is greater than the preset support threshold are added to the frequent itemsets; According to the frequent item sets, determine the association between key fields and calculate the confidence of each association; The association relationship with a confidence level greater than a preset confidence level threshold is determined as an association rule between key fields.

6. The method for analyzing network data packets of unknown transmission protocol according to claim 5, characterized in that: After determining the correlation between the fields of the network data packet to be analyzed according to the correlation coefficient matrix and the association rule, the method further includes: The reduction of the Gini index of each key field when it is used as a split node in the random forest model is accumulated, and the weight of each key field is calculated; According to the key field matrix and the weights of the key fields, the weighted Euclidean distance between the network data packets to be analyzed is calculated to generate a distance matrix; According to the distance matrix, the network data packets to be analyzed are clustered to generate several clusters of network data packets to be analyzed; According to the network data packet clusters to be analyzed, the correlation coefficient matrix and the association rules, transmission protocol parsing and traffic analysis are performed.

7. The method for analyzing network data packets of unknown transmission protocol according to claim 6, characterized in that: The weighted Euclidean distance between network packets is calculated using the following formula: Among them, d weighted (x i , x j ) is the network data packet x i and network packet x j The weighted Euclidean distance between them; n is the number of key features; w k is the weight of the kth key feature; f ik For network data packet x i The kth key feature; f jk For network data packet x j The kth key feature.

8. A network data packet correlation analysis device for an unknown transmission protocol, characterized in that: include: Data acquisition module, key field matrix generation module, correlation coefficient matrix generation module, association rule generation module and field correlation determination module; The data acquisition module is used to acquire a number of network data packets to be analyzed of an unknown transmission protocol and field types of key fields to be extracted from the network data packets to be analyzed; wherein the field types of the key fields to be extracted are generated by random forest model decision; The key field matrix generation module is used to extract the key fields in each network data packet to be analyzed according to the field type of the key field to be extracted, and generate a key field matrix; wherein each row of the key field matrix corresponds to a network data packet to be analyzed, and each column corresponds to a key field; The correlation coefficient matrix generation module is used to calculate the correlation coefficients between the key fields according to the key field matrix and generate the correlation coefficient matrix; The association rule generation module is used to count the occurrence frequency of each key field combination in the key field matrix according to the key field matrix, and generate association rules between the key fields; The field correlation determination module is used to determine the correlation between the fields of the network data packet to be analyzed according to the correlation coefficient matrix and the association rule.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, it can implement the network data packet correlation analysis method of the unknown transmission protocol as described in any one of claims 1 to 7.

10. A storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, it can implement the network data packet correlation analysis method of the unknown transmission protocol described in any one of claims 1 to 7.