Network transmission data detection method and device, computer device and storage medium

By combining semantic representation models and frequency domain feature analysis, the problem of low detection efficiency of encrypted network transmission data is solved, and efficient and accurate network transmission data detection is achieved.

CN116865996BActive Publication Date: 2026-08-04PENG CHENG LAB +1
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
PENG CHENG LAB
Filing Date
2023-06-01
Publication Date
2026-08-04

AI Technical Summary

Technical Problem

In existing technologies, the encryption of network transmission data by SSL and TLS protocols increases the difficulty of detecting abnormal network transmission data and reduces the efficiency of network transmission data detection.

Method used

A semantic representation model is used to extract time-domain feature data of network transmission data. By combining time-frequency conversion and frequency-domain feature analysis with statistical features, a representation image is generated. The detection is performed by integrating time-domain, first frequency domain and second frequency domain feature data, so as to detect network transmission data without decryption.

Benefits of technology

It improves the efficiency and accuracy of network transmission data detection, and can effectively identify abnormal network transmission data without decryption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116865996B_ABST
    Figure CN116865996B_ABST
Patent Text Reader

Abstract

The embodiment of the application provides a kind of network transmission data detection method and device, computer equipment and storage medium, belong to network security technical field.The method comprises: obtaining the network transmission data to be detected;Based on semantic representation model, feature extraction is carried out to network transmission data, and time domain feature data is obtained;Network transmission data is converted into time sequence data, and time-frequency conversion is carried out to time sequence data, and first frequency domain feature data is obtained;The statistical characteristics of network transmission data are extracted, and representation image is generated based on statistical characteristics;Frequency domain conversion is carried out to representation image, and second frequency domain feature data is obtained;Based on time domain feature data, first frequency domain feature data and second frequency domain feature data, data detection is carried out, and the detection result of the network transmission data to be detected is obtained.The embodiment of the application can improve the detection efficiency and detection accuracy of detecting network transmission data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of network security technology, and in particular to a method, apparatus, computer equipment, and storage medium for detecting network transmission data. Background Technology

[0002] Currently, with the continuous development of internet technology, people's production and daily life are becoming increasingly intertwined with the internet, making cybersecurity issues increasingly serious. Among these, security testing of transmitted data is a crucial means of maintaining network data security.

[0003] Furthermore, with the emergence and continuous development of encryption technologies (such as the Secure Socket Layer (SSL) protocol and the Transport Layer Security (TLS) protocol), plaintext data in network transmission has gradually decreased. This has largely protected the data privacy and integrity of network users, and has also improved the security of network data to a certain extent.

[0004] However, while SSL and TLS protocols encrypt network transmission data using public-key or symmetric-key technologies to prevent the leakage of plaintext data, they also encrypt abnormal network transmission data, which increases the difficulty of detecting abnormal network transmission data and thus reduces the efficiency of network transmission data detection. Summary of the Invention

[0005] The main objective of this application is to provide a method, apparatus, computer device, and storage medium for detecting network transmission data, aiming to improve the detection efficiency of network transmission data.

[0006] To achieve the above objectives, a first aspect of this application proposes a method for detecting network transmission data, the method comprising:

[0007] Acquire the network transmission data to be detected;

[0008] Based on the semantic representation model, feature extraction is performed on the network transmission data to obtain temporal feature data;

[0009] The network transmission data is converted into time-series data, and the time-series data is subjected to time-frequency conversion to obtain the first frequency domain feature data;

[0010] Extract statistical features from the network transmission data, and generate a characterization image based on the statistical features;

[0011] The representation image is transformed in the frequency domain to obtain the second frequency domain feature data;

[0012] Data detection is performed based on the time-domain feature data, the first frequency-domain feature data, and the second frequency-domain feature data to obtain the detection result of the network transmission data to be detected.

[0013] In some embodiments, converting the network transmission data into time-series data and performing time-frequency conversion on the time-series data to obtain first frequency domain feature data includes:

[0014] Convert the network transmission data into time-series data;

[0015] Wavelet analysis was performed on the time series to obtain high-frequency and low-frequency features;

[0016] The high-frequency features are concatenated with the low-frequency features to obtain the first frequency domain feature data.

[0017] In some embodiments, converting the network transmission data into time-series data includes:

[0018] The data sequence corresponding to the network transmission data is cut according to a preset step size to obtain multiple data frames of the same length;

[0019] Time sequence data is generated based on the frame sequence composed of multiple data frames of the same length.

[0020] In some embodiments, extracting statistical features of the network transmission data and generating a characterization image based on the statistical features includes:

[0021] Feature extraction is performed on the statistical data of the transport layer in the network transmission data to obtain statistical features;

[0022] The statistical features are mapped to a preset numerical range to obtain a set of values;

[0023] The numerical values ​​in the set of values ​​are used as pixel values ​​to construct a representation image.

[0024] In some embodiments, constructing a representation image using values ​​from the set of values ​​as pixel values ​​includes:

[0025] The values ​​in the numerical set are expanded to obtain the target numerical set;

[0026] The target numerical values ​​are used as pixel values ​​to construct a representation image.

[0027] In some embodiments, the step of data expansion of the values ​​in the value set to obtain the target value set includes:

[0028] Calculate the mean and variance of multiple values ​​in the dataset;

[0029] The mean and variance are added to the set of values ​​to obtain the target set of values.

[0030] In some embodiments, performing frequency domain transformation on the representation image to obtain second frequency domain feature data includes:

[0031] Obtain orthogonal wavelet basis functions;

[0032] Multiple image frequency domain features representing the image are generated based on the orthogonal wavelet basis functions;

[0033] The multiple image frequency domain features are concatenated to obtain the second frequency domain feature data.

[0034] In some embodiments, the network transmission data includes multiple network data packets, and the feature extraction of the network transmission data based on a semantic representation model to obtain temporal feature data includes:

[0035] Based on the semantic representation model, feature extraction is performed on the multiple network data packets to obtain multiple sub-time domain feature data;

[0036] The multiple sub-time domain feature data are concatenated to obtain time domain feature data.

[0037] In some embodiments, after extracting features from the plurality of network data packets based on a semantic representation model to obtain multiple sub-time domain feature data, the method further includes:

[0038] Obtain a target number of sub-time domain features from the plurality of sub-time domain feature data;

[0039] The target number of sub-temporal features are combined to obtain a feature matrix;

[0040] The third frequency domain feature data of the feature matrix is ​​generated based on orthogonal wavelet basis functions;

[0041] The step of performing data detection based on the time-domain feature data, the first frequency-domain feature data, and the second frequency-domain feature data to obtain the detection result of the network transmission data to be detected includes:

[0042] Data detection is performed based on the time-domain feature data, the first frequency-domain feature data, the second frequency-domain feature data, and the third frequency-domain feature data to obtain the detection result of the network transmission data to be detected.

[0043] In some embodiments, the step of performing data detection based on the time-domain feature data, the first frequency-domain feature data, the second frequency-domain feature data, and the third frequency-domain feature data to obtain the detection result of the network transmission data to be detected includes:

[0044] Obtain the weight coefficients corresponding to the second frequency domain feature data and the third frequency domain feature data respectively;

[0045] Based on the obtained corresponding weight coefficients, the second frequency domain feature data and the third frequency domain feature data are fused to obtain fused frequency domain features;

[0046] The time-domain feature data, the first frequency-domain feature data, and the fused frequency-domain feature data are concatenated to obtain the target feature.

[0047] Data detection is performed on the target features to obtain the detection results of the network transmission data to be detected.

[0048] In some embodiments, the semantic representation model includes an encoder and a decoder, and the semantic representation model is trained using the following method:

[0049] Obtain a first training sample set, which includes multiple first sample data packets;

[0050] The first sample data packet is vector-transformed to obtain the sample feature vector;

[0051] The sample feature vector is input into the encoder for feature extraction to obtain the latent vector;

[0052] The latent vector is input into the decoder for decoding to obtain the output vector;

[0053] The parameters of the encoder and the parameters of the decoder are adjusted based on the difference between the output vector and the sample feature vector.

[0054] The following steps are repeated until the semantic representation model is detected to have converged:

[0055] The first sample data packet is reacquired from the first training sample set, and a vector transformation is performed on the first sample data packet to obtain a sample feature vector. The sample feature vector is input into the encoder for feature extraction to obtain a latent vector, and the latent vector is input into the decoder for decoding to obtain an output vector. The parameters of the encoder and the parameters of the decoder are adjusted based on the difference between the output vector and the sample feature vector, and the convergence detection is performed on the semantic representation model.

[0056] In some embodiments, adjusting the parameters of the encoder and the decoder based on the difference between the output vector and the first sample feature vector includes:

[0057] Calculate the paradigm distance between the sample feature vector and the output vector to obtain the self-supervised learning loss;

[0058] The parameters of the encoder and the parameters of the decoder are adjusted based on the self-supervised learning loss.

[0059] In some embodiments, the step of performing data detection based on the time-domain feature data, the first frequency-domain feature data, and the second frequency-domain feature data to obtain the detection result of the network transmission data to be detected includes:

[0060] The fused feature obtained by fusing the time-domain feature data, the first frequency-domain feature data, and the second frequency-domain feature data is input into the network transmission data detection model for detection, and the network transmission data to be detected is obtained.

[0061] The network transmission data detection model is trained using the following method:

[0062] Obtain a second training sample set, which includes multiple sets of sample network transmission data and label data corresponding to each sample network transmission data;

[0063] Based on the semantic representation model, feature extraction is performed on each sample of network transmission data to obtain sample temporal feature data corresponding to each sample of network transmission data.

[0064] Each sample network transmission data is converted into sample time-series data, and the sample time-series data is converted into time-frequency data to obtain the first sample frequency domain feature data corresponding to each sample network transmission data.

[0065] Extract the sample statistical features of each sample network transmission data, generate a sample representation image based on the sample statistical features, and perform frequency domain transformation on the sample representation image to obtain the second sample frequency domain feature data.

[0066] The time-domain feature data, the first sample frequency-domain feature data, and the second sample frequency-domain feature data corresponding to each sample of network transmission data are fused to obtain the fused feature corresponding to each sample of network transmission data.

[0067] The network transmission data detection model is obtained by using the fusion features corresponding to each sample of network transmission data as the model input and the label data corresponding to each sample of network transmission data as the model training label training preset neural network model.

[0068] In some embodiments, after acquiring the network transmission data to be detected, the method further includes:

[0069] The network transmission data is cleaned to obtain the target network transmission data;

[0070] The feature extraction of the network transmission data based on the semantic representation model to obtain temporal feature data includes:

[0071] Based on the semantic representation model, feature extraction is performed on the target network transmission data to obtain temporal feature data;

[0072] The step of converting the network transmission data into time-series data includes:

[0073] Convert the target network transmission data into time-series data;

[0074] The extraction of statistical features from the network transmission data includes:

[0075] Extract the statistical characteristics of the target network transmission data.

[0076] To achieve the above objectives, a second aspect of this application provides a network transmission data detection device, the device comprising:

[0077] The acquisition module is used to acquire the network transmission data to be detected.

[0078] The extraction module is used to extract features from the network transmission data based on a semantic representation model to obtain temporal feature data;

[0079] The first conversion module is used to convert the network transmission data into time-series data, and to perform time-frequency conversion on the time-series data to obtain first frequency domain feature data.

[0080] A generation module is used to extract statistical features of the network transmission data and generate a representation image based on the statistical features;

[0081] The second conversion module is used to perform frequency domain conversion on the representation image to obtain second frequency domain feature data;

[0082] The detection module is used to perform data detection based on the time-domain feature data, the first frequency-domain feature data, and the second frequency-domain feature data to obtain the detection result of the network transmission data to be detected.

[0083] To achieve the above objectives, a third aspect of the present application provides a computer device, the computer device including a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the method described in the first aspect.

[0084] To achieve the above objectives, a fourth aspect of the present application provides a storage medium storing a computer program that, when executed by a processor, implements the method described in the first aspect.

[0085] The network transmission data detection method, apparatus, computer equipment, and storage medium proposed in this application extract time-domain feature data of network transmission data through a semantic representation model, convert the network transmission data into time-series data, and extract the first frequency-domain feature data of the network transmission data in the time-series dimension through time-frequency transformation of the time-series data. Furthermore, it extracts statistical features of the network transmission data to generate a representation image, and performs frequency-domain transformation on the representation image to extract the second frequency-domain feature data in the semantic dimension of the network transmission data. Finally, it integrates the time-domain feature data, the first frequency-domain feature data, and the second frequency-domain feature data of the network transmission data to perform data anomaly detection. This allows for the detection of network transmission data through multi-dimensional indicators at the traffic level, achieving detection without decrypting the network transmission data, thereby improving the efficiency of network transmission data detection. Attached Figure Description

[0086] Figure 1 This is a flowchart of a network transmission data detection method provided in an embodiment of this application;

[0087] Figure 2 This is a flowchart of training a semantic representation model;

[0088] Figure 3 yes Figure 2 The flowchart of step S205 in the document;

[0089] Figure 4 yes Figure 1 The flowchart of step S102 in the document;

[0090] Figure 5 yes Figure 1 The flowchart of step S103 in the process;

[0091] Figure 6 yes Figure 5 The flowchart of step S501 in the text;

[0092] Figure 7 yes Figure 1 The flowchart of step S104 in the process;

[0093] Figure 8 yes Figure 1 The flowchart of step S105 in the process;

[0094] Figure 9This is a flowchart of training a network transmission data detection model;

[0095] Figure 10 This is a schematic diagram of the network transmission data detection model;

[0096] Figure 11 This is a schematic diagram of the network transmission data detection device provided in the embodiments of this application;

[0097] Figure 12 This is a schematic diagram of the hardware structure of the computer device provided in the embodiments of this application. Detailed Implementation

[0098] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0099] It should be noted that although functional modules are divided in the device schematic diagram and a logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than the module division in the device or the order in the flowchart. The terms "first," "second," etc., in the specification, claims, and the aforementioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence.

[0100] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.

[0101] First, let's analyze some of the terms used in this application:

[0102] Artificial intelligence (AI) is a new branch of computer science that studies, develops, and applies theories, methods, technologies, and systems to simulate, extend, and expand human intelligence. It aims to understand the essence of intelligence and produce intelligent machines that can react in a way similar to human intelligence. Research in this field includes robotics, speech recognition, image recognition, natural language processing, and expert systems. AI can simulate the information processes of human consciousness and thought. Furthermore, AI utilizes digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceiving the environment, acquiring knowledge, and using that knowledge to achieve optimal results.

[0103] Network traffic: refers to the amount of data transmitted over a network.

[0104] In related technologies, with the emergence and continuous development of encrypted data, the amount of plaintext content transmitted over the network has gradually decreased, which largely protects the data privacy and integrity of network users. However, some abnormal network transmission data (such as malicious data) can bypass abnormal network transmission data detection modules during transmission due to the use of encryption technology, thereby launching attacks on network users. Decrypting network transmission data to its original state before detection requires enormous computational costs; therefore, how to achieve anomaly detection in network transmission data without decryption is a pressing problem. Currently, commonly used methods for anomaly detection in network transmission data without decryption generally include traditional machine learning methods and deep learning-based methods. These two methods manually extract features such as data packet fields, packet length, and transmission time from network transmission data, and then combine these features with expert knowledge before inputting them into a pre-built machine learning model or deep neural network model for learning. However, these methods extract features from network transmission data in a relatively single dimension, leading to low accuracy in detecting network transmission data based on a single dimension.

[0105] Based on this, embodiments of this application provide a method and apparatus for detecting network transmission data, a computer device and a storage medium, aiming to improve the accuracy of network transmission data detection.

[0106] The network transmission data detection method, apparatus, computer equipment, and storage medium provided in this application are specifically described through the following embodiments. First, the network transmission data detection method in this application is described.

[0107] The embodiments of this application can acquire and process relevant data based on artificial intelligence technology. Artificial intelligence (AI) refers to the theories, methods, technologies, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results.

[0108] Foundational artificial intelligence technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, detection technologies for large-scale network data transmission, operating / interactive systems, and mechatronics. AI software technologies mainly encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning.

[0109] The network transmission data detection method provided in this application relates to the field of network security technology. This method can be applied to a terminal, a server, or software running on either a terminal or a server. In some embodiments, the terminal can be a smartphone, tablet, laptop, desktop computer, etc.; the server can be configured as an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms; the software can be an application implementing the network transmission data detection method, but is not limited to the above forms.

[0110] This application can be used in a wide variety of general-purpose or special-purpose computer system environments or configurations. Examples include: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer computer devices, network PCs, minicomputers, mainframe computers, and distributed computing environments including any of the above systems or devices. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform specific tasks or implement specific abstract data types. This application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.

[0111] It should be noted that in all specific embodiments of this application, when processing data related to user identity or characteristics, such as user information, user behavior data, user historical data, and user location information, user permission or consent is obtained first. Furthermore, the collection, use, and processing of this data comply with relevant laws, regulations, and standards. In addition, when embodiments of this application require access to sensitive personal information of users, separate permission or consent from the user is obtained through pop-ups or redirection to confirmation pages. Only after obtaining the user's separate permission or consent is the necessary user-related data required for the proper functioning of these embodiments acquired.

[0112] Figure 1 This is an optional flowchart of the network transmission data detection method provided in the embodiments of this application. Figure 1 The method may include, but is not limited to, steps S101 to S106.

[0113] Step S101: Obtain the network transmission data to be detected.

[0114] The network transmission data to be detected can be network transmission data directly obtained from the network interface card (NIC) node, or network transmission data acquired and saved in advance for unified detection. The network transmission data to be detected can include one network data packet or multiple network data packets. When the network transmission data to be detected includes multiple network data packets, these multiple network data packets can form a network data packet sequence according to the time order in which they were acquired. In this embodiment, the acquired network transmission data to be detected can specifically be network transmission data encrypted using encryption technology. In other embodiments, the network transmission data to be detected can also be network transmission data that is not encrypted using encryption technology.

[0115] The acquired network transmission data can be in pcap format. pcap is short for Packet Capture, an industry-standard network packet capture format. Network developers typically use various network analyzers to capture TCP / IP packets, and the captured packets are saved in pcap format.

[0116] Step S102: Based on the semantic representation model, feature extraction is performed on the network transmission data to obtain temporal feature data.

[0117] In this embodiment, a semantic representation model corresponding to the network transmission data can be pre-constructed. The pre-constructed semantic representation model can extract semantic features from the network transmission data. Since the acquired network transmission data is data that changes over time, the semantic features of the network transmission data can change over time. Therefore, the semantic features of the network transmission data are also referred to as the temporal feature data of the network transmission data.

[0118] Step S103: Convert the network transmission data into time-series data, and perform time-frequency conversion on the time-series data to obtain the first frequency domain feature data.

[0119] While temporal feature data of network transmission data is extracted as a characteristic of how network transmission data changes over time, and regular network transmission data exhibits certain regularities over time, it can, to some extent, detect anomalies in network transmission data. However, in some cases, abnormal network transmission data can mimic the time-series transmission patterns of normal network transmission data, thus evading detection by network transmission data detection modules and launching attacks against network users. Therefore, relying solely on temporal feature data for anomaly detection can lead to inaccurate results.

[0120] In this embodiment, the acquired network transmission data can be further converted into time-series data. The time-series data can be one-dimensional numerical data that varies over time, such as converting the network transmission data into a curve that varies over time. Then, time-frequency conversion can be performed on the time-series data to convert the characteristics of the network transmission data in the time dimension into characteristics in the frequency dimension, thereby obtaining the frequency domain feature data of the network transmission data. To distinguish it from the frequency domain feature data mentioned below, it can be referred to here as the first frequency domain feature data.

[0121] Step S104: Extract statistical features of network transmission data and generate a representation image based on the statistical features.

[0122] Furthermore, in this embodiment, the acquired network transmission data can also be visualized at the pixel level. Specifically, statistical features of the network transmission data can be extracted, and then a representation image can be generated based on these statistical features. These statistical features can specifically be statistical information about the network transmission data at the transport layer, such as "total number of forward packets," "total number of reverse packets," "total size of forward data packets," and "maximum size of forward data packets." Generally, the statistical features of network transmission data have approximately 80 dimensions, which are not listed here. In this embodiment, a specific feature extraction tool can be used to extract the statistical features of the network transmission data, thereby outputting statistical features represented by numerical data. These numerical data can then be mapped to pixel space to obtain multiple pixel values. Thus, a corresponding representation image can be constructed based on the obtained multiple pixel values.

[0123] Step S105: Perform frequency domain transformation on the representation image to obtain the second frequency domain feature data.

[0124] Furthermore, after generating a representational image based on the statistical characteristics of the network transmission data, the generated representational image can be further transformed in the frequency domain to obtain second frequency domain feature data. Specifically, the frequency domain transformation of the image involves a mathematical transformation from the spatial domain to the frequency domain. Utilizing the properties of orthogonal transformations, complex calculations in the spatial domain can be simplified after conversion to the frequency domain; and the analysis of frequency domain characteristics will also facilitate the acquisition of various image properties and the performance of special processing.

[0125] Step S106: Data detection is performed based on time-domain feature data, first frequency-domain feature data, and second frequency-domain feature data to obtain the detection result of the network transmission data to be detected.

[0126] After acquiring the time-domain, first-frequency-domain, and second-frequency-domain feature data of the network transmission data to be detected, a fusion detection can be performed based on these data to obtain the detection result. The detection result can indicate that the network transmission data is normal or abnormal. When the detection result determines that the network transmission data is abnormal, an anomaly alert can be issued, and related data can be investigated to prevent the abnormal network transmission data from impacting network user data security.

[0127] Steps S101 to S106 of this embodiment involve: acquiring network transmission data to be detected; extracting features from the network transmission data based on a semantic representation model to obtain time-domain feature data; converting the network transmission data into time-series data and performing time-frequency conversion on the time-series data to obtain first frequency-domain feature data; extracting statistical features from the network transmission data and generating a representation image based on the statistical features; performing frequency-domain conversion on the representation image to obtain second frequency-domain feature data; and performing data detection based on the time-domain feature data, the first frequency-domain feature data, and the second frequency-domain feature data to obtain the detection result of the network transmission data to be detected. The network transmission data processing method provided in this embodiment only requires extracting time-domain features, the first frequency-domain features, and the second frequency-domain features from the network transmission data to be detected, and performing detection based on the extracted time-domain features, the first frequency-domain features, and the second frequency-domain features. This eliminates the need to decrypt the network transmission data to be detected, significantly improving the efficiency of network transmission data detection. Furthermore, this method not only extracts the temporal features of the network transmission data to be detected, but also extracts the frequency domain features of the network transmission data to be detected from both temporal and semantic dimensions, thereby extracting multi-level joint feature data of the network transmission data to be detected. Detecting network transmission data based on this multi-level joint feature data can significantly improve the accuracy of network transmission data detection.

[0128] In some embodiments, after step 101, the following may also be included:

[0129] Data cleaning is performed on the network transmission data to obtain the target network transmission data;

[0130] Feature extraction of network transmission data based on semantic representation models yields temporal feature data, including:

[0131] Based on the semantic representation model, feature extraction is performed on the target network transmission data to obtain temporal feature data;

[0132] Converting network-transmitted data into time-series data includes:

[0133] Convert the target network transmission data into time-series data;

[0134] Extract statistical features from network transmission data, including:

[0135] Extract statistical characteristics of data transmitted over the target network.

[0136] In this embodiment, after acquiring the network transmission data to be detected, the acquired network transmission data can first be cleaned to obtain the target network transmission data. Then, time-domain features, first frequency-domain features, and second frequency-domain features are extracted based on the target network transmission data.

[0137] The specific process of cleaning the network transmission data to be detected can involve first performing operations such as deleting the Ethernet header, hiding the IP address, aligning the transport layer packet header, and aligning the data packets on the pcap file corresponding to the acquired network transmission data. Then, the data obtained after the above processing steps undergoes numerical conversion. Since the data in the pcap data packets is presented in binary sequence form, the numerical conversion specifically involves mapping these binary sequences to the set [0, 255] to obtain the target network transmission data. The target network transmission data can specifically be a vector x0.

[0138] In this embodiment, after obtaining the network transmission data to be detected, the network transmission data can be cleaned first to remove redundant data in the obtained network transmission data, so as to avoid the interference of redundant data on the network transmission data detection process, thereby further improving the accuracy of network transmission data detection.

[0139] In some embodiments, the semantic representation model provided in step 102 may specifically include an encoder and a decoder. (See [link to relevant documentation]). Figure 2 , Figure 2 This is a flowchart illustrating the method for training the semantic representation model in step 102. This method may include, but is not limited to, steps S201 to S206:

[0140] Step S201: Obtain the first training sample set, which includes multiple first sample data packets;

[0141] Step S202: Perform vector transformation on the first sample data packet to obtain the sample feature vector;

[0142] Step S203: Input the sample feature vector into the encoder for feature extraction to obtain the latent vector;

[0143] Step S204: Input the latent vector into the decoder and decode it to obtain the output vector;

[0144] Step S205: Adjust the parameters of the encoder and the decoder based on the difference between the output vector and the sample feature vector;

[0145] Step S206: Repeat the following steps until the semantic representation model is detected to have converged:

[0146] The first sample data packet is reacquired from the first training sample set. The first sample data packet is vector transformed to obtain the sample feature vector. The sample feature vector is input into the encoder for feature extraction to obtain the latent vector. The latent vector is input into the decoder for decoding to obtain the output vector. The parameters of the encoder and the decoder are adjusted based on the difference between the output vector and the sample feature vector. The convergence detection of the semantic representation model is then performed.

[0147] In step S201 of some embodiments, training the semantic representation model provided in step S102 of this application requires first obtaining a training sample set for training the semantic representation model. The forest patrol sample set contains multiple training samples, and each training sample can be a sample data packet. To distinguish it from the training samples of the network transmission data detection model below, the training sample set here can be called the first training sample set, which includes multiple first sample data packets. Since the semantic representation model is trained using a self-supervised training method in this embodiment, the first sample data packets in the first training sample set do not contain corresponding label data.

[0148] In this embodiment, the encoder and decoder of the semantic representation model can both be composed of multiple transformer-based encoding modules (a transformer is a typical deep learning model, specifically a deep learning model based entirely on a self-attention mechanism) stacked together. Each transformer encoding module includes a multi-head self-attention layer, a feedforward network layer, and two normalization layers. After determining the model structure of the semantic representation model and the first training sample set for training the semantic representation model, the semantic representation model can be trained based on the first training sample set.

[0149] In step S202 of some embodiments, the first sample data packet can be vector-converted to obtain a sample feature vector. Specifically, before performing vector conversion on the first sample data packet, data cleaning can be performed on the first sample data packet, followed by vector conversion on the cleaned first sample data packet to obtain the sample feature vector.

[0150] In step S203 of some embodiments, the sample feature vector can be further input into the encoder of the semantic representation model for feature extraction to obtain the latent vector output by the encoder. Specifically, the latent vector can be a 256-dimensional vector.

[0151] In step S204 of some embodiments, the latent vectors extracted by the encoder can be further input into the decoder of the semantic representation model for decoding to obtain the output vector output by the decoder. Specifically, this output vector can be a vector with the same dimension as the latent vectors, i.e., it can also be a 256-dimensional vector.

[0152] In step S205 of some embodiments, since a self-supervised training method is used to train the semantic representation model, the training objective is to minimize the difference between the output and input of the semantic representation model. Therefore, after the decoder outputs the output vector, the difference between the output vector and the input sample feature vector can be further calculated, and the parameters of the encoder and decoder in the semantic representation model can be adjusted based on this difference.

[0153] In step S206 of some embodiments, the steps 201 to 205, which involve adjusting the parameters of the encoder and decoder using the first sample data packet, can be executed repeatedly. During this loop, convergence detection of the semantic representation model can be performed simultaneously. If convergence of the semantic representation model is detected, the loop execution stops. Otherwise, the above steps are executed repeatedly to iteratively update the parameters of the semantic representation model. Specifically, convergence detection of the semantic representation model can be performed by detecting the number of loop executions. When the number of loops reaches a preset number, the semantic representation model can be determined to have converged. Alternatively, the degree of change in the parameters of the semantic representation model can be compared; when the degree of change in the parameters of the semantic representation model is less than a preset level, the semantic representation model can also be determined to have converged.

[0154] The specific process of extracting temporal feature data from the acquired network transmission data to be detected using the semantic representation model after training can be described as follows: the encoder using the semantic representation model encodes the network transmission data, obtaining the latent vector output by the encoder. The latent vector output by the encoder is the extracted temporal feature data. .

[0155] In this embodiment, the semantic representation model is trained using self-supervised training, which, compared to supervised training, eliminates the need for acquiring large amounts of labeled data. This reduces the time and effort lost due to manual data annotation, thereby improving the training efficiency of the semantic representation model. Consequently, it enhances the detection efficiency of network transmission data.

[0156] Please see Figure 3 In some embodiments, step S205 may include, but is not limited to, steps S301 to S302:

[0157] Step S301: Calculate the norm distance between the sample feature vector and the output vector to obtain the self-supervised learning loss;

[0158] Step S302: Adjust the parameters of the encoder and the decoder based on the self-supervised learning loss.

[0159] In step S301 of some embodiments, the self-supervised learning loss can be calculated by calculating the normative distance between the sample feature vector and the output vector. Specifically, the self-supervised learning loss can be calculated using the following formula:

[0160]

[0161] Where self_loss is the self-supervised loss, and x is the sample feature vector. This is the output vector. To calculate the normal distance between the sample feature vector and the output vector.

[0162] In step S302 of some embodiments, after the self-supervised loss is determined, the backpropagation gradient can be calculated based on the self-supervised loss, and then the parameters of the encoder and decoder in the semantic representation model can be adjusted according to the backpropagation gradient.

[0163] In this embodiment, the parameters of the semantic representation model are updated by calculating the paradigm distance between the sample feature vector and the output vector as a self-supervised loss, which can improve the accuracy of the trained semantic representation model.

[0164] Please see Figure 4 In some embodiments, network transmission data may include multiple network data packets, and step S102 may include, but is not limited to, steps S401 to S402:

[0165] Step S401: Based on the semantic representation model, feature extraction is performed on the multiple network data packets to obtain multiple sub-time domain feature data.

[0166] Step S402: The multiple sub-time domain feature data are concatenated to obtain time domain feature data.

[0167] In step S401 of some embodiments, when the acquired network transmission data to be detected includes multiple network data packets, feature extraction of the network transmission data based on the semantic representation model can specifically involve using the semantic representation model to extract features from each network data packet separately. Specifically, the encoder in the semantic representation model can be used to extract features from the multiple network data packets separately, and the encoder outputs multiple latent vectors to obtain multiple sub-temporal domain feature data.

[0168] In step S402 of some embodiments, multiple sub-temporal feature data extracted from multiple network data packets can be further concatenated. As mentioned above, the multiple sub-temporal feature data are all latent vectors output by the encoder. Specifically, concatenating the multiple sub-temporal feature data can be a vector concatenation process, resulting in a final vector that represents the temporal feature data corresponding to the network transmission data.

[0169] In this embodiment, when the network transmission data to be detected includes multiple network data packets, an encoder using a semantic representation model can be used to extract features from each of the multiple network data packets, and the extracted sub-temporal feature data can be concatenated to obtain the temporal feature data of the network transmission data. This improves the accuracy of the extracted temporal feature data.

[0170] Please see Figure 5 In some embodiments, step S103 may include, but is not limited to, steps S501 to S503:

[0171] Step S501: Convert the network transmission data into time-series data.

[0172] Step S502: Perform wavelet analysis on the time series to obtain high-frequency and low-frequency features.

[0173] Step S503: The high-frequency features and low-frequency features are concatenated to obtain the first frequency domain feature data.

[0174] In step S501 of some embodiments, before extracting frequency domain feature data from the network transmission data, the network transmission data can be converted into time-series data. The time-series data can be data that changes over time. For example, the network transmission data can be converted into time-series data f(t).

[0175] In step S502 of some embodiments, after converting the network transmission data into time-series data, the time-series data can be further converted to time-frequency data. In the embodiments of this application, wavelet transform can be used to perform time-frequency conversion on the time-series data. The basic wavelet function of the time-series data f(t) can be expressed as:

[0176]

[0177] in, It is the basic wavelet function. It is a normalization factor that ensures that wavelet functions have the same energy at different scales. It is a scaling and translation transform of the basic wavelet function, representing the scaling... The next Sub-wavelet functions.

[0178] The discrete wavelet transform can be expressed as:

[0179]

[0180] in, It is the time series data f(t) at scale and displacement The binary wavelet coefficients below.

[0181] Then, we can use The algorithm performs wavelet decomposition. The algorithm is a classic discretization method of continuous wavelet transform. It decomposes a signal into approximate signals at different frequency channels and detail signals at each scale; these detail signals can be called wavelet surfaces. This algorithm can be used to perform multi-level wavelet decomposition on time-series data; for example, the decomposition level can be set to 5. Wavelet decomposition yields multiple wavelet components, such as high-frequency components and low-frequency components. The high-frequency components can be called high-frequency features, and the low-frequency components can be called low-frequency features.

[0182] In step S503 of some embodiments, after performing wavelet analysis on the time series to obtain high-frequency and low-frequency features, the high-frequency and low-frequency features obtained from the wavelet analysis can be further concatenated to obtain first frequency domain feature data, which can be represented as V. f .

[0183] This application employs wavelet transform to convert network transmission data from time-series data into frequency domain features. Wavelet transform effectively highlights certain characteristics of the problem, enabling localized analysis of time (space) and frequency. Through scaling and translation operations, it progressively refines the time-series data at multiple scales, ultimately achieving time subdivision at high frequencies and frequency subdivision at low frequencies. This automatically adapts to the requirements of time-frequency signal analysis, allowing focus on any detail of the signal. Consequently, the accuracy of the extracted frequency domain features of the network transmission data is significantly improved, thereby enhancing the accuracy of network transmission data detection.

[0184] Please see Figure 6 In some embodiments, step S501 may include, but is not limited to, steps S601 to S602:

[0185] Step S601: Cut the data sequence corresponding to the network transmission data according to the preset step size to obtain multiple data frames of the same length.

[0186] Step S602: Generate time sequence data based on a frame sequence composed of multiple data frames of the same length.

[0187] In step S601 of some embodiments, the network transmission data may include multiple network data packets, wherein each network data packet may use a vector x i This means that a data sequence consisting of multiple network data packets can be represented as: ,in Let X be the total number of network data packets. Then, the data sequence X corresponding to the network transmission data can be segmented according to a preset step size to obtain multiple data frames with the same frame length.

[0188] In step S602 of some embodiments, after the network transmission data is divided into multiple data frames of the same length, each data frame can be numerically converted to a corresponding value. This converts the network transmission data into a one-dimensional numerical sequence, and these one-dimensional numerical sequences with consistent step sizes constitute time-series data. When the step size is sufficiently small, the time-series data can be waveform data that changes over time.

[0189] In this embodiment, the data sequence consisting of network transmission data is divided into multiple data frames of equal length according to a preset step size. Then, based on these data frames of equal length, the network transmission data is converted into a time-series data, thereby facilitating time-frequency conversion and frequency domain feature extraction of the network transmission data. This significantly improves the efficiency of frequency domain feature extraction from network transmission data, and consequently, the efficiency of network transmission data detection.

[0190] Please see Figure 7 In some embodiments, step S104 may also include, but is not limited to, steps S701 to S703:

[0191] Step S701: Extract features from the statistical data of the transport layer in the network transmission data to obtain statistical features;

[0192] Step S702: Map the statistical features to a preset numerical range to obtain a set of values;

[0193] Step S703: Construct a representation image using the values ​​in the numerical set as pixel values.

[0194] In step S701 of some embodiments, statistical features of the network transmission data are extracted. Specifically, a statistical feature extraction tool (such as CICFlowMeter, a network transmission data feature extraction tool that takes a pcap file as input and outputs the feature information of the data packets contained in the pcap file, which has more than 80 dimensions) can be used to extract the statistical features. The feature data extracted by the statistical feature extraction tool is output in the form of a CSV table, that is, the extracted more than 80-dimensional features are numerical data.

[0195] In step S702 of some embodiments, since the feature values ​​obtained by statistical feature extraction of network transmission data have a large variance, feature values ​​greater than 255 can be further mapped to the interval [0, 255]. Specifically, the feature values ​​of the extracted statistical features can be mapped to the above interval using a modulo operation.

[0196] In step S703 of some embodiments, after mapping the feature values ​​of the statistical characteristics of the network transmission data to a range of 0 to 255, the mapped feature values ​​can be used as pixel values ​​to construct the corresponding grayscale image. .

[0197] In some embodiments, constructing a representation image using numerical values ​​from a set of values ​​as pixel values ​​may further include:

[0198] Expand the values ​​in the numerical set to obtain the target numerical set;

[0199] The representation image is constructed using the values ​​in the target numerical set as pixel values.

[0200] In this embodiment, before constructing the grayscale image, the size (dimensions) of the grayscale image to be constructed can be determined first, for example, a 32*32 dimensional grayscale image. After determining the size of the image to be constructed, the required amount of data can be further clarified. In this embodiment, since the statistical features obtained by extracting statistical features from each network data packet have more than 80 dimensions, these 80+ dimensional statistical features can be expanded through a certain data expansion method, that is, the values ​​in the numerical set constituted by the statistical features can be expanded. Specifically, the values ​​in the numerical set can be expanded to 128, and these 128 values ​​constitute the target numerical set corresponding to a network data packet.

[0201] Furthermore, eight network data packets can be randomly selected from the multiple network data packets contained in the network transmission data. The 128-dimensional statistical features of each network data packet can be extracted, and finally, a 32*32-dimensional grayscale image can be generated based on the statistical features of these eight network data packets. .

[0202] In some embodiments, data expansion is performed on the numerical values ​​in the numerical set to obtain a target numerical set, including:

[0203] Calculate the mean and variance of multiple values ​​in a set of values;

[0204] The mean and variance are added to the set of values ​​to obtain the target set of values.

[0205] In this embodiment of the application, the specific method for expanding the numerical values ​​in the numerical set can be to calculate the mean or variance of some or all of the numerical values ​​in the numerical set, obtaining multiple means and multiple variances. Then, the calculated multiple means and multiple variances are added to the numerical set to form the target numerical set.

[0206] Thus, this embodiment of the application expands the statistical features by extending the statistical features of multiple dimensions into a target numerical set of preset dimensions, thereby facilitating the construction of a representation image of the corresponding size.

[0207] Please see Figure 8 In some embodiments, step S105 includes, but is not limited to, steps S801 to S803:

[0208] Step S801: Obtain the orthogonal wavelet basis functions.

[0209] Step S802: Generate multiple image frequency domain features representing the image based on the orthogonal wavelet basis functions.

[0210] Step S803: The multiple image frequency domain features are spliced ​​together to obtain the second frequency domain feature data.

[0211] In step S801 of some embodiments, the representation image can be time-frequency converted using wavelet transform in this application embodiment. Specifically, the orthogonal wavelet basis function Haar can be used to perform time-frequency conversion on the representation image to generate frequency domain features of the representation image, that is, to generate second frequency domain feature data. The orthogonal wavelet basis function Haar can be a single rectangular wave on the interval [0, 1], specifically defined as follows:

[0212] .

[0213] In step S802 of some embodiments, after obtaining the orthogonal wavelet basis function Haar, the representation image can be time-frequency transformed based on the orthogonal wavelet basis function Haar. Specifically, the orthogonal wavelet basis function Haar can be used to generate multiple image frequency domain features containing high-frequency and low-frequency features of the representation image.

[0214] In step S803 of some embodiments, multiple image frequency domain features representing the image can be further concatenated to obtain second frequency domain feature data. .

[0215] In some embodiments, after extracting features from multiple network data packets based on a semantic representation model to obtain multiple sub-time domain feature data, the method further includes:

[0216] Obtain the target number of sub-time domain features from multiple sub-time domain feature data;

[0217] The feature matrix is ​​obtained by combining the number of sub-time-domain features of the target.

[0218] Third frequency domain feature data based on generating feature matrices using orthogonal wavelet basis functions;

[0219] Data detection is performed based on time-domain feature data, first-frequency-domain feature data, and second-frequency-domain feature data to obtain the detection results of the network transmission data to be detected, including:

[0220] Data detection is performed based on time-domain feature data, first frequency-domain feature data, second frequency-domain feature data, and third frequency-domain feature data to obtain the detection results of the network transmission data to be detected.

[0221] In this embodiment, to further improve the accuracy of detecting network transmission data, corresponding frequency domain features can be extracted from the temporal features of the network transmission data, i.e., the third frequency domain feature data of the network transmission data can be extracted. Then, fusion detection is performed based on the temporal features of the network transmission data and the three aspects of frequency domain features. Specifically, after extracting temporal features from multiple network data packets in the network transmission data using a semantic representation model to obtain multiple sub-temporal features, a target number of sub-temporal features can be obtained from the extracted sub-temporal features. Here, the target number can be calculated based on the dimension of the aforementioned representation image and the dimension of the sub-temporal features. For example, if the dimension of the representation image is 32*32, and the dimension of the sub-temporal features is 256, then four sub-temporal features can be randomly obtained.

[0222] After obtaining the target number of sub-time-domain features, these features can be combined to obtain a feature matrix. This feature matrix can be a 32*32 dimensional feature matrix.

[0223] Furthermore, similar to the time-frequency transformation of the representation image, orthogonal wavelet basis functions (Haar) can be used to extract frequency domain features from the feature matrix to obtain the third frequency domain feature data. .

[0224] In some embodiments, data detection is performed based on time-domain feature data, first frequency-domain feature data, second frequency-domain feature data, and third frequency-domain feature data to obtain the detection result of the network transmission data to be detected, including:

[0225] Obtain the weight coefficients corresponding to the second frequency domain feature data and the third frequency domain feature data respectively;

[0226] Based on the obtained corresponding weight coefficients, the second frequency domain feature data and the third frequency domain feature data are fused to obtain the fused frequency domain features;

[0227] The target features are obtained by concatenating the time-domain feature data, the first frequency-domain feature data, and the fused frequency-domain features.

[0228] Data detection is performed on the target features to obtain the detection results of the network transmission data to be detected.

[0229] In this embodiment, the second and third frequency domain feature data can be fused first based on the weighting coefficients corresponding to the second and third frequency domain feature data to obtain the fused frequency domain features, as expressed by the following formula:

[0230]

[0231] in, These are the weighting coefficients corresponding to the second frequency domain feature data. These are the weighting coefficients corresponding to the third frequency domain feature data.

[0232] Furthermore, the time-domain feature data V can be... t First frequency domain feature data V f and fused frequency domain features V fe The target feature V is obtained by concatenating the features. tf .

[0233] In this way, by detecting network transmission data, target features can be detected directly, and the detection results of network transmission data can be obtained.

[0234] In some embodiments, target features are detected, specifically using a network transmission data detection model.

[0235] Please see Figure 9 In some embodiments, the training process of the network transmission data detection model may include, but is not limited to, steps S901 to S906:

[0236] Step S901: Obtain the second training sample set, which includes multiple sets of sample network transmission data and label data corresponding to each sample network transmission data.

[0237] Step S902: Based on the semantic representation model, feature extraction is performed on each sample of network transmission data to obtain sample time-domain feature data corresponding to each sample of network transmission data.

[0238] Step S903: Convert each sample network transmission data into sample time-series data, and perform time-frequency conversion on the sample time-series data to obtain the first sample frequency domain feature data corresponding to each sample network transmission data.

[0239] Step S904: Extract the sample statistical features of each sample network transmission data, generate a sample representation image based on the sample statistical features, and perform frequency domain transformation on the sample representation image to obtain the second sample frequency domain feature data.

[0240] Step S905: The time-domain feature data, the first sample frequency-domain feature data, and the second sample frequency-domain feature data corresponding to each sample of network transmission data are fused to obtain the fused feature corresponding to each sample of network transmission data.

[0241] Step S906: Using the fusion features corresponding to each sample of network transmission data as model input, and using the label data corresponding to each sample of network transmission data as model training labels, a preset neural network model is trained to obtain the network transmission data detection model.

[0242] In step S901 of some embodiments, before training the network transmission data detection model, the model structure and training sample data of the network transmission data detection model can be determined first. In the embodiments of this application, the model structure of the network transmission data detection model can specifically be a bidirectional long short-term memory (BiLSTM) network, such as... Figure 10 The figure shows a schematic diagram of the network transmission data detection model in this application. As shown, the network transmission data detection model includes an input layer, a forward transmission layer, a backward transmission layer, and an output layer. The input layer is responsible for sequence encoding of the input data to ensure it meets the model's input requirements; the forward transmission layer extracts forward features from the input sequence; the backward transmission layer extracts backward features from the input sequence; and the output layer integrates the data output from the forward and backward transmission layers. As illustrated in the figure, the BiLSTM model structure has unique advantages in studying the correlation between features.

[0243] The training sample data used to train the network transmission data detection model can be a second training sample set, which includes multiple sets of sample network transmission data, each of which corresponds to a specific label. In other words, training the network transmission data detection model is supervised training.

[0244] In step S902 of some embodiments, for each sample network transmission data in the second training sample set, a semantic representation model can be used to extract features to obtain sample temporal feature data corresponding to each sample network transmission data. Here, the semantic representation model can be the same as the semantic representation model in step 102.

[0245] In step S903 of some embodiments, each sample network transmission data can be further converted into sample time-series sequence data, and time-frequency conversion can be performed on the sample time-series sequence data to obtain the first sample frequency domain feature data corresponding to each sample network transmission data. This step is consistent with the process of extracting the first frequency domain feature data of the network transmission data in step 103, and will not be described again here.

[0246] In step S904 of some embodiments, sample statistical features of each sample of network transmission data can be further extracted, and second sample frequency domain feature data of each sample of network transmission data can be generated accordingly. This process is consistent with the process of extracting second frequency domain feature data of network transmission data, and will not be described again here.

[0247] In step S905 of some embodiments, the sample time-domain feature data, the first sample frequency-domain feature data, and the second sample frequency-domain feature data of each sample of network transmission data can be fused to obtain the fused feature corresponding to each sample of network transmission data. At this point, the data pair for training the network transmission data detection model is constructed.

[0248] In step S906 of some embodiments, the fusion features corresponding to the sample network transmission data can be used as the model input of the network transmission data detection model, and the corresponding label data can be used as training labels to train the BiLSTM model to obtain the network transmission data detection model.

[0249] Specifically, during training, cross-entropy can be calculated as the loss function for training the model:

[0250]

[0251] Where yi represents the label of sample network transmitted data i, with 1 for the abnormal class and 0 for the normal class. pi represents the probability that the sample network transmitted data is predicted to be of the abnormal class.

[0252] In some embodiments, after training the network transmission data detection model, a test set can be obtained to test the generalization ability of the trained model. Evaluation metrics may include accuracy (Acc), recall (R), and F1 score. The specific calculation formulas are as follows:

[0253]

[0254]

[0255]

[0256]

[0257] Wherein, TP represents the number of samples where both actual and predicted network transmission data are abnormal; FP represents the number of samples where actual network transmission data is normal but predicted as abnormal; TN represents the number of samples where both actual and predicted network transmission data are normal; and FN represents the number of samples where actual network transmission data is abnormal but predicted as normal. If the test results are good, the model is saved as the final network transmission data detection model. If the test results are poor, feedback is used for iteration, redesigning the network structure and adjusting the hyperparameters.

[0258] The network transmission data detection method provided in this application involves: acquiring the network transmission data to be detected; extracting features from the network transmission data based on a semantic representation model to obtain time-domain feature data; converting the network transmission data into time-series data and performing time-frequency conversion on the time-series data to obtain first frequency-domain feature data; extracting statistical features from the network transmission data and generating a representation image based on the statistical features; performing frequency-domain conversion on the representation image to obtain second frequency-domain feature data; and performing data detection based on the time-domain feature data, the first frequency-domain feature data, and the second frequency-domain feature data to obtain the detection result of the network transmission data to be detected. The network transmission data processing method provided in this embodiment only requires extracting time-domain features, the first frequency-domain features, and the second frequency-domain features from the network transmission data to be detected, and performing detection based on the extracted time-domain features, the first frequency-domain features, and the second frequency-domain features. This eliminates the need to decrypt the network transmission data to be detected, significantly improving the efficiency of network transmission data detection. Furthermore, this method not only extracts the temporal features of the network transmission data to be detected, but also extracts the frequency domain features of the network transmission data to be detected from both temporal and semantic dimensions, thereby extracting multi-level joint feature data of the network transmission data to be detected. Detecting network transmission data based on this multi-level joint feature data can significantly improve the accuracy of network transmission data detection.

[0259] Please see Figure 11 This application also provides a network transmission data detection device, which can implement the above-described network transmission data detection method. The device includes:

[0260] The acquisition module is used to acquire the network transmission data to be detected.

[0261] The extraction module is used to extract features from the network transmission data based on a semantic representation model to obtain temporal feature data;

[0262] The first conversion module is used to convert the network transmission data into time-series data, and to perform time-frequency conversion on the time-series data to obtain first frequency domain feature data.

[0263] A generation module is used to extract statistical features of the network transmission data and generate a representation image based on the statistical features;

[0264] The second conversion module is used to perform frequency domain conversion on the representation image to obtain second frequency domain feature data;

[0265] The detection module is used to perform data detection based on the time-domain feature data, the first frequency-domain feature data, and the second frequency-domain feature data to obtain the detection result of the network transmission data to be detected.

[0266] The specific implementation of the network transmission data detection device is basically the same as the specific implementation of the network transmission data detection method described above, and will not be repeated here.

[0267] This application also provides a computer device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the aforementioned method for detecting network transmission data. This computer device can be any smart terminal, including tablet computers, in-vehicle computers, etc.

[0268] Please see Figure 12 , Figure 12 The hardware structure of a computer device according to another embodiment is illustrated. The computer device includes:

[0269] The processor 1201 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this application.

[0270] The memory 1202 can be implemented as a read-only memory (ROM), static storage device, dynamic storage device, or random access memory (RAM). The memory 1202 can store the operating system and other applications. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 1202 and is called and executed by the processor 1201 to execute the network transmission data detection method of the embodiments of this application.

[0271] The input / output interface 1203 is used to implement information input and output;

[0272] The communication interface 1204 is used to enable communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.).

[0273] Bus 1205 transmits information between various components of the device (e.g., processor 1201, memory 1202, input / output interface 1203, and communication interface 1204);

[0274] The processor 1201, memory 1202, input / output interface 1203 and communication interface 1204 are connected to each other within the device via bus 1205.

[0275] This application also provides a storage medium storing a computer program that, when executed by a processor, implements the above-described method for detecting network transmission data.

[0276] Memory, as a non-transitory storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. Furthermore, memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, memory may optionally include memory remotely located relative to the processor, and these remote memories can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.

[0277] The network transmission data detection method, apparatus, computer equipment, and storage medium provided in this application embodiment acquire the network transmission data to be detected; extract features from the network transmission data based on a semantic representation model to obtain time-domain feature data; convert the network transmission data into time-series data and perform time-frequency conversion on the time-series data to obtain first frequency-domain feature data; extract statistical features from the network transmission data and generate a representation image based on the statistical features; perform frequency-domain conversion on the representation image to obtain second frequency-domain feature data; and perform data detection based on the time-domain feature data, the first frequency-domain feature data, and the second frequency-domain feature data to obtain the detection result of the network transmission data to be detected. The network transmission data processing method provided in this embodiment only requires extracting time-domain features, the first frequency-domain features, and the second frequency-domain features from the network transmission data to be detected, and then performing detection based on the extracted time-domain features, the first frequency-domain features, and the second frequency-domain features. This eliminates the need to decrypt the network transmission data to be detected, significantly improving the efficiency of network transmission data detection. Furthermore, this method not only extracts the temporal features of the network transmission data to be detected, but also extracts the frequency domain features of the network transmission data to be detected from both temporal and semantic dimensions, thereby extracting multi-level joint feature data of the network transmission data to be detected. Detecting network transmission data based on this multi-level joint feature data can significantly improve the accuracy of network transmission data detection.

[0278] The embodiments described in this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided by the embodiments of this application. As those skilled in the art will know, with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of this application are also applicable to similar technical problems.

[0279] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of this application, and may include more or fewer steps than shown, or combine certain steps, or different steps.

[0280] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.

[0281] Those skilled in the art will understand that all or some of the steps in the methods disclosed above, as well as the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, or suitable combinations thereof.

[0282] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0283] It should be understood that in this application, "at least one (item)" means one or more, and "more than" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.

[0284] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of the units described above is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0285] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0286] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0287] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing programs, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0288] The preferred embodiments of the present application have been described above with reference to the accompanying drawings, but this does not limit the scope of the claims of the present application. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and substance of the embodiments of the present application shall be within the scope of the claims of the present application.

Claims

1. A method of detecting network transmission data, characterized by, The method includes: Acquire the network transmission data to be detected; Based on the semantic representation model, feature extraction is performed on the network transmission data to obtain temporal feature data; The network transmission data is converted into time-series data, and the time-series data is subjected to time-frequency conversion to obtain the first frequency domain feature data; Extract statistical features from the network transmission data, and generate a characterization image based on the statistical features; The representation image is transformed in the frequency domain to obtain the second frequency domain feature data; Data detection is performed based on the time-domain feature data, the first frequency-domain feature data, and the second frequency-domain feature data to obtain the detection result of the network transmission data to be detected. The step of converting the network transmission data into time-series data and performing time-frequency conversion on the time-series data to obtain first frequency domain feature data includes: Convert the network transmission data into time-series data; Wavelet analysis was performed on the time series data to obtain high-frequency and low-frequency features; The high-frequency features are concatenated with the low-frequency features to obtain the first frequency domain feature data.

2. The method according to claim 1, wherein converting the network transmission data into time-series data comprises: The data sequence corresponding to the network transmission data is cut according to a preset step size to obtain multiple data frames of the same length; Time sequence data is generated based on the frame sequence composed of multiple data frames of the same length.

3. The method of claim 1, wherein, The step of extracting statistical features from the network transmission data and generating a representation image based on the statistical features includes: Feature extraction is performed on the statistical data of the transport layer in the network transmission data to obtain statistical features; The statistical features are mapped to a preset numerical range to obtain a set of values; The numerical values ​​in the set of values ​​are used as pixel values ​​to construct a representation image.

4. The method of claim 3, wherein, The step of constructing a representation image using the values ​​in the set of values ​​as pixel values ​​includes: The values ​​in the numerical set are expanded to obtain the target numerical set; The target numerical values ​​are used as pixel values ​​to construct a representation image.

5. The method of claim 4, wherein, The step of expanding the values ​​in the numerical set to obtain the target numerical set includes: Calculate the mean and variance of multiple values ​​in the numerical set; The mean and variance are added to the set of values ​​to obtain the target set of values.

6. The method of claim 3, wherein, The step of performing frequency domain transformation on the representation image to obtain second frequency domain feature data includes: Obtain orthogonal wavelet basis functions; Multiple image frequency domain features representing the image are generated based on the orthogonal wavelet basis functions; The multiple image frequency domain features are concatenated to obtain the second frequency domain feature data.

7. The method of claim 1, wherein, The network transmission data includes multiple network data packets. The feature extraction of the network transmission data based on the semantic representation model to obtain temporal feature data includes: Based on the semantic representation model, feature extraction is performed on the multiple network data packets to obtain multiple sub-time domain feature data; The multiple sub-time domain feature data are concatenated to obtain time domain feature data.

8. The method of claim 7, wherein, After extracting features from the multiple network data packets based on the semantic representation model to obtain multiple sub-time domain feature data, the method further includes: Obtain a target number of sub-time domain features from the plurality of sub-time domain feature data; The target number of sub-temporal features are combined to obtain a feature matrix; The third frequency domain feature data of the feature matrix is ​​generated based on orthogonal wavelet basis functions; The step of performing data detection based on the time-domain feature data, the first frequency-domain feature data, and the second frequency-domain feature data to obtain the detection result of the network transmission data to be detected includes: Data detection is performed based on the time-domain feature data, the first frequency-domain feature data, the second frequency-domain feature data, and the third frequency-domain feature data to obtain the detection result of the network transmission data to be detected.

9. The method of claim 8, wherein, The step of performing data detection based on the time-domain feature data, the first frequency-domain feature data, the second frequency-domain feature data, and the third frequency-domain feature data to obtain the detection result of the network transmission data to be detected includes: Obtain the weight coefficients corresponding to the second frequency domain feature data and the third frequency domain feature data respectively; Based on the obtained corresponding weight coefficients, the second frequency domain feature data and the third frequency domain feature data are fused to obtain fused frequency domain features; The time-domain feature data, the first frequency-domain feature data, and the fused frequency-domain feature data are concatenated to obtain the target feature. Data detection is performed on the target features to obtain the detection results of the network transmission data to be detected.

10. The method of claim 7, wherein, The semantic representation model includes an encoder and a decoder, and the semantic representation model is trained using the following method: Obtain a first training sample set, which includes multiple first sample data packets; The first sample data packet is vector-transformed to obtain the sample feature vector; The sample feature vector is input into the encoder for feature extraction to obtain the latent vector; The latent vector is input into the decoder for decoding to obtain the output vector; The parameters of the encoder and the parameters of the decoder are adjusted based on the difference between the output vector and the sample feature vector. The following steps are repeated until the semantic representation model is detected to have converged: The first sample data packet is reacquired from the first training sample set. The first sample data packet is vector-transformed to obtain the sample feature vector. The sample feature vector is input into the encoder for feature extraction to obtain the latent vector. The latent vector is input into the decoder for decoding to obtain the output vector. The parameters of the encoder and the decoder are adjusted based on the difference between the output vector and the sample feature vector. The convergence detection of the semantic representation model is then performed.

11. The method according to claim 10, characterized in that, The adjustment of the encoder parameters and the decoder parameters based on the difference between the output vector and the first sample feature vector includes: Calculate the paradigm distance between the sample feature vector and the output vector to obtain the self-supervised learning loss; The parameters of the encoder and the parameters of the decoder are adjusted based on the self-supervised learning loss.

12. The method of claim 1, wherein, The step of performing data detection based on the time-domain feature data, the first frequency-domain feature data, and the second frequency-domain feature data to obtain the detection result of the network transmission data to be detected includes: The fused feature obtained by fusing the time-domain feature data, the first frequency-domain feature data, and the second frequency-domain feature data is input into the network transmission data detection model for detection, and the network transmission data to be detected is obtained. The network transmission data detection model is trained using the following method: Obtain a second training sample set, which includes multiple sets of sample network transmission data and label data corresponding to each sample network transmission data; Based on the semantic representation model, feature extraction is performed on each sample of network transmission data to obtain sample temporal feature data corresponding to each sample of network transmission data. Each sample network transmission data is converted into sample time-series data, and the sample time-series data is converted into time-frequency data to obtain the first sample frequency domain feature data corresponding to each sample network transmission data. Extract the sample statistical features of each sample network transmission data, generate a sample representation image based on the sample statistical features, and perform frequency domain transformation on the sample representation image to obtain the second sample frequency domain feature data. The time-domain feature data, the first sample frequency-domain feature data, and the second sample frequency-domain feature data corresponding to each sample of network transmission data are fused to obtain the fused feature corresponding to each sample of network transmission data. The network transmission data detection model is obtained by using the fusion features corresponding to each sample of network transmission data as the model input and the label data corresponding to each sample of network transmission data as the model training label training preset neural network model.

13. The method of claim 1, wherein, After acquiring the network transmission data to be detected, the process further includes: The network transmission data is cleaned to obtain the target network transmission data; The feature extraction of the network transmission data based on the semantic representation model to obtain temporal feature data includes: Based on the semantic representation model, feature extraction is performed on the target network transmission data to obtain temporal feature data; The step of converting the network transmission data into time-series data includes: Convert the target network transmission data into time-series data; The extraction of statistical features from the network transmission data includes: Extract the statistical characteristics of the target network transmission data.

14. A network transmission data detecting apparatus characterized by comprising: The device includes: The acquisition module is used to acquire the network transmission data to be detected. The extraction module is used to extract features from the network transmission data based on a semantic representation model to obtain temporal feature data; The first conversion module is used to convert the network transmission data into time-series data, and to perform time-frequency conversion on the time-series data to obtain first frequency domain feature data. A generation module is used to extract statistical features of the network transmission data and generate a representation image based on the statistical features; The second conversion module is used to perform frequency domain conversion on the representation image to obtain second frequency domain feature data; The detection module is used to perform data detection based on the time-domain feature data, the first frequency-domain feature data, and the second frequency-domain feature data to obtain the detection result of the network transmission data to be detected. The step of converting the network transmission data into time-series data and performing time-frequency conversion on the time-series data to obtain first frequency domain feature data includes: Convert the network transmission data into time-series data; Wavelet analysis was performed on the time series data to obtain high-frequency and low-frequency features; The high-frequency features are concatenated with the low-frequency features to obtain the first frequency domain feature data.

15. A computer device, comprising: The computer device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the network transmission data detection method according to any one of claims 1 to 13.

16. A storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the method for detecting network transmission data as described in any one of claims 1 to 13.