Game live video traffic identification method and system based on deep Gaussian random convolution
Through the deep Gaussian random convolution method, the problems of inefficiency and high computing resource consumption in network live video traffic recognition are solved, and efficient and robust game live video traffic recognition is achieved, capturing the characteristics of local details and global patterns.
Patent Information
- Application Number
- CN202510574714.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-06
- Publication Date
- 2025-08-08
AI Technical Summary
The existing online live video traffic recognition technology is inefficient when facing encrypted traffic, has high computing resource consumption, and is difficult to achieve efficient online real-time recognition. The traditional method has gradually failed, and emerging deep learning models have problems such as large amount of parameters, poor generalization and poor interpretability.
The method based on deep Gaussian random convolution is adopted, by collecting network traffic data, extracting five-tuple and statistical feature information, constructing histogram features, using deep Gaussian random convolution neural network for training, designing three-layer convolutional layer and fully connected layer, and using binary cross entropy loss function for model optimization to realize the identification of game live video traffic.
The training parameters are reduced, the robustness and interpretability of the model are improved, and the traffic characteristics can be extracted at different levels, local details and global patterns are captured, and efficient game live video traffic recognition is achieved.
Smart Images

Figure CN120455647A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of cyberspace security technology, and in particular to a method and system for identifying game live video traffic based on deep Gaussian random convolution. Background Art
[0002] Live streaming, as a crucial channel for information dissemination in the video streaming sector, has become widely integrated into various sectors, including entertainment, education, and commerce. With the rapid growth in user numbers and the increasing variety of live streaming content, the quality of live streaming content varies greatly. For example, some minors are easily exposed to inappropriate content, such as live gaming broadcasts, which can have a serious negative impact on their development. To protect the physical and mental health and safety of minors, the analysis and identification of encrypted live video traffic has become a research hotspot in academia and industry in recent years.
[0003] In today's network environment, the use of dynamic ports and the frequent adoption of port masquerading technology have rendered traditional port-based identification methods increasingly ineffective. Currently, existing traffic analysis and identification methods fall into two main categories: content-based and content-independent.
[0004] Specifically, content-based identification methods, such as deep packet inspection, require in-depth analysis of the content carried by each data packet and matching it with a pre-built fingerprint library to achieve identification. However, this type of identification model is relatively inefficient, especially when dealing with encrypted traffic, and has obvious limitations, making it difficult to be effective.
[0005] Non-content-based identification methods primarily utilize traffic statistics to identify and classify traffic. These methods demonstrate relatively high accuracy in identifying encrypted traffic and effectively address the challenges of identifying encrypted traffic. With the rapid development of deep learning technology, traditional methods that rely on manually extracting traffic statistics are gradually being replaced by emerging deep learning methods. Deep learning methods offer significant advantages in processing large-scale traffic data. However, traditional deep learning models have a large number of parameters and poor generalization. Training requires a large number of data samples and significant computing power, resulting in poor model interpretability. Deep neural networks are often viewed as "black box" models, making their decision-making processes difficult to understand and explain.
[0006] In summary, current network traffic analysis and identification technologies still face numerous challenges. On the one hand, with the widespread adoption of encryption technology, traditional plaintext-based analysis methods have become increasingly ineffective. On the other hand, while emerging traffic analysis methods have somewhat addressed the shortcomings of traditional methods, they still have numerous limitations, such as excessive computational resource consumption, large memory footprint, low processing efficiency, poor interpretability, and difficulty achieving efficient online, real-time identification. Summary of the Invention
[0007] In order to address the deficiencies of the prior art, the present invention provides a method and system for identifying game live video traffic based on deep Gaussian random convolution; In the present invention, first, the game live video traffic is collected and the histogram features of the video traffic are extracted. Then, the one-dimensional vector of the histogram features is input into the designed deep Gaussian random convolutional neural network for training. Finally, the trained model is used to reliably classify game and non-game live videos.
[0008] The technical solution of the present invention is: In a first aspect, the present invention provides a method for identifying game live video traffic based on deep Gaussian random convolution; A method for identifying game live video traffic based on deep Gaussian random convolution, comprising: Collect raw network traffic data; Extract the five-tuple information and statistical feature information of the original network traffic data. The five-tuple information includes the source IP address, source port, destination IP address, destination port and protocol type of the data packet; the statistical feature information includes the load size and arrival time; obtain the flow based on the five-tuple information; Based on the statistical characteristics of the flow, the duration and flow rate of the flow are used to perform coarse-grained identification of the flow, extract the video flow, and obtain the game and non-game live video flow training set; The histogram features are extracted from the training set video traffic, and the one-dimensional vector of the histogram features is input into the designed deep Gaussian random convolution game live video traffic recognition model for training; Collect the live video to be tested, extract features from the live video to be tested, and filter the extracted features to obtain a feature data set to be tested; The feature dataset to be tested is input into the trained recognition model, and the recognition results of Internet game live video traffic are output.
[0009] Preferably, according to the present invention, data packets with the same quintuple information constitute one flow.
[0010] Preferably, according to the present invention, the flow is coarsely identified based on the statistical characteristic information of the flow, that is, the duration of the flow and the flow rate, that is, the number of bytes per second, to extract the video flow; including: Get the duration of each stream , and calculate the duration of each flow The number of bytes arriving in , as shown below: ; in, Indicates the obtained Flow rate, The weight corresponding to the duration of the flow and the number of bytes per second, represents the threshold, if The value is greater than or equal to , it is judged as video traffic, otherwise it is judged as non-video traffic.
[0011] According to the preferred embodiment of the present invention, if multiple streams in the live broadcast room session reach the threshold The final judgment is made through the domain name resolved in the DNS response package; the information resolved by the DNS response package includes IP, protocol, and domain name, and multiple streams in the live broadcast room are marked; and whether it is video traffic is determined from these domain names.
[0012] Preferably, according to the present invention, the histogram features are extracted from the training set video traffic; including: The histogram feature is constructed by extracting the arrival time and payload of the data packets from the statistical feature information of the flow; the time interval is defined, and the payload arriving within the time interval is counted to obtain a one-dimensional histogram feature vector of the payload arriving within the time interval; and a training data set of games and non-games including the histogram features is formed.
[0013] Preferably, according to the present invention, a recognition model for live game video traffic using deep Gaussian random convolution is designed; the recognition model includes three convolutional layers and a fully connected layer, each convolutional layer uses three convolution kernels of different sizes, and the size of each convolution kernel is randomly selected from the three different sizes of convolution kernels; the parameters of the three different sizes of convolution kernels are randomly initialized based on Gaussian distribution, and all convolution kernel parameters are fixed and do not participate in backpropagation updates; Further preferably, in the recognition model, the convolution kernel size increases layer by layer as the number of convolution layers increases; the convolution kernel size of the first convolution layer is 3, the convolution kernel size of the second convolution layer is 5, and the convolution kernel size of the third convolution layer is 7; when the input one-dimensional vector enters the three convolution layers, each convolution layer generates One-dimensional vector output; then for this The one-dimensional vectors are assigned initial weights and weighted summation is performed to generate a new one-dimensional input, which is then put into the next convolutional layer; finally, the extracted features are input into the fully connected layer to obtain the final classification result.
[0014] Preferably, according to the present invention, the one-dimensional vector of the histogram feature is input into the designed deep Gaussian random convolution game live video traffic recognition model for training; comprising: The dataset is divided into a training set and a test set. The training set is input into the recognition model. The feature map generated by each convolution operation is passed to the next layer through weighted summation. After multiple convolutions, the extracted features are finally input into the fully connected layer to output the classification results. Specifically, it includes: Assume the input data is ,in, It is Video traffic samples, is the time step; each convolutional layer has convolution kernels, where ; K is the number of convolution kernels in each layer; the size of each convolution kernel is Randomly select from, where is the length of the convolution kernel; For the convolution kernels , the corresponding bias term is ; pass Padding The operation keeps the output length of the recognition model unchanged as follows: ; in, is the input data length; is the convolution kernel size; is the length that needs to be filled; Is the step size of the convolution kernel movement, so: ; in, is the input data length; is the output data length; Output of the convolutional layer Calculated by the following formula: ; in, is the output data at time step Hedi m Layer The elements corresponding to the convolution kernels, is to proceed padding The input data after that is at time step The corresponding elements, It is m Layer The convolution kernel is at position Corresponding elements; the internal parameters of the convolution kernel are randomly initialized and fixed based on the Gaussian distribution and do not participate in the back propagation update; thus we get: ; ; in, is the output data calculated by the kth convolution kernel; is the value obtained by convolution at the Tth time step; It is i In the convolution layer, k The set of output data calculated by the convolution kernel; It is i In the convolution layer, k The output data calculated by the convolution kernel; initialization The weight corresponding to each output data ; From this we get: ; ; in, It is i The set of weights corresponding to each output data in the layer convolution; Is the calculated input data of the next layer of convolution; Then, the feature maps generated by each convolution operation are weighted and summed as the input of the next convolution layer; After multiple convolutions, the extracted features are finally input into the fully connected layer to output the classification results.
[0015] According to the present invention, preferably, the recognition model Loss The function uses binary cross entropy loss, also known as log loss; it is shown below: ; in: is the value of the loss function, is the total number of samples, It is The true label of each sample is 0 or 1; It is The predicted probability of a sample, that is, the probability that the recognition model predicts that the current sample is a positive class, and the value range is (0,1).
[0016] The feature dataset to be tested is input into the trained classification model, and the recognition results of Internet game live video traffic are output.
[0017] In a second aspect, the present invention provides a system for identifying game live video traffic based on deep Gaussian random convolution; A system for identifying live game video traffic based on deep Gaussian random convolution, including: Traffic collection module: collects live broadcast traffic and marks game and non-game live broadcasts; Feature extraction module: Extracts video traffic from the collected live streaming traffic to obtain training and test sets of game and non-game live streaming video traffic; extracts histogram features from the game and non-game live streaming video traffic to be tested and trained to obtain the feature datasets to be tested and trained; Classification model training module: input the features of the optimized training set into the classification model, train the classification model, and obtain a trained classification model; Identification module: Inputs the feature dataset to be tested into the trained classification model and outputs the identification results of the live game video traffic.
[0018] In a third aspect, the present invention provides an electronic device comprising a memory and a processor, and computer instructions stored in the memory and executed on the processor, wherein the computer instructions complete the method in the first aspect when executed by the processor.
[0019] In a fourth aspect, the present invention provides a computer-readable storage medium for storing computer instructions, which, when executed by a processor, complete the steps of the method in the first aspect.
[0020] Compared with the prior art, the present invention has the following beneficial effects: The present invention relates to a method and system for identifying live game video traffic based on deep Gaussian random convolution. Network traffic is viewed as a one-dimensional time series signal. From this perspective, the convolution operation in a convolutional neural network (CNN) can be viewed as a filtering process for the network signal. The parameters of the convolution kernel (filter) are initially randomly initialized based on a Gaussian distribution and remain fixed during training, making the convolution sum analogous to a traditional signal filter. This method generates diverse and rich convolutional features, helping to capture different patterns and characteristics in traffic signals. This reduces training parameters, improves interpretability, and enhances robustness. Furthermore, different convolution kernel sizes (i.e., different receptive fields) allow the model to capture contextual information of varying ranges. This hierarchical convolution kernel design enables the network to extract traffic features at different levels while maintaining sensitivity to local details. Specifically, smaller convolution kernels help capture local patterns and fine-grained features in traffic signals, while larger convolution kernels integrate broader contextual information and extract global patterns. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 This is a flow chart of a method for identifying live game video traffic based on deep Gaussian random convolution provided by Example 1 of the present invention; Figure 2 This is an architecture diagram of a system for identifying live game video traffic based on deep Gaussian random convolution, provided in Example 2 of the present invention.
[0022] Figure 3 This is a framework diagram of three convolutional layers in the recognition model provided in Example 1 of the present invention.
[0023] Figure 4 This is a framework diagram of the fully connected layer in the recognition model provided in Example 1 of the present invention.
[0024] Figure 5 This is the F1 score change graph of the recognition model training.
[0025] Figure 6 This is the confusion matrix diagram generated by the recognition model test set. DETAILED DESCRIPTION
[0026] The present invention will be further defined below with reference to the accompanying drawings and embodiments, but is not limited thereto.
[0027] It should be noted that the following detailed descriptions are exemplary and intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which the present invention belongs.
[0028] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present invention. As used herein, unless the context clearly indicates otherwise, the singular form is intended to include the plural form. In addition, it should be understood that when the terms "comprise" and / or "include" are used in this specification, they indicate the presence of features, steps, operations, devices, components and / or their combinations;
[0029] Example 1 A method for identifying live game video traffic based on deep Gaussian random convolution, such as Figure 1 Shown, including: Collect original network traffic data; collect original game and non-game live broadcast traffic within the campus network.
[0030] Specifically, the collection and analysis device is a host running a Windows system with its network card operating in local mode. Using the NPcap development toolkit, we developed a new tool to collect traffic. This tool can also capture traffic remotely, separating the collection and analysis ends and improving system efficiency.
[0031] More specifically, code was written using selenium and webdriver-manager libraries to simulate the user's behavior of opening a web live broadcast room, opening game and non-game live broadcast rooms respectively, and running a distributed data acquisition application (WinPcap remote capture server program) on the captured network device to transmit the captured traffic data to the capture device; 30 seconds of data traffic was collected for each live broadcast and saved in the form of a Pcap file to form a data set for game and non-game live broadcasts.
[0032] Extracting quintuple information and statistical feature information from raw network traffic data. The quintuple information includes the source IP address, source port, destination IP address, destination port, and protocol type of the data packet; the statistical feature information includes payload size and arrival time; obtaining the flow based on the quintuple information; the payload size is the size of the effective payload at the transport layer of the data packet, and the arrival time is the kernel time of the network device when the data packet is recorded by the network device on its storage medium; More specifically, the arrival time and total length of the data packet are obtained; the data link layer header information of the data packet is parsed according to the Ethernet data frame format, the network layer information of the data packet is parsed, the source IP address, destination IP address and transport layer protocol type information of the data packet are obtained, and the IP datagram header length and the total length including its effective load are obtained to calculate the payload size; More specifically, the transport layer information of the data packet is parsed to obtain the source port, destination port and transport layer header length information. Data packets with the same five-tuple information are a stream, and a live broadcast room will generate multiple such flows.
[0033] Based on the statistical characteristics of the flow, the duration and flow rate of the flow are used to perform coarse-grained identification of the flow, extract the video flow, and obtain the game and non-game live video flow training set; The histogram features are extracted from the training set video traffic, and the one-dimensional vector of the histogram features is input into the designed deep Gaussian random convolution game live video traffic recognition model for training; Collect the live video to be tested, extract features from the live video to be tested, and filter the extracted features to obtain a feature data set to be tested; The feature dataset to be tested is input into the trained recognition model, and the recognition results of Internet game live video traffic are output.
[0034] Data packets with the same five-tuple information constitute a flow.
[0035] Based on the statistical characteristics of the flow, namely the duration and flow rate (bytes per second), the flow is coarsely identified and the video traffic is extracted. This includes: Get the duration of each stream , and calculate the duration of each flow The number of bytes arriving in , as shown below: ; in, Indicates the obtained Flow rate, The weight corresponding to the duration of the flow and the number of bytes per second, represents the threshold, if The value is greater than or equal to , it is judged as video traffic, otherwise it is judged as non-video traffic. The threshold is set according to task requirements and experience. Obtained through sample training;
[0036] If multiple streams in a live broadcast session reach the threshold The final judgment is made through the domain name resolved in the DNS response package; the information resolved by the DNS response package includes IP, protocol, and domain name, and multiple streams in the live broadcast room are marked; and whether it is video traffic is determined from these domain names.
[0037] By capturing and parsing the DNS response packets related to live broadcast services, we extract a feature set of three elements: the target IP address cluster, the transport layer protocol type (TCP / UDP), and the associated domain name information. When the following conditions are met simultaneously in the network traffic, the traffic is determined to belong to a specific live broadcast service: (1) The target IP address is exactly the same as the DNS response record; (2) The transport layer protocol type matches the DNS preset protocol; (3) The time when the traffic occurs is within the TTL validity period of the DNS record.
[0038] Domain names for live video streams often contain characteristic video service fields such as "video," "flv," "live," and "cdn." Domain names are semantically parsed and traffic is marked when they contain characteristic fields (for example, the core word "video" is identified in "d1--ov-gotcha207.bliblivideo.com").
[0039] Because when accessing domestic web platforms, the DNS response content is usually resolvable. By parsing the information in the DNS response package, the resolved domain name and corresponding IP address can be used to mark each flow, thereby realizing the association identification of the flow and the domain name.
[0040] Extract histogram features from the training set video traffic; including: The histogram feature is constructed by extracting the arrival time and payload of the data packets from the statistical feature information of the flow. 1s is defined as the time interval, and the payload arriving within the 1s time interval is counted, with the total time length being 30s. A one-dimensional histogram feature vector of the time interval for payload arrival is obtained. This forms a training dataset for both gaming and non-gaming that includes the histogram features. The pseudo-random code is as follows:
[0041] Input: packet_times (a list of packet arrival timestamps in seconds) Output: packet_payload (a list of the number of packets arriving per second) 1. Initialize an array packet_payload of length 30, with all elements initially set to 0 / / packet_payload[i] represents the number of packets arriving in the i-th second 2. For each timestamp t in packet_times: a. If t is between 0 and 30 seconds: i. Calculate the seconds of the current timestamp t index = floor(t) / / Round down ii. If index is between 0 and 29: packet_payload[index] += 1 0. Output packet_payload Finish Design a deep Gaussian random convolution recognition model for live game video traffic; the recognition model consists of three convolutional layers and a fully connected layer, and each convolutional layer uses three different sizes of convolution kernels ( ), the size of each convolution kernel is randomly selected from three different sizes of convolution kernels; the parameters of the three different sizes of convolution kernels are randomly initialized based on Gaussian distribution, and all convolution kernel parameters are fixed and do not participate in backpropagation updates; In the recognition model, Figure 3As shown in the figure, the convolution kernel size of the three convolution layers increases layer by layer as the number of convolution layers increases; the convolution kernel size of the first convolution layer is 3, the convolution kernel size of the second convolution layer is 5, and the convolution kernel size of the third convolution layer is 7; when the input one-dimensional vector enters the three convolution layers, each convolution layer generates One-dimensional vector output; then for this The one-dimensional vectors are assigned initialization weights and weighted summation is performed to generate a new one-dimensional input, which is then put into the next convolution layer. Since the size of the convolution kernel increases layer by layer, unique data features can be extracted from each convolution layer. Finally, the extracted features are input into the fully connected layer to obtain the final classification result.
[0042] The one-dimensional vector of histogram features is input into the designed deep Gaussian random convolution game live video traffic recognition model for training; including: The dataset is divided into a training set and a test set (8:2). The training set is input into the recognition model. The feature map generated by each convolution operation is passed to the next layer through weighted summation. After multiple convolutions, the extracted features are finally input into the fully connected layer to output the classification results. Specifically, Assume the input data is ,in, It is Video traffic samples, is the time step; each convolutional layer has convolution kernels (or filters), where ; K is the number of convolution kernels in each layer; the size of each convolution kernel is Randomly select from, where is the length of the convolution kernel; For the convolution kernels , the corresponding bias term is ; In order to better extract the local and global features of the input signal, Padding The operation keeps the output length of the recognition model unchanged and extracts features at different scales without losing features due to boundary issues; as shown below: ; in, is the input data length; is the convolution kernel size; is the length that needs to be padded to ensure that the output length is equal to the input length; is the step size of the convolution kernel movement, which is set to 1 here; so: ; in, is the input data length; is the output data length; Output of the convolutional layer Calculated by the following formula: ; in, is the output data at time step Hedi m Layer The elements corresponding to the convolution kernels, is to proceed padding The input data after that is at time step The corresponding elements, It is m Layer The convolution kernel is at position Corresponding elements; the internal parameters of the convolution kernel are randomly initialized and fixed based on the Gaussian distribution and do not participate in the back propagation update; thus we get: ; ; in, is the output data calculated by the kth convolution kernel; is the value obtained by convolution at the Tth time step; It is i In the convolution layer, k The set of output data calculated by the convolution kernel; It is i In the convolution layer, k The output data calculated by the convolution kernel; initialization The weight corresponding to each output data ; From this we get: ; ; in, It is i The set of weights corresponding to each output data in the layer convolution; Is the calculated input data of the next layer of convolution; Then, the feature maps generated by each convolution operation are weighted and summed as the input of the next convolution layer; After multiple convolutions, Figure 4 As shown in the figure, the feature data obtained in the first three convolutional layers are flattened and then put into the fully connected layer for classification to output the classification results.
[0043] Identification model Loss The function uses binary cross entropy loss ( Binary Cross-Entropy Loss ), also known as logarithmic loss ( Log Loss ); as follows:
[0044] in: is the value of the loss function, is the total number of samples, It is The true label of each sample is 0 or 1; It is The predicted probability of a sample, that is, the probability that the recognition model predicts that the current sample is a positive class (label is 1), and the value range is (0,1).
[0045] The training hyper parameters are shown in Table 1 below; Table 1 Introduction to the parameters of the deep Gaussian random convolution model;
[0046] The feature dataset to be tested was input into the trained classification model, which then output the identification results for internet game live streaming video traffic. The classification model outputs the category of the stream to be tested, resulting in a predicted label: 1 for game live streaming and 0 for non-game live streaming. A total of 3,000 live video streams were collected, including 1,500 game live streams and 1,500 non-game live streams. The model achieved a classification accuracy of 93.44% and an F1 score of 0.9319.
[0047] Example 2 A recognition system for live game video traffic based on deep Gaussian random convolution, such as Figure 2 Shown, including: Traffic collection module: collects live broadcast traffic and marks game and non-game live broadcasts; Feature extraction module: Extracts video traffic from the collected live streaming traffic to obtain training and test sets of game and non-game live streaming video traffic; extracts histogram features from the game and non-game live streaming video traffic to be tested and trained to obtain the feature datasets to be tested and trained; Classification model training module: Input the features of the optimized training set into the classification model, train the classification model, and obtain a trained classification model; the F1 score is as follows Figure 5 As shown in Figure 3, with the increase in the number of iterations, the F1 score gradually increases and reaches about 0.94.
[0048] Identification module: Input the feature data set to be tested into the trained classification model and output the identification results of the game live video traffic. The confusion matrix of the test results is as follows Figure 6As shown, among 581 game live broadcast test samples, 551 were successfully predicted, and among 380 non-game live broadcast samples, 347 were successfully predicted, with an overall accuracy rate of 93.44%.
[0049] Example 3 An electronic device includes a memory and a processor, as well as computer instructions stored in the memory and executed on the processor. When the computer instructions are executed by the processor, the steps of a method for identifying live game video traffic based on deep Gaussian random convolution as described in Example 1 are completed.
[0050] Example 4 A computer-readable storage medium is used to store computer instructions. When the computer instructions are executed by a processor, the steps of the method for identifying game live video traffic based on deep Gaussian random convolution described in Example 1 are completed.
Claims
1. A method for identifying game live video traffic based on deep Gaussian random convolution, characterized in that: include: Collect raw network traffic data; Extract the five-tuple information and statistical feature information of the original network traffic data. The five-tuple information includes the source IP address, source port, destination IP address, destination port and protocol type information of the data packet; Statistical characteristic information includes load size and arrival time; Get the flow based on the five-tuple information; Based on the statistical characteristics of the flow, the duration and flow rate of the flow are used to perform coarse-grained identification of the flow, extract the video flow, and obtain the game and non-game live video flow training set; The histogram features are extracted from the training set video traffic, and the one-dimensional vector of the histogram features is input into the designed deep Gaussian random convolution game live video traffic recognition model for training; Collect the live video to be tested, extract features from the live video to be tested, and filter the extracted features to obtain a feature data set to be tested; Input the feature dataset to be tested into the trained recognition model and output the recognition results of Internet game live video traffic; Further preferably, data packets with the same quintuple information constitute one flow.
2. A method for identifying game live video traffic based on deep Gaussian random convolution according to claim 1, characterized in that: Based on the statistical characteristics of the flow, namely the duration and flow rate (bytes per second), the flow is coarsely identified and the video traffic is extracted. This includes: Get the duration of each stream , and calculate the duration of each flow The number of bytes arriving in , as shown below: ; in, Indicates the obtained Flow rate, The weight corresponding to the duration of the flow and the number of bytes per second, represents the threshold, if The value is greater than or equal to , it is judged as video traffic, otherwise it is judged as non-video traffic.
3. The method for identifying live game video traffic based on deep Gaussian random convolution according to claim 1 is characterized in that: If multiple streams in a live broadcast session reach the threshold , the final judgment is made through the domain name resolved in the DNS response package; the information resolved by the DNS response package includes IP, protocol, and domain name, and multiple streams in the live broadcast room are marked; Determine whether the traffic is video traffic from these domain names.
4. The method for identifying live game video traffic based on deep Gaussian random convolution according to claim 1 is characterized in that: Extract histogram features from the training set video traffic; including: The histogram feature is constructed by extracting the arrival time and payload of the data packets from the statistical feature information of the flow; the time interval is defined, and the payload arriving within the time interval is counted to obtain a one-dimensional histogram feature vector of the payload arriving within the time interval; and a training data set of games and non-games including the histogram features is formed.
5. The method for identifying live game video traffic based on deep Gaussian random convolution according to claim 1 is characterized in that: A recognition model for live game video traffic using deep Gaussian random convolution was designed. The recognition model consists of three convolutional layers and a fully connected layer. Each convolutional layer uses three convolution kernels of different sizes, and the kernel size is randomly selected from the three sizes. The parameters of the three kernels are randomly initialized based on a Gaussian distribution, and all kernel parameters are fixed and not updated through backpropagation. Further preferably, in the recognition model, the convolution kernel size increases layer by layer as the number of convolution layers increases; the convolution kernel size of the first convolution layer is 3, the convolution kernel size of the second convolution layer is 5, and the convolution kernel size of the third convolution layer is 7; when the input one-dimensional vector enters the three convolution layers, each convolution layer generates One-dimensional vector output; then for this The one-dimensional vectors are assigned initial weights and weighted summation is performed to generate a new one-dimensional input, which is then put into the next convolutional layer; finally, the extracted features are input into the fully connected layer to obtain the final classification result.
6. The method for identifying live game video traffic based on deep Gaussian random convolution according to claim 1 is characterized in that: The one-dimensional vector of histogram features is input into the designed deep Gaussian random convolution game live video traffic recognition model for training; including: The dataset is divided into a training set and a test set. The training set is input into the recognition model. The feature map generated by each convolution operation is passed to the next layer through weighted summation. After multiple convolutions, the extracted features are finally input into the fully connected layer to output the classification results. Specifically, it includes: Assume the input data is ,in, It is Video traffic samples, is the time step; each convolutional layer has convolution kernels, where ; K is the number of convolution kernels in each layer; the size of each convolution kernel is Randomly select from, where is the length of the convolution kernel; For the convolution kernels , the corresponding bias term is ; pass Padding The operation keeps the output length of the recognition model unchanged as follows: ; in, is the input data length; is the convolution kernel size; is the length that needs to be filled; Is the step size of the convolution kernel movement, so: ; in, is the input data length; is the output data length; Output of the convolutional layer Calculated by the following formula: ; in, is the output data at time step Hedi m Layer The elements corresponding to the convolution kernels, is to proceed padding The input data after that is at time step The corresponding elements, It is m Layer The convolution kernel is at position Corresponding elements; the internal parameters of the convolution kernel are randomly initialized and fixed based on the Gaussian distribution and do not participate in the back propagation update; thus we get: ; ; in, is the output data calculated by the kth convolution kernel; is the value obtained by convolution at the Tth time step; It is i In the convolution layer, k The set of output data calculated by the convolution kernel; It is i In the convolution layer, k The output data calculated by the convolution kernel; initialization The weight corresponding to each output data ; From this we get: ; ; in, It is i The set of weights corresponding to each output data in the layer convolution; Is the calculated input data of the next layer of convolution; Then, the feature maps generated by each convolution operation are weighted and summed as the input of the next convolution layer; After multiple convolutions, the extracted features are finally input into the fully connected layer to output the classification results.
7. A method for identifying live game video traffic based on deep Gaussian random convolution according to claims 1-6, characterized in that: Identification model Loss The function uses binary cross entropy loss, also known as log loss; it is shown below: ; in: is the value of the loss function, is the total number of samples, It is The true label of each sample is 0 or 1; It is The predicted probability of a sample, that is, the probability that the recognition model predicts that the current sample is a positive class, with a value range of (0,1); The feature dataset to be tested is input into the trained classification model, and the recognition results of Internet game live video traffic are output.
8. A system for identifying live game video traffic based on deep Gaussian random convolution, characterized in that: include: Traffic collection module: collects live broadcast traffic and marks game and non-game live broadcasts; Feature extraction module: Extracts video traffic from the collected live streaming traffic to obtain training and test sets of game and non-game live streaming video traffic; extracts histogram features from the game and non-game live streaming video traffic to be tested and trained to obtain the feature datasets to be tested and trained; Classification model training module: input the features of the optimized training set into the classification model, train the classification model, and obtain a trained classification model; Identification module: Inputs the feature dataset to be tested into the trained classification model and outputs the identification results of the live game video traffic.
9. An electronic device comprising a memory and a processor and computer instructions stored in the memory and executed on the processor, characterized in that When the computer instructions are executed by the processor, a method for identifying game live video traffic based on deep Gaussian random convolution as described in any one of claims 1-7 is completed.
10. A computer-readable storage medium for storing computer instructions, characterized in that: When the computer instructions are executed by the processor, a method for identifying game live video traffic based on deep Gaussian random convolution as described in any one of claims 1-7 is completed.