Model training method, code type identification method, electronic equipment and computer program product

Through pre-training and fine-tuning methods, the encoding identification model is trained using unlabeled and marked traffic data, which solves the problem that the model relies on labeled data in the prior art, and realizes the accurate identification and generalization of audio and video data encoding types.

CN120279892APending Publication Date: 2025-07-08HEFEI IFLY DIGITAL TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510236259.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

When identifying the encoding types of audio data and video data in network traffic data, the prior art highly relies on the number and distribution of labeled data, resulting in poor model deviation and generalization capabilities.

Method used

By pre-training using traffic data of unlabeled encoding type, a pre-trained model is obtained, and fine-tuning is performed through traffic data of marked encoding type, adding a coding recognition layer to form an encoded recognition model.

Benefits of technology

The generalization ability of the model is improved, the dependence on the number and distribution of training data is reduced, and the encoding type of audio data or video data in the traffic data can be accurately identified.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279892A_ABST
    Figure CN120279892A_ABST
Patent Text Reader

Abstract

The invention provides a model training method, a code type identification method, electronic equipment and a computer program product, which are applied to the technical field of code type identification. The code recognition model training method comprises the following steps: pre-training an initial network model by using traffic data of which a code type is not labeled to obtain a pre-trained model; and performing fine tuning on the pre-training model by using traffic data with a marked coding type to obtain a coding recognition model, the coding type comprising a coding format of audio data or video data in the traffic data. According to the method, pre-training is performed by means of the traffic data of which the coding type is not labeled, and the pre-training model is used for outputting the feature representation of the audio data or the video data and is irrelevant to the coding type of the audio data or the video data, so that the pre-training model has relatively high generalization ability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of coding type recognition, and in particular, to a method for training a model, a method for recognizing a coding type, an electronic device, and a computer program product. Background Art

[0002] Traffic data in a network can be understood as data transmitted in the network at a certain level. In some scenarios, such as audio call and video call scenarios, traffic data usually includes audio data and video data. These audio data and video data are usually data generated after encoding raw data. For example, audio data will be encoded into coding formats such as alaw (A-Law Encoding), mulaw (Mu-Law Encoding), opus (Opus Interactive Audio Codec), amrnb (Adaptive Multi-Rate Narrowband), etc., and video data will be encoded into coding formats such as H264, VP8, etc. Moreover, to enhance data security, these encoded data will undergo encryption processing. Thus, after obtaining the traffic data, its coding type / coding format cannot be directly obtained through the header file. However, in certain specific tasks, it is necessary to identify the coding types of the audio data and video data in these traffic data to restore the original voice and video.

[0003] Currently, most use deep learning recognition methods to identify the coding types of audio data or video data in traffic data. For example, a certain amount of training data is labeled, and the training data is used to train a neural network model so that it can capture discriminative and robust traffic representations from content-invisible and unbalanced traffic data for accurate classification, thereby determining the coding type.

[0004] However, the above method highly depends on the quantity and distribution of the labeled training data, which easily leads to model bias and poor generalization ability. Summary of the Invention

[0005] Based on the above technical status quo, the present application proposes a method for training a model, a method for recognizing a coding type, an electronic device, and a computer program product.

[0006] According to the first aspect of the embodiments of the present application, a method for training a coding recognition model is provided, and the method includes:

[0007] Pre-training an initial network model using traffic data with unlabeled coding types to obtain a pre-trained model, where the pre-trained model is used to output a feature representation of audio data or video data;

[0008] Fine-tune the pre-trained model using the traffic data with the labeled encoding type to obtain an encoding recognition model, where the encoding type includes: the encoding format of the audio data or video data in the traffic data.

[0009] Optionally, pre-train the initial network model using the traffic data with the unlabeled encoding type to obtain a pre-trained model capable of extracting semantic features, including:

[0010] Extract the payload data from the traffic data with the unlabeled encoding type and convert the payload data into character encoding;

[0011] Use the character encoding as a sentence pair to pre-train the initial network model to obtain a pre-trained model capable of extracting semantic features, and / or use the character encoding in units of payloads to pre-train the initial network model to obtain a pre-trained model capable of extracting payload features.

[0012] Optionally, use the character encoding as a sentence pair to pre-train the initial network model to obtain a pre-trained model capable of extracting semantic features, including:

[0013] Use the character encoding corresponding to each payload as a sentence, mask and predict part of the content of the sentence;

[0014] Calculate the first model loss based on the prediction results and update the model parameters in the initial network model according to the first model loss.

[0015] Optionally, use the character encoding as a sentence pair to pre-train the initial network model to obtain a pre-trained model capable of extracting semantic features, including:

[0016] Determine paired character encodings, where each pair of character encodings includes the character encodings corresponding to two payloads that are adjacent or non-adjacent in time series;

[0017] Use the paired character encodings as the first sentence and the second sentence, and predict whether the second sentence is the next sentence adjacent in time series according to the first sentence; where the next sentence adjacent in time series includes the character encoding corresponding to the next payload adjacent in time series to the target payload; the target payload includes the payload corresponding to the first sentence;

[0018] Calculate the second model loss based on the prediction results and update the model parameters in the initial network model according to the second model loss.

[0019] Optionally, use the character encoding in units of payloads to pre-train the initial network model to obtain a pre-trained model capable of extracting payload features, including:

[0020] Determine paired first character encodings and second character encodings, where the first character encoding and the second character encoding correspond to different payloads;

[0021] Predict whether the first character encoding and the second character encoding are from the same video file or audio file;

[0022] Calculate a third model loss based on the prediction result, and update the model parameters in the initial network model according to the third model loss.

[0023] Optionally, pre-train the initial network model with the character encodings in units of payloads to obtain a pre-trained model capable of extracting payload features, including:

[0024] Determine paired third character encodings and fourth character encodings, where the third character encoding corresponds to a payload and the fourth character encoding corresponds to a payload or a non-payload;

[0025] Predict whether both the third character encoding and the fourth character encoding correspond to payloads;

[0026] Calculate a fourth model loss based on the prediction result, and update the model parameters in the initial network model according to the fourth model loss.

[0027] Optionally, fine-tune the pre-trained model with traffic data with labeled encoding types to obtain an encoding recognition model, including:

[0028] Add an encoding recognition layer after the output layer of the pre-trained model;

[0029] Input the traffic data with labeled encoding types into the pre-trained model, obtain a target feature representation through the pre-trained model, and obtain a predicted encoding type based on the target feature representation through the encoding recognition layer; where the target feature representation includes: the feature representation of the audio data or video data in the traffic data with labeled encoding types;

[0030] Calculate a fifth model loss based on the predicted encoding type and the labeled encoding type, and update the model parameters of the encoding recognition layer according to the fifth model loss.

[0031] Optionally, obtain a predicted encoding type based on the target feature representation through the encoding recognition layer, including:

[0032] Process the frame-level target feature representation into a traffic data packet-level feature representation through the encoding recognition layer, and obtain a predicted encoding type based on the traffic data packet-level feature representation.

[0033] Optionally, the encoding recognition layer includes: an average pooling layer and a fully connected layer;

[0034] The average pooling layer is used to process the target feature representation at the frame level into a feature representation at the traffic data packet level;

[0035] The fully connected layer is used to obtain the predicted coding type based on the feature representation at the traffic data packet level.

[0036] According to the second aspect of the embodiments of the present application, a method for identifying a coding type is provided. The method includes:

[0037] Obtain target data transmitted in the network, where the target data includes: audio traffic data or video traffic data;

[0038] Input the target data into the coding recognition model to obtain the coding type of the audio data or video data in the target data output by the coding recognition model;

[0039] Among them, the coding recognition model is trained by the training method of the coding recognition model described in the first aspect above.

[0040] According to the third aspect of the embodiments of the present application, a training device for a coding recognition model is provided. The device includes:

[0041] A pre-training module is used to pre-train an initial network model with traffic data whose coding type is not labeled to obtain a pre-trained model, where the pre-trained model is used to output the feature representation of audio data or video data;

[0042] A fine-tuning module is used to fine-tune the pre-trained model with traffic data whose coding type has been labeled to obtain a coding recognition model, where the coding type includes: the coding format of the audio data or video data in the traffic data.

[0043] According to the fourth aspect of the embodiments of the present application, a device for identifying a coding type is provided. The device includes:

[0044] An acquisition module is used to acquire target data transmitted in the network, where the target data includes: audio traffic data or video traffic data;

[0045] An identification module is used to input the target data into the coding recognition model to obtain the coding type of the audio data or video data in the target data output by the coding recognition model;

[0046] Among them, the coding recognition model is trained by the training method of the coding recognition model described in the first aspect above.

[0047] According to a fifth aspect of the embodiments of the present application, an electronic device is provided, including a memory and a processor; the memory is connected to the processor and is used for storing programs; the processor is used for implementing the training method of the encoding recognition model as described in the first aspect or the method for recognizing the encoding type as described in the second aspect by running the programs in the memory.

[0048] According to a sixth aspect of the embodiments of the present application, a storage medium is provided. A computer program is stored on the storage medium, and when the computer program is run by a processor, the training method of the encoding recognition model as described in the first aspect or the method for recognizing the encoding type as described in the second aspect is implemented.

[0049] According to a seventh aspect of the embodiments of the present application, a computer program product is provided, including: a computer program, and when the computer program is executed by a processor, the training method of the encoding recognition model as described in the first aspect or the method for recognizing the encoding type as described in the second aspect is implemented.

[0050] In the embodiments of the present application, pre-training is performed by means of traffic data without labeled encoding types. The pre-training model is used to output feature representations of audio data or video data, regardless of the encoding types of the audio data or video data. Therefore, the pre-training model has strong generalization ability. Finally, in a fine-tuning manner, the pre-training model is fine-tuned by using traffic data with labeled encoding types, and an encoding recognition model can be obtained. The encoding recognition model can not only recognize the encoding types of audio data or video data in traffic data, but also has strong generalization ability. At the same time, since the training process of the pre-training model is actually a fine-tuning process of the model, the dependence on the quantity and distribution of training data can be reduced. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.

[0052] Figure 1 It is one of the schematic flowcharts of a method for training an encoding recognition model provided by an embodiment of the present application;

[0053] Figure 2 It is another schematic flowchart of a method for training an encoding recognition model provided by an embodiment of the present application;

[0054] Figure 3 It is the schematic flowchart of a method for recognizing an encoding type provided by an embodiment of the present application;

[0055] Figure 4 A structural schematic diagram of a training device for an encoding recognition model provided by an embodiment of the present application;

[0056] Figure 5 A structural schematic diagram of a device for recognizing an encoding type provided by an embodiment of the present application;

[0057] Figure 6 A structural schematic diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0058] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0059] Overview

[0060] As described in the background art, for the audio data and video data in the traffic data, their encoding formats usually cannot be directly obtained from the header file. For example, the audio data and video data in the traffic data are usually transmitted in the form of binary data, and it is impossible to intuitively distinguish their encoding types. Therefore, as a method for recognizing the encoding type, using a neural network model to automatically learn complex patterns from the original traffic can achieve a good recognition effect. However, this requires a certain amount of labeled training data.

[0061] However, the above method has a strong dependence on the quantity and distribution of the training data, which is very likely to cause model bias and is difficult to generalize to unseen data. Therefore, in order to ensure that the model has a good recognition effect, it takes a long time and effort to prepare the training data. However, the problem of its poor generalization ability still cannot be solved.

[0062] In view of the above technical status quo, an embodiment of the present application proposes a training method for an encoding recognition model. Pre-training is performed by means of traffic data without labeled encoding types. The pre-training model is used to output the feature representation of the audio data or video data in the model input, which has nothing to do with the encoding type of the audio data or video data. Therefore, the pre-training model has strong generalization ability. Finally, in a fine-tuning manner, the pre-training model is fine-tuned by using the traffic data with labeled encoding types, and an encoding recognition model can be obtained. The encoding recognition model can not only recognize the encoding type of the audio data or video data in the traffic data, but also has strong generalization ability. At the same time, since the training process of the pre-training model is actually a fine-tuning process of the model, the dependence on the quantity and distribution of the training data can be reduced.

[0063] Exemplary Method

[0064] In an exemplary embodiment, a method for training an encoding recognition model is provided. This method can be applied to any electronic device. For example, the electronic device can include a computer, a server, a smart phone, etc. As Figure 1 shown, the method for training the encoding recognition model can include:

[0065] S101: Pre-train an initial network model using traffic data with unlabeled encoding types to obtain a pre-trained model.

[0066] This step is the pre-training process. The traffic data with unlabeled encoding types is used as training data. Through the training data, pre-training can be carried out to obtain a pre-trained model that can output feature representations. Specifically, the pre-trained model is used to output the feature representations of audio data or video data. For example, if the input of the pre-trained model is traffic data containing audio data or video data, then its output is the feature representation of the audio data or video data in the traffic data. Here, the traffic data can be traffic data that can be transmitted in the network or traffic data used for transmission in the network, and the traffic data carries audio data or video data.

[0067] Regarding the pre-training process of the model and the network architecture of the initial network model, no limitations are made here. The network architecture of the initial network model can be the network architecture of any pre-trained model. For example, the pre-trained model can be a BERT (Bidirectional Encoder Representations from Transformers) network model, RoBERTa, a GPT network model, etc.

[0068] Since the traffic data used in the pre-training process does not need to be labeled, therefore, in this embodiment, a large amount of traffic data can be obtained for pre-training. For example, a large amount of traffic data can be obtained in the network, and the traffic data carrying audio data or video data can be screened out and used as traffic data with unlabeled encoding types for pre-training. Through pre-training with traffic data with unlabeled encoding types, the pre-trained model can well learn the characteristics of the audio data or video data in the traffic data.

[0069] S102: Fine-tune the pre-trained model using traffic data with labeled encoding types to obtain an encoding recognition model.

[0070] It should be noted that the coding type includes: the coding format of audio data or video data in the traffic data. Regarding the coding formats of audio data and video data, they can be any coding formats, whether commonly used or not commonly used at present, which will not be elaborated here.

[0071] It can be understood that the pre-trained model can extract the features of audio data or video data in the traffic data, but does not have the ability to identify the coding type. On this basis, by using the labeled training data (traffic data with the coding type labeled), fine-tuning the model can obtain a model that can identify the coding type while retaining the capabilities of the pre-trained model. Among them, the model fine-tuning process is the process of training the model with the labeled training data to enable it to have the ability to identify the coding type. In some embodiments, since the pre-trained model already has good capabilities after the pre-training stage. Therefore, when performing model fine-tuning, a small amount of traffic data with the coding type labeled can be used to fine-tune the pre-trained model.

[0072] In the embodiments of the present application, pre-training is performed by means of traffic data without the coding type labeled. The pre-trained model is used to output the feature representation of audio data or video data, which is independent of the coding type of the audio data or video data. Therefore, the pre-trained model has strong generalization ability. Finally, in a fine-tuning manner, the pre-trained model is fine-tuned by using the traffic data with the coding type labeled to obtain a coding recognition model. The coding recognition model can not only identify the coding type of audio data or video data in the traffic data, but also has strong generalization ability. At the same time, since the training process of the pre-trained model is actually the model fine-tuning process, the dependence on the quantity and distribution of the training data can be reduced.

[0073] In some embodiments of the present application, pre-training the initial network model by using traffic data without the coding type labeled to obtain a pre-trained model capable of extracting semantic features includes:

[0074] Extracting the payload data in the traffic data without the coding type labeled and converting the payload data into character encoding;

[0075] Using the character encoding as a sentence pair to pre-train the initial network model to obtain a pre-trained model capable of extracting semantic features, and / or using the character encoding in units of payload to pre-train the initial network model to obtain a pre-trained model capable of extracting payload features.

[0076] It should be noted that traffic data is usually transmitted over the network in the form of traffic data packets (packets). Among them, a traffic data packet is usually composed of a five-tuple, which includes the source IP address, source port, destination IP address, destination port, and the 4-layer communication protocol. The encoded and encrypted data is mainly included in the transport layer payload in the communication protocol. That is, audio data and video data are located in the payload. For the convenience of pre-training, first extract the required payload data (payload data of each traffic data packet) from the traffic data, and then convert it into character encoding, and based on this, perform pre-training.

[0077] In some embodiments, the payload data in the traffic data is usually binary data. In the process of converting the payload data into character encoding, it can be first converted into hexadecimal data, and then every four hexadecimal numbers are regarded as one character encoding, and finally the character encoding of the traffic data is obtained. For example, if the payload data is 0110101010111010001, when performing hexadecimal processing, every 4-bit binary number is converted into one hexadecimal number, so that each byte can be converted into two hexadecimal numbers. Then combine the data of two bytes to obtain a 4-bit hexadecimal number, such as "0A4B 812DC52A". Finally, every 4-bit hexadecimal number is regarded as one character encoding, and its numerical range can be from 0 to 65535, that is, the encoding (character encoding) of all data samples is completed and limited within a limited range.

[0078] It can be understood that the encoding characteristics or features of the audio data or video data in the traffic data can be directly extracted from the character encoding. If the character encoding corresponding to a payload is regarded as a statement / sentence. Then, with the help of natural language processing solutions, the process of extracting encoding characteristics can be regarded as the process of extracting semantic features. Therefore, the character encoding can be used as a statement to pre-train the initial network model. Among them, the character encoding corresponding to each payload can be used as a statement.

[0079] It can be understood that there are also temporal or contextual relationships between the payloads in the traffic data. For example, adjacent speech frames or video frames in time are contextually related, and this information is also of great significance for payload modeling and encoding recognition. For this reason, the character encoding can also be used to pre-train the initial network model in units of payloads, and train the model to understand and learn the association relationships, contextual features, and features that distinguish payloads from non-payloads.

[0080] In the embodiments of the present application, a natural language processing solution can be used to regard the process of extracting encoding features as the process of extracting semantic features for pre-training. It is also possible to consider the correlation relationship between the payloads in the traffic data, and learn the relevant information between the payloads during the pre-training process, so that the pre-trained model can learn the representation of the traffic data from multiple aspects.

[0081] In some embodiments of the present application, character encoding is used to pre-train the initial network model of the sentence pair to obtain a pre-trained model capable of extracting semantic features, including:

[0082] Regarding the character encoding corresponding to each payload as a sentence, masking and predicting part of the content of the sentence;

[0083] Calculating a first model loss based on the prediction result, and updating the model parameters in the initial network model according to the first model loss.

[0084] It should be noted that the natural language processing solution of masking prediction can well extract semantic features. Therefore, in this embodiment, the character encoding corresponding to each payload is regarded as a sentence, part of the character encoding is masked and predicted, and the model parameters are updated according to the first model loss calculated from the prediction result.

[0085] In some embodiments, when using character encoding to implement masking prediction, part of the character encoding (which can also be regarded as a word) can be masked, and prediction is performed by aligning with other character encodings. It can be regarded as a character-level classification task to learn the relationship between each character encoding. For example, all character encodings have a 15% probability of being randomly replaced with a mask. Among these 15% of the character encodings selected as masks, 80% of the probability is that they are really replaced with a special mask symbol ([MASK]), 10% of the probability is replaced with a random token, and finally 10% of the probability is not replaced at all. The loss function is designed as shown in Formula 1 below:

[0086] Formula 1: m i ∈[1,2,…|V|];

[0087] In Formula 1, L mask (θ,θ1) represents the model loss, θ represents the parameters of the Encoder part when using the BERT model as the pre-trained model; θ1 represents the parameters in the fully connected output layer connected after the Encoder; the set of masked words is M. Since it is a multi-classification problem on a vocabulary size |V|, the loss function can be called the negative log-likelihood function. m represents the prediction result, m i represents the true label of the i-th masked character, and p represents the probability that m is the same as m i the same.

[0088] In the embodiments of the present application, pre-training is carried out by means of a mask prediction scheme, so that the pre-trained model can extract semantic features well.

[0089] In some embodiments of the present application, character encoding is used to pre-train the initial network model of the sentence pair to obtain a pre-trained model capable of extracting semantic features, including:

[0090] Determine paired character encodings, where each pair of character encodings includes character encodings corresponding to two payloads that are adjacent or non-adjacent in time series;

[0091] Use the paired character encodings as the first sentence and the second sentence, and predict whether the second sentence is the next sentence adjacent in time series according to the first sentence; among them, the next sentence adjacent in time series includes the character encoding corresponding to the next payload adjacent in time series to the target payload; the target payload includes the payload corresponding to the first sentence;

[0092] Calculate the second model loss based on the prediction result, and update the model parameters in the initial network model according to the second model loss.

[0093] It should be noted that the natural language processing scheme for next sentence prediction can extract semantic features well. Therefore, in this embodiment, the character encoding corresponding to each payload is used as a sentence / statement. Thus, multiple payloads in a piece of traffic data are regarded as multiple statements in a paragraph. Whether different payloads belong to adjacent data in time series can be obtained through time information, and the time information in the payload is discarded. Between two payloads, a prediction target of whether it belongs to the "next sentence" is constructed, and it can be learned whether two payloads belong to adjacent data in time series. This method can enable the model to learn the time series relationship, causal relationship, etc. between payloads. In some embodiments, when determining paired character encodings, the character encodings corresponding to two adjacent payloads in time series can be combined as positive examples, and the character encodings corresponding to two non-adjacent payloads in time series can be combined as negative examples.

[0094] In some embodiments, when using character encoding to implement next sentence prediction, whether different data packets in the same piece of traffic data belong to adjacent data in time series can be obtained through time information, the payload in the data packet is extracted, and the time information is discarded. Between two payloads, a prediction target of whether it belongs to the "next sentence" is constructed. For example, there are two payloads, a and b, in an input sequence. There is a 50% probability that b is really after a, and there is also a 50% probability that b is a payload randomly selected from other places. Then, this means that 50% of the samples are positive examples (the two payloads are adjacent), and 50% of the samples are negative examples (the two payloads have no great relationship). The loss function is designed as shown in Formula 2, Formula 2:

[0095] n i∈[Is the next sentence, not the next sentence];

[0096] In Equation (2), Lnp(θ,θ2) represents the model loss, θ represents the parameters of the Encoder part when using the BERT model as the pre-trained model; θ2 represents the parameters in the fully connected output layer connected after the Encoder; the set of payloads for pre-training is N. Since it is a binary classification problem of whether it is the next sentence, the loss function used can be the negative log-likelihood function, n represents the prediction result, and n j represents the true label of the j-th payload pair. p represents the probability that n is the same as n j Same probability.

[0097] In the embodiments of the present application, the temporal relationship between payloads can be modeled, enabling the pre-trained model to learn payload-level statistical and semantic information, and the feature representation output by the pre-trained model to be more comprehensive.

[0098] In some embodiments of the present application, the initial network model is pre-trained with character encodings in units of payloads to obtain a pre-trained model capable of extracting payload features, including:

[0099] Determine paired first character encodings and second character encodings, where the first character encodings and the second character encodings correspond to different payloads;

[0100] Predict whether the first character encoding and the second character encoding come from the same video file or audio file;

[0101] Calculate a third model loss based on the prediction result, and update the model parameters in the initial network model according to the third model loss.

[0102] It should be noted that by distinguishing whether different payloads are from the same source, that is, whether they belong to the same video file or the same audio file, the distinguishability between different audio or videos can be learned, similar to the distinguishability of voiceprints in audio and scenes in videos, so as to model the attributes of traffic data from a higher level, making the feature representation output by the pre-trained model more comprehensive. In some embodiments, when determining paired first character encodings and second character encodings, the character encodings corresponding to two payloads from the same source can be combined as positive examples, and the character encodings corresponding to two payloads from different sources can be combined as negative examples.

[0103] In some embodiments, when implementing homology prediction using character encoding, an audio file (such as an audio-video call) or a video file consists of several traffic data packets, and the payload is contained in the traffic data packets. When constructing pre-training data, information on whether two payloads are homologous can be obtained. Since a large number of different audio files or video files can be collected during pre-training, a training objective for such homology prediction can be constructed. For example, in an input sequence, there are two payloads: a and b. There is a 50% probability that b and a belong to the same audio file or video file, and there is also a 50% probability that b is a payload selected from other audio files or video files. This means that 50% of the samples are positive examples and 50% of the samples are negative examples. The designed loss function is shown in Formula 3 below:

[0104] Formula 3: x k ∈ [Homologous, Non-homologous];

[0105] In Formula 3, Lss(θ,θ3) represents the model loss. θ represents the parameters of the Encoder part when using the BERT model as the pre-training model; θ3 represents the parameters in the fully connected output layer connected after the Encoder. The set of payloads for pre-training homology prediction is U. Since this is a binary classification problem of whether two payloads are homologous, the loss function used can be the negative log-likelihood function. x represents the prediction result, and x k represents the true label of the kth payload pair. p represents the probability that x and x k are the same.

[0106] In the embodiments of the present application, the traffic attributes between different audio data or video data can be modeled, enabling the pre-training model to learn the distinguishability between different audio data or video data, and the feature representation output by the pre-training model is more comprehensive.

[0107] In some embodiments of the present application, the initial network model is pre-trained using character encoding in units of payloads to obtain a pre-training model capable of extracting payload features, including:

[0108] Determine paired third character encoding and fourth character encoding, where the third character encoding corresponds to the payload, and the fourth character encoding corresponds to the payload or non-payload;

[0109] Predict whether both the third character encoding and the fourth character encoding correspond to the payload;

[0110] Calculate the fourth model loss based on the prediction result, and update the model parameters in the initial network model according to the fourth model loss.

[0111] It should be noted that in order to better learn the coding characteristics of the payload level and distinguish it from other non-payload data, a training objective for payload detection and prediction can be constructed. That is, to learn whether both of two payloads are normal payloads in traffic data containing audio data or video data. In some embodiments, when determining the paired third character encoding and fourth character encoding, the character encodings corresponding to two normal traffic payloads can be combined as positive examples, and the character encoding corresponding to one normal traffic payload and the character encoding corresponding to one abnormal traffic payload can be combined as negative examples. Among them, the traffic payload is the payload in traffic data. In some embodiments, the normal traffic payload includes the payload in traffic data, and the abnormal traffic payload includes the non-payload part in traffic data or randomly generated character data.

[0112] In some embodiments, when using character encoding to implement payload detection and prediction, assume that there are two payloads, a and b, in an input sequence. Then there is a 50% probability that both a and b are normal traffic payloads; there is also a 50% probability that one of a and b is not a normal traffic payload. Among them, when making abnormal traffic payloads, 70% of the probability can be selected to randomly intercept the non-payload part from traffic data packets, and 30% of the probability can be selected to generate random characters. The designed loss function is shown in Formula 4, Formula 4:

[0113] y q ∈ [both are payloads, non-payload];

[0114] In Formula 4, Lss(θ,θ4) represents the model loss, θ represents the parameters of the Encoder part when using the BERT model as the pre-training model; θ4 represents the parameters in the fully connected output layer connected after the Encoder. The payload set for pre-training for payload detection and prediction is V. It is a binary classification problem of whether two data are both payloads. The loss function used can be the negative log-likelihood function, y represents the prediction result, and y q represents the true label of the qth payload pair. p represents the probability that y is the same as y q the same.

[0115] In the embodiments of the present application, the statistical characteristics and attribute features of the payload can be modeled, so that the pre-training model can learn the distinguishability between the payload and the non-payload, and thus the feature representation output by the pre-training model is more comprehensive.

[0116] In some embodiments, multiple of the above four pre-training purposes can be used for pre-training. For example, Figure 2 As shown, a large amount of network traffic data can obtain binary payload data after payload data extraction, and the binary payload data is processed to generate character encoding. Among them, Figure 2The binary data therein: 00001010 1100100 111101000000000, and the character encoding: 1b4d2b8c a3b5 are only examples. The process of generating character encoding using the binary payload data can be referred to the above relevant description and will not be elaborated here.

[0117] After obtaining the character encoding, pre-training is performed using the character encoding. Among them, the pre-training process includes predictions for four pre-training purposes: masked prediction, next sentence prediction, homology prediction, and payload detection prediction. Regarding the processes of masked prediction, next sentence prediction, homology prediction, and payload detection prediction, reference can be made to the processes of masked prediction, next sentence prediction, homology prediction, and payload detection prediction in the above embodiments and will not be elaborated here. The model loss in the pre-training process is calculated through the following Formula Five:

[0118] Loss = L mask (θ, θ1) + Lnp(θ, θ2) + Lss(θ, θ3) + Lss(θ, θ4);

[0119] In Formula Five, Loss represents the model loss, and L mask (θ, L1) represents the model loss in Formula One, Lnp(θ, θ2) represents the model loss in Formula Two, Lss(θ, θ3) represents the model loss in Formula Three, and Lss(θ, θ4) represents the model loss in Formula Four. For the explanations of relevant parameters, refer to Formulas One to Four above.

[0120] In some embodiments, the BERT network model can be used as the pre-training model. Then, during the pre-training and fine-tuning training using the character encoding, it can be first converted into embedding vectors, namely token embedding and position embedding. Among them, the token embedding can convert the character encoding into an embedding vector, and the position embedding records the position information of each character encoding.

[0121] In the embodiments of the present application, by processing traffic data to generate character encoding, and then through the methods of BERT pre-training and fine-tuning training, learning of massive Internet data is realized, the generalization of encoding type recognition is improved, and the requirement for labeled data is reduced. At the same time, when constructing the pre-training objective, for the payload data, masked prediction and next sentence prediction methods similar to BERT are adopted. At the same time, combined with the characteristics of traffic data, homology pre-training and payload detection pre-training are proposed. Through homology pre-training, the distinguishability between different calls can be learned, and the traffic attributes can be modeled from a higher level; through payload detection pre-training, the distinguishability between traffic payload and other non-payloads can be learned, and the statistical characteristics and attribute features of the payload can be modeled. Finally, the efficiency and accuracy of encoding type recognition are improved.

[0122] In some embodiments of the present application, a pre-trained model is fine-tuned using traffic data with labeled coding types to obtain a coding recognition model, including:

[0123] Add a coding recognition layer after the output layer of the pre-trained model;

[0124] Input the traffic data with labeled coding types into the pre-trained model to obtain a target feature representation through the pre-trained model, and obtain a predicted coding type based on the target feature representation through the coding recognition layer; wherein, the target feature representation includes: the feature representation of the audio data or video data in the traffic data with labeled coding types;

[0125] Calculate a fifth model loss based on the predicted coding type and the labeled coding type, and update the model parameters of the coding recognition layer according to the fifth model loss.

[0126] It should be noted that a large amount of traffic data is learned through the above pre-training. Therefore, only a small amount of labeled data is required in the fine-tuning stage to obtain good recognition results, which can greatly reduce the demand for labeled data. In the process of model fine-tuning in this embodiment, the model architecture will be changed, that is, a coding recognition layer for identifying coding types will be added, and the original pre-trained model will be used as the feature extraction unit of the overall model. In some embodiments, in the process of fine-tuning, only the model parameters of the coding recognition layer can be adjusted, or the model parameters of the feature extraction unit and the coding recognition layer can be adjusted simultaneously.

[0127] It can be understood that the feature representation output by the feature extraction unit can well represent the audio data or video data in the traffic data. Furthermore, based on this, the coding recognition layer performs the recognition or classification of coding types, and the obtained result is the predicted coding type. The predicted coding type includes: the coding format of the audio data or video data in the predicted traffic data.

[0128] In the embodiments of the present application, in the process of model fine-tuning, the original pre-trained model is used as the feature extraction unit of the overall model, and a coding recognition layer for identifying coding types is added to achieve model fine-tuning.

[0129] In some embodiments of the present application, obtaining a predicted coding type based on the target feature representation through the coding recognition layer includes:

[0130] The coding recognition layer processes the frame-level target feature representation into a traffic data packet-level feature representation, and obtains a predicted coding type based on the traffic data packet-level feature representation.

[0131] It should be noted that for audio data or video data, it is usually composed of data frames. For example, video data is composed of video frames. After processing the audio data or video data into the target feature representation at the frame level, the amount of data of this level of target feature representation is usually very large. If directly using the target feature representation at the frame level to carry out classification tasks or coding type recognition tasks, a very complex network layer needs to be designed. To simplify the design of the subsequent network layer, it can be converted into the feature representation at the data traffic packet level, which is used as the basis for subsequent classification tasks. Compared with the target feature network at the frame level, the amount of data of the target feature network at the data traffic packet level has a significant decrease, so that a simple network layer can be designed.

[0132] In some embodiments, the model parameters obtained by pre-training are used to initialize the bottom layer of the entire network model. An average pooling layer is used to transform the frame-level feature representation output by the bottom layer into the feature representation at the data traffic packet level. After the average pooling layer, a fully connected layer is added as the class classifier for outputting the coding type. Among them, the final output of the model can be regularized by the softmax function, and the cross-entropy loss function can be used to guide the training to fine-tune the entire network parameters. Specifically, the coding recognition layer includes: an average pooling layer and a fully connected layer; the average pooling layer is used to process the target feature representation at the frame level into the feature representation at the data traffic packet level; the fully connected layer is used to obtain the predicted coding type based on the feature representation at the data traffic packet level.

[0133] In the embodiments of the present application, by processing the target feature representation at the frame level into the feature representation at the data traffic packet level through the coding recognition layer, the network architecture complexity of the coding recognition part in the coding recognition model can be simplified.

[0134] In another exemplary embodiment, a method for identifying a coding type is provided, as Figure 3 shown. The method for identifying a coding type includes:

[0135] S301: Obtain the target data transmitted in the network, where the target data includes: audio traffic data or video traffic data;

[0136] S302: Input the target data into the coding recognition model to obtain the coding type of the audio data or video data in the target data output by the coding recognition model;

[0137] Among them, the coding recognition model is trained by the training method of the coding recognition model provided in the above embodiments.

[0138] It should be noted that the audio traffic data includes traffic data carrying audio data, and the video traffic data includes traffic data carrying video data. The encoding type is the encoding format of audio data or video data. For the process of training the encoding recognition model through the training method of the encoding recognition model, refer to the various embodiments of the above-mentioned training method of the encoding recognition model, which will not be elaborated here.

[0139] In the embodiments of the present application, pre-training is performed by means of traffic data without labeled encoding types. The pre-training model is used to output the feature representation of audio data or video data in the model input, regardless of the encoding type of the audio data or video data. Therefore, the pre-training model has strong generalization ability. Finally, in a fine-tuning manner, the pre-training model is fine-tuned by using traffic data with labeled encoding types, and an encoding recognition model can be obtained. The encoding recognition model can not only identify the encoding type of audio data or video data in the traffic data, but also has strong generalization ability. At the same time, since the training process of the pre-training model is actually the fine-tuning process of the model, the dependence on the quantity and distribution of the training data can be reduced.

[0140] Exemplary Apparatus

[0141] Correspondingly, the embodiments of the present application also provide a training device for an encoding recognition model. Refer to Figure 4 As shown, the training device for the encoding recognition model includes:

[0142] A pre-training module 401, configured to pre-train an initial network model by using traffic data without labeled encoding types to obtain a pre-training model, where the pre-training model is used to output the feature representation of audio data or video data;

[0143] A fine-tuning module 402, configured to fine-tune the pre-training model by using traffic data with labeled encoding types to obtain an encoding recognition model, where the encoding type includes: the encoding format of audio data or video data in the traffic data.

[0144] In some embodiments, the pre-training module 401 includes:

[0145] An extraction unit, configured to extract the payload data in the traffic data without labeled encoding types and convert the payload data into character encoding;

[0146] A pre-training unit, configured to use the character encoding as a statement to pre-train the initial network model to obtain a pre-training model capable of extracting semantic features, and / or use the character encoding in units of payloads to pre-train the initial network model to obtain a pre-training model capable of extracting payload features.

[0147] In some embodiments, the pre-training unit is specifically configured to use the character encoding corresponding to each payload as a sentence, mask and predict part of the content of the sentence; calculate a first model loss based on the prediction result, and update the model parameters in the initial network model according to the first model loss.

[0148] In some embodiments, the pre-training unit is specifically configured to:

[0149] Determine paired character encodings, where each pair of character encodings includes character encodings corresponding to two payloads that are adjacent or non-adjacent in time series;

[0150] Use the paired character encodings as the first sentence and the second sentence, and predict whether the second sentence is the next sentence adjacent in time series according to the first sentence; among them, the next sentence adjacent in time series includes the character encoding corresponding to the next payload adjacent in time series to the target payload; the target payload includes the payload corresponding to the first sentence;

[0151] Calculate a second model loss based on the prediction result, and update the model parameters in the initial network model according to the second model loss.

[0152] In some embodiments, the pre-training unit is specifically configured to:

[0153] Determine paired first character encoding and second character encoding, where the first character encoding and the second character encoding correspond to different payloads;

[0154] Predict whether the first character encoding and the second character encoding come from the same video file or audio file;

[0155] Calculate a third model loss based on the prediction result, and update the model parameters in the initial network model according to the third model loss.

[0156] In some embodiments, the pre-training unit is specifically configured to:

[0157] Determine paired third character encoding and fourth character encoding, where the third character encoding corresponds to a payload, and the fourth character encoding corresponds to a payload or a non-payload;

[0158] Predict whether both the third character encoding and the fourth character encoding correspond to payloads;

[0159] Calculate a fourth model loss based on the prediction result, and update the model parameters in the initial network model according to the fourth model loss.

[0160] In some embodiments, the fine-tuning module 402 includes:

[0161] An architecture adjustment unit, configured to add an encoding recognition layer after the output layer of the pre-trained model;

[0162] The first fine-tuning unit is configured to input the traffic data with the labeled encoding type into a pre-trained model, obtain a target feature representation through the pre-trained model, and obtain a predicted encoding type based on the target feature representation through an encoding recognition layer; wherein, the target feature representation includes: the feature representation of the audio data or video data in the traffic data with the labeled encoding type.

[0163] The second fine-tuning unit is configured to calculate a fifth model loss based on the predicted encoding type and the labeled encoding type, and update the model parameters of the encoding recognition layer according to the fifth model loss.

[0164] In some embodiments, the first fine-tuning unit obtains the predicted encoding type based on the target feature representation through the encoding recognition layer, including:

[0165] Processing the frame-level target feature representation into a traffic data packet-level feature representation through the encoding recognition layer, and obtaining the predicted encoding type based on the traffic data packet-level feature representation.

[0166] In some embodiments, the encoding recognition layer includes: an average pooling layer and a fully connected layer;

[0167] The average pooling layer is configured to process the frame-level target feature representation into a traffic data packet-level feature representation; the fully connected layer is configured to obtain the predicted encoding type based on the traffic data packet-level feature representation.

[0168] The training device of the encoding recognition model provided in this embodiment belongs to the same inventive concept as the training method of the encoding recognition model provided in the above embodiments of the present application, and can execute the training method of the encoding recognition model provided in any of the above embodiments of the present application, and has the corresponding functional modules and beneficial effects of the execution method. For the technical details not described in detail in this embodiment, reference can be made to the specific processing content of the training method of the encoding recognition model provided in the above embodiments of the present application, which will not be elaborated here.

[0169] Correspondingly, an embodiment of the present application further provides a device for identifying an encoding type. Refer to Figure 5 As shown, the device for identifying an encoding type includes:

[0170] An acquisition module 501, configured to acquire target data transmitted in a network, where the target data includes: audio traffic data or video traffic data;

[0171] An identification module 502, configured to input the target data into an encoding recognition model, and obtain the encoding type of the audio data or video data in the target data output by the encoding recognition model;

[0172] Wherein, the encoding recognition model is trained by the training method of the encoding recognition model provided in the above embodiments.

[0173] The device for identifying the coding type provided in this embodiment belongs to the same inventive concept as the method for identifying the coding type provided in the above embodiments of the present application. It can execute the method for identifying the coding type provided in any of the above embodiments of the present application, and has the corresponding functional modules and beneficial effects for executing the method. For the technical details not described in detail in this embodiment, reference can be made to the specific processing content of the method for identifying the coding type provided in the above embodiments of the present application, which will not be elaborated here.

[0174] It should be understood that the modules in the above device can be implemented in the form of a processor calling software. For example, the device includes a processor, the processor is connected to a memory, and instructions are stored in the memory. The processor calls the instructions stored in the memory to implement any of the above methods or the functions of each unit of the device. The processor can be a general-purpose processor, such as a CPU or a microprocessor, etc., and the memory can be a memory inside the device or a memory outside the device. Alternatively, the units in the device can be implemented in the form of hardware circuits. By designing the hardware circuits, the functions of some or all of the units can be realized. The hardware circuits can be understood as one or more processors. For example, in one implementation, the hardware circuit is an ASIC, and the functions of some or all of the above units are realized by designing the logical relationships of the components in the circuit. Another example, in another implementation, the hardware circuit can be implemented by a PLD. Taking an FPGA as an example, it can include a large number of logic gate circuits, and the connection relationships between the logic gate circuits are configured through a configuration file, so as to realize the functions of some or all of the above units. All units of the above device can be all implemented in the form of a processor calling software, or all implemented in the form of hardware circuits, or some implemented in the form of a processor calling software, and the remaining part implemented in the form of hardware circuits.

[0175] In the embodiments of the present application, a processor is a circuit with the ability to process signals. In one implementation, the processor can be a circuit with the ability to read and execute instructions, such as a CPU, a microprocessor, a GPU, or a DSP, etc. In another implementation, the processor can realize certain functions through the logical relationships of hardware circuits, and the logical relationships of the hardware circuits are fixed or can be reconstructed. For example, the processor is a hardware circuit implemented by an ASIC or a PLD, such as an FPGA. In a reconfigurable hardware circuit, the process of the processor loading a configuration document to realize the configuration of the hardware circuit can be understood as the process of the processor loading instructions to realize the functions of some or all of the above units. In addition, it can also be a hardware circuit designed for artificial intelligence, which can be understood as a kind of ASIC, such as an NPU, a TPU, a DPU, etc.

[0176] It can be seen that each unit in the above device can be one or more processors (or processing circuits) configured to implement the above method. For example: CPU, GPU, NPU, TPU, DPU, microprocessor, DSP, ASIC, FPGA, or a combination of at least two of these processor forms.

[0177] In addition, each unit in the above device can be integrated in whole or in part, or can be independently implemented. In one implementation, these units are integrated together and implemented in the form of an SOC. The SOC can include at least one processor for implementing any of the above methods or the functions of each unit of the device. The types of the at least one processor can be different. For example, it includes a CPU and an FPGA, a CPU and an artificial intelligence processor, a CPU and a GPU, etc.

[0178] Exemplary Electronic Device

[0179] An embodiment of the present application provides an electronic device. Refer to Figure 6 As shown, the device includes:

[0180] A memory 600 and a processor 610;

[0181] Wherein, the memory 600 is connected to the processor 610 and is used for storing programs;

[0182] The processor 610 is configured to implement the training method of the encoding recognition model or the method of identifying the encoding type disclosed in any of the above embodiments by running the program stored in the memory 600.

[0183] Specifically, the above electronic device may further include: a bus, a communication interface 620, an input device 630, and an output device 640.

[0184] The processor 610, the memory 600, the communication interface 620, the input device 630, and the output device 640 are interconnected through the bus. Among them:

[0185] The bus may include a path for transmitting information between various components of the computer system.

[0186] The processor 610 can be a general-purpose processor, such as a general-purpose central processing unit (CPU), a microprocessor, etc., or an application-specific integrated circuit (ASIC), or one or more integrated circuits for controlling the execution of the program of the present invention. It can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0187] The processor 610 may include a main processor, and may also include a baseband chip, a modem, etc.

[0188] The memory 600 stores a program for implementing the technical solution of the present invention, and may also store an operating system and other key services. Specifically, the program may include program code, and the program code includes computer operation instructions. More specifically, the memory 600 may include a read-only memory (ROM), other types of static storage devices that can store static information and instructions, a random access memory (RAM), other types of dynamic storage devices that can store information and instructions, a disk memory, a flash memory, etc.

[0189] The input device 630 may include a device for receiving user input data and information, such as a keyboard, a mouse, a camera, a scanner, a light pen, a voice input device, a touch screen, a pedometer, or a gravity sensor, etc.

[0190] The output device 640 may include a device for allowing information to be output to the user, such as a display screen, a printer, a speaker, etc.

[0191] The communication interface 620 may include a device of any transceiver type for communicating with other devices or communication networks, such as Ethernet, a radio access network (RAN), a wireless local area network (WLAN), etc.

[0192] The processor 610 executes the program stored in the memory 600 and calls other devices, which can be used to implement each step of any one of the training methods of the encoding recognition model or the method of identifying the encoding type provided in the above embodiments of the present application.

[0193] An embodiment of the present application also proposes a chip, which includes a processor and a data interface. The processor reads and runs a program stored on a memory through the data interface to execute the training method of the encoding recognition model or the method of identifying the encoding type introduced in any of the above embodiments. The specific processing process and its beneficial effects can be referred to the embodiments of the training method of the encoding recognition model or the method of identifying the encoding type described above.

[0194] Exemplary Computer Program Product and Storage Medium

[0195] In addition to the above methods and devices, an embodiment of the present application may also be a computer program product, which includes computer program instructions. When the computer program instructions are run by a processor, the processor is caused to execute the steps in the training method of the encoding recognition model or the method of identifying the encoding type according to various embodiments of the present application described in any of the above embodiments of this specification.

[0196] The computer program product may be written in any combination of one or more programming languages for programming code to perform the operations of the embodiments of the present application. The programming languages include object-oriented programming languages such as Java, C++, etc., and also include conventional procedural programming languages such as the "C" language or similar programming languages. The programming code may be executed entirely on the user's computing device, partially on the user's device, executed as a stand-alone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.

[0197] In addition, an embodiment of the present application may also be a storage medium on which a computer program is stored, and the computer program is executed by a processor to perform the steps in the training method of the encoding recognition model or the method of recognizing the encoding type according to various embodiments of the present application described in any of the above embodiments of this specification.

[0198] For the foregoing method embodiments, for simplicity of description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present application is not limited by the described order of actions, because according to the present application, certain steps may be performed in other orders or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the present application.

[0199] It should be noted that the embodiments in this specification are all described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The same or similar parts between the embodiments can be referred to each other. For the device embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiments.

[0200] The steps in the methods of the embodiments of the present application can be adjusted, combined, and deleted according to actual needs. The technical features recorded in each embodiment can be replaced or combined.

[0201] The modules and sub-modules in the devices and terminals in the embodiments of the present application can be combined, divided, and deleted according to actual needs.

[0202] In several embodiments provided by the present application, it should be understood that the disclosed terminals, devices, and methods can be implemented in other ways. For example, the terminal embodiments described above are merely illustrative. For example, the division of modules or sub-modules is only a logical function division. In actual implementation, there may be other division methods. For example, multiple sub-modules or modules can be combined or integrated into another module, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling, direct coupling, or communication connection to each other can be an indirect coupling or communication connection through some interfaces, devices, or modules, and can be in electrical, mechanical, or other forms.

[0203] The modules or sub-modules described as separate components may or may not be physically separated. The components as modules or sub-modules may or may not be physical modules or sub-modules, that is, they can be located in one place, or can be distributed to multiple network modules or sub-modules. Some or all of the modules or sub-modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0204] In addition, each functional module or sub-module in various embodiments of the present application can be integrated in a processing module, or each module or sub-module can exist physically alone, or two or more modules or sub-modules can be integrated in one module. The above-mentioned integrated modules or sub-modules can be implemented in the form of hardware or in the form of software functional modules or sub-modules.

[0205] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed in this article can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.

[0206] The steps of the methods or algorithms described in combination with the embodiments disclosed in this article can be directly implemented by hardware, software units executed by a processor, or a combination of the two. The software units can be placed in a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium well-known in the technical field.

[0207] Finally, it should also be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.

[0208] The foregoing description of the disclosed embodiments enables those skilled in the art to implement or use the present application. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A training method for a coding recognition model, characterized in that, The method includes: Pre-training an initial network model using traffic data with unlabeled encoding types to obtain a pre-trained model, where the pre-trained model is used to output feature representations of audio data or video data; Fine-tuning the pre-trained model using traffic data with labeled encoding types to obtain an encoding recognition model, where the encoding types include: the encoding formats of audio data or video data in the traffic data.

2. The method according to claim 1, wherein Pre-training an initial network model using traffic data with unlabeled encoding types to obtain a pre-trained model capable of extracting semantic features, including: Extracting payload data from traffic data with unlabeled encoding types and converting the payload data into character encodings; Using the character encodings as sentences to pre-train the initial network model to obtain a pre-trained model capable of extracting semantic features, and / or using the character encodings in units of payloads to pre-train the initial network model to obtain a pre-trained model capable of extracting payload features.

3. The method according to claim 2, wherein Using the character encodings as sentences to pre-train the initial network model to obtain a pre-trained model capable of extracting semantic features, including: Using the character encoding corresponding to each payload as a sentence, masking and predicting part of the content of the sentence; Calculating a first model loss based on the prediction results and updating the model parameters in the initial network model according to the first model loss.

4. The method according to claim 2, wherein Using the character encodings as sentences to pre-train the initial network model to obtain a pre-trained model capable of extracting semantic features, including: Determining paired character encodings, where each pair of character encodings includes character encodings corresponding to two payloads that are adjacent or non-adjacent in time series; Using the paired character encodings as a first sentence and a second sentence, and predicting whether the second sentence is the next sentence adjacent in time series according to the first sentence; where the next sentence adjacent in time series includes the character encoding corresponding to the next payload adjacent in time series to the target payload; the target payload includes the payload corresponding to the first sentence; Calculating a second model loss based on the prediction results and updating the model parameters in the initial network model according to the second model loss.

5. The method according to claim 2, wherein Using the character encodings in units of payloads to pre-train the initial network model to obtain a pre-trained model capable of extracting payload features, including: Determining paired first character encodings and second character encodings, where the first character encoding and the second character encoding correspond to different payloads; Predicting whether the first character encoding and the second character encoding come from the same video file or audio file; Calculating a third model loss based on the prediction results and updating the model parameters in the initial network model according to the third model loss.

6. The method according to claim 2, characterized in that Using the character encodings in units of payloads to pre-train the initial network model to obtain a pre-trained model capable of extracting payload features, including: Determining paired third character encodings and fourth character encodings, where the third character encoding corresponds to a payload and the fourth character encoding corresponds to a payload or a non-payload; Predicting whether both the third character encoding and the fourth character encoding correspond to payloads; Calculate the fourth model loss based on the prediction result, and update the model parameters in the initial network model according to the fourth model loss.

7. The method according to claim 1, characterized in that, Fine-tune the pre-trained model using the traffic data with the labeled encoding type to obtain an encoding recognition model, including: Add an encoding recognition layer after the output layer of the pre-trained model; Input the traffic data with the labeled encoding type into the pre-trained model, obtain the target feature representation through the pre-trained model, and obtain the predicted encoding type based on the target feature representation through the encoding recognition layer; wherein, the target feature representation includes: the feature representation of the audio data or video data in the traffic data with the labeled encoding type; Calculate the fifth model loss based on the predicted encoding type and the labeled encoding type, and update the model parameters of the encoding recognition layer according to the fifth model loss.

8. The method according to claim 7, wherein Obtain the predicted encoding type based on the target feature representation through the encoding recognition layer, including: Process the frame-level target feature representation into a traffic data packet-level feature representation through the encoding recognition layer, and obtain the predicted encoding type based on the traffic data packet-level feature representation.

9. The method according to claim 8, wherein The encoding recognition layer includes: an average pooling layer and a fully connected layer; The average pooling layer is used to process the frame-level target feature representation into a traffic data packet-level feature representation; The fully connected layer is used to obtain the predicted encoding type based on the traffic data packet-level feature representation.

10. A method for identifying an encoding type, characterized in that, The method includes: Obtain the target data transmitted in the network, wherein the target data includes: audio traffic data or video traffic data; Input the target data into the encoding recognition model to obtain the encoding type of the audio data or video data in the target data output by the encoding recognition model; Wherein, the encoding recognition model is trained by the training method of the encoding recognition model according to any one of claims 1 to 9.

11. An electronic device, characterized in that, Includes a memory and a processor; The memory is connected to the processor and is used to store programs; The processor is used to implement the method according to any one of claims 1 to 10 by running the programs in the memory.

12. A computer program product, characterized in that, Includes: A computer program, which implements the method according to any one of claims 1 to 10 when executed by a processor.

Citation Information

Cited By

  • Hybrid coded data stream decoding method and related device

    CN121708946A

  • Hybrid encoded data stream decoding method and related apparatus

    CN121708946B