Encrypted call audio flow identification method based on spectral analysis
Through the encrypted call audio traffic recognition method based on spectrum analysis, packet length division and frequency domain feature extraction are performed on the encrypted voice traffic data, which solves the problems of poor universality and high training cost of encrypted call audio recognition methods in the prior art, and achieves high-accuracy voice traffic recognition.
Patent Information
- Application Number
- CN202510316605.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-17
- Publication Date
- 2025-07-08
AI Technical Summary
The existing encrypted call audio recognition methods are poor in universality among different call applications, the speech recognition model is expensive to train, and complex feature input leads to increased noise, and simple features require a large number of labeled samples.
Using a spectrum analysis method, encrypted voice traffic data is used to divide packet lengths and extract frequency domain features, accurately classify the audio traffic recognition model and classification recognition model, and use frequency characteristics and packet length attributes for speech traffic recognition.
It improves the universality and recognition accuracy of encrypted voice recognition methods, reduces the cost of model training, and realizes the precise classification of encrypted voice traffic data.
Smart Images

Figure CN120281512A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of encrypted audio recognition, and particularly relates to a method for identifying encrypted call audio traffic based on spectrum analysis, a method for training an audio traffic recognition model, a method for training a classification recognition model, a device, a storage medium, a device, and a computer program product. Background Art
[0002] With the rapid development of Internet technology, voice transmission over Internet protocol (VoIP) call applications for instant messaging are widely sought after by users due to their convenience and speed. To protect user call privacy, the traffic of these VoIPs is usually encrypted. However, based on a large number of market and user requirements, how to identify non-privacy content without decrypting the call content has become a concerned topic.
[0003] The existing encrypted call audio recognition methods mainly fall into three categories: language recognition methods based on statistical analysis, phrase recognition methods based on hidden Markov chains, and user identity recognition methods based on statistical machine learning. These call audio recognition software and hardware usually use a dedicated audio recognition network and are designed and developed based on open-source coding technology or specific call applications.
[0004] However, the above methods have poor universality among different call applications due to the cross-over between coding strategies. At the same time, in different application scenarios, higher noise will be brought by complex feature inputs, while simple features respectively require a large number of labeled samples of different types, which increases the difficulty of training the speech recognition model. Summary of the Invention
[0005] This application aims to provide a method for identifying encrypted call audio traffic based on spectrum analysis, a method for training an audio traffic recognition model, a method for training a classification recognition model, a device, a storage medium, a device, and a computer program product, which at least solves the problems of poor universality of existing encrypted speech recognition methods and high training cost of speech recognition models.
[0006] In a first aspect, an embodiment of this application discloses a method for identifying encrypted call audio traffic based on spectrum analysis, including: Performing a first partitioning of all the voice data packets according to the packet length of each voice data packet in the encrypted voice traffic data to generate a plurality of first voice data packet sequences; each of the first voice data packet sequences has a corresponding first packet length range; the range width of each of the first packet length ranges is the same; Input the generated multiple first voice data packet sequences into an audio traffic recognition model to generate multiple second voice data packet sequences through the second division of all the voice data packets by the audio traffic recognition model; each of the second voice data packet sequences has a corresponding second data packet length range; each of the second data packet length ranges is used to characterize the corresponding frequency range. Determine the classification result of the encrypted voice traffic data according to the extraction of the data characteristics of each second voice data packet sequence in the frequency domain.
[0007] In a second aspect, an embodiment of the present application also discloses an encrypted call audio traffic recognition method based on spectrum analysis, including: According to the data packet length of each voice data packet in the encrypted voice traffic data, perform a first division of all the voice data packets to generate multiple first voice data packet sequences; each of the first voice data packet sequences has a corresponding first data packet length range; the range widths of each of the first data packet length ranges are the same. Input the generated multiple first voice data packet sequences into an audio traffic recognition model to generate multiple second voice data packet sequences through the second division of all the voice data packets by the audio traffic recognition model; each of the second voice data packet sequences has a corresponding second data packet length range; each of the second data packet length ranges is used to characterize the corresponding frequency range. Determine the classification result of the encrypted voice traffic data according to the extraction of the data characteristics of each second voice data packet sequence in the frequency domain; The determining the classification result of the encrypted voice traffic data according to the extraction of the data characteristics of each second voice data packet sequence in the frequency domain includes: Extract multiple statistical metrics of the voice data packets in each second voice data packet sequence respectively to obtain the statistical feature data of the encrypted voice traffic data in each of the frequency ranges; Determine a target classification recognition model corresponding to the target classification requirement among multiple classification recognition models, and input all the statistical feature data of the obtained encrypted voice traffic data into the target classification recognition model to obtain the target classification result of the encrypted voice traffic data.
[0008] In a third aspect, an embodiment of the present application also discloses a training method for an audio traffic recognition model, which is used to train the audio traffic recognition model in the encrypted call audio traffic recognition method based on spectrum analysis as described in the first aspect or the second aspect, including: Obtain a first original training speech data set for training, and a first encrypted training speech data set corresponding to the first original training speech data set; the first original training speech data in the first original training speech data set corresponds one-to-one with the first encrypted training speech data in the first encrypted training speech data set; Determine the first audio spectrum feature of the first original training speech data set according to the first original training speech data in the first original training speech data set, and determine the first data packet length statistical feature of the first encrypted training speech data set according to the first encrypted training speech data in the first encrypted training speech data set; Train the audio traffic recognition model according to the first audio spectrum feature and the first data packet length statistical feature to obtain the trained audio traffic recognition model.
[0009] In a fourth aspect, an embodiment of the present application also discloses a training method for a classification recognition model, which is used to train the classification recognition model in the encrypted call audio traffic recognition method based on spectrum analysis as described in the second aspect, including: Obtain a second original training speech data set for training, and a second encrypted training speech data set corresponding to the second original training speech data set; the second original training speech data in the second original training speech data set corresponds one-to-one with the second encrypted training speech data in the second encrypted training speech data set; Determine the second audio spectrum feature of the second original training speech data set according to the second original training speech data in the second original training speech data set, and determine the second data packet length feature and data packet distribution feature of the second encrypted training speech data set according to the second encrypted training speech data in the second encrypted training speech data set; Train the classification recognition model according to the second audio spectrum feature, the second data packet length statistical feature, and the data packet distribution feature to obtain the trained classification recognition model.
[0010] In a fifth aspect, an embodiment of the present application also discloses an encrypted call audio traffic recognition device based on spectrum analysis, including: A first division module, configured to perform a first division on all the voice data packets according to the data packet length of each voice data packet in the encrypted voice traffic data, to generate a plurality of first voice data packet sequences; each of the first voice data packet sequences has a corresponding first data packet length range; the range width of each of the first data packet length ranges is the same; A second partitioning module, configured to input the generated multiple first voice data packet sequences into an audio traffic recognition model, so as to perform a second partitioning of all the voice data packets through the audio traffic recognition model, and generate multiple second voice data packet sequences; each of the second voice data packet sequences has a corresponding second data packet length range; each of the second data packet length ranges is respectively used to characterize the respective corresponding frequency range; A classification module, configured to determine a classification result of the encrypted voice traffic data according to the extraction of data characteristics of each of the second voice data packet sequences in the frequency domain.
[0011] In a sixth aspect, an embodiment of the present application further discloses a training device for an audio traffic recognition model, configured to train the audio traffic recognition model in the encrypted call audio traffic recognition method based on spectrum analysis as described in the first aspect or the second aspect, including: A first data set module, configured to obtain a first original training voice data set for training, and a first encrypted training voice data set corresponding to the first original training voice data set; the first original training voice data in the first original training voice data set corresponds one-to-one with the first encrypted training voice data in the first encrypted training voice data set; A first extraction module, configured to determine a first audio spectrum feature of the first original training voice data set according to the first original training voice data in the first original training voice data set, and determine a first data packet length statistical feature of the first encrypted training voice data set according to the first encrypted training voice data in the first encrypted training voice data set; A first training module, configured to train the audio traffic recognition model according to the first audio spectrum feature and the first data packet length statistical feature, so as to obtain the trained audio traffic recognition model.
[0012] In a seventh aspect, an embodiment of the present application further discloses a training device for a classification recognition model, configured to train the classification recognition model in the encrypted call audio traffic recognition method based on spectrum analysis as described in the second aspect, including: A second data set module, configured to obtain a second original training voice data set for training, and a second encrypted training voice data set corresponding to the second original training voice data set; the second original training voice data in the second original training voice data set corresponds one-to-one with the second encrypted training voice data in the second encrypted training voice data set; A second extraction module, configured to determine second audio spectrum features of the second original training speech dataset according to the second original training speech data in the second original training speech dataset, and determine second packet length features and packet distribution features of the second encrypted training speech dataset according to the second encrypted training speech data in the second encrypted training speech dataset; A second training module, configured to train the classification and recognition model according to the second audio spectrum features, the second packet length statistical features, and the packet distribution features, so as to obtain the trained classification and recognition model.
[0013] In a eighth aspect, an embodiment of the present application further discloses a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps described in the first aspect or the second aspect are implemented.
[0014] In a ninth aspect, an embodiment of the present application further discloses an electronic device, including a processor, a memory, and a computer program stored on the memory and executable on the processor, and when the computer program is executed by the processor, the steps described in the first aspect or the second aspect are implemented..
[0015] In a tenth aspect, an embodiment of the present application further discloses a computer program product, on which a computer program is stored, and when the computer program is executed by a processor, the steps described in the first aspect or the second aspect are implemented.
[0016] In summary, in the embodiment of the present application, through the second division of all speech data packets by the audio traffic recognition model, the frequency feature representation of the speech data packets is refined, so that the speech data packets in different frequency ranges can more accurately represent the frequency information of the speech traffic; furthermore, by extracting the frequency features in the second speech data packet sequence, the statistical feature data in different frequency ranges can be efficiently obtained, and two related data attributes that the occurrence features in different scenarios have different frequency features and the speech data packets under different frequency features have different packet lengths are utilized, thereby realizing the accurate classification of encrypted speech traffic data. Therefore, based on the method of the embodiment of the present application, the universality and recognition accuracy of the model are improved, and the problem of poor universality of the existing encrypted speech recognition method is solved. In addition, the training method of the audio traffic recognition model and the training method of the classification and recognition model provided based on the embodiment of the present application solve the problem of high training cost of related models for speech recognition. Description of the Drawings
[0017] By reading the following detailed description of the preferred embodiments, various other advantages and benefits will become apparent to those of ordinary skill in the art. The drawings are only for the purpose of showing the preferred embodiments and are not considered to be a limitation of the present application. Moreover, throughout the drawings, the same reference symbols are used to represent the same components. In the drawings: Figure 1 is a statistical chart of the data packet length under multiple encrypted call software based on gender; Figure 2 is a flowchart of the steps of a method for identifying encrypted call audio traffic based on spectrum analysis provided by an embodiment of the present application; Figure 3 is a comparison chart of the data packet length and frequency distribution; Figure 4 is a flowchart of the steps of another method for identifying encrypted call audio traffic based on spectrum analysis provided by an embodiment of the present application; Figure 5 is a flowchart of the steps of a method for training an audio traffic recognition model provided by an embodiment of the present application; Figure 6 is a flowchart of the steps of a method for training a classification recognition model provided by an embodiment of the present application; Figure 7 is a complete audio traffic recognition process under an embodiment of the present application; Figure 8 is a schematic structural diagram of a device for identifying encrypted call audio traffic based on spectrum analysis provided by an embodiment of the present application; Figure 9 is a schematic structural diagram of a device for training an audio traffic recognition model provided by an embodiment of the present application; Figure 10 is a schematic structural diagram of a device for training a classification recognition model provided by an embodiment of the present application; Figure 11 is a block diagram of an electronic device provided by an embodiment of the present application; Figure 12 is a block diagram of another electronic device provided by an embodiment of the present application. Detailed Embodiments
[0018] The exemplary embodiments of the present application will be described in more detail below with reference to the drawings. Although the exemplary embodiments of the present application are shown in the drawings, it should be understood that the present application can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present application can be more thoroughly understood and the scope of the present application can be completely conveyed to those skilled in the art.
[0019] As Figure 1As shown, 1-a, 1-b, 1-c, and 1-d in the figure are respectively examples of the gender of the callers. They are the statistical results of the packet length distributions of several existing call audio on the market for the same call audio traffic under different call applications based on open-source coding technology or specific call application designs, without fully considering black-box coding technology and the coding strategies of non-open-source commercial call applications. It can be seen from the statistical results that the packet length can, to a certain extent, characterize the classification results of data frequencies.
[0020] Based on the above assumption, as Figure 2 shown, considering the impact of different call applications on bandwidth, this application retains the characteristics of variable bit rate (VBR) coding. According to the relationship between audio complexity and packet length, it filters call audio traffic to divide VoIP traffic data packets into different packet length intervals, thereby constructing a traffic frequency domain distribution of call audio and proposing an encrypted call audio traffic recognition method based on spectrum analysis. This is aimed at making the traffic frequency domain distribution of the same call audio content smoother in different call applications to address the traffic distribution differences under black-box coding technology, and further improving the portability of the model among various call applications. It specifically includes the following steps: Step 101: Generate multiple first voice data packet sequences by making a first division of all voice data packets according to the packet length of each voice data packet in the encrypted voice traffic data.
[0021] Among them, each first voice data packet sequence respectively has a corresponding first data packet length range; the range widths of each first data packet length range are the same.
[0022] In some embodiments of this application, in order to initially simplify the complexity of encrypted voice data and facilitate the extraction of subsequent frequency domain features, the packet lengths of all voice data packets are first statistically analyzed, and then they are divided into multiple first voice data packet sequences according to the statistical results. Each sequence respectively has a corresponding first data packet length range. This can independently process voice data packets with different length ranges. The data packet length range refers to the length value of each data packet. Through this division method, the voice traffic can be initially classified. In this way, multiple first voice data packet sequences are generated, enabling more accurate frequency domain feature extraction and classification recognition in subsequent steps.
[0023] In a specific example, based on the actually captured encrypted voice traffic data, the length of each voice data packet can be recorded and divided into multiple first voice data packet sequences. For example, the data packet lengths can be divided into ranges such as 40 - 50 bytes, 50 - 60 bytes, 60 - 70 bytes, etc., and each range corresponds to a first voice data packet sequence. After dividing in this way, it is possible to more clearly distinguish data packets in different length ranges and prepare for subsequent frequency domain feature extraction and classification recognition.
[0024] Step 102: Input the generated multiple first voice data packet sequences into the audio traffic recognition model to generate multiple second voice data packet sequences through the second division of all the voice data packets by the audio traffic recognition model.
[0025] Among them, each second voice data packet sequence has a corresponding second data packet length range respectively; each second data packet length range is respectively used to characterize the corresponding frequency range.
[0026] In some embodiments of the present application, in order to more precisely characterize the frequency characteristics of voice data packets and thus improve the classification accuracy of encrypted voice traffic data, multiple first voice data packet sequences will first be input into the audio traffic recognition model, and the model will perform a second division on these data packets to generate multiple second voice data packet sequences. Each second voice data packet sequence has a corresponding second data packet length range respectively, and these ranges are used to characterize the corresponding frequency ranges. Through such a division method, voice data packets in different frequency ranges can more precisely characterize the frequency information of voice traffic. The multiple second voice data packet sequences generated in this way can effectively distinguish voice characteristics of different frequencies, thus providing more accurate data support for subsequent classification recognition.
[0027] In a specific example, as Figure 3 shown, input the generated multiple first voice data packet sequences into the audio traffic recognition model. Assuming that gender is still used as a reference, the division of the data packet length for the frequency domain interval under different genders can be compared: for example, it can be divided into four ranges: high (H), relatively high (BH), relatively low (BL), and low (L), and four second voice data packet sequences are generated correspondingly. These second data packet length ranges respectively characterize different frequency ranges, enabling each second voice data packet sequence to more precisely reflect the voice characteristics within that frequency range.
[0028] Step 103: Determine the classification result of the encrypted voice traffic data according to the extraction of the data characteristics of each second voice data packet sequence in the frequency domain.
[0029] In some embodiments of the present application, in order to accurately classify voice data packets in different frequency ranges, thereby improving the accuracy of encrypted voice traffic recognition, the data features of the second voice data packet sequence in the frequency domain are extracted to obtain statistical feature data in different frequency ranges. These statistical feature data are used to determine the classification result of the encrypted voice traffic data, making the recognition process more accurate and efficient. The data features refer to various information of the voice data packet within the frequency range. By extracting and analyzing these features, accurate classification of the voice traffic data can be achieved. After performing this step, the classification result can effectively distinguish the voice features of different frequencies, thereby achieving accurate recognition of the encrypted voice traffic.
[0030] In a specific example, the data features of multiple second voice data packet sequences in the frequency domain are extracted. For example, the corresponding frequency ranges in each second voice data packet sequence are extracted, such as low frequency (40 - 60 Hz), medium frequency (60 - 80 Hz), high frequency (80 - 100 Hz), as well as the mean length, length deviation value, etc. of the data packets contained therein. Next, according to the extracted data features in the frequency domain, the statistical feature data in each frequency range are calculated. Finally, these statistical feature data are input into the classification and recognition model to determine the classification result of the encrypted voice traffic data.
[0031] In summary, in the embodiments of the present application, through the second division of all voice data packets by the audio traffic recognition model, the frequency feature representation of the voice data packets is refined, enabling the voice data packets in different frequency ranges to more accurately represent the frequency information of the voice traffic; furthermore, by extracting the frequency features in the second voice data packet sequence, the statistical feature data in different frequency ranges are efficiently obtained, utilizing the two associated data attributes that the occurrence features in different scenarios have different frequency features and the voice data packets under different frequency features have different packet lengths, thereby achieving accurate classification of the encrypted voice traffic data. Thus, based on the method of the embodiments of the present application, the universality and recognition accuracy of the model are improved, and the problem of poor universality of existing encrypted voice recognition methods is solved. In addition, the training method of the audio traffic recognition model and the training method of the classification and recognition model provided based on the embodiments of the present application solve the problem of high training cost of related models for voice recognition.
[0032] Figure 4 This is another encrypted call audio traffic recognition method based on spectrum analysis provided in this embodiment, specifically including the following steps: Step 201, denoise the encrypted voice traffic data and extract all data packets in the denoised encrypted voice traffic data.
[0033] In some embodiments of the present application, in order to reduce the interference of noise on voice data packets and thus improve the accuracy of subsequent frequency feature extraction and classification, a denoising technique will be first applied to process the encrypted voice traffic data to remove the noise and interference signals in the data. Then, all data packets in the denoised encrypted voice traffic data are extracted to ensure that the data packets are analyzed and processed in the denoised state. The denoising technique includes using filters or other signal processing methods to reduce or eliminate unnecessary noise in the data. The obtained denoised encrypted voice traffic data packets will provide a cleaner data basis for frequency domain feature extraction and classification recognition in the subsequent steps.
[0034] In a specific example, denoising processing is performed on the actually captured encrypted voice traffic data. For example, a low-pass filter can be used to process the data to remove high-frequency noise and interference signals. The denoised data packets contain clearer voice information, and then these denoised encrypted voice traffic data packets are extracted. The denoising processing effectively improves the quality of the data, thereby enhancing the accuracy and robustness of the entire speech recognition process.
[0035] Step 202: According to the packet lengths of each voice data packet in the encrypted voice traffic data, perform a first division on all the voice data packets to generate multiple first voice data packet sequences.
[0036] Among them, each first voice data packet sequence respectively has a corresponding first packet length range; the range widths of each first packet length range are the same.
[0037] The method shown in this step has been described in step 101 and will not be elaborated here.
[0038] Optionally, step 202 includes the following sub-steps: Sub-step 2021: According to the matching situation between multiple first packet length ranges and the packet lengths of each voice data packet, perform a first division on all the voice data packets to obtain multiple groups of voice data packets.
[0039] In some embodiments of the present application, in order to preliminarily classify and organize voice data packets for subsequent processing and analysis, the corresponding first packet length range will be determined first according to the length value of each voice data packet; then, according to these length ranges, all voice data packets are assigned to multiple groups, and each group contains voice data packets that meet a specific length range. The packet length range refers to the length value of each data packet, and the matching situation refers to assigning the data packets to the groups within the corresponding length ranges. In this way, multiple groups of voice data packets divided according to the packet length range are obtained, enabling more efficient processing and analysis in subsequent steps.
[0040] In a specific example, the lengths of the actually captured voice data packets can be counted, and they can be matched and divided according to multiple data packet length ranges. For example, the data packet lengths can be divided into ranges such as 40 - 50 bytes, 50 - 60 bytes, 60 - 70 bytes, etc., and the data packets within each length range are assigned to corresponding groups. These divided data packet groups will be used for subsequent merging and processing, improving the efficiency and accuracy of data processing.
[0041] Sub - step 2022: Merge each group of voice data packets to obtain multiple first voice data packet sequences.
[0042] In some embodiments of the present application, in order to organize the preliminarily classified voice data packets into continuous sequences for subsequent analysis and processing, each already divided voice data packet group will first be sorted and checked, and then each group of voice data packets will be merged according to its corresponding first data packet length range to form continuous first voice data packet sequences. These voice data packet sequences retain the order and structure of the original data packets. Through this merging method, it can be ensured that the data packet length ranges in each voice data packet sequence are the same, so that frequency domain features can be more efficiently extracted and classification recognition can be performed in subsequent steps. The multiple first voice data packet sequences obtained in this way enable more accurate processing and analysis in subsequent steps.
[0043] In a specific example, the voice data packet groups previously divided according to the data packet length range can be merged. For example, the voice data packets with a length range of 40 - 50 bytes can be merged into a first voice data packet sequence, and the voice data packets with a length range of 50 - 60 bytes can be merged into another first voice data packet sequence. In this way, the preliminarily classified voice data packets can be organized into continuous sequences, obtaining multiple first voice data packet sequences. These merged voice data packet sequences will be used for subsequent frequency feature extraction and classification recognition, improving the efficiency and accuracy of data processing.
[0044] Step 203: Input the generated multiple first voice data packet sequences into the audio traffic recognition model to generate multiple second voice data packet sequences through the second division of all the voice data packets by the audio traffic recognition model.
[0045] Among them, each second voice data packet sequence has a corresponding second data packet length range respectively; each second data packet length range is respectively used to represent the corresponding frequency range.
[0046] The method shown in this step has been described in step 102 and will not be elaborated here.
[0047] Step 204: Determine the classification result of the encrypted voice traffic data according to the extraction of the data characteristics of each second voice data packet sequence in the frequency domain.
[0048] The method shown in this step has been described in step 103 and will not be elaborated here.
[0049] Optionally, step 204 includes the following sub-steps: Sub-step 2041: Extract multiple statistical metrics of the voice data packets in each second voice data packet sequence respectively to obtain the statistical feature data of the encrypted voice traffic data in each frequency range.
[0050] In some embodiments of the present application, in order to perform quantitative analysis on voice data packets in different frequency ranges and thus provide accurate feature information for subsequent classification and recognition, the voice data packets in each second voice data packet sequence will be analyzed first to calculate multiple statistical metrics. These statistical metrics include the mean, median, variance, standard deviation, total packet number, packet number per unit time, etc. of the packet length. These metrics reflect various feature information of the voice data packets in the frequency domain. Statistical metrics refer to the numerical features used to describe and analyze a set of data, and through statistical metrics, the distribution of the data can be summarized. The statistical feature data of the encrypted voice traffic data obtained in this way in each frequency range provides key input features for subsequent classification and recognition.
[0051] In a specific example, statistical analysis can be performed on the voice data packets in multiple second voice data packet sequences. For example, the following statistical metrics can be extracted: the mean, median, variance, standard deviation, total packet number, packet number per unit time, etc. of the packet length. Through these statistical metrics, the packet features in each frequency range can be obtained. For example, the mean packet length in a certain frequency range is 50 bytes, the median is 48 bytes, the variance is 10 bytes, the total packet number is 2000, and the packet number per unit time is 100 packets / second. The statistical feature data of the encrypted voice traffic data obtained in different frequency ranges in this way will provide detailed and accurate input features for the subsequent classification and recognition model, improving the accuracy and robustness of the recognition.
[0052] Optionally, sub-step 2041 includes the following sub-steps: Sub-step 20411: Determine one or more of the variance of the packet length sequence, standard deviation of the packet length sequence, mean of the packet length sequence, median of the packet length sequence, variance in the front and back directions of the packet length sequence, decile percentile of the packet length, total number of data packets, and number of data packets per unit time determined from the voice data packets in each second voice data packet sequence respectively as a set of statistical feature data components corresponding to the second voice data packet sequence.
[0053] In some embodiments of the present application, in order to extract important statistical features from each sequence of voice data packets for subsequent analysis and classification, the voice data packets in each second sequence of voice data packets are first analyzed to calculate a plurality of statistical metrics, such as the variance of the packet length sequence, the standard deviation of the packet length sequence, the mean of the packet length sequence, the median of the packet length sequence, the variance of the packet length sequence in the forward and backward directions, the decile percentile of the packet length, the total number of data packets, and the number of data packets per unit time. The statistical feature data component refers to a plurality of numerical features extracted from the voice data packets for describing the attributes of the data packets. Such a set of statistical feature data components for each second sequence of voice data packets provides key data for subsequent feature storage and classification.
[0054] In a specific example, statistical analysis can be performed on the second sequence of voice data packets. For example, the following statistical metrics can be calculated for the voice data packets in each sequence: the variance of the packet length sequence is 15, the standard deviation of the packet length sequence is 4, the mean of the packet length sequence is 50, the median of the packet length sequence is 48, the variance of the packet length sequence in the forward and backward directions is 3, the decile percentile distribution of the packet length is 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, the total number of data packets is 2000, and the number of data packets per unit time is 100 packets per second.
[0055] Sub-step 20412: Store each set of statistical feature data components respectively according to a preset statistical feature data structure to obtain statistical feature data corresponding to each frequency range respectively.
[0056] In some embodiments of the present application, in order to systematically store and manage the statistical feature data within each frequency range for subsequent analysis and processing, each set of statistical feature data components is first stored according to a preset statistical feature data structure. Through this structured storage method, it can be ensured that the statistical feature data within each frequency range can be stored systematically and orderly. The statistical feature data obtained in this way corresponding to each frequency range respectively provides a key data basis for subsequent classification and recognition.
[0057] In a specific example, each set of statistical feature data components can be stored according to a preset statistical feature data structure. For example, the statistical features of voice data packets with a length range of 40 - 50 bytes can be stored in a data structure, including the variance of the packet length sequence, the standard deviation, the mean, the median, the variance in the forward and backward directions, the decile percentile of the packet length, the total number of data packets, and the number of data packets per unit time. Similarly, the statistical features of data packets in other length ranges are also stored respectively according to the preset data structure.
[0058] Optionally, in sub-step 20412, 20-dimensional features can be extracted from the packet length sequences in four frequency domain intervals respectively to construct an 80-dimensional multi-frequency domain feature vector and encoded into a latent space through a convolutional neural network to capture more accurate call audio features and optimize the identification of encrypted call audio traffic. Among the 20-dimensional features: the 1-dimensional components included are: the variance of the packet length sequence, the standard deviation of the packet length sequence, the mean of the packet length sequence, the median of the packet length sequence, the total number of data packets, and the number of data packets per unit time; the 3-dimensional components are: the variance of the packet length sequence in the front and back directions, and the remaining dimensions are used to store the 10-equal-part percentile of the packet length.
[0059] Sub-step 2042: Determine the target classification and recognition model corresponding to the target classification requirement among multiple classification and recognition models, and input all the statistical feature data of the obtained encrypted voice traffic data into the target classification and recognition model to obtain the target classification result of the encrypted voice traffic data.
[0060] In some embodiments of the present application, in order to select the most suitable model for a specific classification requirement among multiple classification and recognition models, thereby improving the accuracy and effectiveness of the classification result, first, according to the target classification requirement, multiple available classification and recognition models are evaluated and compared to determine the target classification and recognition model that best meets the requirement. Then, all the statistical feature data of the previously extracted encrypted voice traffic data is input into the target classification and recognition model to generate the final classification result. The target classification and recognition model refers to the best model selected from multiple classification models to meet a specific classification requirement. In this way, an accurate classification result of the encrypted voice traffic data can be effectively obtained, improving the accuracy and efficiency of recognition.
[0061] In a specific example, two classification tasks can be performed on the encrypted voice traffic data: call content recognition and call user identity recognition. Suppose there are three classification and recognition models available for selection: Support Vector Machine (SVM), Random Forest (RF), and Convolutional Neural Networks (CNN). Then, according to the specific requirements of each classification task, these three models can be evaluated and compared to determine that the SVM model is most suitable for call content recognition, while the CNN model is most suitable for call user identity recognition. Next, the previously extracted statistical feature data is input into the corresponding target classification and recognition model, and the classification results of call content recognition and call user identity recognition are obtained respectively. Through this method, the encrypted voice traffic data can be classified efficiently and accurately, improving the accuracy and robustness of recognition.
[0062] In summary, in the embodiments of the present application, through the second division of all voice data packets by the audio traffic recognition model, the frequency feature representation of the voice data packets is refined, so that the voice data packets in different frequency ranges can more accurately represent the frequency information of the voice traffic; and then by extracting the frequency features in the second voice data packet sequence, the statistical feature data in different frequency ranges can be efficiently obtained. By using the two related data attributes that the occurrence features in different scenarios have different frequency features and the voice data packets with different frequency features have different packet lengths, the accurate classification of the encrypted voice traffic data is realized. Therefore, based on the method of the embodiments of the present application, the universality and recognition accuracy of the model are improved, and the problem of poor universality of the existing encrypted voice recognition methods is solved. In addition, the training method of the audio traffic recognition model and the training method of the classification recognition model provided based on the embodiments of the present application solve the problem of high training cost of the related models for voice recognition.
[0063] Figure 5 This is a training method of an audio traffic recognition model provided by an embodiment of the present application, which is used to train the audio traffic recognition model involved in the above embodiment, and specifically includes the following steps: Step 301, obtain a first original training voice data set for training, and a first encrypted training voice data set corresponding to the first original training voice data set.
[0064] Among them, the first original training voice data in the first original training voice data set corresponds one-to-one with the first encrypted training voice data in the first encrypted training voice data set.
[0065] In some embodiments of the present application, in order to ensure the diversity and integrity of the training data, and thus improve the training effect of the audio traffic recognition model, a first original training voice data set containing unencrypted voice data and a first encrypted training voice data set corresponding to these original voice data will be obtained first. Each first original training voice data corresponds one-to-one with its corresponding first encrypted training voice data. This can ensure that the model can learn the relationship between encrypted and unencrypted voice data during the training process, thereby improving the recognition ability of the model. Through this matching data set, the model can better understand the characteristics of encrypted voice data and improve the accuracy and robustness of recognition.
[0066] In a specific example, a large amount of voice data can be collected, and the voice data can be encrypted to obtain an unencrypted first original training voice data set and an encrypted first encrypted training voice data set, and make the records in the two data sets correspond one-to-one.
[0067] Step 302: Determine the first audio spectrum feature of the first original training speech dataset according to the first original training speech data in the first original training speech dataset, and determine the first packet length statistical feature of the first encrypted training speech dataset according to the first encrypted training speech data in the first encrypted training speech dataset.
[0068] In some embodiments of the present application, in order to extract the key features of the original and encrypted speech datasets for subsequent model training, spectral analysis is first performed on the first original training speech data in the first original training speech dataset to determine its audio spectrum feature; then, the first encrypted training speech data in the first encrypted training speech dataset is analyzed to determine its packet length statistical feature. The audio spectrum feature refers to various characteristic information of speech data within the frequency range, and the packet length statistical feature includes the statistical distribution, mean, median, variance, etc. of the packet length. The obtained audio spectrum feature of the first original training speech dataset and the packet length statistical feature of the first encrypted training speech dataset can provide key input features for subsequent model training.
[0069] In a specific example, a spectral analysis tool can be used to perform spectral analysis on the speech data in the first original training speech dataset. For example, the frequency components of each segment of speech can be extracted, such as low frequency (20 - 200 Hz), medium frequency (200 - 2000 Hz), high frequency (2000 - 20000 Hz), etc. Next, the encrypted speech packets in the first encrypted training speech dataset are statistically analyzed for packet length, and the length distribution features of each packet are calculated, such as mean, median, variance, etc. In this way, the obtained audio spectrum feature of the first original training speech dataset and the packet length statistical feature of the first encrypted training speech dataset will be used for the training of the subsequent audio traffic recognition model, improving the recognition ability and accuracy of the model.
[0070] Step 303: Train the audio traffic recognition model according to the first audio spectrum feature and the first packet length statistical feature to obtain a trained audio traffic recognition model.
[0071] In some embodiments of the present application, in order to utilize the extracted audio spectrum feature and packet length statistical feature to enable the model to accurately identify encrypted speech traffic, the first audio spectrum feature and the first packet length statistical feature are first used as input features and input into the audio traffic recognition model; then, the model is trained through a training algorithm so that it can learn the relationship between the input features and the speech traffic classification result, thereby improving the recognition ability of the model. The obtained trained audio traffic recognition model will be able to efficiently and accurately identify encrypted speech traffic data.
[0072] In a specific example, the previously extracted first audio spectrum features and first data packet length statistical features can be used as training data and input into the audio traffic recognition model. During the training process, the model continuously adjusts its parameters to accurately map the input features to the classification results. After multiple training iterations, an audio traffic recognition model that can utilize the input audio spectrum features and data packet length statistical features will ultimately be obtained to efficiently identify encrypted voice traffic data.
[0073] In summary, in the embodiments of the present application, through the second division of all voice data packets by the audio traffic recognition model, the frequency feature representation of the voice data packets is refined, enabling voice data packets in different frequency ranges to more accurately represent the frequency information of the voice traffic; furthermore, by extracting the frequency features in the second voice data packet sequence, statistical feature data in different frequency ranges can be efficiently obtained. By leveraging the two associated data attributes that the occurrence features have different frequency features in different scenarios and the voice data packets under different frequency features have different packet lengths, the accurate classification of encrypted voice traffic data is thus achieved. Therefore, based on the method of the embodiments of the present application, the universality and recognition accuracy of the model are improved, and the problem of poor universality of existing encrypted voice recognition methods is solved. In addition, the training method of the audio traffic recognition model and the training method of the classification recognition model provided based on the embodiments of the present application solve the problem of high training costs for related models in voice recognition.
[0074] Figure 6 This is a training method for a classification recognition model provided by an embodiment of the present application, used to train the classification recognition model involved in the above embodiments, and specifically includes the following steps: Step 401, obtain a second original training voice data set for training, and a second encrypted training voice data set corresponding to the second original training voice data set.
[0075] Among them, the second original training voice data in the second original training voice data set corresponds one-to-one with the second encrypted training voice data in the second encrypted training voice data set.
[0076] In some embodiments of the present application, to ensure the diversity and integrity of the training data and thus improve the training effect of the classification recognition model, a second original training voice data set containing unencrypted voice data and a second encrypted training voice data set corresponding to these original voice data will first be obtained. Each second original training voice data corresponds one-to-one with its corresponding second encrypted training voice data. This can ensure that the model can learn the relationship between encrypted and unencrypted voice data during the training process, thereby improving the model's recognition ability. With this matching data set, the model can better understand the characteristics of encrypted voice data and improve the accuracy and robustness of recognition.
[0077] In a specific example, a large amount of voice data can be collected and the voice data can be encrypted to obtain an unencrypted second original training voice data set and an encrypted second encrypted training voice data set, and the records in the two data sets are made to correspond one by one.
[0078] Step 402: Determine the second audio spectrum feature of the second original training voice data set according to the second original training voice data in the second original training voice data set, and determine the second data packet length feature and data packet distribution feature of the second encrypted training voice data set according to the second encrypted training voice data in the second encrypted training voice data set.
[0079] In some embodiments of the present application, in order to extract the key features of the original and encrypted voice data sets to facilitate subsequent model training, the second original training voice data in the second original training voice data set will first be subjected to spectrum analysis to determine its audio spectrum feature; then, the second encrypted training voice data in the second encrypted training voice data set will be analyzed to determine its data packet length feature and data packet distribution feature. The audio spectrum feature refers to various characteristic information of the voice data within the frequency range, and the data packet length feature includes the statistical distribution, mean, median, variance, etc. of the data packet length, and the data packet distribution feature includes the time and space distribution of the data packet during transmission. The audio spectrum feature of the second original training voice data set and the data packet length feature and data packet distribution feature of the second encrypted training voice data set obtained in this way can provide key input features for subsequent model training.
[0080] In a specific example, a spectrum analysis tool can be used to perform spectrum analysis on the voice data in the second original training voice data set. For example, the frequency components of each segment of voice can be extracted, such as low frequency (20 - 200 Hz), medium frequency (200 - 2000 Hz), high frequency (2000 - 20000 Hz), etc. Next, the data packet lengths of the encrypted voice data packets in the second encrypted training voice data set are counted, and the length distribution features of each data packet are calculated, such as mean, median, variance, etc. In this way, the audio spectrum feature of the second original training voice data set and the data packet length statistical feature of the second encrypted training voice data set will be used for the training of the subsequent classification and recognition model, improving the recognition ability and accuracy of the model.
[0081] Step 403: Train the classification and recognition model according to the second audio spectrum feature, the second data packet length statistical feature, and the data packet distribution feature to obtain a trained classification and recognition model.
[0082] In some embodiments of the present application, in order to utilize a comprehensive feature set and enable the model to more accurately identify and classify encrypted voice traffic data, the second audio spectrum feature, the second packet length statistical feature, and the packet distribution feature are first used as input features and input into the classification and recognition model; then, the model is trained through a training algorithm so that it can learn the relationship between the input features and the voice traffic classification result, thereby improving the recognition ability of the model. In this way, a trained classification and recognition model is obtained to efficiently and accurately identify encrypted voice traffic data.
[0083] In a specific example, the previously extracted second audio spectrum feature, second packet length statistical feature, and packet distribution feature can be used as training data and input into the classification and recognition model. Assume that the classification and recognition model uses a Long Short-Term Memory (LSTM) model. During the training process, the model continuously adjusts its parameters to accurately map the input features to the classification result. After multiple training iterations, the finally obtained classification and recognition model can utilize the input audio spectrum feature, packet length statistical feature, and packet distribution feature to efficiently identify encrypted voice traffic data, thereby improving the accuracy and robustness of the recognition.
[0084] As Figure 7 shown, it is a complete audio traffic recognition process under the embodiments of the present application, which specifically includes the following processes: Step S0, obtaining training data through extraction of audio frequency domain distribution features: Use the method of extracting frequency domain distribution features to obtain training data. By performing frequency domain analysis on the audio signal, its frequency distribution features are extracted.
[0085] Step S1, encrypted call traffic capture: Perform data capture of encrypted call traffic. The capture process involves collecting actual call packets to ensure that the model can process and identify encrypted voice traffic in a real environment.
[0086] Step S2, extraction of packet length distribution and subsequence division under variable-length coding: Statistically analyze the length distribution of each packet, and divide the packets into multiple traffic subsequences according to these distributions.
[0087] Step S3, extraction of frequency domain feature fingerprints: Perform frequency domain analysis on the packets of each subsequence, and extract the frequency domain feature fingerprints by reorganizing the packet sequence.
[0088] Step S4, obtaining classification results through processing by the recognition model: Input the extracted frequency domain feature fingerprints into the recognition model, and the model generates the final classification results by learning and analyzing these features.
[0089] In summary, in the embodiment of the present application, through the second division of all voice data packets by the audio traffic recognition model, the frequency feature representation of the voice data packets is refined, so that the voice data packets in different frequency ranges can more accurately represent the frequency information of the voice traffic; furthermore, by extracting the frequency features in the second voice data packet sequence, the statistical feature data in different frequency ranges can be efficiently obtained. By using the two associated data attributes that the occurrence features in different scenarios have different frequency features and the voice data packets with different frequency features have different packet lengths, the accurate classification of the encrypted voice traffic data is realized. Therefore, based on the method of the embodiment of the present application, the universality and recognition accuracy of the model are improved, and the problem of poor universality of the existing encrypted voice recognition method is solved. In addition, the training method of the audio traffic recognition model and the training method of the classification recognition model provided based on the embodiment of the present application solve the problem of high training cost of the related models for voice recognition.
[0090] As Figure 8 shown, the embodiment of the present application also discloses an encrypted call audio traffic recognition device 50 based on spectrum analysis, including: A first division module 501, configured to perform a first division on all voice data packets according to the packet length of each voice data packet in the encrypted voice traffic data, and generate a plurality of first voice data packet sequences; each first voice data packet sequence has a corresponding first packet length range; the range width of each first packet length range is the same; A second division module 502, configured to input the generated plurality of first voice data packet sequences into the audio traffic recognition model, and perform a second division on all voice data packets through the audio traffic recognition model to generate a plurality of second voice data packet sequences; each second voice data packet sequence has a corresponding second packet length range; each second packet length range is respectively used to represent the corresponding frequency range; A classification module 503, configured to determine the classification result of the encrypted voice traffic data according to the extraction of the data features in the frequency domain of each second voice data packet sequence.
[0091] Optionally, the first division module 501 includes: A first division sub-module, configured to perform a first division on all voice data packets according to the matching situation between a plurality of first packet length ranges and the packet length of each voice data packet, so as to obtain multiple groups of voice data packets; A merging sub-module, configured to merge each group of voice data packets to obtain a plurality of first voice data packet sequences.
[0092] Optionally, the classification module 503 includes: A statistical feature sub-module, configured to respectively extract multiple statistical metrics of the voice data packets in each second voice data packet sequence, so as to obtain statistical feature data of the encrypted voice traffic data in each frequency range; A sub-component module, configured to determine a target classification and recognition model corresponding to the target classification requirement from multiple classification and recognition models, and input all the statistical feature data of the obtained encrypted voice traffic data into the target classification and recognition model, so as to obtain a target classification result of the encrypted voice traffic data.
[0093] Optionally, the statistical feature sub-module includes: A component extraction unit, configured to respectively determine one or more of the variance of the packet length sequence, the standard deviation of the packet length sequence, the mean of the packet length sequence, the median of the packet length sequence, the variance in the front and back directions of the packet length sequence, the decile percentile of the packet length, the total number of data packets, and the number of data packets per unit time determined from the voice data packets in each second voice data packet sequence as a set of statistical feature data components corresponding to the corresponding second voice data packet sequence; A feature generation unit, configured to respectively store each set of statistical feature data components according to a preset statistical feature data structure, so as to obtain statistical feature data corresponding to each frequency range respectively.
[0094] Optionally, the encrypted call audio traffic recognition device 50 based on spectrum analysis further includes: A denoising module, configured to denoise the encrypted voice traffic data and extract all the data packets in the denoised encrypted voice traffic data.
[0095] As Figure 9 shown, an embodiment of the present application also discloses a training device 60 for an audio traffic recognition model, configured to train the audio traffic recognition model involved in the above embodiment, including: A first data set module 601, configured to obtain a first original training voice data set for training and a first encrypted training voice data set corresponding to the first original training voice data set; the first original training voice data in the first original training voice data set corresponds one-to-one with the first encrypted training voice data in the first encrypted training voice data set; A first extraction module 602, configured to determine the first audio spectrum features of the first original training voice data set according to the first original training voice data in the first original training voice data set, and determine the first packet length statistical features of the first encrypted training voice data set according to the first encrypted training voice data in the first encrypted training voice data set; A first training module 603, configured to train the audio traffic recognition model according to the first audio spectrum features and the first packet length statistical features, so as to obtain a trained audio traffic recognition model.
[0096] As shown in Figure 10 the figure, an embodiment of the present application also discloses a training device 61 for a classification and recognition model, which is used to train the classification and recognition model involved in the above embodiment, and includes: A second data set module 611, configured to obtain a second original training speech data set for training, and a second encrypted training speech data set corresponding to the second original training speech data set; the second original training speech data in the second original training speech data set corresponds one-to-one with the second encrypted training speech data in the second encrypted training speech data set; A second extraction module 612, configured to determine the second audio spectrum feature of the second original training speech data set according to the second original training speech data in the second original training speech data set, and determine the second data packet length feature and the data packet distribution feature of the second encrypted training speech data set according to the second encrypted training speech data in the second encrypted training speech data set; A second training module 613, configured to train the classification and recognition model according to the second audio spectrum feature, the second data packet length statistical feature, and the data packet distribution feature, so as to obtain a trained classification and recognition model.
[0097] In summary, in the embodiment of the present application, through the second division of all speech data packets by the audio traffic recognition model, the frequency feature representation of the speech data packets is refined, so that the speech data packets in different frequency ranges can more accurately represent the frequency information of the speech traffic; furthermore, by extracting the frequency features in the second speech data packet sequence, the statistical feature data in different frequency ranges can be efficiently obtained, and two related data attributes that the occurrence features in different scenarios have different frequency features and the speech data packets with different frequency features have different packet lengths are utilized, thereby realizing the accurate classification of encrypted speech traffic data. Therefore, based on the method of the embodiment of the present application, the universality and recognition accuracy of the model are improved, and the problem of poor universality of the existing encrypted speech recognition method is solved. In addition, the training method of the audio traffic recognition model and the training method of the classification and recognition model provided based on the embodiment of the present application solve the problem of high training cost of the related models for speech recognition.
[0098] The embodiment of the present application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements each process of the above-mentioned method for recognizing encrypted call audio traffic based on spectrum analysis, the training method of the audio traffic recognition model, and the training method of the classification and recognition model, and can achieve the same technical effect. To avoid repetition, it will not be elaborated here. Among them, the computer-readable storage medium is, for example, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc, etc.
[0099] Figure 11 It is a block diagram of an electronic device 700 provided by an embodiment of the present application. For example, the electronic device 700 may be a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.
[0100] Referring to Figure 11 , the electronic device 700 may include one or more of the following components: a processing component 702, a memory 704, a power component 706, a multimedia component 708, an audio component 710, an input / output (I / O) interface 712, a sensor component 714, and a communication component 716.
[0101] The processing component 702 generally controls the overall operation of the electronic device 700, such as operations associated with display, telephone calls, data communication, camera operations, and recording operations. The processing component 702 may include one or more processors 720 to execute instructions to complete all or part of the steps of the above-mentioned encrypted call audio traffic recognition method based on spectrum analysis, the training method of the audio traffic recognition model, and the training method of the classification recognition model. In addition, the processing component 702 may include one or more modules to facilitate the interaction between the processing component 702 and other components. For example, the processing component 702 may include a multimedia module to facilitate the interaction between the multimedia component 708 and the processing component 702.
[0102] The memory 704 is used to store various types of data to support the operation of the electronic device 700. Examples of these data include instructions for any application or method operating on the electronic device 700, contact data, phone book data, messages, pictures, multimedia, etc. The memory 704 may be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disk.
[0103] The power component 706 provides power to various components of the electronic device 700. The power component 706 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the electronic device 700.
[0104] The multimedia component 708 includes a screen that provides an output interface between the electronic device 700 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors can sense not only the boundaries of touch or swipe actions but also detect the duration and pressure associated with the touch or swipe operations. In some embodiments, the multimedia component 708 includes a front camera and / or a rear camera. When the electronic device 700 is in an operating mode, such as a shooting mode or a multimedia mode, the front camera and / or the rear camera can receive external multimedia data. Each of the front camera and the rear camera can be a fixed optical lens system or have a focal length and optical zoom capabilities.
[0105] The audio component 710 is used to output and / or input audio signals. For example, the audio component 710 includes a microphone (MIC) that is used to receive external audio signals when the electronic device 700 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signals can be further stored in the memory 704 or sent via the communication component 716. In some embodiments, the audio component 710 further includes a speaker for outputting audio signals.
[0106] The I / O interface 712 provides an interface between the processing component 702 and a peripheral interface module, which can be a keyboard, a click wheel, buttons, etc. These buttons can include but are not limited to: a home button, a volume button, a power button, and a lock button.
[0107] The sensor component 714 includes one or more sensors for providing status assessments of various aspects of the electronic device 700. For example, the sensor component 714 can detect the on / off state of the electronic device 700, the relative positioning of components, such as the display and keypad of the electronic device 700. The sensor component 714 can also detect a change in the position of the electronic device 700 or a component of the electronic device 700, the presence or absence of user contact with the electronic device 700, the orientation or acceleration / deceleration of the electronic device 700, and the temperature change of the electronic device 700. The sensor component 714 can include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor component 714 can also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor component 714 can further include an acceleration sensor, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.
[0108] The communication component 716 is used to facilitate communication between the electronic device 700 and other devices in a wired or wireless manner. The electronic device 700 can access a communication standard-based wireless network, such as WiFi, a carrier network (such as 2G, 3G, 4G, or 7G), or a combination thereof. In an exemplary embodiment, the communication component 716 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 716 further includes a Near Field Communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on Radio Frequency Identification (RFID) technology, Infrared Data Association (IrDA) technology, Ultra-Wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.
[0109] In an exemplary embodiment, the electronic device 700 can be implemented by one or more Application Specific Integrated Circuits (ASICs), Digital Signal Processors (DSPs), Digital Signal Processing Devices (DSPDs), Programmable Logic Devices (PLDs), Field Programmable Gate Arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components, for implementing the encrypted call audio traffic identification method based on spectrum analysis, the training method of the audio traffic identification model, and the training method of the classification and identification model provided in the embodiments of the present application.
[0110] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is further provided, such as a memory 704 including instructions. The above instructions can be executed by a processor 720 of the electronic device 700 to complete the above-mentioned encrypted call audio traffic identification method based on spectrum analysis, the training method of the audio traffic identification model, and the training method of the classification and identification model. For example, the non-transitory storage medium can be a ROM, a Random Access Memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device, etc.
[0111] Figure 12 is a block diagram of an electronic device 800 shown according to an exemplary embodiment. For example, the electronic device 800 can be provided as a server. Referring to Figure 12 , the electronic device 800 includes a processing component 822, which further includes one or more processors, and memory resources represented by a memory 832 for storing instructions executable by the processing component 822, such as application programs. The application programs stored in the memory 832 can include one or more modules each corresponding to a set of instructions. In addition, the processing component 822 is configured to execute instructions to perform the encrypted call audio traffic identification method based on spectrum analysis, the training method of the audio traffic identification model, and the training method of the classification and identification model provided in the embodiments of the present application.
[0112] The electronic device 800 may further include a power supply component 826 configured to perform power management of the electronic device 800, a wired or wireless network interface 850 configured to connect the electronic device 800 to a network, and an input / output (I / O) interface 858. The electronic device 800 may operate based on an operating system stored in the memory 832, such as Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSD TM or the like.
[0113] An embodiment of the present application also provides a computer program product, including a computer program, which when executed by a processor, implements an encrypted call audio traffic recognition method based on spectrum analysis, a training method for an audio traffic recognition model, and a training method for a classification recognition model.
[0114] Those skilled in the art will readily conceive of other embodiments of the present application after considering the specification and practicing the application disclosed herein. The present application is intended to cover any variations, uses, or adaptations of the present application, which follow the general principles of the present application and include known common general knowledge or conventional technical means in the technical field not disclosed in the present disclosure. The specification and embodiments are only to be considered as exemplary, and the true scope and spirit of the present application are pointed out by the following claims.
[0115] It should be understood that the present application is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present application is only limited by the appended claims.
[0116] Each embodiment in this specification is described in a progressive manner, with the key point of each embodiment being the differences from other embodiments. The same or similar parts among the various embodiments can be referred to each other.
[0117] It is easy for those skilled in the art to think that any combination application of the above various embodiments is feasible. Therefore, any combination of the above various embodiments is an embodiment of the present application. However, due to space limitations, this specification does not elaborate on them one by one here.
[0118] The method for identifying encrypted call audio traffic based on spectrum analysis, the method for training an audio traffic identification model, and the method for training a classification and identification model provided herein are not inherently related to any specific computer, virtual system, or other device. Various general-purpose systems can also be used in conjunction with the teachings provided herein. Based on the above description, it will be apparent to those skilled in the art how to construct a system having the structure required by the solution of the present application. In addition, the present application is not directed to any specific programming language. It should be understood that the content of the present application described herein can be implemented using various programming languages, and the descriptions made above with respect to specific languages are for the purpose of disclosing the best mode of the present application.
[0119] In the specification provided herein, a large number of specific details are set forth. However, it will be understood that embodiments of the present application may be practiced without these specific details. In some instances, well-known methods, structures, and techniques have not been shown in detail so as not to obscure the understanding of this specification.
[0120] Similarly, it should be understood that, in order to streamline the present application and assist in understanding one or more of the various aspects of the application, in the above description of the exemplary embodiments of the present application, the various features of the present application are sometimes grouped together into a single embodiment, figure, or description thereof. However, the disclosed method should not be construed as reflecting an intention that the claimed subject matter of the present application requires more features than are expressly recited in each claim. Rather, as the claims reflect, the aspects of the application lie in less than all of the features of the single foregoing disclosed embodiment. Thus, the claims following the detailed description are hereby expressly incorporated into the detailed description, with each claim standing on its own as a separate embodiment of the present application.
[0121] Those skilled in the art will appreciate that the modules in the devices in the embodiments can be adaptively changed and disposed in one or more devices different from those of the embodiments. The modules or units or components in the embodiments can be combined into one module or unit or component, and in addition, they can be divided into multiple sub-modules or sub-units or sub-components. Except for the fact that at least some of such features and / or processes or units are mutually exclusive, any combination can be adopted for all the features disclosed in this specification (including the accompanying claims, abstract, and drawings) and all the processes or units of any method or device so disclosed. Unless otherwise expressly stated, each feature disclosed in this specification (including the accompanying claims, abstract, and drawings) can be replaced by an alternative feature providing the same, equivalent, or similar purpose.
[0122] In addition, those skilled in the art can understand that although some of the embodiments described herein include certain features included in other embodiments rather than other features, the combination of features of different embodiments is meant to be within the scope of the present application and forms different embodiments. For example, in the claims, any one of the claimed embodiments can be used in any combination.
[0123] Each component embodiment of the present application can be implemented in hardware, or in software modules running on one or more processors, or in a combination thereof. Those skilled in the art should understand that a microprocessor or a digital signal processor (DSP) can be used in practice to implement some or all of the functions of some or all of the components in the method for identifying encrypted call audio traffic based on spectrum analysis, the method for training an audio traffic recognition model, and the method for training a classification recognition model according to the embodiments of the present application. The present application can also be implemented as a device or apparatus program (for example, a computer program and a computer program product) for executing part or all of the methods described herein. Such a program implementing the present application can be stored on a computer-readable medium, or can be in the form of one or more signals. Such signals can be downloaded from an Internet website, or provided on a carrier signal, or in any other form.
[0124] In yet another embodiment provided by the present invention, there is also provided a computer program product containing instructions, which when run on a computer, causes the computer to execute the method for identifying encrypted call audio traffic based on spectrum analysis, the method for training an audio traffic recognition model, and the method for training a classification recognition model according to the embodiments of the present application.
[0125] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present invention are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer, or a data storage device such as a server or data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)).
[0126] It should be noted that the above embodiments are illustrative of the present application rather than restrictive of the present application, and those skilled in the art can design alternative embodiments without departing from the scope of the appended claims. In the claims, any reference signs placed between parentheses shall not be construed as limiting the claim. The word "comprising" does not exclude the presence of elements or steps not listed in the claim. The word "a" or "an" preceding an element does not exclude the presence of a plurality of such elements. The present application can be implemented by means of hardware including several different elements and by means of a suitably programmed computer. In the unit claims listing several devices, several of these devices can be embodied by the same item of hardware. The use of the words first, second, and third, etc. does not denote any order. These words can be interpreted as names.
[0127] It should be noted that for the method embodiments of the present application, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the embodiments of the present application are not limited by the described order of actions, because according to the embodiments of the present application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily essential to the embodiments of the present application.
[0128] Each embodiment in this specification is described in a related manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the embodiments of the system or device, since they are basically similar to the method embodiments, the description is relatively simple, and for the relevant parts, reference can be made to the corresponding descriptions in the method embodiments.
[0129] The above are only the preferred embodiments of the present invention, and are not intended to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention are all included in the protection scope of the present invention.
Claims
1. An encrypted call audio traffic recognition method based on spectrum analysis, characterized in that Including: Performing a first division of all the voice data packets according to the data packet length of each voice data packet in the encrypted voice traffic data, to generate a plurality of first voice data packet sequences; Each of the first voice data packet sequences respectively has a corresponding first data packet length range; the range width of each of the first data packet length ranges is the same; Inputting the generated plurality of first voice data packet sequences into an audio traffic recognition model, to perform a second division of all the voice data packets through the audio traffic recognition model, to generate a plurality of second voice data packet sequences; each of the second voice data packet sequences respectively has a corresponding second data packet length range; each of the second data packet length ranges is respectively used to characterize the corresponding frequency range; Determining a classification result of the encrypted voice traffic data according to the extraction of data features of each of the second voice data packet sequences in the frequency domain.
2. The encrypted call audio traffic recognition method based on spectrum analysis according to claim 1, wherein The performing a first division of all the voice data packets according to the data packet length of each voice data packet in the encrypted voice traffic data, to generate a plurality of first voice data packet sequences, includes: Performing a first division of all the voice data packets according to the matching situation between a plurality of the first data packet length ranges and the data packet length of each voice data packet, to obtain multiple groups of voice data packets; Merging each group of the voice data packets, to obtain a plurality of the first voice data packet sequences.
3. The method for identifying encrypted call audio traffic based on spectrum analysis according to claim 1, wherein The determining a classification result of the encrypted voice traffic data according to the extraction of data features of each of the second voice data packet sequences in the frequency domain, includes: Respectively extracting a plurality of statistical metrics of the voice data packets in each of the second voice data packet sequences, to obtain statistical feature data of the encrypted voice traffic data in each of the frequency ranges; Determining a target classification recognition model corresponding to a target classification requirement among a plurality of classification recognition models, and inputting all the statistical feature data of the obtained encrypted voice traffic data into the target classification recognition model, to obtain a target classification result of the encrypted voice traffic data.
4. The method for identifying encrypted call audio traffic based on spectrum analysis according to claim 3, wherein The respectively extracting a plurality of statistical metrics of the voice data packets in each of the second voice data packet sequences, to obtain statistical feature data of the encrypted voice traffic data in each of the frequency ranges, includes: Determining, respectively, one or more of the variance of the packet length sequence, the standard deviation of the packet length sequence, the mean of the packet length sequence, the median of the packet length sequence, the variance in the front-back direction of the packet length sequence, the percentile of the decile packet length, the total number of data packets, and the number of data packets per unit time determined from the voice data packets in each of the second voice data packet sequences, as a set of statistical feature data components corresponding to the respective second voice data packet sequences; Respectively storing each set of statistical feature data components according to a preset statistical feature data structure, to obtain statistical feature data respectively corresponding to each of the frequency ranges.
5. The method for identifying encrypted call audio traffic based on spectrum analysis according to claim 1, wherein The encrypted call audio traffic recognition method based on spectrum analysis further includes: Denosing the encrypted voice traffic data, and extracting all the data packets in the denoised encrypted voice traffic data.
6. A training method for an audio traffic recognition model, characterized in that For training the audio traffic recognition model in the encrypted call audio traffic recognition method based on spectrum analysis as described in any one of claims 1 to 5, including: Obtaining a first original training speech data set for training, and a first encrypted training speech data set corresponding to the first original training speech data set; the first original training speech data in the first original training speech data set corresponds one-to-one with the first encrypted training speech data in the first encrypted training speech data set; Determining the first audio spectrum feature of the first original training speech data set according to the first original training speech data in the first original training speech data set, and determining the first packet length statistical feature of the first encrypted training speech data set according to the first encrypted training speech data in the first encrypted training speech data set; Training the audio traffic recognition model according to the first audio spectrum feature and the first packet length statistical feature to obtain the trained audio traffic recognition model.
7. A training method for a classification and recognition model, characterized in that, For training the classification recognition model in the encrypted call audio traffic recognition method based on spectrum analysis as described in claim 3 or 4, including: Obtaining a second original training speech data set for training, and a second encrypted training speech data set corresponding to the second original training speech data set; the second original training speech data in the second original training speech data set corresponds one-to-one with the second encrypted training speech data in the second encrypted training speech data set; Determining the second audio spectrum feature of the second original training speech data set according to the second original training speech data in the second original training speech data set, and determining the second packet length feature and packet distribution feature of the second encrypted training speech data set according to the second encrypted training speech data in the second encrypted training speech data set; Training the classification recognition model according to the second audio spectrum feature, the second packet length statistical feature, and the packet distribution feature to obtain the trained classification recognition model.
8. An encrypted call audio traffic recognition device based on spectrum analysis, characterized in that, Including: A first partitioning module for partitioning all the speech packets according to the packet length of each speech packet in the encrypted speech traffic data to generate a plurality of first speech packet sequences; each of the first speech packet sequences has a corresponding first packet length range; the range width of each of the first packet length ranges is the same; A second partitioning module for inputting the generated plurality of first speech packet sequences into the audio traffic recognition model to perform a second partitioning of all the speech packets through the audio traffic recognition model to generate a plurality of second speech packet sequences; each of the second speech packet sequences has a corresponding second packet length range; each of the second packet length ranges is used to characterize the respective corresponding frequency range; A classification module for determining the classification result of the encrypted speech traffic data according to the extraction of the data features of each of the second speech packet sequences in the frequency domain.
9. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium, and when the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.
10. An electronic device, characterized in that, It includes a processor, a memory, and a computer program stored on the memory and executable on the processor. When the computer program is executed by the processor, the steps of the method according to any one of claims 1 to 7 are implemented.