Method, apparatus, and computer program product for authenticating synthetic audio
By employing different sampling rates and deep learning models to analyze audio features, synthetic audio can be accurately and efficiently distinguished from natural audio, addressing the challenge of synthetic audio proliferation and security threats.
Patent Information
- Application Number
- CN202211179264.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-27
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2042-09-27
AI Technical Summary
The number of synthetic audio has increased explosively, and it takes huge manpower to identify it paragraph by paragraph, sentence by sentence, and it is difficult for the existing technology to identify synthetic audio efficiently and accurately.
By sampling the target audio using different sampling rates, the constant Q transform cepspectral coefficient, linear frequency cepspectral coefficient, Mel frequency cepspectral coefficient, fundamental frequency and short-time energy characteristics are extracted, and a pre-trained identification model is input to automatically identify whether the target audio is synthetic audio.
The identification efficiency and recognition accuracy of synthetic audio are improved, and the features of synthetic audio can be captured from four dimensions: time domain, frequency domain, pitch and volume. The efficiency is improved when some features are recognized, and the robustness is enhanced when all features are recognized.
Smart Images

Figure CN115881083B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of audio processing, and in particular, to a method, device, and computer program product for authenticating synthetic audio. Background Art
[0002] With the development of audio synthesis technology, more and more realistic synthetic audio similar to human voices has emerged. Although synthetic audio has facilitated people's work, life, and entertainment, it has also posed a great threat to information security. Therefore, it is necessary to authenticate synthetic audio. However, the number of synthetic audio has shown an explosive growth, and authenticating it sentence by sentence requires a huge amount of manpower. Summary of the Invention
[0003] Based on this, in view of the above technical problems, it is necessary to provide a method, computer device, and computer program product for authenticating synthetic audio.
[0004] This application provides a method for authenticating synthetic audio, and the method includes:
[0005] Sampling a target audio at different sampling rates to obtain multiple audios corresponding to different sampling rates;
[0006] Extracting at least one of the constant Q transform cepstral coefficients, linear frequency cepstral coefficients, mel frequency cepstral coefficients, fundamental frequency, and short-time energy features of each of the audios;
[0007] Inputting at least one of the constant Q transform cepstral coefficients, linear frequency cepstral coefficients, mel frequency cepstral coefficients, fundamental frequency, and short-time energy features of each of the audios into a pre-trained authentication model to obtain the result output by the authentication model;
[0008] Determining whether the target audio is synthetic audio according to the result output by the authentication model.
[0009] This application provides a computer device, including a memory and a processor, where the memory stores a computer program, and the processor executes the following steps:
[0010] Sampling a target audio at different sampling rates to obtain multiple audios corresponding to different sampling rates;
[0011] Extracting at least one of the constant Q transform cepstral coefficients, linear frequency cepstral coefficients, mel frequency cepstral coefficients, fundamental frequency, and short-time energy features of each of the audios;
[0012] Inputting at least one of the constant Q transform cepstral coefficients, linear frequency cepstral coefficients, mel frequency cepstral coefficients, fundamental frequency, and short-time energy features of each of the audios into a pre-trained authentication model to obtain the result output by the authentication model;
[0013] Determine whether the target audio is a synthetic audio according to the result output by the discrimination model.
[0014] This application provides a computer program product, on which a computer program is stored, and the computer program is executed by a processor to perform the following steps:
[0015] Sample the target audio at different sampling rates to obtain multiple audio corresponding to different sampling rates;
[0016] Extract at least one of the constant Q transform cepstral coefficients, linear frequency cepstral coefficients, mel frequency cepstral coefficients, fundamental frequency, and short-time energy features of each audio;
[0017] Input at least one of the constant Q transform cepstral coefficients, linear frequency cepstral coefficients, mel frequency cepstral coefficients, fundamental frequency, and short-time energy features of each audio into a pre-trained discrimination model to obtain the result output by the discrimination model;
[0018] Determine whether the target audio is a synthetic audio according to the result output by the discrimination model.
[0019] In the above discrimination method, computer device, and computer program product for synthetic audio, the discrimination model is used to automatically discriminate whether the target audio is a synthetic audio, improving the discrimination efficiency of synthetic audio; moreover, this application samples the target audio at different sampling rates, combines the features of the audio of the target audio at different sampling rates, and determines whether the target audio is a synthetic audio, improving the recognition accuracy; in addition, compared with non-synthetic audio, some features of synthetic audio are different. For example, some synthetic audio is different in the performance of constant Q transform cepstral coefficients, and some synthetic audio is different in the performance of linear frequency cepstral coefficients. Based on this characteristic of synthetic audio, in the features of constant Q transform cepstral coefficients, linear frequency cepstral coefficients, mel frequency cepstral coefficients, fundamental frequency, and short-time energy features of the audio, some features or all features are used to identify whether the target audio is a synthetic audio; when using some features, since the used partial features belong to the above features, the recognition accuracy can be guaranteed to a certain extent, and when using partial features, only partial features can be extracted, thereby improving the recognition efficiency; when using all features, the constant Q transform cepstral coefficients, linear frequency cepstral coefficients, mel frequency cepstral coefficients, fundamental frequency, and short-time energy features of each audio can be extracted, covering more features, and no matter which aspect the features of the audio are different, they can be identified as much as possible, with strong robustness; in addition, starting from four dimensions of time domain (reflected in multi-sampling rate), frequency domain (reflected in constant Q transform cepstral coefficients, linear frequency cepstral coefficients, mel frequency cepstral coefficients), pitch (reflected in fundamental frequency), and volume (reflected in short-time energy features), the features of synthetic audio can be better captured and discriminated. Brief Description of the Drawings
[0020] Figure 1 It is a schematic flowchart of a method for authenticating synthesized audio in an embodiment;
[0021] Figure 2 It is a schematic diagram of the processing of an authentication model in an embodiment;
[0022] Figure 3 It is a schematic diagram of forming multiple sampling rate groups in an embodiment;
[0023] Figure 4 It is an internal structure diagram of a computer device in an embodiment. Detailed Description of the Embodiment
[0024] In order to make the objectives, technical solutions and advantages of the present application clearer and more understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0025] Referring to "embodiment" in the present application means that specific features, structures or characteristics described in combination with the embodiment may be included in at least one embodiment of the present application. The occurrence of this phrase in various positions in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described in the present application can be combined with other embodiments.
[0026] In one embodiment, as Figure 1 shown, a method for authenticating synthesized audio is provided, which can be applied to a computer device. The method includes the following steps:
[0027] Step S101: Sample the target audio at different sampling rates to obtain multiple audios corresponding to different sampling rates.
[0028] Exemplarily, the sampling rates may include sampling rates selected from a high sampling rate range, sampling rates selected from a medium sampling rate range, and sampling rates selected from a low sampling rate range, such as 16 kHz, 8 kHz, and 4 kHz. The target audio may be singing, speech, etc. After determining different sampling rates, the target audio can be sampled to obtain multiple audios, and different audios correspond to different sampling rates.
[0029] In the above processing method, the sampling rates used to sample the target audio cover high, medium, and low sampling rates, which can better capture the characteristics of the target audio and are beneficial to authenticating whether the target audio is synthesized audio.
[0030] Step S102: Extract at least one of the constant Q transform cepstral coefficients, linear frequency cepstral coefficients, Mel frequency cepstral coefficients, fundamental frequency, and short-time energy features of each audio.
[0031] For synthesized audio, compared with non-synthesized audio, the performance of some features is different. Some synthesized audio has different performances in terms of constant Q transform cepstral coefficients, and some synthesized audio has different performances in terms of linear frequency cepstral coefficients. Based on this, in this application, among the features of constant Q transform cepstral coefficients (CQCC), linear frequency cepstral coefficients (LFCC), Mel frequency cepstral coefficients (MFCC), and fundamental frequency (f0) of audio, some features or all features are used to determine whether the target audio is synthesized audio.
[0032] When using some features, since the used partial features belong to the above-mentioned features, the accuracy of recognition can be guaranteed to a certain extent. Moreover, when using partial features for recognition, only some features need to be extracted, thereby improving the recognition efficiency.
[0033] In the case of using partial features for recognition, when extracting partial features from audio with different sampling rates, the extracted partial features can be the same or different. Exemplarily: If the target audio is sampled at 16 kHz, 8 kHz, and 4 kHz to obtain Audio_1, Audio_2, and Audio_3; if the extracted partial features are the same, all being CQCC, then the CQCC of Audio_1, the CQCC of Audio_2, and the CQCC of Audio_3 can be extracted; if the extracted partial features are different, then the CQCC of Audio_1, the LFCC of Audio_2, and the f0 of Audio_3 can be extracted.
[0034] After obtaining the partial features of audio with different sampling rates, enter step S103: Input the partial features of audio with different sampling rates into a pre-trained discrimination model, and obtain the result output by the discrimination model to determine whether the target audio is synthesized audio.
[0035] When using all features, the CQCC, LFCC, MFCC, f0, and Energy of each audio can be extracted, and then enter step S103: Input all the features of each audio into a pre-trained discrimination model, and obtain the result output by the discrimination model to determine whether the target audio is synthesized audio; when using all features, more features are covered, and no matter what the differences in the features of the audio are, they can be identified as much as possible, and the robustness is stronger.
[0036] Step S103: Input at least one of the constant Q transform cepstral coefficients, linear frequency cepstral coefficients, mel frequency cepstral coefficients, fundamental frequency, and short-time energy features of each audio into a pre-trained discrimination model, and obtain the result output by the discrimination model.
[0037] Among them, the pre-trained discrimination model can identify whether an audio is a synthetic audio based on the features of the audio. To enable the discrimination model to have this ability, a corresponding training set can be constructed and the discrimination model can be trained using the corresponding training set. Taking the extraction of CQCC, LFCC, MFCC, f0, and Energy of the audio as an example, the method for constructing the training set is introduced as follows: After obtaining multiple sample audios through different sampling rates, obtain the CQCC, LFCC, MFCC, f0, and Energy of each sample audio, and use them as the data input to the discrimination model during training. In addition, label each sample audio, and the label can represent the true probability that the sample audio is a synthetic audio. Based on the CQCC, LFCC, MFCC, f0, and Energy of each sample audio and the label given to the sample audio, a training set is formed.
[0038] When training the discrimination model using the above training set, the CQCC, LFCC, MFCC, f0, and Energy of the sample audio can be input into the discrimination model. The discrimination model processes the input features according to the model parameters and outputs the corresponding result. The output result can represent the predicted probability that the sample audio is a synthetic audio. Based on the difference between the predicted probability represented by the output result and the true probability represented by the label, a loss value is obtained. Among them, the greater the difference, the greater the loss value. After obtaining the loss value, use the loss value to adjust the model parameters in the discrimination model until the loss value is less than the set value. When the loss value is less than the set value, it indicates that the discrimination model can accurately identify whether the sample audio is a synthetic audio.
[0039] In addition, to improve the recognition accuracy of the discrimination model, the sampling rate used in the training stage can be used as the sampling rate in the application stage. For example, if the sampling rates used in the training stage are 16 kHz, 8 kHz, and 4 kHz, then the sampling rates used in the application stage (corresponding to the above step S101) can be 16 kHz, 8 kHz, and 4 kHz.
[0040] Step S104: Determine whether the target audio is a synthetic audio according to the result output by the discrimination model.
[0041] The above discrimination model can be obtained through deep learning, and the result it outputs can be a probability. When the probability output by the discrimination model is greater than the preset value, it is determined that the target audio is a synthetic audio. When the probability output by the discrimination model is less than or equal to the preset value, it is determined that the target audio is a non-synthetic audio.
[0042] In the above method for identifying synthesized audio, through an identification model, it automatically identifies whether the target audio is synthesized audio, improving the identification efficiency of synthesized audio. Moreover, this application samples the target audio at different sampling rates, synthesizes the features of the audio at different sampling rates of the target audio, and determines whether the target audio is synthesized audio, improving the recognition accuracy. Additionally, compared with non-synthesized audio, synthesized audio has different manifestations in some features. For example, some synthesized audio has different manifestations in constant Q transform cepstral coefficients, and some synthesized audio has different manifestations in linear frequency cepstral coefficients. Based on this characteristic of synthesized audio, in the features of constant Q transform cepstral coefficients, linear frequency cepstral coefficients, mel frequency cepstral coefficients, fundamental frequency, and short-time energy features of the audio, some features or all features are used to determine whether the target audio is synthesized audio. When using some features, since the some features used belong to the above features, the recognition accuracy can be guaranteed to a certain extent. Moreover, when using some features, only some features need to be extracted, thus improving the recognition efficiency. When using all features, the constant Q transform cepstral coefficients, linear frequency cepstral coefficients, mel frequency cepstral coefficients, fundamental frequency, and short-time energy features of each audio can be extracted. The covered features are more, and no matter which aspect the features of the audio are different, they can be identified as much as possible, with strong robustness. Additionally, starting from four dimensions: time domain (reflected in multi-sampling rate), frequency domain (reflected in constant Q transform cepstral coefficients, linear frequency cepstral coefficients, mel frequency cepstral coefficients), pitch (reflected in fundamental frequency), and volume (reflected in short-time energy features), the features of synthesized audio can be better captured and identified.
[0043] When the sampling rate is different, the influence on the fundamental frequency and short-time energy features of the audio is small, while the influence on the constant Q transform cepstral coefficients, linear frequency cepstral coefficients, and mel frequency cepstral coefficients of the audio is large. Therefore, when extracting the constant Q transform cepstral coefficients, linear frequency cepstral coefficients, mel frequency cepstral coefficients, fundamental frequency, and short-time energy features of each audio, the following steps can be included: extract for multiple audios to obtain the constant Q transform cepstral coefficients, linear frequency cepstral coefficients, and mel frequency cepstral coefficients of each audio respectively; extract the fundamental frequency from any one of the audios as the fundamental frequency of each audio; extract the short-time energy feature from any one of the audios as the short-time energy feature of each audio.
[0044] Exemplarily, the target audio is sampled using 16 kHz, 8 kHz, and 4 kHz to obtain Audio_1, Audio_2, and Audio_3. The constant Q transform cepstral coefficients, linear frequency cepstral coefficients, and Mel frequency cepstral coefficients of each audio can be extracted, obtaining the constant Q transform cepstral coefficients, linear frequency cepstral coefficients, and Mel frequency cepstral coefficients of Audio_1 itself, the constant Q transform cepstral coefficients, linear frequency cepstral coefficients, and Mel frequency cepstral coefficients of Audio_2 itself, and the constant Q transform cepstral coefficients, linear frequency cepstral coefficients, and Mel frequency cepstral coefficients of Audio_3 itself.
[0045] In addition, any one of Audio_1, Audio_2, and Audio_3 is selected, and the fundamental frequency and short-time energy features are extracted from the selected audio. The obtained fundamental frequency and short-time energy features are shared by Audio_1, Audio_2, and Audio_3.
[0046] In the above processing method, the fundamental frequency and short-time energy features that can be shared by multiple audios are extracted from any one audio, and the constant Q transform cepstral coefficients, linear frequency cepstral coefficients, and Mel frequency cepstral coefficients of multiple audios are extracted. While improving the audio feature extraction efficiency, it is also possible to extract unique features of audios with different sampling rates.
[0047] Furthermore, in the case where the fundamental frequency and short-time energy features are shared by multiple audios, the computer device can splice the constant Q transform cepstral coefficients, linear frequency cepstral coefficients, Mel frequency cepstral coefficients, fundamental frequency, and short-time energy features of the same audio to obtain the splicing features of each audio; and input the splicing features of each audio into a pre-trained discrimination model.
[0048] Exemplarily, the target audio is sampled using 16 kHz, 8 kHz, and 4 kHz to obtain Audio_1, Audio_2, and Audio_3, and the CQCC, LFCC, MFCC, f0, and Energy of each audio are obtained. Then, the CQCC, LFCC, MFCC, f0, and Energy of the same audio are spliced to obtain the splicing features of the audio. For example, the CQCC, LFCC, MFCC, f0, and Energy of Audio_1 are spliced to obtain the splicing features of Audio_1; after obtaining the splicing features of Audio_1, the splicing features of Audio_2, and the splicing features of Audio_3, since these audios are sampled from the same target audio, the splicing features of these audios can be input into a pre-trained discrimination model, and the discrimination model outputs the probability that the above target audio is a synthetic audio based on the splicing features of these audios.
[0049] Furthermore, for multiple audios sampled from the same target audio, the splicing features of these audios can form a matrix, which can be called the first matrix. Then, the first matrix is input into a pre-trained discrimination model. Using multiple convolutional kernels of the discrimination model, convolutional calculations are respectively performed on the first matrix to obtain the convolutional calculation results corresponding to each convolutional kernel, and the convolutional calculation results corresponding to each convolutional kernel are spliced to obtain a second matrix; using the linear layer and rectified linear unit of the discrimination model, dimensionality reduction processing is performed on the second matrix; using the bidirectional recurrent unit and linear layer of the discrimination model, global modeling is performed on the result of the dimensionality reduction processing and dimensionality reduction processing is performed on the global modeling result; according to the result of the dimensionality reduction processing of the global modeling result, the result output by the discrimination model is obtained.
[0050] The architecture of the discrimination model used in this application is as Figure 2 shown. The following combines Figure 2 to introduce the processing process inside the discrimination model.
[0051] For T audios sampled from the same target audio, after obtaining the splicing features of the T audios, the splicing features of the T audios can be formed into a first matrix, and then the first matrix is input into a pre-trained discrimination model. The dimension of the first matrix can be denoted as [T, N], where N represents the number of features included in the splicing features.
[0052] 1) To obtain local and global features in the time-frequency domain, different convolutional kernels (the convolutional kernel sizes can be 1, 3, 5, 7, 9 respectively, and the number of channels is 64 for all) are used to perform convolutional calculations on the above first matrix respectively to obtain multiple convolutional calculation results, and then the convolutional calculation results are spliced to obtain a second matrix; since the dimension of the above first matrix is [T, N], therefore, the dimension of the second matrix obtained after being processed by the above convolutional kernels is [T, 64 * 5];
[0053] 2) To enhance the non-linear ability of the discrimination model, through the linear layer (Linear) and rectified linear unit (ReLU) of the discrimination model, dimensionality reduction processing is performed on the second matrix, and the dimension of the second matrix is reduced from [T, 64 * 5] to [T, 64].
[0054] 3) To extract key information in the time domain, a bidirectional gated recurrent unit (Bi-GRU) and a linear layer (Linear) are used to perform global modeling on the result of the dimensionality reduction processing obtained in 2), and dimensionality reduction processing is performed on the global modeling result, and the dimension of the dimensionality reduction processing result is [T, 1].
[0055] 4) Use the sigmoid activation function and mean operation of the discrimination model to process the dimensionality reduction result of dimension [T, 1] obtained in 3). The processed result is the result output by the discrimination model, and the result output by this discrimination model can represent the probability that the target audio is a synthetic audio, and the value range is [0, 1].
[0056] In the above processing method, convolution, multiple dimensionality reduction, and global modeling processing are performed using the discrimination model to improve the accuracy of the discrimination result of the synthetic audio.
[0057] In one embodiment, the computer device can train the discrimination model through the following steps: obtain an un-augmented data set; the un-augmented data set includes non-synthetic audio and synthetic audio; perform noise addition processing and reverberation addition processing on the un-augmented data set to obtain an augmented data set; based on the augmented data set, obtain a training set, and use the training set to train the discrimination model.
[0058] The non-synthetic audio belongs to real audio that is not synthesized, and can include speech and singing. The speech can be the recording data commonly used in speech recognition, and the singing can be the singing data in karaoke.
[0059] The synthetic audio can come from the ASVspoof2021 data set (this data set includes audio obtained by speech synthesis and speech conversion) and audio generated by a singing synthesis tool.
[0060] For example, the durations of the non-synthetic audio and the synthetic audio are both 5 hours. In order to augment the un-augmented data set and improve the generalization ability of the discrimination model, noise addition and reverberation addition processing can be performed on the audio of the un-augmented data set. Finally, the durations of the obtained non-synthetic audio and synthetic audio are both 7 hours, forming an augmented data set, and this augmented data set covers a variety of timbres, a variety of voices, and a wide pitch range.
[0061] Further, after obtaining the augmented data set, frame windowing processing can be performed on each audio in the augmented data set; extract the constant Q transform cepstral coefficients, linear frequency cepstral coefficients, mel frequency cepstral coefficients, fundamental frequency, and short-time energy features of each audio after frame windowing processing to form a training set.
[0062] Before feature extraction, frame windowing processing is first performed on the audio of the augmented data set. The frame windowing processing is exemplary: for an audio with a duration of D, T frames can be obtained, and the duration of each frame is D / T.
[0063] The features extracted in this application include CQCC, LFCC, MFCC, f0, and Energy. Among them, CQCC, LFCC, MFCC, and the pYin algorithm used to extract f0 can all be implemented by the open-source librosa library. Energy is the mean square value of the Mel spectrogram of a single-frame audio. After extracting features from the augmented dataset to obtain the training set, the discriminative model can be trained using the training set; if the architecture of the discriminative model is Figure 2 the architecture shown, then, during the training phase, the processing process inside the discriminative model is similar to that introduced in other embodiments and will not be elaborated here.
[0064] It should be noted that the loss function used in the training phase is the cross-entropy loss function. Based on the loss value, gradient backpropagation is performed to update the parameters. The learning rate is 0.001, and the optimizer is the Adam optimizer. When the result output by the discriminative model is greater than 0.5, it indicates that the audio to be discriminated is a non-synthetic audio. When the result output by the discriminative model is less than or equal to 0.5, it indicates that the audio to be discriminated is a synthetic audio. When the accuracy of the validation set reaches the maximum value and the discriminative model is basically convergent, the training ends.
[0065] In one embodiment, the steps of obtaining different sampling rates include: determining a preset high sampling rate interval, a medium sampling rate interval, and a low sampling rate interval; selecting at least one sampling rate from the high sampling rate interval, at least one sampling rate from the medium sampling rate interval, and at least one sampling rate from the low sampling rate interval to form different sampling rates.
[0066] When selecting sampling rates from different intervals, the difference between the sampling rate selected from the high sampling rate interval and the sampling rate selected from the medium sampling rate interval, and the difference between the sampling rate selected from the medium sampling rate interval and the sampling rate selected from the low sampling rate interval can be equal and can be a fixed value. That is, after selecting a sampling rate from any sampling rate interval, according to this fixed value, the corresponding sampling rates are selected from the other two sampling rate intervals. For example, after selecting 8 kHz from the medium sampling rate interval, if the fixed value is 4, then a sampling rate of 4 kHz can be selected from the low sampling rate interval, and a sampling rate of 12 kHz can be selected from the high sampling rate interval.
[0067] Furthermore, the computer device can also determine multiple sampling rate groups; the sampling rates within the same group are respectively selected from the high sampling rate interval, the medium sampling rate interval, and the low sampling rate interval. As Figure 3 shown, a sampling rate a1 can be selected from the preset low sampling rate interval, a sampling rate b1 can be selected from the medium sampling rate interval, and a sampling rate c1 can be selected from the high sampling rate interval. The sampling rates a1, b1, and c1 are used as a sampling rate group to obtain the first sampling rate group. Similarly, the second sampling rate group can be obtained in the above manner.
[0068] After obtaining multiple sampling rate groups, each sampling rate group can be used to sample the target audio to obtain the audio corresponding to the same sampling rate group.
[0069] For example, use a1, b1, and c1 in the first sampling rate group to sample the target audio to obtain the audio corresponding to the first sampling rate group. Another example is to use a2, b2, and c2 in the second sampling rate group to sample the target audio to obtain the audio corresponding to the second sampling rate group.
[0070] Then, extract the CQCC, LFCC, MFCC, f0, and Energy of each audio, and input the CQCC, LFCC, MFCC, f0, and Energy of the audio corresponding to the same sampling rate group into a pre-trained discrimination model, and the discrimination model outputs the result corresponding to this sampling rate group.
[0071] For example, input the CQCC, LFCC, MFCC, f0, and Energy of the audio corresponding to the first sampling rate group into the discrimination model, and the discrimination model outputs the result corresponding to the first sampling rate group; another example is to input the CQCC, LFCC, MFCC, f0, and Energy of the audio corresponding to the second sampling rate group into the discrimination model, and the discrimination model outputs the result corresponding to the second sampling rate group. Among them, f0 and Energy can be shared by the audio of all sampling rate groups.
[0072] After obtaining the results corresponding to each sampling rate group output by the discrimination model, these results are integrated to determine whether the target audio is a synthetic audio.
[0073] For example, the result corresponding to the first sampling rate group output by the discrimination model is 0.3, and the result corresponding to the second sampling rate group output by the discrimination model is 0.4. Then, these results can be integrated by taking the average to obtain a comprehensive value of 0.35. If the comprehensive value is greater than the preset value, it is determined that the target audio is a synthetic audio; if the comprehensive value is greater than the preset value, it is determined that the target audio is a non-synthetic audio.
[0074] In the above processing method, by integrating the results corresponding to different sampling rate groups output by the discrimination model for discrimination, the accuracy of synthetic audio discrimination can be improved.
[0075] It should be understood that although the steps in the flowcharts involved in the above-described embodiments are shown sequentially in the direction of the arrows, these steps are not necessarily executed sequentially in the order indicated by the arrows. Unless there is a clear indication in this article, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-described embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.
[0076] In one embodiment, a computer device is provided, and its internal structure diagram can be as Figure 4 shown. The computer device includes a processor, a memory, and a network interface connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store the identification data of the synthesized audio. The network interface of the computer device is used to communicate with an external terminal through a network connection. The computer device also includes an input / output interface, which is a connection circuit for exchanging information between the processor and external devices. They are connected to the processor through a bus, abbreviated as the I / O interface. When the computer program is executed by the processor, it implements a method for identifying synthesized audio.
[0077] Those skilled in the art can understand that Figure 4 the structure shown in
[0078] is only a block diagram of a part of the structure related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have a different component layout.
[0079] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, it implements the steps in the above-mentioned method embodiments.
[0080] In one embodiment, a computer program product is provided, on which a computer program is stored, and the computer program is executed by a processor to perform the steps in the above-mentioned method embodiments.
[0081] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties.
[0082] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through a computer program. The above computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above method embodiments. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in this application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, or optical memory, etc. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.
[0083] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0084] The above embodiments only represent several implementation manners of this application, and their descriptions are relatively specific and detailed, but they should not be construed as limitations on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of this application, several modifications and improvements can still be made, and these all belong to the protection scope of this application. Therefore, the protection scope of this application patent should be subject to the appended claims.
Claims
1. A method for identifying synthesized audio, characterized in that, The method includes: Sampling the target audio using multiple sampling rate groups to obtain the audio corresponding to each sampling rate within the multiple sampling rate groups; Extracting at least two of the constant Q transform cepstral coefficients, linear frequency cepstral coefficients, mel frequency cepstral coefficients, fundamental frequency, and short-time energy features of each of the audio; For each of the sampling rate groups, according to at least two of the constant Q transform cepstral coefficients, linear frequency cepstral coefficients, mel frequency cepstral coefficients, fundamental frequency, and short-time energy features of the audio corresponding to each sampling rate within the sampling rate group, obtaining the splicing features of the audio corresponding to each sampling rate within the sampling rate group to form a first matrix corresponding to the sampling rate group; For each of the sampling rate groups, using a pre-trained discrimination model to perform convolution calculations on the first matrix corresponding to the sampling rate group to obtain multiple convolution calculation results, and obtaining the result corresponding to the sampling rate group output by the discrimination model according to a second matrix obtained by splicing the multiple convolution calculation results; Based on the results corresponding to each of the multiple sampling rate groups, determining whether the target audio is a synthetic audio.
2. The method according to claim 1, wherein The extracting at least two of the constant Q transform cepstral coefficients, linear frequency cepstral coefficients, mel frequency cepstral coefficients, fundamental frequency, and short-time energy features of each of the audio includes: Respectively extracting the constant Q transform cepstral coefficients, linear frequency cepstral coefficients, and mel frequency cepstral coefficients of each of the audio from each of the audio; Extracting the fundamental frequency from any one of the audio as the fundamental frequency of each of the audio; Extracting the short-time energy feature from any one of the audio as the short-time energy feature of each of the audio.
3. The method according to claim 1, characterized in that According to at least two of the constant Q transform cepstral coefficients, linear frequency cepstral coefficients, mel frequency cepstral coefficients, fundamental frequency, and short-time energy features of the audio corresponding to each sampling rate within the sampling rate group, obtaining the splicing features of the audio corresponding to each sampling rate within the sampling rate group includes: Splicing at least two of the constant Q transform cepstral coefficients, linear frequency cepstral coefficients, mel frequency cepstral coefficients, fundamental frequency, and short-time energy features of the audio corresponding to the same sampling rate within the sampling rate group to obtain the splicing features of the audio corresponding to each sampling rate within the sampling rate group.
4. The method according to claim 1, characterized in that, Using a pre-trained discrimination model to perform convolution calculations on the first matrix corresponding to the sampling rate group to obtain multiple convolution calculation results includes: Using multiple convolution kernels of a pre-trained discrimination model to perform convolution calculations on the first matrix corresponding to the sampling rate group respectively to obtain the convolution calculation results corresponding to each convolution kernel; According to a second matrix obtained by splicing the multiple convolution calculation results, obtaining the result corresponding to the sampling rate group output by the discrimination model includes: Splicing the convolution calculation results corresponding to each of the convolution kernels to obtain a second matrix; Using the linear layer and rectified linear unit of the discrimination model to perform dimensionality reduction processing on the second matrix; Using the bidirectional recurrent neural network and linear layer of the discrimination model to perform global modeling on the result obtained by the dimensionality reduction processing and perform dimensionality reduction processing on the global modeling result; According to the result of the dimensionality reduction processing of the global modeling result, obtaining the result corresponding to the sampling rate group output by the discrimination model.
5. The method according to any one of claims 1 to 4, characterized in that, The method further includes: Obtain an unexpanded dataset; the unexpanded dataset includes non-synthetic audio and synthetic audio; Perform noise addition processing and reverberation addition processing on the unexpanded dataset to obtain an expanded dataset; Based on the expanded dataset, obtain a training set, and use the training set to train a discrimination model.
6. The method according to claim 5, characterized in that, The obtaining of the training set based on the expanded dataset includes: Perform frame addition and windowing processing on each audio in the expanded dataset; Extract at least two of the constant Q transform cepstral coefficients, linear frequency cepstral coefficients, mel frequency cepstral coefficients, fundamental frequency, and short-time energy features of each audio after frame addition and windowing processing to obtain a training set.
7. The method according to claim 1, wherein The method further includes: Determine a preset high sampling rate interval, a medium sampling rate interval, and a low sampling rate interval; Select at least one sampling rate from the high sampling rate interval, select at least one sampling rate from the medium sampling rate interval, and select at least one sampling rate from the low sampling rate interval to form a sampling rate group.
8. The method according to claim 1, characterized in that, Integrate the results corresponding to each of the multiple sampling rate groups to determine whether the target audio is synthetic audio, including: Through an averaging method, comprehensively process the results corresponding to each of the multiple sampling rate groups to obtain a comprehensive result; Determine whether the target audio is synthetic audio according to the comprehensive result.
9. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, the method according to any one of claims 1 to 8 is implemented.
10. A computer program product comprising a computer program, characterized in that, When the computer program is executed by the processor, the method according to any one of claims 1 to 8 is implemented.
Citation Information
Patent Citations
Snoring sound signal identification method
CN110570880A