Method and System for Detecting and Locating Stitched Audio Based on Spectrogram Segmentation
By adopting a full convolutional network and an encoder-decoder architecture in audio splicing detection and positioning, extracting the MFCC characteristics of the audio segments and calculating the binary prediction mask and element ratio, the problem of insufficient audio splicing detection and positioning performance in the prior art is solved, and an audio splicing detection and positioning method with high accuracy and scalability is realized.
Patent Information
- Application Number
- CN202210368335.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-08
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2042-04-08
AI Technical Summary
The existing audio splicing detection and positioning methods have shortcomings in detection and positioning performance, especially when the signal-to-noise ratio of the splicing segment is close, the performance of the noise level-based method is degraded, the ENF-based method is restricted by law, and the neural network method can only infer whether it is spliced but cannot be positioned.
Using a spliced audio detection and positioning method based on spectral image segmentation, the Mel spectral coefficient (MFCC) characteristics of the audio segment are extracted through the full convolutional network (FCN) and the encoder-decoder architecture, the binary prediction mask and element ratio are calculated, and whether the audio segment is a spliced segment is determined, and the minimum positioning area is located.
It improves the accuracy of audio stitching detection and positioning, can intuitively display the positioning area to a certain extent, effectively alleviate the problem of data set mismatch, is scalable, and is suitable for large-scale and long-term audio detection and positioning scenarios.
Smart Images

Figure CN114819067B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for detecting tampered audio, and more particularly to a method for detecting and locating spliced audio based on spectrogram segmentation and its application in the field of digital audio forensics. This method belongs to the field of multimedia privacy protection in the field of information security technology. Background Art
[0002] The methods of audio tampering are mainly divided into Copy-Move and Splicing. Among them, the audio Copy-Move operation is to intercept a small audio segment from a piece of audio and paste the intercepted audio segment to another position in the same piece of audio to construct a new audio segment; the audio splicing technology is to insert another piece of audio at the beginning, middle or end of the audio to form a new audio segment; the two methods of audio tampering are mainly used to change the original content of the audio to achieve the purpose of making false audio; audio tampering detection and positioning mainly use characteristics such as background noise inconsistency and speech feature similarity combined with machine learning methods to judge whether the audio file to be tested has been tampered with.
[0003] With the wide use of online audio editing and processing tools, it has become easier to create tampered audio without perceivable traces. Audio splicing can be divided into inserting another piece of audio in the middle or at the end of the audio to construct a new audio segment: the former inserts a small piece of audio at a certain point in the middle of the complete audio to form a new audio; the latter splices a small piece of audio to the end of the complete audio to form a new audio. Audio Copy-Move and audio splicing reduce the reliability of audio as judicial evidence and are not conducive to the protection of intellectual property rights. In addition, these spliced audios can be used to spread fake news, having a negative impact on society. Therefore, the ability to detect whether an audio recording has been spliced is a task of great interest to the audio forensics community.
[0004] In the past few decades, various studies have been conducted on audio splicing detection and localization. According to the detection principle, the detection of spliced audio can be roughly divided into three categories: detection based on background noise, detection based on ENF, and detection based on deep learning. First, due to the inconsistency of the noise level caused by audio splicing operations, researchers have developed audio splicing detection methods based on the local noise level of audio signals. For example, the spectral entropy method (SE) is used to determine the length of each syllable, calculate the variance of the background noise of each syllable, and then judge whether there is an audio with heterologous splicing forgery by comparing the similarity of the variances of the background noise of each syllable (Reference: Meng, X., Li, C., Tian, L.: Detecting audio splicing forgery algorithm based on local noise level estimation. In: 2018 5th international conference on systems and informatics (ICSAI). pp. 861-865. IEEE 2018); a noise estimation algorithm with parameter optimization is used to extract the noise signal of the suspicious speech, and the statistic of the Mel frequency feature of the estimated noise signal is calculated to determine the detection splicing trace (Reference: Yan, D., Dong, M., Gao, J.: Exposing speech transsplicing forgery with noise level inconsistency. Security and Communication Networks 2021). However, when the signal-to-noise ratio between the spliced audio segments is close to or even the same, the performance of the audio splicing detection method based on the noise level will drop sharply. In addition, based on the fact that inserting one audio segment into another audio recording will cause abnormal changes in the electric network frequency (ENF) signal, it is a good method to detect spliced audio by analyzing the ENF signal.Some researchers have proposed a method that applies a wavelet filter to the ENF signal to highlight abnormal ENF variations and uses autoregressive coefficients to train a classifier under a supervised learning framework to detect spliced audio segments (Reference: Lin, X., Kang, X.: Supervised audio tampering detection using an autoregressive model. In: 2017 IEEE international conference on acoustics, speech and signal processing (ICASSP). pp. 2142-2146. IEEE 2017); subsequently, some researchers have used multiple ENF features as input feature vectors of a convolutional neural network to detect spliced audio (Reference: Mao, M., Xiao, Z., Kang, X., Li, X., Xiao, L.: Electric network frequency based audio forensics using convolutional neural networks. In: IFIP International Conference on Digital Forensics. pp. 253-270. Springer 2020). However, due to legal restrictions, obtaining a concurrent reference dataset of the power system has been severely limited, which challenges the practicality of ENF-based audio splicing detection methods. Currently, Convolutional Neural Networks (CNNs) have been introduced into the field of audio splicing detection. When CNNs are introduced into audio splicing detection, the spectrogram of the audio segment is directly input into the convolutional neural network, and a classifier based on the convolutional neural network is trained to detect spliced speech (Reference: Yan, D., Dong, M., Gao, J.: Exposing speech trans-splicing forgery with noise level inconsistency. Security and Communication Networks 2021). However, neural network-based methods can only infer whether a given audio has been spliced and cannot locate the spliced segments. Based on the above overview, although some audio splicing detection and localization methods have achieved effective performance, new technologies are still needed to improve detection and localization performance. To the best of our knowledge and literature review, the encoder-decoder architecture has not been used in the research of audio splicing detection and localization.
[0005] After a patent search, the relevant patent applications in the field of the present invention are as follows:
[0006] The Chinese patent "A method for detecting multiple forgery operations of speech based on RNN" with the patent application number CN111564163A discloses a method for detecting multiple forged speeches based on a recurrent neural network. This method is based on the dependency relationship between linear spectral coefficients and audio frames, and uses a recurrent neural network (RNN) to learn the inherent characteristics of spectral coefficients, thereby effectively improving the accuracy of forged speech detection. Since this invention does not involve the operation of detecting splicing tampering of audio, this method is significantly different from the design concept and specific implementation method of the present invention. Summary of the Invention
[0007] The purpose of the present invention is to accurately segment the spectrogram of audio, and then accurately determine whether a specific length block of audio is a spliced audio block. On this basis, a method for detecting and locating audio splicing with high accuracy is finally designed.
[0008] Compared with other methods for detecting and locating spliced audio, the present invention adopts a technology in the image scene style of the visual field, and defines three variables: the minimum localization region length (The smallest localization region, L slr ), the element ratio (The ratio, ρ), and the final determination threshold (Threshold, T). Finally, according to the binary mask image (Binary output mask) output by the fully convolutional network (FCN), it is calculated whether a certain minimum localization region block is a spliced audio block. It can be seen that the method proposed by the present invention is different from any previous method for detecting and locating audio splicing, and is particularly suitable for the detection and localization scenarios of large-scale and long-duration audio.
[0009] According to research, the existing methods for detecting and locating audio splicing have the following three limitations: First, due to the inconsistency of the noise level caused by audio splicing operations, when the signal-to-noise ratio between splicing segments is close to or even the same, the performance of the audio splicing detection method based on the noise level will drop sharply; Second, for the method of detecting spliced audio based on analyzing the power grid frequency signal, due to legal restrictions, it is greatly restricted to obtain the concurrent reference data set of the power system, which challenges the practicality of the ENF-based audio splicing detection method; Finally, for the method of detecting spliced audio based on neural networks, this type of method can only infer whether a given audio is spliced and cannot locate the spliced audio segment.
[0010] Specifically, the technical solution adopted by the present invention is as follows:
[0011] A method for detecting and locating spliced audio based on spectrogram segmentation, comprising the following steps:
[0012] 1) Detection segment division: Divide the audio to be tested into several audio segments S to be detected according to the minimum positioning area length L for audio splicing tampering positioning slr , where each audio segment is composed of continuous sampling points, and the length of each audio segment is L g . slr .
[0013] 2) Preprocessing: Extract the spectrogram feature F of the audio segment S g , and form a spliced audio segment S′ by combining t audio segments S according to the size of the network input g , and splice the corresponding spectrogram features F g into a spectrogram feature F′ to be input into the network g . g . g .
[0014] 3) Calculate the binary prediction mask: Input the spliced spectrogram feature F′ g into the trained ASLNet (Audio Splicing Detection and Localization Network) network, and calculate the binary prediction mask corresponding to the spliced audio segment S′ g , where 1 in the binary prediction mask represents the spliced sample points, and 0 represents the original sample points.
[0015] 4) Calculate the element ratio ρ: According to the calculation result of step 3), calculate the ratio ρ of the number of sample points that are spliced sample points (element is 1) in the binary prediction mask of each audio segment S g to the total number of sample points.
[0016] 5) Spliced segment judgment: According to the calculation result of step 4), compare the size of ρ and the preset judgment threshold T value, and then judge whether the segment S g is a spliced segment, where when ρ > T, the segment is a spliced segment, otherwise, it is an original segment.
[0017] 6) For the N divided audio segments S′ g , execute steps 2) to 5) to sequentially judge whether all segments S of the audio to be tested g are spliced segments.
[0018] Now, the design and training of the ASLNet network proposed by the present invention, the extraction of the spectrogram feature F g and the definition and calculation of the ratio ρ will be described in detail as follows.
[0019] [1] Spectrogram feature F g and extraction of the binary true mask:
[0020] The present invention extracts Mel-Frequency Cepstral Coefficients (MFCC) features from the audio segment S g as the spectrogram feature. The extraction process of the spectrogram feature is as Figure 1 shown. The detailed MFCC extraction process is as follows: First, the pre-emphasis module is used to enhance the energy of the signal at high frequencies, and the short-time Fourier transform (STFT) of the pre-emphasized signal is calculated using a periodic Hamming window with a length of 2048 samples and an overlap of 512 samples. Then, the Mel-Filter is used to map the energy to the Mel frequency scale, and the logarithm is taken to create a power map. Finally, the discrete cosine transform is used to calculate the transformed coefficients containing important energy, that is, the Mel-spectrum coefficients.
[0021] For an audio segment with a minimum localization unit of 16000 sampling points (i.e., L slr = 16000), the first 24 coefficients are selected as the static MFCC features, the dynamic coefficients and acceleration coefficients are calculated, and they are concatenated after the static coefficients to form 72 feature vectors. Therefore, the shape of the MFCC feature matrix is 72×32, where 72 is the number of coefficients and 32 is the number of frames. In addition, in order to train the decoder network, a binary true mask (Ground Truth Mask) is designed for each MFCC feature matrix. The binary true mask consists of 0 or 1 elements, and its size is 72×32. For the original audio segment, each element in the corresponding binary true mask is 0, while for the spliced audio segment, each element in the corresponding binary true mask is 1. In the present invention, the length L slr can be set according to the length of the audio segment that the user wants to localize in the actual application.
[0022] [2] Design and training of the ASLNet network:
[0023] The overall flowchart of the method for detecting and localizing spliced audio based on spectrogram segmentation is as Figure 2As shown in the figure, the fully convolutional network is the core of the entire process; the fully convolutional network structure is a commonly used network structure in current semantic segmentation algorithms, which consists of an encoder-decoder. The encoder performs convolution and downsampling to capture context information, while the decoder is responsible for deconvolution and upsampling to predict pixel-level class labels. Many encoder-decoder architectures have been proposed (FCN, U-Net, and SegNet) and successfully applied to the field of image pixel segmentation. The basic network architecture of the ASLNet of the present invention is a modified FCN-VGG16, which consists of a VGG16 encoder and a decoder with a residual structure. The goal of the VGG16 encoder is to capture the context representation of acoustic features, while the goal of the decoder is to convert the intermediate feature map into a binary prediction mask.
[0024] As Figure 3 shown, the VGG blocks are stacked to construct the VGG16 encoder, where each VGG block consists of two to three convolutional blocks followed by a max-pooling layer, with a total of 13 convolutional layers and 5 max-pooling layers; the convolutional block consists of a convolutional layer, a batch normalization layer, and a linear unit (ReLU) activation function. For all convolutional layers, the same kernel size of 3×3 is used, and the convolution stride is 1; in addition, the padding size is 1 to keep the output size the same after each convolutional layer. The size of the max-pooling layer is 2×2, and the stride is 2, which is used to halve the resolution after each VGG block. The purpose of the decoder is to reconstruct the binary ground truth mask using the basic information extracted by the VGG16 encoder. The decoder consists of two transposed convolutional layers and a SoftMax activation function. The kernel size of the first transposed convolution is 4×4, and the stride is 2; the kernel size of the second transposed convolution is 32×32, and the stride is 16. In addition, skip connections from the fourth VGG block to the first transposed convolution are used to aggregate the features learned at the lower layer to the higher layer. Finally, the SoftMax activation function is used to calculate the probability that the elements come from the concatenated audio segments. To train the ASLNet network, first determine the data size of the input network, and then concatenate the MFCC matrices of multiple audio segments and the corresponding binary ground truth masks together as the input and label of the ASLNet network.
[0025] [3] Definition and calculation of the ratio ρ:
[0026] This method determines whether a small piece of audio is concatenated audio based on the binary prediction mask output by the ASLNet network. Just as the elements in the binary ground truth mask are defined, for the original audio segment, each element in the corresponding binary ground truth mask is 0, while for the concatenated audio segment, each element in the corresponding binary ground truth mask is 1. Therefore, this method calculates the ratio ρ of the number of concatenated sample points in the mask to the total number of elements according to the binary prediction mask. The specific formula for ρ is as follows:
[0027]
[0028] Among them, Num represents the number of elements in the set. ρ calculated according to the above formula is compared with a preset threshold T to determine whether the audio block S g is a spliced sample. The specific formula is as follows:
[0029]
[0030] A spliced audio detection and positioning system based on spectrogram segmentation using the above method, which includes:
[0031] A detection segment division module, which is used to divide the audio to be tested into several audio segments S to be detected according to the minimum positioning area length L of audio splicing tampering positioning slr ; g ;
[0032] A preprocessing module, which is used to extract the spectrogram feature F of the audio segment S g , and splice t audio segments S g into an audio segment S' according to the size of the network input g , and splice the corresponding spectrogram features F g into a spectrogram feature F' to be input into the network g ; g ;
[0033] A binary prediction mask calculation module, which is used to input the spliced spectrogram feature F' g into the trained spliced audio detection and positioning network to calculate the binary prediction mask corresponding to the spliced audio segment S' g ;
[0034] An element ratio calculation module, which is used to calculate the ratio ρ of the number of sample points that are spliced sample points in the binary prediction mask of each audio segment S g to the total number of sample points;
[0035] A spliced segment judgment module, which is used to compare the size of ρ and the preset judgment threshold T value, and then judge whether the audio segment S g is a spliced segment; for the N audio segments S' divided g , it is judged in turn whether all the audio segments S of the audio to be tested g are spliced segments.
[0036] The beneficial effects of the spliced audio detection and positioning method based on spectrogram segmentation of the present invention on the related technical fields are as follows:
[0037] 1) It can improve the accuracy of detection and positioning. Since the image target segmentation technology in the visual field has been well-developed, there are many advanced network structures that can achieve a high recognition accuracy; applying these advanced network structures to the spectrogram of audio for positioning can also achieve a high accuracy.
[0038] 2) It can visually display the positioning area to a certain extent. Through the binary prediction mask map output by the ASLNet network, it can well display the audio block S g The proportion of the elements spliced in the spectrogram features. Since there is a certain corresponding relationship between the proportion of the elements and the audio sampling points, it can be located in the original waveform diagram of the audio, and the two-dimensional image positioning display can better show the possibility of tampering.
[0039] 3) It can effectively alleviate the problem of dataset mismatch to a certain extent. There are its own characteristics within the same audio segment. The trained ASLNet network can learn the internal characteristics of a complete audio, and its implementation does not depend on a specific training dataset. Therefore, the present invention has a large applicable range and can effectively analyze unknown audio datasets.
[0040] 4) It has scalability. In the present invention, for small segments with a length of L slr parameters such as the fully convolutional network structure and the final decision threshold (Threshold, T) can be adjusted according to the actual environmental requirements, so as to customize different splicing audio detection and positioning methods based on spectrogram segmentation for application in different speech splicing detection and positioning analysis scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] Figure 1 is the flowchart of audio spectrogram coefficient feature extraction in the present invention;
[0042] Figure 2 is the flowchart of splicing audio detection and positioning in the present invention;
[0043] Figure 3 is the schematic diagram of the Encoder-Decoder fully convolutional network in the present invention;
[0044] Figure 4 is the schematic diagram of the detection result after network iteration in the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0045] The following is a further description of the present invention through specific implementation examples in combination with the attached Figure 2 drawings.
[0046] The splicing audio detection and positioning method based on spectrogram segmentation proposed by the present invention has the following specific operation details:
[0047] 1) Detection segment division: Divide the audio to be tested into several audio segments S to be detected according to the minimum positioning area length L for audio splicing tampering positioning slr , where each audio segment is composed of consecutive sampling points, and the length of each audio segment is L g . slr .
[0048] 2) Preprocessing: Extract the spectrogram features F of the audio segment S g , and form t audio segments S into a spliced audio segment S' according to the size of the network input g , and splice the corresponding spectrogram features F g into a spectrogram feature F' to be input into the network g . g . g .
[0049] 3) Calculate the binary prediction mask: Input the spliced spectrogram feature F' into the trained ASLNet network, and calculate the binary prediction mask corresponding to the spliced audio segment S', where 1 in the binary prediction mask represents the spliced sample points and 0 represents the original sample points. g , g .
[0050] 4) Calculate the element ratio ρ: According to the calculation result of step 3), calculate the ratio ρ of the number of sample points that are spliced sample points (element is 1) in the binary prediction mask of each audio segment S to the total number of sample points. g .
[0051] 5) Spliced segment judgment: According to the calculation result of step 4), compare the size of ρ and the preset judgment threshold T value, and then judge whether the segment S is a spliced segment. When ρ > T, the segment is a spliced segment, otherwise, it is an original segment. g .
[0052] 6) For the N divided audio segments S', execute steps 2) to 5) to sequentially judge whether all segments S of the audio to be tested are spliced segments. g . g .
[0053] As can be seen from the above specific implementation manners: First, the present invention mainly realizes the calibration of the forgery area of the audio spectrogram through the fully convolutional network in the visual field, so as to realize the function of splicing and positioning the original waveform of the audio, and its implementation does not depend on specific background noise and ENF signals; Second, according to different actual application scenarios, the minimum positioning area creation length L can be changed slrBy using the preset value of T, detection results with different lengths of positioning intervals and different confidence levels are generated. Therefore, the present invention has a relatively wide range of applications and strong flexibility.
[0054] To highlight that the method proposed by the present invention is an effective spliced audio detection and positioning method, the following experimental configuration is adopted for the spliced audio detection and positioning experiment:
[0055] 1) Produce two spliced sample data sets: a 2-second audio data set (CNSet2s) and a 3-second audio data set (CNSet3s); First, cut the audio segments of the FMFCC-A corpus into 1-second, 2-second, and 3-second audio segments. Among them, the 2-second audio clips and 3-second audio clips are used as the original samples of the CNSet2s and CNSet3s data sets, with 44,727 and 44,669 audio segments respectively. Then, randomly select two non-homologous 1-second audio segments and connect them into a spliced audio segment (i.e., the splicing position is at the end of another audio segment). Randomly produce 86,073 2-second spliced audio segments to construct CNSet2s. In addition, randomly select a 1-second audio segment and a 2-second audio segment, insert the 1-second audio segment into the middle of the 2-second audio segment, and after randomly producing 85,865 3-second spliced audio segments, CNSet3s is completely produced.
[0056] 2) Data set division: Divide CNSet2s and CNSet3s according to the ratio of 6:2:2 into a training data set, a validation data set, and a test data set, which are used for the training of the network model, the selection of the model, and the prediction of the model performance respectively.
[0057] 3) Parameter extraction: First, in this experiment, define the minimum positioning region creation length L slr = 16000, that is, 16000 sampling points are used as the minimum positioning interval for each extraction of MFCC feature coefficients, obtain the spectrogram matrix of a 72×32 audio segment, and prepare the binary true mask matrix corresponding to each audio segment in the training set and validation set;
[0058] 4) Comparative splicing audio detection method: Since Jadhav is currently a neural network-based splicing audio detection method, it is compared with the present invention (Reference: Jadhav, S., Patole, R., Rege, P.: Audio splicing detection using convolutional neural network. In: 2019 10th International Conference on Computing, Communication and Networking Technologies (ICCCNT). pp. 1-5. IEEE 2019).
[0059] 5) Training and detection: Use the training set and validation set to train and optimize the ASLNet network of this method and the network in the Jadhav method. Set each batch to 64 audio files. After the model has gone through one epoch on the training set, it is validated once on the validation set. This loop runs for 200 epochs, and the model that achieves the best result on the validation set is selected as the final model for testing on the test set. In the testing phase, the MFCC matrix of the spectrogram features of the audio segments on the validation set, which have been extracted, are input into the trained network model to obtain the corresponding predicted binary mask.
[0060] 6) Calculate the proportion ρ of the number of elements equal to 1 in the obtained binary prediction mask to the total number of elements, and compare ρ with the predefined threshold T to determine whether the audio segment is a spliced segment. Thus, calculate the true positive rate, true negative rate, and accuracy rate of the model. Repeat the experiment 10 times and take the average of the obtained data as the final result of the model.
[0061] According to the above experimental configuration, the obtained splicing audio detection and localization results are shown in Table 1. It can be seen that the present invention can effectively detect the segments of spliced audio. When the threshold T is increased, the true positive rate of the present invention increases significantly, thereby reducing the number of samples missed by detection. In addition, the splicing audio detection and localization results of the present invention and the Jadhav method are shown in Table 2. It can be seen that the detection effect of the present invention is significantly better than that of the Jadhav method. Therefore, the present invention is very suitable for splicing audio detection and localization scenarios with high security requirements.
[0062] Table 1. Detection results of the present invention with different thresholds T
[0063]
[0064] Table 2. Splicing audio detection results using the Jadhav method and the present invention
[0065]
[0066] Based on the same inventive concept, another embodiment of the present invention provides a splicing audio detection and positioning system based on spectrogram segmentation using the above method, which includes:
[0067] A detection segment division module, configured to divide the audio to be tested into several audio segments S to be detected according to the minimum positioning region length L for audio splicing tampering positioning slr , g ;
[0068] A preprocessing module, configured to extract the spectrogram feature F of the audio segment S g , and splice t audio segments S g into an audio segment S' according to the size of the network input g , and splice the corresponding spectrogram features F g into a spectrogram feature F' to be input into the network g ; g ;
[0069] A binary prediction mask calculation module, configured to input the spliced spectrogram feature F' g into the trained splicing audio detection and positioning network, and calculate the binary prediction mask corresponding to the spliced audio segment S' g ;
[0070] An element ratio calculation module, configured to calculate the ratio ρ of the number of sample points with splicing sample points in the binary prediction mask of each audio segment S g to the total number of sample points;
[0071] A splicing segment judgment module, configured to compare the size of ρ and the preset judgment threshold T value, and further judge whether the audio segment S g is a spliced segment; for the N divided audio segments S' g , sequentially judge whether all the audio segments S g of the audio to be tested are spliced segments.
[0072] Based on the same inventive concept, another embodiment of the present invention provides an electronic device (such as a computer, a server, a smart phone, etc.), which includes a memory and a processor, the memory stores a computer program, the computer program is configured to be executed by the processor, and the computer program includes instructions for executing each step in the method of the present invention.
[0073] Based on the same inventive concept, another embodiment of the present invention provides a computer-readable storage medium (such as ROM / RAM, disk, optical disc), and the computer-readable storage medium stores a computer program, and when the computer program is executed by a computer, each step of the method of the present invention is implemented.
[0074] Other embodiments of the present invention: For the preprocessing in step 2), the spectrogram features involved can be replaced by audio waveforms, statistical features of spectrograms, or any acoustic features of audio.
[0075] For the binary prediction mask in step 3), the ASLNet network involved can be replaced by any network with an encoder-decoder structure, such as U-net and SegNet, etc.
[0076] For calculating the element ratio ρ in step 4), the ratio ρ does not necessarily have to be the ratio of the number of sample points and can be replaced by a weighted ratio or any quantity representing a proportion.
[0077] For the determination of the spliced segment in step 5), the final result does not necessarily have to be obtained by comparing with a threshold and can be replaced by any decision-making method, such as training a binary classifier.
[0078] The above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Those of ordinary skill in the art can modify or equivalently replace the technical solutions of the present invention without departing from the spirit and scope of the present invention. The protection scope of the present invention shall be subject to what is described in the claims.
Claims
1. A method for detecting and locating spliced audio based on spectrogram segmentation, characterized in that, it includes the following steps: The minimum positioning area length L for audio splicing tampering positioning slr , divide the audio to be measured into several audio segments S to be detected g ; Extract the audio segment S g of the spectrogram feature F g , and splice t audio segments S g into an audio segment S' according to the size of the network input g , and splice the corresponding spectrogram features F g into a spectrogram feature F' to be input into the network g ; The spliced spectrogram feature F′ g is input into the trained spliced audio detection and localization network to calculate the corresponding binary prediction mask for the spliced audio segment S′ g ; Calculate the number ratio ρ of the number of stitching samples to the total number of samples in the binary prediction mask of each audio segment S g ; Compare the magnitudes of ρ and a pre-set determination threshold value T, and then determine whether the audio segment S g is a spliced segment; For the N divided audio segments S' g , successively determine whether all audio segments S g of the audio to be measured are spliced segments.
2. The method according to claim 1, characterized in that, The extracted audio segment S g of the spectrogram feature F g is to extract Mel spectrogram coefficient features as the spectrogram features from the audio segment S g .
3. The method according to claim 2, characterized in that, The extraction process of the spectrogram features includes: First, use a pre-weighting module to enhance the energy of the signal at high frequencies, and calculate the short-time Fourier transform of the pre-weighted signal using a periodic Hamming window with a length of 2048 samples and an overlap of 512 samples; Then, use a Mel filter bank to map the energy to the Mel frequency scale and take the logarithm to produce a power spectrogram; Finally, use the discrete cosine transform to calculate the transformed coefficients containing important energy, that is, Mel-frequency cepstral coefficients.
4. The method according to claim 1, characterized in that, The basic network architecture of the spliced audio detection and localization network is a modified FCN-VGG16, which consists of a VGG16 encoder and a decoder with a residual structure. The goal of the VGG16 encoder is to capture the context representation of acoustic features, and the goal of the decoder is to convert the intermediate feature map into a binary prediction mask.
5. The method according to claim 4, characterized in that, The VGG16 encoder is composed of stacking VGG blocks, where each VGG block consists of two to three convolutional blocks followed by a max-pooling layer; The convolutional block consists of a convolutional layer, a batch normalization layer, and a linear unit activation function; The decoder consists of two transposed convolutional layers and a SoftMax activation function. Using the skip connection from the fourth VGG block to the first transposed convolution, the features learned at the lower layer are aggregated to the higher layer, and finally the SoftMax activation function is used to calculate the probability that the element comes from the spliced audio segment.
6. The method according to claim 1, characterized in that, In the binary prediction mask, 1 represents the spliced sample points, and 0 represents the original sample points; The calculation formula of the ratio ρ is as follows: where Num represents the number of elements in the set.
7. The method according to claim 1, characterized in that, Compare the magnitudes of ρ and a preset determination threshold value T, and then determine whether the audio segment S g is a spliced segment, including: when ρ > T, the audio segment S g is a spliced segment; otherwise, the audio segment S g is an original segment.
8. A system for detecting and locating spliced audio based on spectrogram segmentation using the method according to any one of claims 1 to 7, characterized in that, it includes: A detection segment division module, which is used to divide the audio to be tested into a plurality of audio segments S to be detected according to the minimum positioning region length L for audio splicing tampering positioning slr , g ; A preprocessing module for extracting the spectrogram feature F g of the audio segment S g , and concatenating t audio segments S g into an audio segment S' according to the size of the network input g , and concatenating the corresponding spectrogram features F g into a spectrogram feature F' to be input into the network g ; The binary prediction mask calculation module is used to input the spliced spectrogram feature F′ g into the trained spliced audio detection and localization network to calculate the spliced audio segment S′ g and the corresponding binary prediction mask; A calculation element ratio module for calculating each audio segment S g The ratio ρ of the number of splicing sample points to the total number of sample points among the sample points in the binary prediction mask The splicing segment judgment module is used to compare the magnitudes of ρ and a preset judgment threshold T value, and further judge whether the audio segment S g is a spliced segment; for the N divided audio segments S′ g , it is sequentially judged whether all the audio segments S of the audio to be measured g are spliced segments.
9. An electronic device, characterized in that, it includes a memory and a processor, the memory stores a computer program, the computer program is configured to be executed by the processor, and the computer program includes instructions for executing the method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, the computer-readable storage medium stores a computer program, and when the computer program is executed by the computer, it implements the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
RNN-based voice detection method for multiple counterfeit operations
CN111564163A
Audio anti-splicing detection method and system based on GRU
CN110942776A
Video content integrity identification method and device, equipment and storage medium
CN112418011A