Speech speed attack detection method based on prosodic features and random forest classifier

Through the speech speed attack detection method based on pronunciation characteristics and random forest classifiers, the problem of difficulty in detecting speech speed attacks in the existing technology is solved, and effective protection of speech recognition systems is achieved, with the advantages of low cost and high accuracy.

CN114550751BActive Publication Date: 2025-05-02ZHEJIANG UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210127689.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-02-11
Publication Date
2025-05-02
Estimated Expiration
2042-02-11

AI Technical Summary

Technical Problem

Existing voice recognition systems are susceptible to voice speed attacks, especially difficult to detect voice speed attacks without adding noise.

Method used

The speech speed attack detection method based on pronunciation characteristics and random forest classifier is adopted, and the attack detection is detected using the random forest classifier by extracting features such as jitter, tremolo and harmonic noise ratio.

Benefits of technology

It can accurately and effectively detect voice speed attacks against voice recognition systems, which have the characteristics of low cost and high attack detection accuracy, and is suitable for the security protection of voice recognition systems on smart devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114550751B_ABST
    Figure CN114550751B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for detecting speech speed-up attacks based on rhythmic features and random forest classifiers, and belongs to the fields of speech recognition technology and security technology. An audio data set is obtained, including normal audio and speed-up adversarial audio; jitter features, vibrato features, and harmonic noise ratio features of all audio in the audio data set are extracted to form a feature vector; a random forest classifier is trained using the feature vectors of normal audio and speed-up adversarial audio, and speech speed-up attack detection is performed using the trained random forest classifier. The present invention can efficiently detect speech speed-up deception attacks through the existing microphone and speech hardware of the speech recognition system, has the characteristics of low cost and high attack detection accuracy, can be used for security protection of speech recognition systems on smart devices such as mobile phones, and has broad demand and application prospects.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of speech recognition technology and security technology, and in particular relates to a speech speed attack detection method based on prosodic features and a random forest classifier. Background Art

[0002] Automatic Speech Recognition (ASR) systems can recognize speech and output speech recognition text. Existing popular ASR systems include open source systems (such as Kaldi and DeepSpeech) and commercial systems (such as Google Cloud Speech-to-Text, Baidu ASR, and iFlytek). For the input audio, the ASR system first performs signal processing to reduce noise and remove irrelevant frequency components; then divides the processed audio signal into short segments and extracts features such as Mel-frequency cepstral coefficients (MFCC); finally, using the extracted features, the most likely word sequence is inferred through a pre-trained speech recognition model.

[0003] Time-scale Modification (TSM) refers to the operation of speeding up or slowing down the playback speed of an audio clip. Common audio players or audio editing software use time-scale modification to change the playback speed of audio without changing the audio pitch. Time-scale modification mainly includes three steps: 1) signal decomposition, 2) frame repositioning and change, and 3) signal reconstruction. It first decomposes the input audio into shorter and overlapping analysis frames. The length of the analysis frame is usually in milliseconds, and it retains the pitch content of the frame in the original audio. According to the processing algorithm of frame transformation, there are many implementation schemes for time-scale modification, such as FFmpeg, SoundTouch, Waveform Similarity Overlap-Add (WSOLA) and Phase Vocoder (PV-TSM). Among the existing implementation schemes, FFmpeg is the most commonly used time-scale modification in commercial audio players and audio editing software.

[0004] However, studies have shown that existing speech recognition systems are vulnerable to speech speed-up attacks. Speech speed-up attacks refer to the process of causing ASR recognition errors by simply speeding up or slowing down the original audio. Untargeted attacks can be achieved by simply modifying the speed of a segment (for example, 20ms), while targeted attacks only require a few rounds of optimization to achieve malicious targets with good audibility (such as opening a door). With the development of voice technology and electronic devices, the threshold for speech speed-up attacks is getting lower, the effect is getting better, and the harm is getting greater. Therefore, in this case, it is urgent to propose an efficient and low-cost speech speed-up attack detection method.

[0005] At present, there have been many related studies that protect against voice attacks by detecting the noise and distortion introduced in the process of generating voice adversarial samples. However, this type of detection method is difficult to detect voice speed attacks without adding noise. Summary of the invention

[0006] In order to solve the technical problems existing in the above-mentioned background technology, the present invention provides a method for detecting speech speed-up attacks based on rhythmic features and random forest classifiers. Speed-up adversarial audio is obtained through speech speed-up attacks based on a particle swarm algorithm, and three types of features that can effectively and truly reflect the rhythmic differences between normal audio and speed-up adversarial audio are extracted. Attack detection is performed based on a random forest classifier, which can accurately and effectively detect speech deception attacks represented by speed-up attacks against speech recognition systems.

[0007] In order to achieve the above object, the technical solution adopted by the present invention is:

[0008] A method for detecting speech speed-up attacks based on prosodic features and a random forest classifier, comprising:

[0009] Obtain an audio dataset, including normal audio and double-speed adversarial audio;

[0010] Extract the jitter features, vibrato features and harmonic noise ratio features of all audio in the audio data set to form a feature vector;

[0011] The feature vectors of normal audio and speed-up adversarial audio are used to train a random forest classifier, and the trained random forest classifier is used to detect speech speed-up attacks.

[0012] Furthermore, the double-speed countermeasure audio is obtained by double-speeding normal audio without adding additional noise.

[0013] Furthermore, the jitter characteristics include jitt characteristics, jitta characteristics, rap characteristics, and ppq5 characteristics, and the calculation formula is:

[0014]

[0015]

[0016]

[0017]

[0018] Among them, T i represents the duration of the i-th jitter in the audio, N represents the total number of jitters in the audio, jitta, jitt, jitt rap , jitt ppq5They are jitt feature, jitta feature, rap feature, and ppq5 feature.

[0019] Furthermore, the vibrato feature includes Shim feature, ShdB feature, apq5 feature, and apq11 feature, and the calculation formula is:

[0020]

[0021]

[0022]

[0023]

[0024] Among them, A i represents the duration of the i-th vibrato in the audio, M represents the total number of vibratos in the audio, shim, ShdB, apq5, and apq11 represent the Shim feature, ShdB feature, apq5 feature, and apq11 feature, respectively.

[0025] Furthermore, the harmonic-to-noise ratio characteristic calculation formula is:

[0026]

[0027] Among them, sig per is the ratio of the audio period signal, sig noise It is the ratio of noise in the audio signal, hnr represents the harmonic noise ratio feature, and by designing different analysis window lengths, two harmonic noise ratio features hnr05 and hnr15 are obtained. The analysis window length of hnr15 is 3 times that of hnr05.

[0028] Furthermore, when using the trained random forest classifier to detect speech speed attack, the jitter features, vibrato features and harmonic noise ratio features of the audio to be detected are extracted to form a feature vector, which is used as the input of the trained random forest classifier to obtain the detection result.

[0029] The beneficial effects of the present invention are:

[0030] For the double-speed attack audio without adding noise that is difficult to identify by traditional methods, the present invention proposes to use 10-dimensional features such as jitter, vibrato, and harmonic noise ratio to provide effective feature data for attack detection based on the differences in pronunciation rhythm features between double-speed attack audio and normal audio, and combines the random forest algorithm for replay attack detection. During a voice deception attack, even if the attacker produces a sound that is very similar to the real user's voice, the sound will produce differences in jitter, vibrato, and harmonic noise ratio when it is accelerated, so this method can be used to detect voice double-speed deception attacks.

[0031] The present invention can efficiently detect voice speed spoofing attacks through the existing microphone and voice hardware of the voice recognition system, has the characteristics of low cost and high attack detection accuracy, can be used for security protection of voice recognition systems on smart devices such as mobile phones, and has broad demand and application prospects. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] Figure 1 It is a flow chart of a method for detecting speech speed-up attacks based on prosodic features and a random forest classifier, shown in an embodiment of the present invention. DETAILED DESCRIPTION

[0033] The following is a detailed description of the technical solutions provided by various embodiments of the present invention in conjunction with the accompanying drawings. The flowcharts shown in the accompanying drawings are only exemplary and do not necessarily include all the steps. For example, some steps can be decomposed, while some steps can be combined or partially combined, so the actual execution order may change according to the actual situation.

[0034] A speech speed-up attack is a means of speech adversarial attack that simply speeds up or slows down the original audio instead of adding interference. In order to achieve the recognition of both normal audio and speed-up adversarial audio without adding noise, the present invention utilizes the "unnatural" distortion caused by the speed-up operation, and proposes a speech speed-up attack detection method based on rhythmic features and random forest classifiers. The specific features of the speech signal received by the microphone of the electronic device are extracted and marked, and the random forest classifier is trained using the marked features. The trained classifier is used to perform speech speed-up attack detection on the speech signal to be tested, and the output is whether it is a speed-up speech adversarial result. It mainly includes the following steps:

[0035] 1) Feature extraction:

[0036] The overall process is as follows Figure 1 As shown, the input audio first goes through the feature extraction stage. The change in speech speed will cause the pronunciation difference between the double-speed adversarial audio and the normal audio. The key to using rhythmic features for attack detection is to extract features with large differences from the normal audio and the double-speed adversarial audio. Rhythmic features are one of the important forms of human language and emotional expression, including pitch, intonation, voice jitter, etc. The present invention adopts 4 jitter features, namely jitta, jitt, rap, ppq5; 4 vibrato features, namely Shim, ShdB, apq5, apq11; and 2 harmonic noise ratio features, namely hnr05 and hnr15 as a rhythmic feature combination to distinguish normal audio from double-speed adversarial audio, specifically:

[0037] 1.1) Jitter

[0038] Jitter represents the frequency change between signal cycles, which is mainly caused by the lack of control over the vibration of the vocal cords. The high or low jitter value represents the degree of hoarseness of the specific subject's voice, and the jitter of the patient is generally stronger. The present invention screened four jitter characteristics, including jitt, jitta, rap, and ppq5. i represents the duration of the i-th jitter in the audio, and N represents the total number of jitters in the audio.

[0039] The jitta feature is expressed as the mean absolute difference between consecutive jitter cycles and is calculated as:

[0040]

[0041] The JITT feature is expressed as the mean absolute difference between consecutive cycles divided by the average jitter period, calculated as:

[0042]

[0043] The rap feature is the relative average jitter, expressed as the average absolute difference between a jitter and the average of its two adjacent jitters divided by the average jitter period, denoted as jitt rap , the calculation formula is:

[0044]

[0045] The PPQ5 feature is expressed as the average absolute difference between a jitter and the average of its four adjacent jitters divided by the average jitter period, denoted as JITT. ppq5 , the calculation formula is:

[0046]

[0047] 1.2) Vibrato

[0048] Vibrato represents the amplitude fluctuation of the sound wave signal, which is related to human breathing and the noise it makes. It is generally regarded as a pathological rhythmic feature. The present invention screened four vibrato features, including Shim, ShdB, apq5, and apq11, and used A i represents the duration of the i-th vibrato in the audio, and M represents the total number of vibratos in the audio.

[0049] The Shim feature represents the mean absolute difference between the amplitudes of the vibrato divided by the average amplitude of the vibrato, and is calculated as:

[0050]

[0051] The ShdB feature is expressed as the average absolute logarithm of the vibrato amplitude and is calculated as:

[0052]

[0053] APQ5 is similar to the PPQ5 feature, which represents the average absolute difference between the amplitude of a vibrato and the average of the amplitudes of its four adjacent vibratos, divided by the average amplitude of the vibrato. The calculation formula is:

[0054]

[0055] apq11 is the average absolute difference between the amplitude of a vibrato and the average of its ten adjacent vibrato amplitudes divided by the average amplitude of the vibrato. The calculation formula is:

[0056]

[0057] 1.3) Harmonic noise ratio

[0058] The harmonic noise ratio characterizes the ratio of periodic components to non-periodic components in an audio signal, and reflects the texture of human speech, such as a soft or hard tone. The present invention selects two harmonic noise ratio features, including hnr05 and hnr15.

[0059] HNR, or harmonic noise ratio, indicates the degree of periodicity of sound waves in dB. The difference between HNR05 and HNR15 is the length of the analysis window. The calculation formula is:

[0060]

[0061] Among them, sig per is the ratio of the periodic signal, sig noise The analysis window length indicates the length of data used to calculate the harmonic noise ratio. The analysis window length for HNR15 is 3 times that for HNR05.

[0062] 2) Attack Detection:

[0063] For the problem of detecting speech speed-up attacks, prosodic features can better characterize the pronunciation distortion caused by changes in speech speed, thereby detecting speed-up adversarial audio.

[0064] The present invention uses the interface of Google TTS to convert text into speech, generates one hundred normal audios, and uses the particle swarm optimization algorithm to generate one hundred double-speed adversarial audios based on normal audios. As described in the above part, the present invention extracts a total of 10 rhythmic features, including 4 jitter features, namely jitt, jitta, rap, ppq5; 4 vibrato features, namely Shim, ShdB, apq5, apq11; 2 harmonic noise ratio features, namely hnr05, hnr15. Then, the 10-dimensional features of normal audio and double-speed adversarial audio are learned using a random forest classifier, and the probability that each decision tree outputs the detection result as normal audio or double-speed adversarial audio is determined by the random forest based on the voting selection of the decision tree.

[0065] In this embodiment, the detection results of random forest are compared with those of decision tree and SVM classifier, as shown in Table 1:

[0066] Table 1 Detection results of different classifiers for double-speed attacks

[0067] Accuracy Recall Accuracy F1 score Equal error rate AUC Random Forest 0.921 0.897 0.903 0.916 0.192 0.838 Decision Tree 0.865 0.866 0.864 0.872 0.14 0.866 SVM 0.783 0.741 0.864 0.842 0.241 0.849

[0068] It can be seen that random forest, decision tree and SVM can all achieve recognition of normal audio and double-speed adversarial audio, among which the accuracy of random forest reaches 92.1%.

[0069] The present invention obtains speed-up adversarial audio through a speech speed-up attack based on a particle swarm algorithm, and extracts three types of features that can effectively and truly reflect the rhythmic differences between normal audio and speed-up adversarial audio. According to the regular differences in pronunciation rhythm between normal audio and speed-up adversarial audio, the features are input into a random forest classifier, and then normal audio and speed-up adversarial audio are detected. During a speech deception attack, even if the attacker produces a sound that is very similar to the real user's voice, the sound will inevitably cause a certain degree of pronunciation difference when it is accelerated. Although the difference is extremely small, it is difficult to distinguish between the two using traditional feature extraction methods. However, the speech speed-up attack detection method based on rhythmic features and a random forest classifier proposed in the present invention can effectively detect this speed-up adversarial audio without adding noise, and provides guidance for the protection of speech recognition systems.

[0070] It should be understood by those skilled in the art that the embodiments of the present invention described above and shown in the accompanying drawings are only examples and do not limit the present invention. The purpose of the present invention has been fully and effectively achieved. The functional and structural principles of the present invention have been demonstrated and explained in the embodiments, and the embodiments of the present invention may be deformed or modified in any way without departing from the principles.

Claims

1. A method for detecting speech speed attack based on prosodic features and random forest classifier, characterized in that: include: Obtain an audio dataset, including normal audio and double-speed adversarial audio; The double-speed countermeasure audio is obtained by operating the normal audio at double speed without adding additional noise; Extract the jitter features, vibrato features and harmonic noise ratio features of all audio in the audio data set to form a feature vector; the jitter features include jitt features, jitta features, rap features, and ppq5 features, and the calculation formula is: Among them, T i represents the duration of the i-th jitter in the audio, N represents the total number of jitters in the audio, jitta, jitt, jitt rap , jitt ppq5 They are jitt feature, jitta feature, rap feature, and ppq5 feature; The vibrato features include Shim feature, ShdB feature, apq5 feature, and apq11 feature, and the calculation formula is: Among them, A i represents the duration of the i-th vibrato in the audio, M represents the total number of vibratos in the audio, shim, ShdB, apq5, apq11 represent the Shim feature, ShdB feature, apq5 feature, apq11 feature respectively; The harmonic-to-noise ratio characteristic calculation formula is: Among them, sig per is the ratio of the audio period signal, sig noise is the ratio of noise in the audio signal, hnr represents the harmonic noise ratio feature, and two harmonic noise ratio features hnr05 and hnr15 are obtained by designing different analysis window lengths. The analysis window length of hnr15 is 3 times that of hnr05. The feature vectors of normal audio and speed-up adversarial audio are used to train a random forest classifier, and the trained random forest classifier is used to detect speech speed-up attacks.

2. The method for detecting speech speed attack based on prosodic features and random forest classifier according to claim 1, characterized in that: When using the trained random forest classifier to detect speech speed attack, the jitter features, vibrato features and harmonic noise ratio features of the audio to be detected are extracted to form a feature vector, which is used as the input of the trained random forest classifier to obtain the detection result.

Citation Information

Patent Citations

  • Camouflage voice detection method adopting joint features and random forest

    CN113436646A