Speech anomaly detection method, device, electronic device and storage medium

By comparing the long-term power spectrum characteristics of the speech to be detected with the long-term power spectrum characteristics of the sample of the positive sample speech, and using the likelihood score of the speech abnormality detection model, the problem of low speech abnormality detection accuracy in the prior art is solved, and the accurate detection of complex speech spectrum is achieved.

CN114283790BActive Publication Date: 2025-06-03XIAN XUNFEI SUPER BRAIN INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111437540.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-29
Publication Date
2025-06-03
Estimated Expiration
2041-11-29

AI Technical Summary

Technical Problem

In the prior art, the speech abnormality detection method cannot effectively detect complex and irregular speech spectrum changes caused by abnormal noise reduction system, resulting in low detection accuracy.

Method used

By determining the long-term power spectrum characteristics of the speech to be detected, and comparing and analyzing it based on the sample long-term power spectrum characteristics of the positive sample speech, the speech abnormality detection result is determined using the likelihood score output by the speech abnormality detection model.

Benefits of technology

Accurate abnormality detection of complex and irregular speech spectrum is achieved, and the accuracy of speech abnormality detection is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114283790B_ABST
    Figure CN114283790B_ABST
Patent Text Reader

Abstract

The present invention provides a method, apparatus, electronic device, and storage medium for voice anomaly detection. The method includes: determining the long-term power spectrum feature of the voice to be detected; and performing anomaly detection on the voice to be detected based on the sample long-term power spectrum feature of the positive sample voice and the long-term power spectrum feature of the voice to be detected, so as to obtain the voice anomaly detection result of the voice to be detected. The method, apparatus, electronic device, and storage medium for voice anomaly detection provided by the present invention can perform anomaly detection on any complex voice to be detected and accurately obtain the voice anomaly detection result of the voice to be detected.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of speech recognition, and in particular, to a method, device, electronic device and storage medium for detecting abnormal speech. Background Art

[0002] Speech data acquisition is the basis for many algorithm researches and application systems. In order to adapt to various complex acquisition environments, a noise reduction system is usually built into the acquisition device for noise reduction processing to ensure the quality of speech data. If the operation is improper during the acquisition process, such as the acquisition device is not aligned with the target speaker, the noise reduction system in the acquisition device is likely to incorrectly suppress the target signal, resulting in significant speech distortion, such as the speech being suddenly loud or soft, frequency band missing and other phenomena.

[0003] However, the speech distortion caused by the noise reduction system is not only related to the use environment and method, but also related to the noise reduction algorithm adopted by the noise reduction system. Moreover, the speech spectrum corresponding to the distorted speech collected due to the abnormality of the noise reduction system is constantly changing. The existing speech anomaly detection methods usually detect anomalies for regular speech spectra and cannot detect such complex and changing speech spectra. Summary of the Invention

[0004] The present invention provides a method, device, electronic device and storage medium for detecting abnormal speech, so as to solve the defect of low accuracy of speech anomaly detection in the prior art.

[0005] The present invention provides a method for detecting abnormal speech, including:

[0006] Determine the long-term power spectrum feature of the speech to be detected;

[0007] Based on the sample long-term power spectrum feature of the positive sample speech and the long-term power spectrum feature of the speech to be detected, perform anomaly detection on the speech to be detected to obtain the speech anomaly detection result of the speech to be detected.

[0008] According to a method for detecting abnormal speech provided by the present invention, the performing anomaly detection on the speech to be detected based on the sample long-term power spectrum feature of the positive sample speech and the long-term power spectrum feature of the speech to be detected to obtain the speech anomaly detection result of the speech to be detected includes:

[0009] Input the long-term power spectrum feature of the speech to be detected into a speech anomaly detection model to obtain a likelihood score output by the speech anomaly detection model;

[0010] Based on the likelihood score, determine the speech anomaly detection result of the speech to be detected;

[0011] Among them, the voice anomaly detection model is trained based on the sample long-term power spectrum features of the positive sample voice.

[0012] According to a voice anomaly detection method provided by the present invention, determining the voice anomaly detection result of the voice to be detected based on the likelihood score includes:

[0013] Determining the maximum detection score and the minimum detection score based on the mean supervector of the voice anomaly detection model;

[0014] Determining the long-term power spectrum detection score based on the likelihood score, the maximum detection score, and the minimum detection score;

[0015] Determining the voice anomaly detection result of the voice to be detected based on the long-term power spectrum detection score.

[0016] According to a voice anomaly detection method provided by the present invention, after obtaining the voice anomaly detection result of the voice to be detected, it further includes:

[0017] Updating the voice anomaly detection result of the voice to be detected based on the voice effective duration of the voice to be detected and / or the signal-to-noise ratio of the voice to be detected.

[0018] According to a voice anomaly detection method provided by the present invention, updating the voice anomaly detection result of the voice to be detected based on the voice effective duration of the voice to be detected and / or the signal-to-noise ratio of the voice to be detected includes:

[0019] When the voice anomaly detection result of the voice to be detected is that the voice to be detected is normal, if it is determined that the voice to be detected has an anomaly based on the voice effective duration of the voice to be detected and / or the signal-to-noise ratio of the voice to be detected, then update the voice anomaly detection result of the voice to be detected to the voice to be detected is abnormal.

[0020] According to a voice anomaly detection method provided by the present invention, determining the long-term power spectrum features of the voice to be detected includes:

[0021] Extracting voicing signal data from the voice to be detected;

[0022] Determining the variance of the long-term power spectrum features of each segment of voicing frame data based on the voicing frame data of each segment in the voicing signal data;

[0023] Determining the long-term power spectrum features of the voice to be detected based on the variance of the long-term power spectrum features of each segment of voicing frame data.

[0024] A method for detecting abnormal speech provided by the present invention, determining the long-term power spectrum feature of the speech to be detected according to the variance of the long-term power spectrum features of each segment of voiced frame data, includes:

[0025] Based on the variance of the long-term power spectrum features of each segment of voiced frame data and the number of voiced frames in each segment of voiced frame data, determine the long-term power spectrum feature of the speech to be detected.

[0026] The present invention also provides a device for detecting abnormal speech, including:

[0027] A determination unit for determining the long-term power spectrum feature of the speech to be detected;

[0028] A detection unit for performing abnormal detection on the speech to be detected based on the sample long-term power spectrum feature of the positive sample speech and the long-term power spectrum feature of the speech to be detected, and obtaining the speech abnormal detection result of the speech to be detected.

[0029] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of the speech abnormal detection method as described in any one of the above are implemented.

[0030] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the speech abnormal detection method as described in any one of the above are implemented.

[0031] The present invention also provides a computer program product, including a computer program. When the computer program is executed by a processor, the steps of the speech abnormal detection method as described in any one of the above are implemented.

[0032] For the speech abnormal detection method, device, electronic device, and storage medium provided by the present invention, since the sample long-term power spectrum feature of the positive sample speech has a low-pass characteristic, based on the sample long-term power spectrum feature of the positive sample speech and the long-term power spectrum feature of the speech to be detected, abnormal detection can be performed on any complex speech to be detected, and the speech abnormal detection result of the speech to be detected can be accurately obtained. Description of the Drawings

[0033] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0034] Figure 1It is a schematic flow chart of the voice anomaly detection method provided by the present invention;

[0035] Figure 2 It is a schematic flow chart of the implementation manner of step 120 in the voice anomaly detection method provided by the present invention;

[0036] Figure 3 It is a schematic flow chart of the implementation manner of step 122 in the voice anomaly detection method provided by the present invention;

[0037] Figure 4 It is a schematic flow chart of the implementation manner of step 110 in the voice anomaly detection method provided by the present invention;

[0038] Figure 5 It is a schematic flow chart of another voice anomaly detection method provided by the present invention;

[0039] Figure 6 It is a schematic structural diagram of the voice anomaly detection device provided by the present invention;

[0040] Figure 7 It is a schematic structural diagram of the electronic device provided by the present invention. Specific embodiments

[0041] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below with reference to the accompanying drawings in the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art without making creative efforts based on the embodiments in the present invention fall within the protection scope of the present invention.

[0042] During the process of the acquisition device acquiring voice data, the voice is usually denoised through a built-in noise reduction system to ensure the quality of the voice data. If the operation is improper during the acquisition process, such as the acquisition device not being aimed at the target speaker, the noise reduction system in the acquisition device is likely to wrongly suppress the target signal, resulting in significant voice distortion, such as excessive residual noise, sudden change in signal-to-noise ratio, abnormal sound quality, voice being too loud or too soft, frequency band missing, etc. Since there is a lack of original voice data for comparison, these voice anomaly problems caused by the noise reduction system need to be detected by subjective auditory judgment of the human ear.

[0043] In addition, most of the voice noise reduction algorithms adopted in the noise reduction system are frequency domain algorithms. The distortion caused by this algorithm may occur in different frequency bands, and the degree of distortion may also vary greatly. This change in the spectrum will be mixed with the spectrum change of the voice data itself. The spectrum of the voice data itself changes continuously over time, has a large dynamic range, and has an obvious formant structure. Its spectrum is also affected by the surrounding acoustic environment and the frequency response of the recording system itself. Therefore, the spectra of the normally collected voice data and the distorted voice data caused by the abnormality of the noise reduction system are both changing continuously.

[0044] However, the voice abnormality detection methods in the prior art usually perform abnormality detection on regular voice spectra or voice spectra with specific prior knowledge. For example, for abnormalities such as packet loss and jitter in voice communication, the rule is that no data packet is received within a specified time, resulting in continuous zero signals or inserted noise signals; for abnormal operations of multiple microphone systems, the rule is that significant differences will occur between the original microphone signals, such as a significant abnormality in the sound pressure level difference between the main and auxiliary microphones, or a significant abnormality in the sound pressure level of the signal caused by the blockage of the sound inlet hole of a certain microphone, etc.; for abnormal phenomena such as acoustic echo, howling, clipping, and hum, they have obvious prior physical characteristics and can be detected through specific phenomena occurring in the voice data spectrum; for abnormal phenomena such as too far or too close pick-up distance, or only background noise being picked up, the physical changes in the voice data spectrum generated have strong prior knowledge. However, the above methods cannot accurately perform voice abnormality detection on the complex, irregular, and non-specific prior knowledge voice spectra caused by the abnormality of the noise reduction system.

[0045] In view of this, the present invention provides a voice abnormality detection method. Figure 1 It is a schematic flow chart of the voice abnormality detection method provided by the present invention. As Figure 1 shown, the method includes the following steps:

[0046] Step 110: Determine the long-term power spectrum feature of the voice to be detected.

[0047] Step 120: Based on the sample long-term power spectrum feature of the positive sample voice and the long-term power spectrum feature of the voice to be detected, perform abnormality detection on the voice to be detected to obtain the voice abnormality detection result of the voice to be detected.

[0048] Here, the speech to be detected refers to the speech for which it is to be detected whether there is a speech anomaly, and the positive sample speech refers to the normal speech with qualified speech quality. For example, the positive sample speech can be the sample speech that has passed the manual quality inspection, or the sample speech collected by the noise reduction system of the speech collection device under normal working conditions. The sample long-term power spectrum feature of the positive sample speech is used to characterize the distribution of the signal power of the positive sample speech in the frequency domain, and the long-term power spectrum feature of the speech to be detected is used to characterize the distribution of the signal power of the speech to be detected in the frequency domain.

[0049] Since the normal speech is the normal speech with qualified speech quality, the sample long-term power spectrum feature of the positive sample speech has a low-pass characteristic, that is, the signal power of the positive sample speech gradually decreases as the frequency increases. Since the speech to be detected is usually collected by a speech collection device, during the process of the noise reduction system built in the speech collection device performing noise reduction processing on the speech to be detected, an anomaly in the noise reduction system may cause the speech to be detected to be distorted, and the abnormal state of the noise reduction system usually lasts for a period of time, resulting in obvious fluctuations in the long-term power spectrum feature of the speech to be detected, that is, at this time, the signal power of the speech to be detected does not gradually decrease as the frequency increases.

[0050] Based on this, the sample long-term power spectrum feature of the positive sample speech can be compared and analyzed with the long-term power spectrum feature of the speech to be detected to determine whether the long-term power spectrum feature of the speech to be detected has the low-pass characteristic of the sample long-term power spectrum feature of the positive sample speech. If so, it indicates that the probability of the speech to be detected having an anomaly is relatively low; if not, it indicates that the probability of the speech to be detected having an anomaly is relatively high.

[0051] Optionally, after determining the sample long-term power spectrum feature of the positive sample speech and the long-term power spectrum feature of the speech to be detected, the similarity between the sample long-term power spectrum feature of the positive sample speech and the long-term power spectrum feature of the speech to be detected can be calculated. The higher the similarity between the two, the smaller the difference between the sample long-term power spectrum feature of the positive sample speech and the long-term power spectrum feature of the speech to be detected, that is, the long-term power spectrum feature of the speech to be detected also has the low-pass characteristic, that is, the probability of the speech to be detected having an anomaly is lower; the lower the similarity between the two, the greater the difference between the sample long-term power spectrum feature of the positive sample speech and the long-term power spectrum feature of the speech to be detected, that is, the long-term power spectrum feature of the speech to be detected does not have the low-pass characteristic, that is, the probability of the speech to be detected having an anomaly is higher.

[0052] Optionally, after determining the long-term power spectrum features of the positive sample speech and the long-term power spectrum features of the speech to be detected, the range of signal power corresponding to each frequency can be determined based on the long-term power spectrum features of the positive sample speech. If the signal power corresponding to each frequency of the long-term power spectrum features of the speech to be detected is within the above range, it indicates that the long-term power spectrum features of the speech to be detected have a low-pass feature, that is, the probability that the speech to be detected is abnormal is relatively low; if the signal power corresponding to each frequency of the long-term power spectrum features of the speech to be detected is outside the above range, it indicates that the long-term power spectrum features of the speech to be detected do not have a low-pass feature, that is, the probability that the speech to be detected is abnormal is relatively high.

[0053] In addition, based on the above method, the abnormal point position of the speech to be detected can also be located. For example, based on the long-term power spectrum features of the positive sample speech, the signal power range corresponding to frequency 1 is determined to be (a1, a2), and the signal power range corresponding to frequency 2 is determined to be (b1, b2). The signal power corresponding to frequency 1 in the long-term power spectrum features of the speech to be detected is m, and the signal power corresponding to frequency 2 is n, and a1 < m < a2, n > b2. Therefore, the signal power m corresponding to frequency 1 is within the signal power range (a1, a2), that is, the speech corresponding to frequency 1 in the speech to be detected is normal; the signal power n corresponding to frequency 2 is outside the signal power range (b1, b2), that is, the speech corresponding to frequency 2 in the speech to be detected is abnormal. Thus, according to the method of the above embodiment, not only can the speech to be detected be abnormally detected, but also the abnormal position in the speech to be detected can be accurately located.

[0054] Thus, the embodiment of the present invention utilizes the low-pass characteristic of the long-term power spectrum features of the positive sample speech, so that the abnormal detection of the speech to be detected can be performed without any rules or specific prior knowledge, that is, the embodiment of the present invention can not only realize the abnormal detection of the speech to be detected with rules or specific prior knowledge, but also realize the abnormal detection of the complex speech to be detected without rules and specific prior knowledge.

[0055] The speech abnormal detection method provided by the embodiment of the present invention, due to the low-pass characteristic of the long-term power spectrum features of the positive sample speech, can realize the abnormal detection of any complex speech to be detected based on the long-term power spectrum features of the positive sample speech and the long-term power spectrum features of the speech to be detected, and accurately obtain the speech abnormal detection result of the speech to be detected.

[0056] Based on the above embodiment, Figure 2 is a schematic flowchart of the implementation manner of step 120 in the speech abnormal detection method provided by the present invention, as Figure 2 shown, step 120 includes:

[0057] Step 121: Input the long-term power spectrum features of the speech to be detected into the speech anomaly detection model to obtain the likelihood score output by the speech anomaly detection model;

[0058] Step 122: Determine the speech anomaly detection result of the speech to be detected based on the likelihood score;

[0059] Among them, the speech anomaly detection model is trained based on the sample long-term power spectrum features of positive sample speech.

[0060] Specifically, since the speech anomaly detection model is trained based on the sample long-term power spectrum features of positive sample speech, the speech anomaly detection model can learn the long-term power spectrum features corresponding to normal speech from the sample long-term power spectrum features of positive sample speech. After inputting the long-term power spectrum features of the speech to be detected into the trained speech anomaly detection model, the speech anomaly detection model can determine the similarity between the long-term power features of the speech to be detected and the long-term power spectrum features of normal speech, and then obtain the likelihood score used to characterize this similarity. Among them, the speech anomaly detection model can be a Gaussian Mixture Model (GMM), and the Gaussian components of this model can be 4 - 6, so as to fit the distribution of signal power in the frequency domain in each Gaussian component and accurately obtain the distribution of signal power in the frequency domain of positive sample speech.

[0061] After determining the likelihood score, the smaller the difference between the likelihood score and the score threshold, the higher the similarity between the long-term power spectrum features of the speech to be detected and the long-term power spectrum features of normal speech, that is, the lower the probability that the speech to be detected has an anomaly; the greater the difference between the likelihood score and the score threshold, the lower the similarity between the long-term power spectrum features of the speech to be detected and the long-term power spectrum features of normal speech, that is, the higher the probability that the speech to be detected has an anomaly.

[0062] It can be seen that through the speech anomaly detection model in the embodiments of the present invention, the likelihood score used to characterize the similarity between the long-term power spectrum features of the speech to be detected and the long-term power spectrum features of normal speech can be obtained, and then the speech anomaly detection result can be accurately obtained based on the likelihood score.

[0063] Based on any of the above embodiments, Figure 3 is a schematic flowchart of the implementation manner of step 122 in the speech anomaly detection method provided by the present invention. As Figure 3 shown, step 122 includes:

[0064] Step 1221: Determine the maximum detection score and the minimum detection score based on the mean supervector of the speech anomaly detection model;

[0065] Step 1222: Determine the long-term power spectrum detection score based on the likelihood score, the maximum detection score, and the minimum detection score;

[0066] Step 1223: Determine the speech anomaly detection result of the speech to be detected based on the long-term power spectrum detection score.

[0067] Specifically, the mean supervector is used to characterize the distribution of the signal power corresponding to the long-term power spectrum features of the positive sample speech in each frequency domain. The maximum detection score is used to characterize the signal power distribution at the frequency domain density maximum of the long-term power spectrum features of the positive sample speech. The minimum detection score is used to characterize the signal power distribution at the frequency domain density minimum of the long-term power spectrum features of the positive sample speech.

[0068] Since the likelihood scores obtained based on the speech anomaly detection model are usually less than 0, in order to intuitively determine the speech anomaly detection result, in the embodiments of the present invention, the likelihood scores are converted into values between 0 and 1. Specifically: Determine the maximum detection score and the minimum detection score based on the mean supervector, and then obtain the long-term power spectrum detection score with a numerical range of (0, 1) based on the likelihood score, the maximum detection score, and the minimum detection score, so that the speech anomaly detection result can be determined based on the long-term power spectrum detection score. For example, if the difference between the long-term power spectrum detection score and the score threshold is greater, it indicates that the probability of anomalies in the speech to be detected is higher.

[0069] Optionally, the long-term power spectrum detection score v 3 (l v ) can be determined based on the following formula:

[0070]

[0071]

[0072]

[0073] s min = a·s max

[0074] where s(l v ) represents the likelihood score, p(·|λ) represents the observation probability of the speech anomaly detection model (GMM model), represents the mean supervector of the GMM model, a > 1, which can be an empirical parameter used to control the scale stretching range of the detection score, represents the long-term power spectrum features of the speech to be detected, s max represents the maximum detection score, s min represents the minimum detection score.

[0075] It can be understood that after obtaining the long-term power spectrum detection score, different levels of warnings can be issued based on the difference between the long-term power spectrum score and the score threshold. For example, the greater the difference, the higher the warning level.

[0076] Based on any of the above embodiments, after obtaining the voice anomaly detection result of the voice to be detected, the following steps are further included:

[0077] Update the voice anomaly detection result of the voice to be detected based on the voice effective duration of the voice to be detected and / or the signal-to-noise ratio of the voice to be detected.

[0078] Specifically, during the process of the acquisition device collecting voice, regulations are usually made on the signal-to-noise ratio of the voice, the voice effective duration, the background noise level, etc. If the noise reduction system is abnormal, it usually causes some voice signals of the collected voice to be wrongly suppressed, resulting in a significant decrease in the short-term signal-to-noise ratio, and further affecting the voice effective duration. Therefore, after obtaining the voice anomaly detection result of the voice to be detected, the embodiments of the present invention can update the voice anomaly detection result of the voice to be detected based on the voice effective duration of the voice to be detected and / or the signal-to-noise ratio of the voice to be detected.

[0079] Optionally, the voice effective duration of the voice to be detected is used to represent the cumulative duration of the voice signals existing in the voice to be detected. Based on the voice effective duration of the voice to be detected, it can assist in detecting whether the voice to be detected is abnormal. If the detection result is the same as the voice anomaly detection result, there is no need to update the voice anomaly detection result; if they are different, it is necessary to further determine whether to update the voice anomaly detection result. For example, if the voice effective duration of the voice to be detected exceeds the specified effective duration, it indicates that the probability of some voice signals of the voice to be detected being wrongly suppressed is relatively high, that is, the probability of the voice to be detected being abnormal is relatively high. At this time, if the voice anomaly detection result of the voice to be detected is normal, the voice anomaly detection result can be updated to abnormal to ensure the accuracy of the voice anomaly detection result.

[0080] Optionally, the signal-to-noise ratio of the voice to be detected is used to represent the ratio of the voice signal to the noise signal in the voice to be detected. Based on the signal-to-noise ratio of the voice to be detected, it can assist in detecting whether the voice to be detected is abnormal. If the detection result is the same as the voice anomaly detection result, there is no need to update the voice anomaly detection result; if they are different, it is necessary to further determine whether to update the voice anomaly detection result. For example, if the signal-to-noise ratio of the voice to be detected exceeds the specified signal-to-noise ratio, it indicates that the probability of some voice signals of the voice to be detected being wrongly suppressed is relatively high, that is, the probability of the voice to be detected being abnormal is relatively high. At this time, if the voice anomaly detection result of the voice to be detected is normal, the voice anomaly detection result can be updated to abnormal to ensure the accuracy of the voice anomaly detection result.

[0081] Optionally, based on the voice effective duration of the voice to be detected and the signal-to-noise ratio of the voice to be detected, the voice anomaly detection result can be updated. For example, if the voice anomaly detection result is normal, and it is determined based on the voice effective duration of the voice to be detected that there is no anomaly in the voice to be detected, but it is determined based on the signal-to-noise ratio of the voice to be detected that there is an anomaly in the voice to be detected, then the voice anomaly detection result can be updated to abnormal. If the voice anomaly detection result is normal, and it is determined based on both the voice effective duration and the signal-to-noise ratio of the voice to be detected that there is no anomaly in the voice to be detected, it indicates that the credibility of the voice anomaly detection result being normal is relatively high, that is, there is no need to update the voice anomaly detection result at this time.

[0082] It can be seen that based on the voice effective duration of the voice to be detected and / or the signal-to-noise ratio of the voice to be detected, the embodiments of the present invention can further determine whether there is an anomaly in the voice to be detected, so as to update the voice anomaly detection result of the voice to be detected and ensure the accuracy of the voice anomaly detection result.

[0083] Based on any of the above embodiments, the voice effective duration of the voice to be detected and the signal-to-noise ratio of the voice to be detected are determined based on the following process:

[0084] First, perform a Short-Time Fourier Transform (STFT) on the voice to be detected. Specifically, first frame and window the voice to be detected, and then perform a Fourier transform. It should be noted that since the voice distortion caused by the abnormal noise reduction system is manifested in the channel characteristics, that is, the spectral details of the voice to be detected do not affect the accuracy of the voice anomaly detection result, a broadband spectrogram analysis method can be used here to obtain the spectral envelope information of the voice to be detected.

[0085] Then, perform Voice Activity Detection (VAD) on the voice to be detected after the short-time Fourier transform. The detection basis is the ratio of the signal power to the noise power of the voice to be detected, that is, the signal-to-noise ratio. During the process of collecting the voice to be detected, due to the effect of the noise reduction system, the collected voice to be detected may have a relatively high signal-to-noise ratio. If other noise estimation methods are used, it is easy to have the problem of overestimating the noise, especially when the voice signal continuously appears in the voice to be detected, the overestimation will be more serious. The accurate estimation of the noise level directly affects the accuracy of the signal-to-noise ratio. Therefore, in order to accurately determine the signal-to-noise ratio of the voice to be detected, VAD detection can be performed based on the following method.

[0086] (a) Calculate the initial noise power level p ns (l) of each frame signal in the voice to be detected based on the following formula:

[0087]

[0088] Wherein, l represents the frame number, and k represents the frequency point number. represents the noise power spectrum obtained according to the traditional noise estimation method, and K represents the main frequency range of the speech to be detected, which is also the frequency range concerned by the long-term power spectrum of the speech to be detected. It is related to the specific acquisition specification and acquisition equipment. For example, for a speech signal sampled at 8 kHz, its typical value is from 200 Hz to 3500 Hz.

[0089] (b) Based on the initial noise power level p ns (l) and the power level p y (l) of the speech to be detected, the initial signal-to-noise ratio of each frame can be obtained = p y (l) - p ns (l). VAD detection is performed based on the initial signal-to-noise ratio of each frame. Among them, the power level of the speech to be detected represents the power spectrum of the speech to be detected obtained according to the traditional noise estimation method.

[0090] Here, when performing VAD detection based on the initial signal-to-noise ratio of each frame, the typical four-state method can be adopted. States 0 to 3 respectively represent "silence", "suspected speech start", "speech", and "suspected speech end state". Among them, the initial state of the speech to be detected is 0 (silence).

[0091] In state 1 (silence), if the initial signal-to-noise ratio of each frame suddenly becomes greater than the preset threshold, it can transfer from state 0 (silence) to state 1 (suspected speech start).

[0092] In state 1 (suspected speech start), if the initial signal-to-noise ratio of each frame continuously remains greater than the preset threshold and the duration is sufficient, it can transfer from state 1 to state 2 (speech); if the initial signal-to-noise ratio of each frame continuously remains less than the preset threshold, it can transfer from state 1 (suspected speech start) to state 0 (silence); if the initial signal-to-noise ratio of each frame continuously remains greater than the preset threshold but the duration is insufficient, it remains in state 1 (suspected speech start).

[0093] In state 2 (speech), if the initial signal-to-noise ratio of each frame is greater than the preset threshold, it remains in state 2 (speech); if the initial signal-to-noise ratio of each frame suddenly becomes less than the preset threshold, it transfers from state 2 (speech) to state 3 (suspected speech end).

[0094] In state 3 (suspected speech end), if the initial signal-to-noise ratio of each frame continuously remains lower than the preset threshold and the duration is insufficient, it remains in state 3 (suspected speech end); if the initial signal-to-noise ratio of each frame is greater than the preset threshold, it transfers from state 3 (suspected speech end) to state 2 (speech); if the initial signal-to-noise ratio of each frame continuously remains lower than the preset threshold and the duration is sufficient, it transfers from state 3 (suspected speech end) to state 0 (silence).

[0095] (c) Modify according to the above VAD detection results, specifically: First, for the speech signal segment in the 0 state (silence), directly use the power spectrum of the speech to be detected as the estimate of the modified noise power spectrum; for the speech signal segment in the 3 state (suspected end of speech), take the minimum value of the power spectrum of the speech to be detected and the initially estimated noise power spectrum as the estimate of the modified noise power spectrum; for the modified noise power spectra of the signal segments in the 1 state (suspected start of speech) and 2 state (speech), linearly interpolate using the modified noise power spectra corresponding to the nearby 0 state and 3 state. According to the modified noise power spectrum, calculate the modified noise power level using formula (1).

[0096] (d) According to the modified noise power level and the power level p of the speech to be detected y (l) Perform VAD detection again to correct possible errors in the previous VAD, and finally uniformly mark the 1 state and 2 state as the speech state.

[0097] After completing the VAD detection, add up the durations of all the speech signals marked as the speech state above to obtain the effective speech duration v of the speech to be detected 1 .

[0098] Among the speech distortions caused by the abnormal noise reduction system, one of the common phenomena is that the speech is sometimes loud and sometimes soft, and the signal-to-noise ratio of the corresponding speech signal segment changes greatly. And the signal-to-noise ratio of normal speech also changes greatly. Therefore, the embodiment of the present invention uses the long-term signal-to-noise ratio of the speech segment of the speech to be detected as a feature, that is, the signal-to-noise ratio (SIGNAL-NOISE RATIO, SNR) feature of the speech segment. Among them, a speech segment refers to a set of multiple speech frames marked as the speech state by VAD detection, which can be implemented by a queue composed of speech frames. The SNR feature of the speech segment is the average SNR value of several consecutive speech frames, denoted as v 2 (l s ), where l s is the number of the speech segment. If the noise reduction system completely suppresses the speech signal of the speech to be detected to the level of residual ambient noise, then this part of the speech signal can be directly regarded as residual noise, and the human ear will not obviously perceive it as distorted speech, and it can be considered that it does not affect the quality of the speech to be detected; if the power level of the distorted speech is significantly higher than the ambient residual noise, it will be marked as the speech state by VAD detection, and its sometimes loud and sometimes soft phenomenon will cause the average SNR of the speech segment to be lower than the normal level.

[0099] Based on any of the above embodiments, update the speech anomaly detection result of the speech to be detected based on the effective speech duration of the speech to be detected and / or the signal-to-noise ratio of the speech to be detected, including:

[0100] When the speech anomaly detection result of the speech to be detected is that the speech to be detected is normal, if it is determined that there is an anomaly in the speech to be detected based on the speech effective duration of the speech to be detected and / or the signal-to-noise ratio of the speech to be detected, then update the speech anomaly detection result of the speech to be detected to the speech to be detected is abnormal.

[0101] Specifically, when the speech anomaly detection result of the speech to be detected is that the speech to be detected is normal, and it is determined that there is no anomaly in the speech to be detected based on the speech effective duration of the speech to be detected, but it is determined that there is an anomaly in the speech to be detected based on the signal-to-noise ratio of the speech to be detected, it indicates that some speech signals are wrongly suppressed, resulting in a significant decrease in the short-term signal-to-noise ratio, that is, the probability of the speech to be detected being abnormal is relatively high at this time. Therefore, the speech anomaly detection result can be updated to the speech to be detected is abnormal. Among them, the speech effective duration of the speech to be detected can be represented by the speech effective duration v in the above embodiments. 1 to characterize, and the signal-to-noise ratio of the speech to be detected can be characterized by the segment SNR feature in the above embodiments.

[0102] When the speech anomaly detection result of the speech to be detected is that the speech to be detected is normal, and it is determined that there is no anomaly in the speech to be detected based on the signal-to-noise ratio of the speech to be detected, but it is determined that there is an anomaly in the speech to be detected based on the speech effective duration of the speech to be detected, and it is determined that there is an anomaly in the speech to be detected based on the signal-to-noise ratio of the speech to be detected, it indicates that some speech signals are wrongly suppressed, thereby affecting the speech effective duration, that is, the probability of the speech to be detected being abnormal is relatively high at this time. Therefore, the speech anomaly detection result can be updated to the speech to be detected is abnormal.

[0103] When the speech anomaly detection result of the speech to be detected is that the speech to be detected is normal, and it is determined that there is an anomaly in the speech to be detected based on both the signal-to-noise ratio of the speech to be detected and the speech effective duration of the speech to be detected, it indicates that some speech signals are wrongly suppressed, resulting in a significant decrease in the short-term signal-to-noise ratio and affecting the speech effective duration, that is, the probability of the speech to be detected being abnormal is relatively high at this time. Therefore, the speech anomaly detection result can be updated to the speech to be detected is abnormal.

[0104] If the speech anomaly detection result is normal, and it is determined that there is no anomaly in the speech to be detected based on both the speech effective duration and the signal-to-noise ratio of the speech to be detected, it indicates that the credibility of the speech anomaly detection result being normal is relatively high, that is, there is no need to update the speech anomaly detection result at this time.

[0105] It can be seen that when the voice anomaly detection result of the voice to be detected in the embodiment of the present invention is that the voice to be detected is normal, if it is determined that there is an anomaly in the voice to be detected based on the voice effective duration of the voice to be detected and / or the signal-to-noise ratio of the voice to be detected, the voice anomaly detection result of the voice to be detected is updated to the voice to be detected is abnormal, so that it can be further determined whether there is an anomaly in the voice to be detected according to the voice effective duration and / or the signal-to-noise ratio, thereby ensuring the accuracy of the voice anomaly detection result.

[0106] Based on any of the above embodiments, Figure 4 is a schematic flowchart of the implementation manner of step 110 in the voice anomaly detection method provided by the present invention, as Figure 4 shown, step 110 includes:

[0107] Step 111, extracting voicing signal data from the voice to be detected;

[0108] Step 112, determining the long-term power spectrum feature variance of each segment of voicing frame data based on the voicing frame data in the voicing signal data;

[0109] Step 113, determining the long-term power spectrum feature of the voice to be detected based on the long-term power spectrum feature variance of each segment of voicing frame data.

[0110] Specifically, in the voice to be detected, there may be voicing signal data and unvoiced signal data. Usually, the voicing signal data of normal speech has a low-pass characteristic, while the unvoiced signal data does not have an acoustic pulse and does not have a low-pass characteristic. Therefore, when determining the long-term power spectrum feature of the voice to be detected, it is necessary to filter out the unvoiced signal data from the voice to be detected to obtain the voicing signal data. For example, the voice to be detected can be subjected to voice / unvoice detection, and the voicing signal data can be extracted from the voice to be detected. After obtaining the voicing signal data, the long-term power spectrum feature variance of each segment of voicing frame data is determined, and then the long-term power spectrum feature of the voice to be detected is obtained based on the long-term power spectrum feature variance of each segment of voicing frame data.

[0111] Based on any of the above embodiments, step 113 includes:

[0112] Determining the long-term power spectrum feature of the voice to be detected based on the long-term power spectrum feature variance of each segment of voicing frame data and the number of voicing frames in each segment of voicing frame data.

[0113] Specifically, after obtaining the voiced signal data, based on the voiced frame data of each segment, the long-term power spectrum feature variance of the voiced frame data of each segment can be determined, and the initial long-term power spectrum feature can be determined based on the long-term power spectrum feature variance of the voiced frame data of each segment. Since the volume of the speech to be detected will affect the long-term power spectrum feature of the speech to be detected, after obtaining the initial long-term power spectrum feature, the initial long-term power spectrum feature can be normalized by the power spectrum according to formula (3) to obtain the long-term power spectrum feature of the speech to be detected that eliminates the influence of the volume size.

[0114] Among them, the long-term power spectrum feature of the speech to be detected can be determined by formulas (2) and (3):

[0115]

[0116]

[0117] Among them, represents the initial long-term power spectrum feature, represents the long-term power spectrum feature variance of the voiced frame data of each segment, L v represents the v-th segment in the voiced signal data, l v represents the segment number, and each segment L v includes L voiced frames, and there can be overlap between segments. Among them, in order to eliminate the influence of the formants of different phonemes, the length of the voiced frame needs to meet the preset requirements, for example, the length of the voiced frame is not less than 4s.

[0118] Based on any of the above embodiments, the present invention further provides a method for detecting abnormal speech, Figure 6 is a schematic structural diagram of the speech abnormal detection device provided by the present invention, as Figure 6 shown, the method includes:

[0119] First, frame and window the speech to be detected, then perform short-time Fourier transform on the broadband spectrogram based on the short-time Fourier transform, and then perform speech endpoint detection on the speech to be detected after the short-time Fourier transform, and mark the speech signal corresponding to the speech state on the speech signal to be detected according to the speech endpoint detection result.

[0120] Then, add up the durations of all the speech signals marked as the speech state to obtain the speech effective duration of the speech to be detected, and use the segment long-term signal-to-noise ratio of the speech to be detected as the segment SNR feature, perform voiced / unvoiced detection on the segment SNR feature, and filter out the unvoiced signal data in the speech to be detected to obtain the voiced signal data.

[0121] After obtaining the voiced signal data, based on the variance of the long-term power spectrum features of each segment of voiced frame data in the voiced signal data, the initial long-term power spectrum feature variance is determined, and the initial long-term power spectrum feature variance is subjected to power spectrum normalization processing to obtain the long-term power spectrum features of the speech to be detected.

[0122] The long-term power spectrum features of the speech to be detected are input into the trained Gaussian mixture model to obtain the likelihood score of the speech to be detected. Combining with the mean supervector of the Gaussian mixture model, the likelihood score is converted into a long-term power spectrum detection score with a numerical range of (0, 1), and the long-term power spectrum detection score is compared with the score range to determine the speech anomaly detection result. For example, if it is within the score range, it indicates that the probability of the speech to be detected having an anomaly is small, that is, the speech anomaly detection result is normal; if it exceeds the score range, it indicates that the probability of the speech to be detected having an anomaly is large, that is, the speech anomaly detection result is abnormal. Among them, the score range can be determined by statistically analyzing the score distributions of abnormal samples and normal samples.

[0123] It should be noted that after obtaining the speech anomaly detection result, if the speech anomaly detection result is normal, it is also possible to further determine whether the speech to be detected has an anomaly based on the effective speech duration of the speech to be detected and the signal-to-noise ratio of the speech to be detected. If it is determined that the speech to be detected has an anomaly based on either the effective speech duration of the speech to be detected or the signal-to-noise ratio of the speech to be detected, the speech anomaly detection result can be updated to abnormal.

[0124] Next, the speech anomaly detection device provided by the present invention will be described. The speech anomaly detection device described below can be correspondingly referred to the speech anomaly detection method described above.

[0125] Based on any of the above embodiments, the present invention also provides a speech anomaly detection device. Figure 6 is a schematic structural diagram of the speech anomaly detection device provided by the present invention, as Figure 6 shown, the device includes:

[0126] A determination unit 610, configured to determine the long-term power spectrum features of the speech to be detected;

[0127] A detection unit 620, configured to perform anomaly detection on the speech to be detected based on the sample long-term power spectrum features of the positive sample speech and the long-term power spectrum features of the speech to be detected, and obtain the speech anomaly detection result of the speech to be detected.

[0128] Based on any of the above embodiments, the detection unit 620 includes:

[0129] A likelihood score determination unit, configured to input the long-term power spectrum features of the speech to be detected into the speech anomaly detection model, and obtain the likelihood score output by the speech anomaly detection model;

[0130] A detection result determination unit, configured to determine a voice anomaly detection result of the voice to be detected based on the likelihood score;

[0131] Wherein, the voice anomaly detection model is trained based on the sample long-term power spectrum features of the positive sample voice.

[0132] Based on any of the above embodiments, the detection result determination unit includes:

[0133] A first score determination unit, configured to determine a maximum detection score and a minimum detection score based on the mean supervector of the voice anomaly detection model;

[0134] A second score determination unit, configured to determine a long-term power spectrum detection score based on the likelihood score, the maximum detection score, and the minimum detection score;

[0135] A detection result determination subunit, configured to determine a voice anomaly detection result of the voice to be detected based on the long-term power spectrum detection score.

[0136] Based on any of the above embodiments, it further includes:

[0137] An update unit, configured to update the voice anomaly detection result of the voice to be detected based on the voice effective duration and / or the signal-to-noise ratio of the voice to be detected after obtaining the voice anomaly detection result of the voice to be detected.

[0138] Based on any of the above embodiments, the update unit is configured to:

[0139] When the voice anomaly detection result of the voice to be detected is that the voice to be detected is normal, if it is determined that the voice to be detected has an anomaly based on the voice effective duration and / or the signal-to-noise ratio of the voice to be detected, then update the voice anomaly detection result of the voice to be detected to the voice to be detected is abnormal.

[0140] Based on any of the above embodiments, the determination unit 610 includes:

[0141] An extraction unit, configured to extract unvoiced signal data from the voice to be detected;

[0142] A first calculation unit, configured to determine the long-term power spectrum feature variance of each segment of unvoiced frame data based on each segment of unvoiced frame data in the unvoiced signal data;

[0143] A second calculation unit, configured to determine the long-term power spectrum feature of the voice to be detected based on the long-term power spectrum feature variance of each segment of unvoiced frame data.

[0144] Based on any of the above embodiments, the second calculation unit is configured to:

[0145] Determine the long-term power spectrum feature of the speech to be detected based on the variance of the long-term power spectrum features of each segment of voiced frame data and the number of voiced frames in each segment of voiced frame data.

[0146] Figure 7 FIG. [FIGURE NUMBER] is a schematic structural diagram of an electronic device provided by the present invention. As Figure 7 shown, the electronic device may include: a processor 710, a memory 720, a communication interface 730, and a communication bus 740. Among them, the processor 710, the memory 720, and the communication interface 730 communicate with each other through the communication bus 740. The processor 710 may call the logical instructions in the memory 720 to execute a voice anomaly detection method, which includes: determining the long-term power spectrum feature of the speech to be detected; based on the sample long-term power spectrum feature of the positive sample speech and the long-term power spectrum feature of the speech to be detected, performing anomaly detection on the speech to be detected to obtain the voice anomaly detection result of the speech to be detected.

[0147] In addition, when the logical instructions in the above-mentioned memory 720 are implemented in the form of a software functional unit and sold or used as an independent product, they may be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, may be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk, or an optical disc that can store program codes.

[0148] On the other hand, the present invention also provides a computer program product. The computer program product includes a computer program stored on a non-transitory computer-readable storage medium. The computer program includes program instructions. When the program instructions are executed by a computer, the computer can execute the voice anomaly detection method provided by the above-mentioned various methods, which includes: determining the long-term power spectrum feature of the speech to be detected; based on the sample long-term power spectrum feature of the positive sample speech and the long-term power spectrum feature of the speech to be detected, performing anomaly detection on the speech to be detected to obtain the voice anomaly detection result of the speech to be detected.

[0149] In another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is configured to execute the voice anomaly detection method provided above. The method includes: determining the long-term power spectrum feature of the voice to be detected; based on the sample long-term power spectrum feature of the positive sample voice and the long-term power spectrum feature of the voice to be detected, performing anomaly detection on the voice to be detected to obtain the voice anomaly detection result of the voice to be detected.

[0150] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative work.

[0151] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solution, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disc, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0152] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments or perform equivalent replacements for some of the technical features. These modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for detecting abnormal speech, characterized in that, it includes: Determine the long-term power spectrum feature of the speech to be detected; Based on the sample long-term power spectrum feature of the positive sample speech and the long-term power spectrum feature of the speech to be detected, perform abnormal detection on the speech to be detected to obtain the speech abnormal detection result of the speech to be detected; The performing abnormal detection on the speech to be detected based on the sample long-term power spectrum feature of the positive sample speech and the long-term power spectrum feature of the speech to be detected to obtain the speech abnormal detection result of the speech to be detected includes: Input the long-term power spectrum feature of the speech to be detected into a speech abnormal detection model to obtain the likelihood score output by the speech abnormal detection model; Based on the likelihood score, determine the speech abnormal detection result of the speech to be detected; wherein, the speech abnormal detection model is trained based on the sample long-term power spectrum feature of the positive sample speech; The determining the speech abnormal detection result of the speech to be detected based on the likelihood score includes: Based on the mean supervector of the speech abnormal detection model, determine the maximum detection score and the minimum detection score; Based on the likelihood score, the maximum detection score, and the minimum detection score, determine the long-term power spectrum detection score; Based on the long-term power spectrum detection score, determine the speech abnormal detection result of the speech to be detected.

2. The method for detecting abnormal speech according to claim 1, characterized in that, after obtaining the speech abnormal detection result of the speech to be detected, it further includes: Based on the speech effective duration of the speech to be detected and / or the signal-to-noise ratio of the speech to be detected, update the speech abnormal detection result of the speech to be detected.

3. The method for detecting abnormal speech according to claim 2, characterized in that, The updating the speech abnormal detection result of the speech to be detected based on the speech effective duration of the speech to be detected and / or the signal-to-noise ratio of the speech to be detected includes: When the speech abnormal detection result of the speech to be detected is that the speech to be detected is normal, if it is determined that the speech to be detected is abnormal based on the speech effective duration of the speech to be detected and / or the signal-to-noise ratio of the speech to be detected, then update the speech abnormal detection result of the speech to be detected to the speech to be detected is abnormal.

4. The method for detecting abnormal speech according to any one of claims 1 to 3, characterized in that, The determining the long-term power spectrum feature of the speech to be detected includes: Extract the voiced signal data from the speech to be detected; Based on each segment of voiced frame data in the voiced signal data, determine the variance of the long-term power spectrum feature of each segment of voiced frame data; Based on the variance of the long-term power spectrum feature of each segment of voiced frame data, determine the long-term power spectrum feature of the speech to be detected.

5. The method for detecting abnormal speech according to claim 4, characterized in that, The determining the long-term power spectrum feature of the speech to be detected based on the variance of the long-term power spectrum feature of each segment of voiced frame data includes: Determine the long-term power spectrum feature of the to-be-detected speech based on the variance of the long-term power spectrum features of each segment of voiced frame data and the number of voiced frames in each segment of voiced frame data.

6. A voice anomaly detection device characterized in that it includes: a determination unit for determining the long-term power spectrum feature of the to-be-detected speech; a detection unit for performing anomaly detection on the to-be-detected speech based on the sample long-term power spectrum feature of the positive sample speech and the long-term power spectrum feature of the to-be-detected speech to obtain the voice anomaly detection result of the to-be-detected speech; The performing anomaly detection on the to-be-detected speech based on the sample long-term power spectrum feature of the positive sample speech and the long-term power spectrum feature of the to-be-detected speech to obtain the voice anomaly detection result of the to-be-detected speech includes: inputting the long-term power spectrum feature of the to-be-detected speech into a voice anomaly detection model to obtain the likelihood score output by the voice anomaly detection model; determining the voice anomaly detection result of the to-be-detected speech based on the likelihood score; wherein, the voice anomaly detection model is trained based on the sample long-term power spectrum feature of the positive sample speech; The determining the voice anomaly detection result of the to-be-detected speech based on the likelihood score includes: determining the maximum detection score and the minimum detection score based on the mean supervector of the voice anomaly detection model; determining the long-term power spectrum detection score based on the likelihood score, the maximum detection score and the minimum detection score; determining the voice anomaly detection result of the to-be-detected speech based on the long-term power spectrum detection score.

7. An electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that when the processor executes the program, it implements the steps of the voice anomaly detection method according to any one of claims 1 to 5.

8. A non-transitory computer-readable storage medium, on which a computer program is stored, characterized in that when the computer program is executed by a processor, it implements the steps of the voice anomaly detection method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Voice input exception determining method, apparatus, terminal, and storage medium

    CA2981775A1

  • Voice model training method, voice recognition method, devices, facility and medium

    CN108922515A