Voice detection method, device, equipment and readable storage medium

By separating human voice and background noise from the detected voice, extracting features for similarity calculation, it is solved in the prior art that AI imitates speech and real-person speech splicing, and achieving higher speech detection accuracy.

CN116168725BActive Publication Date: 2025-06-06FIBERHOME TELECOMMUNICATION TECHNOLOGIES CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202310089664.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-02-09
Publication Date
2025-06-06
Estimated Expiration
2043-02-09

AI Technical Summary

Technical Problem

The prior art is difficult to effectively distinguish the situation where AI software imitates voice and live voice splicing, resulting in low accuracy of authenticity recognition.

Method used

By obtaining the voice to be detected, the real target voice and the real background noise, separating the voice and background noise, extracting features and calculating the similarity, and determining the authenticity of the voice to be detected.

Benefits of technology

It improves the accuracy of voice detection, can effectively distinguish between real-person voice and fake voice, and reduces the risk of victims.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116168725B_ABST
    Figure CN116168725B_ABST
Patent Text Reader

Abstract

The present invention provides a speech detection method, device, equipment and readable storage medium. The speech detection method comprises: separating the human voice and background noise of the speech to be detected to obtain the human voice to be detected and the background noise to be detected; determining whether the background noise to be detected is real by calculating the similarity between the first feature of the background noise to be detected and the first feature of the real background noise; if the background noise to be detected is not real, determining that the speech to be detected is not a real person's voice; if the background noise to be detected is real, determining whether the speech to be detected is a real person's voice by calculating the similarity between the second feature of the human voice to be detected and the second feature of the real target human voice. The present invention extracts features and compares the similarity between the background noise to be detected and the real background noise obtained by separating the speech to be detected. Since the real background noise is difficult to obtain and imitate, the accuracy of speech detection can be effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of speech recognition technology, and in particular to a speech detection method, device, equipment and readable storage medium. Background Art

[0002] In daily life, we often receive marketing and fraud calls from disguised human voices. These voice calls automatically generated by computer AI (Artificial Intelligence) are very close to real human voices. Inexperienced people are easily deceived, resulting in economic losses. Some voices generated by using speech synthesis technology or splicing real voices send voice messages to relatives and friends through chat tools to borrow money, resulting in property losses. Existing software with voice intelligence can highly imitate the voice characteristics of real people or generate voices by splicing real human voices. It is difficult for ordinary people to distinguish between the real and the fake, and they may fall into the trap and become victims if they are not careful.

[0003] At present, voiceprint models are usually established to identify whether they are the same person. However, since AI software can highly imitate the voice characteristics of real people, the accuracy of authenticity judgment by analyzing the similarity between the voice characteristics of real people and recorded files is not high. In addition, real people's voices can be obtained through chat tools and recording software, and real people's voices can be forged by splicing voices, which makes it difficult to distinguish the true from the false. Therefore, the current voice detection method based on voiceprint recognition technology is difficult to effectively distinguish the authenticity of some AI software's imitation voices and real people's voices. Summary of the invention

[0004] The main purpose of the present invention is to provide a voice detection method, device, equipment and readable storage medium, aiming to solve the technical problem that the current voice detection method based on voiceprint recognition technology is difficult to effectively distinguish the authenticity of some AI software's imitation voice and real person voice splicing.

[0005] In a first aspect, the present invention provides a speech detection method, the speech detection method comprising:

[0006] Obtain the speech to be detected, the real target voice and the real background noise;

[0007] Separate the human voice and background noise of the speech to be detected to obtain the human voice to be detected and the background noise to be detected;

[0008] Extracting first features from the background noise to be detected and the real background noise respectively, and determining whether the background noise to be detected is real by calculating the similarity between the first feature of the background noise to be detected and the first feature of the real background noise;

[0009] If the background noise to be detected is not real, it is determined that the voice to be detected is not a real person's voice;

[0010] If the background noise to be detected is real, the second feature is extracted from the human voice to be detected and the real target human voice respectively, and the similarity between the second feature of the human voice to be detected and the second feature of the real target human voice is calculated to determine whether the voice to be detected is a real person's voice.

[0011] Optionally, the separating the human voice and the background noise of the speech to be detected to obtain the human voice to be detected and the background noise to be detected includes:

[0012] Divide the speech to be detected into a plurality of continuous speech frames of equal duration;

[0013] Convert each speech frame into the frequency domain using Fourier transform, and divide each speech frame into multiple frequency intervals according to the frequency domain;

[0014] According to the energy value of each frequency interval and the preset energy value of each frequency interval, a probability calculation model is used to determine whether each speech frame is a human voice;

[0015] Merging the speech frames that are determined to be human voices and are continuous in time, and using one or more segments obtained by merging as human voices to be detected;

[0016] The speech frames that are determined not to be human voices and are continuous in time are merged, and one or more segments obtained by merging are used as the background noise to be detected.

[0017] Optionally, extracting the first feature from the background noise to be detected and the real background noise respectively, and determining whether the background noise to be detected is real by calculating the similarity between the first feature of the background noise to be detected and the first feature of the real background noise includes:

[0018] The longest segment in the background noise to be detected is used as the first comparison segment;

[0019] Extracting first features from the first comparison segment, wherein the first features include standard deviation, peak value, single peak mean value and average energy;

[0020] If the standard deviation of the first comparison segment is less than the preset standard deviation, the background noise to be detected is determined to be unreal, otherwise the standard deviation, peak value, single peak mean and average energy are extracted from the real background noise;

[0021] Using a probability calculation model, respectively calculating the similarity between the standard deviation, peak value, single peak mean value and average energy of the real background noise and the standard deviation, peak value, single peak mean value and average energy of the first comparison segment, to obtain a first similarity, a second similarity, a third similarity and a fourth similarity;

[0022] Calculate the product of the first similarity, the second similarity, the third similarity and the fourth similarity to obtain a first joint similarity;

[0023] If the first joint similarity is greater than the first preset joint similarity, the background noise to be detected is determined to be real; otherwise, the background noise to be detected is determined to be unreal.

[0024] Optionally, extracting the second feature from the human voice to be detected and the real target human voice respectively, and determining whether the voice to be detected is a real person's voice by calculating the similarity between the second feature of the human voice to be detected and the second feature of the real target human voice includes:

[0025] Separating the real target human voice from the background noise, and selecting the longest real target human voice segment from the one or more separated real target human voice segments as the second comparison segment;

[0026] Extracting second features from the second comparison segment, wherein the second features include volume mean, pitch period, speech rate, and spectrum distribution;

[0027] Extract the volume mean, pitch period, speech rate and spectrum distribution of each human voice segment to be detected;

[0028] Using a probability calculation model, respectively calculating the volume mean, pitch period, speech rate and spectral distribution of each human voice segment to be detected and the similarity between the volume mean, pitch period, speech rate and spectral distribution of the second comparison segment, to obtain the fifth similarity, sixth similarity, seventh similarity and eighth similarity of each human voice segment to be detected;

[0029] Calculate the product of the fifth similarity, the sixth similarity, the seventh similarity and the eighth similarity of each to-be-detected vocal segment to obtain a second joint similarity of each to-be-detected vocal segment;

[0030] If the second joint similarity of at least one to-be-detected human voice segment is greater than the second preset joint similarity, the to-be-detected voice is determined to be a real person's voice; otherwise, the to-be-detected voice is determined not to be a real person's voice.

[0031] Optionally, after determining that the voice to be detected is a real person's voice, the method further includes:

[0032] Acquire a real third-person voice, separate the real third-person voice from background noise, and select the longest real third-person voice segment from one or more separated segments as a third comparison segment;

[0033] Extracting the volume mean, pitch period, speech rate and spectrum distribution from the third comparison segment;

[0034] Using a probability calculation model, respectively calculating the similarity between the volume mean, pitch period, speech rate and spectrum distribution of each human voice segment to be detected and the volume mean, pitch period, speech rate and spectrum distribution of the third comparison segment, to obtain the ninth similarity, the tenth similarity, the eleventh similarity and the twelfth similarity of each human voice segment to be detected;

[0035] Calculate the product of the ninth similarity, the tenth similarity, the eleventh similarity and the twelfth similarity of each to-be-detected human voice segment to obtain a third joint similarity of each to-be-detected human voice segment;

[0036] If the third joint similarities of all the segments of the human voice to be detected are not greater than the third preset joint similarity, it is determined that the voice to be detected is a real person's voice, but there is no third person's proof, otherwise the ratio between the average volume of the third comparison segment and the average volume of the second comparison segment is calculated as the first ratio;

[0037] Calculating a ratio between an average volume of a third person's vocal segment in the vocal segment to be detected and an average volume of the target vocal segment as a second ratio;

[0038] If the difference between the first ratio and the second ratio is less than the preset difference, the voice to be detected is determined to be a real person's voice and is certified by a third party; otherwise, the voice to be detected is determined to be a real person's voice but is not certified by a third party.

[0039] In a second aspect, the present invention further provides a speech detection device, the speech detection device comprising:

[0040] An acquisition module is used to acquire the speech to be detected, the real target human voice and the real background noise;

[0041] A separation module is used to separate the human voice and background noise of the speech to be detected, so as to obtain the human voice to be detected and the background noise to be detected;

[0042] A determination module, used to extract first features from the background noise to be detected and the real background noise respectively, and determine whether the background noise to be detected is real by calculating the similarity between the first feature of the background noise to be detected and the first feature of the real background noise;

[0043] If the background noise to be detected is not real, it is determined that the voice to be detected is not a real person's voice;

[0044] If the background noise to be detected is real, the second feature is extracted from the human voice to be detected and the real target human voice respectively, and the similarity between the second feature of the human voice to be detected and the second feature of the real target human voice is calculated to determine whether the voice to be detected is a real person's voice.

[0045] Optionally, the separation module is used to:

[0046] Divide the speech to be detected into a plurality of continuous speech frames of equal duration;

[0047] Convert each speech frame into the frequency domain using Fourier transform, and divide each speech frame into multiple frequency intervals according to the frequency domain;

[0048] According to the energy value of each frequency interval and the preset energy value of each frequency interval, a probability calculation model is used to determine whether each speech frame is a human voice;

[0049] Merging the speech frames that are determined to be human voices and are continuous in time, and using one or more segments obtained by merging as human voices to be detected;

[0050] The speech frames that are determined not to be human voices and are continuous in time are merged, and one or more segments obtained by merging are used as the background noise to be detected.

[0051] Optionally, the determination module is used to:

[0052] The longest segment in the background noise to be detected is used as the first comparison segment;

[0053] Extracting first features from the first comparison segment, wherein the first features include standard deviation, peak value, single peak mean value and average energy;

[0054] If the standard deviation of the first comparison segment is less than the preset standard deviation, the background noise to be detected is determined to be unreal, otherwise the standard deviation, peak value, single peak mean and average energy are extracted from the real background noise;

[0055] Using a probability calculation model, respectively calculating the similarity between the standard deviation, peak value, single peak mean value and average energy of the real background noise and the standard deviation, peak value, single peak mean value and average energy of the first comparison segment, to obtain a first similarity, a second similarity, a third similarity and a fourth similarity;

[0056] Calculate the product of the first similarity, the second similarity, the third similarity and the fourth similarity to obtain a first joint similarity;

[0057] If the first joint similarity is greater than the first preset joint similarity, the background noise to be detected is determined to be real; otherwise, the background noise to be detected is determined to be unreal.

[0058] In a third aspect, the present invention also provides a speech detection device, comprising a processor, a memory, and a speech detection program stored in the memory and executable by the processor, wherein when the speech detection program is executed by the processor, the steps of the speech detection method described above are implemented.

[0059] In a fourth aspect, the present invention further provides a readable storage medium, on which a speech detection program is stored, wherein when the speech detection program is executed by a processor, the steps of the speech detection method as described above are implemented.

[0060] In the present invention, a speech to be detected, a real target human voice and real background noise are obtained; the human voice and background noise of the speech to be detected are separated to obtain the human voice to be detected and the background noise to be detected; the first feature is extracted from the background noise to be detected and the real background noise respectively, and the similarity between the first feature of the background noise to be detected and the first feature of the real background noise is calculated to determine whether the background noise to be detected is real; if the background noise to be detected is not real, it is determined that the speech to be detected is not a real person's voice; if the background noise to be detected is real, the second feature is extracted from the human voice to be detected and the real target human voice respectively, and the similarity between the second feature of the human voice to be detected and the second feature of the real target human voice is calculated to determine whether the speech to be detected is a real person's voice. The present invention further extracts features and compares similarities between the background noise to be detected and the real background noise obtained by separating the speech to be detected on the basis of human voice detection. Since the background noise is difficult to obtain and imitate, the accuracy of speech detection can be effectively improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0061] Figure 1 A flow chart of an embodiment of a speech detection method of the present invention;

[0062] Figure 2 for Figure 1 Detailed flow chart of step S20;

[0063] Figure 3 for Figure 1 Detailed flow chart of step S30;

[0064] Figure 4 for Figure 1 A detailed flow chart of step S50;

[0065] Figure 5 A schematic diagram of functional modules of an embodiment of a speech detection device of the present invention;

[0066] Figure 6 The figure is a schematic diagram of the hardware structure of a speech detection device according to an embodiment of the present invention.

[0067] The realization of the purpose, functional features and advantages of the present invention will be further explained in conjunction with embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION

[0068] It should be understood that the specific embodiments described herein are only used to explain the present invention, and are not used to limit the present invention.

[0069] In a first aspect, an embodiment of the present invention provides a speech detection method.

[0070] In order to more clearly demonstrate the speech detection method provided in the embodiment of the present application, the application scenario of the speech detection method provided in the embodiment of the present application is first introduced.

[0071] The voice detection method provided in the embodiment of the present application is applied to voice calls automatically generated by computer AI to search for targets over a large range, as well as some voice synthesis technologies that splice or generate real-person voices, and send voice messages to relatives and friends through chat tools to borrow money, which may cause property losses. It is difficult for ordinary people to distinguish between the true and the false, and they may fall into the trap and become victims if they are not careful. Therefore, it is very necessary to accurately detect the authenticity of the voice.

[0072] In one embodiment, referring to Figure 1 , Figure 1 FIG. 1 is a flow chart of an embodiment of a speech detection method of the present invention. Figure 1 As shown, the voice detection method includes:

[0073] Step S10, obtaining the speech to be detected, the real target human voice and the real background noise.

[0074] In this embodiment, the voice to be detected is a voice that needs to be detected to distinguish its authenticity. The voice to be detected can be obtained from channels such as real-time voice calls or chat tools, and the real target human voice and real background noise are obtained in advance for comparison with the voice to be detected to distinguish the authenticity of the voice to be detected.

[0075] Step S20, separating the human voice and the background noise of the speech to be detected to obtain the human voice to be detected and the background noise to be detected.

[0076] In this embodiment, the human voice to be detected and the background noise to be detected are separated from the speech to be detected for comparison with the real target human voice and the real background noise.

[0077] Step S30, extracting the first feature from the background noise to be detected and the real background noise respectively, and determining whether the background noise to be detected is real by calculating the similarity between the first feature of the background noise to be detected and the first feature of the real background noise.

[0078] In this embodiment, after extracting the first feature from the background noise to be detected and the real background noise respectively, the similarity between the first feature of the background noise to be detected and the first feature of the real background noise is calculated to determine whether the background noise to be detected is real.

[0079] Step S40: If the background noise to be detected is not real, it is determined that the voice to be detected is not a real person's voice.

[0080] In this embodiment, since it is difficult for a voice forger to obtain and imitate real background noise, if the background noise to be detected is not real, it can be directly determined that the voice to be detected is not a real person's voice.

[0081] Step S50, if the background noise to be detected is real, extract the second feature from the human voice to be detected and the real target human voice respectively, and determine whether the voice to be detected is a real person's voice by calculating the similarity between the second feature of the human voice to be detected and the second feature of the real target human voice.

[0082] In this embodiment, on the basis that the background noise to be detected is real, after extracting the second feature from the human voice to be detected and the real target human voice respectively, the similarity between the second feature of the human voice to be detected and the second feature of the real target human voice is calculated to determine whether the voice to be detected is a real person's voice.

[0083] In the present embodiment, in addition to extracting features and performing similarity comparison between the human voice to be detected obtained by separating the voice to be detected and the real target human voice, feature extraction and similarity comparison between the background noise to be detected obtained by separating the voice to be detected and the real background noise are added. Since it is difficult for a voice forger to obtain and imitate the real background noise, the accuracy of voice detection can be effectively improved. After detecting a non-human voice, the user can be reminded of the risk in real time, and computer-generated marketing and fraud calls can be automatically blocked.

[0084] Further, in one embodiment, referring to Figure 2 , Figure 2 for Figure 1 The detailed flow chart of step S20 is as follows: Figure 2 As shown, step S20 includes:

[0085] Step S201, dividing the speech to be detected into a plurality of continuous speech frames of equal duration;

[0086] Step S202, converting each speech frame into a frequency domain using Fourier transform, and dividing each speech frame into a plurality of frequency intervals according to the frequency domain;

[0087] Step S203, according to the energy value of each frequency interval and the preset energy value of each frequency interval, a probability calculation model is used to determine whether each speech frame is a human voice;

[0088] Step S204, merging the speech frames that are determined to be human voices and are continuous in time, and using the merged one or more segments as human voices to be detected;

[0089] Step S205: merge the speech frames that are determined not to be human voices and are continuous in time, and use one or more merged segments as background noise to be detected.

[0090] In this embodiment, the speech to be detected can be divided into a plurality of continuous speech frames of equal duration every 20 milliseconds, and each speech frame can be divided into six frequency intervals of 3000Hz-4000Hz, 2000Hz-3000Hz, 1000Hz-2000Hz, 500Hz-1000Hz, 250Hz-500Hz and 80Hz-250Hz according to the frequency domain. The low-frequency part of 0Hz-80Hz that is inaudible to the human ear can be filtered out, and the energy values ​​of the six frequency intervals are extracted respectively. Combined with the preset energy value of each frequency interval, a probability calculation model is used to determine whether each speech frame is a human voice or a non-human voice. The probability calculation model mainly adopts probability analysis. The statistical method of the cloth is used to calculate the similarity. In the specific implementation process, Gaussian probability model, Markov model and logistic regression model can be used. Non-human voice is background noise. Then, the speech frames that are continuous in time and judged to be human voice or background noise are merged, that is, the human voice to be detected and the background noise to be detected are obtained. Since the divided speech frames are temporal, and the same speech frame will be judged as human voice or background noise, therefore, the speech frames that are continuous in time and judged to be human voice or background noise are merged, which may cause the speech frames to be discontinuous in time. Therefore, the speech frames that are both human voice or background noise and continuous in time are merged to obtain one or more fragments.

[0091] Further, in one embodiment, referring to Figure 3 , Figure 3 for Figure 1 The detailed flow chart of step S30 is as follows: Figure 3 As shown, step S30 includes:

[0092] Step S301, taking the longest segment in the background noise to be detected as the first comparison segment;

[0093] Step S302, extracting a first feature from the first comparison segment, wherein the first feature includes a standard deviation, a peak value, a single peak mean value, and an average energy;

[0094] Step S303, if the standard deviation of the first comparison segment is less than the preset standard deviation, the background noise to be detected is determined to be unreal, otherwise the standard deviation, peak value, single peak mean and average energy are extracted from the real background noise;

[0095] Step S304, using a probability calculation model, respectively calculating the similarity between the standard deviation, peak value, single peak mean value and average energy of the real background noise and the standard deviation, peak value, single peak mean value and average energy of the first comparison segment, to obtain a first similarity, a second similarity, a third similarity and a fourth similarity;

[0096] Step S305, calculating the product of the first similarity, the second similarity, the third similarity and the fourth similarity to obtain a first joint similarity;

[0097] Step S306: if the first joint similarity is greater than the first preset joint similarity, the background noise to be detected is determined to be real; otherwise, the background noise to be detected is determined to be unreal.

[0098] In this embodiment, the longest segment in the background noise to be detected is the most representative, so it is used as the first comparison segment, and the first feature is extracted for similarity comparison with the first feature extracted from the real background noise. If the standard deviation of the first comparison segment is less than the preset standard deviation, it means that the first comparison segment in the background noise to be detected is silent. Since the speech in the real scene is noisy, it can be directly determined that the background noise to be detected is not real. If the background noise to be detected is not silent, a probability calculation model is further used to calculate the similarity between the standard deviation, peak value, single peak mean and average energy of the real background noise and the standard deviation, peak value, single peak mean and average energy of the first comparison segment, and compare the joint similarity to determine whether the background noise to be detected is real.

[0099] Further, in one embodiment, referring to Figure 4 , Figure 4 for Figure 1 The detailed flow chart of step S50 is as follows: Figure 4 As shown, step S50 includes:

[0100] Step S501, separating the real target human voice from the background noise, and selecting the longest real target human voice segment from one or more real target human voice segments obtained by separation as the second comparison segment;

[0101] Step S502, extracting second features from the second comparison segment, where the second features include volume mean, pitch period, speech rate, and spectrum distribution;

[0102] Step S503, extracting the volume mean, pitch period, speech rate and spectrum distribution of each human voice segment to be detected;

[0103] Step S504, using a probability calculation model, respectively calculating the volume mean, pitch period, speech rate and spectral distribution of each human voice segment to be detected and the similarity between the volume mean, pitch period, speech rate and spectral distribution of the second comparison segment, to obtain the fifth similarity, sixth similarity, seventh similarity and eighth similarity of each human voice segment to be detected;

[0104] Step S505, calculating the product of the fifth similarity, the sixth similarity, the seventh similarity and the eighth similarity of each to-be-detected vocal segment, to obtain a second joint similarity of each to-be-detected vocal segment;

[0105] Step S506: if the second joint similarity of at least one of the to-be-detected human voice segments is greater than the second preset joint similarity, it is determined that the to-be-detected voice is a real person's voice; otherwise, it is determined that the to-be-detected voice is not a real person's voice.

[0106] In this embodiment, the same method of separating human voice and background noise as that of the speech to be detected is used to separate the human voice and background noise of the real target human voice, and the most representative real target human voice segment with the longest time is selected as the second comparison segment. Among the one or more human voice segments to be detected separated from the speech to be detected, as long as the second joint similarity of one human voice segment to be detected is greater than the second preset joint similarity, it means that the human voice to be detected in the speech to be detected is consistent with the real target human voice, and it can be determined that the speech to be detected is a real person's voice, otherwise it is determined that the speech to be detected is not a real person's voice.

[0107] Further, in one embodiment, after determining that the voice to be detected is a real person's voice, the following steps are included:

[0108] Acquire a real third-person voice, separate the real third-person voice from background noise, and select the longest real third-person voice segment from one or more separated segments as a third comparison segment;

[0109] Extracting the volume mean, pitch period, speech rate and spectrum distribution from the third comparison segment;

[0110] Using a probability calculation model, respectively calculating the similarity between the volume mean, pitch period, speech rate and spectrum distribution of each human voice segment to be detected and the volume mean, pitch period, speech rate and spectrum distribution of the third comparison segment, to obtain the ninth similarity, the tenth similarity, the eleventh similarity and the twelfth similarity of each human voice segment to be detected;

[0111] Calculate the product of the ninth similarity, the tenth similarity, the eleventh similarity and the twelfth similarity of each to-be-detected human voice segment to obtain a third joint similarity of each to-be-detected human voice segment;

[0112] If the third joint similarities of all the segments of the human voice to be detected are not greater than the third preset joint similarity, it is determined that the voice to be detected is a real person's voice, but there is no third person's proof, otherwise the ratio between the average volume of the third comparison segment and the average volume of the second comparison segment is calculated as the first ratio;

[0113] Calculating a ratio between an average volume of a third person's vocal segment in the vocal segment to be detected and an average volume of the target vocal segment as a second ratio;

[0114] If the difference between the first ratio and the second ratio is less than the preset difference, the voice to be detected is determined to be a real person's voice and is certified by a third party; otherwise, the voice to be detected is determined to be a real person's voice but is not certified by a third party.

[0115] In this embodiment, the use of a third person's voice as proof is mainly used in some application scenarios that require higher accuracy and security of voice detection, such as the authenticity of a judicial authentication recording file, etc., to obtain the real third person's voice in advance, use the same voice and background noise separation method as the voice to be detected, separate the real third person's voice from the background noise, and also select the most representative and longest real third person voice segment as the third comparison segment, detect the second feature of each segment in the voice to be detected, and whether it is consistent with the second feature of the third comparison segment, to determine whether there is a third person in the voice to be detected as proof. If the similarity of all the voice segments to be detected is not verified, it means that the real third person's voice has not been verified, that is, the real third person's voice is not detected in the voice segment to be detected, that is, there is no third person proof. Whether there is a third person as proof can assist in verifying whether the voice to be detected is real, thereby improving the accuracy and security of voice detection. It should be noted that the use of a third person's voice to assist in verification does not necessarily require the third person and the target to be at the same scene, such as using a voice connection between the third person and the target to assist in verification and other application scenarios. The validity of the third-party proof is further determined by the fusion characteristics of the target voice and the third-party voice, that is, the ratio of the average volume, making it difficult for the voice generation software to imitate. The third comparison segment is the longest segment of the real third-party voice, and the second comparison segment is the longest segment of the real target voice. The ratio of their average volumes is the first ratio, that is, the ratio of the average volume of the real third-party voice to the average volume of the real target voice. The second ratio is the ratio between the average volume of the third-party voice segment in the voice segment to be detected and the average volume of the target voice segment. Under normal circumstances, the first ratio and the second ratio should be close, that is, the difference between the two should be less than the preset difference. If the difference between the two is too large, it means that the third-party voice may be forged, that is, there is no third-party proof.

[0116] In a second aspect, an embodiment of the present invention further provides a speech detection device.

[0117] Reference Figure 5 , Figure 5 Schematic diagram of functional modules of a speech detection device according to an embodiment of the present invention.

[0118] In this embodiment, the speech detection device includes:

[0119] An acquisition module 10 is used to acquire the speech to be detected, the real target human voice and the real background noise;

[0120] The separation module 20 is used to separate the human voice and the background noise of the speech to be detected, so as to obtain the human voice to be detected and the background noise to be detected;

[0121] The determination module 30 is used to extract the first feature from the background noise to be detected and the real background noise respectively, and determine whether the background noise to be detected is real by calculating the similarity between the first feature of the background noise to be detected and the first feature of the real background noise;

[0122] If the background noise to be detected is not real, it is determined that the voice to be detected is not a real person's voice;

[0123] If the background noise to be detected is real, the second feature is extracted from the human voice to be detected and the real target human voice respectively, and the similarity between the second feature of the human voice to be detected and the second feature of the real target human voice is calculated to determine whether the voice to be detected is a real person's voice.

[0124] Furthermore, in one embodiment, the separation module 20 is used to:

[0125] Divide the speech to be detected into a plurality of continuous speech frames of equal duration;

[0126] Convert each speech frame into the frequency domain using Fourier transform, and divide each speech frame into multiple frequency intervals according to the frequency domain;

[0127] According to the energy value of each frequency interval and the preset energy value of each frequency interval, a probability calculation model is used to determine whether each speech frame is a human voice;

[0128] Merging the speech frames that are determined to be human voices and are continuous in time, and using one or more segments obtained by merging as human voices to be detected;

[0129] The speech frames that are determined not to be human voices and are continuous in time are merged, and one or more segments obtained by merging are used as the background noise to be detected.

[0130] Furthermore, in one embodiment, the determination module 30 is used to:

[0131] The longest segment in the background noise to be detected is used as the first comparison segment;

[0132] Extracting first features from the first comparison segment, wherein the first features include standard deviation, peak value, single peak mean value and average energy;

[0133] If the standard deviation of the first comparison segment is less than the preset standard deviation, the background noise to be detected is determined to be unreal, otherwise the standard deviation, peak value, single peak mean and average energy are extracted from the real background noise;

[0134] Using a probability calculation model, respectively calculating the similarity between the standard deviation, peak value, single peak mean value and average energy of the real background noise and the standard deviation, peak value, single peak mean value and average energy of the first comparison segment, to obtain a first similarity, a second similarity, a third similarity and a fourth similarity;

[0135] Calculate the product of the first similarity, the second similarity, the third similarity and the fourth similarity to obtain a first joint similarity;

[0136] If the first joint similarity is greater than the first preset joint similarity, the background noise to be detected is determined to be real; otherwise, the background noise to be detected is determined to be unreal.

[0137] Furthermore, in one embodiment, the determination module 30 is further configured to:

[0138] Separating the real target human voice from the background noise, and selecting the longest real target human voice segment from the one or more separated real target human voice segments as the second comparison segment;

[0139] Extracting second features from the second comparison segment, wherein the second features include volume mean, pitch period, speech rate, and spectrum distribution;

[0140] Extract the volume mean, pitch period, speech rate and spectrum distribution of each human voice segment to be detected;

[0141] Using a probability calculation model, respectively calculating the volume mean, pitch period, speech rate and spectral distribution of each human voice segment to be detected and the similarity between the volume mean, pitch period, speech rate and spectral distribution of the second comparison segment, to obtain the fifth similarity, sixth similarity, seventh similarity and eighth similarity of each human voice segment to be detected;

[0142] Calculate the product of the fifth similarity, the sixth similarity, the seventh similarity and the eighth similarity of each to-be-detected vocal segment to obtain a second joint similarity of each to-be-detected vocal segment;

[0143] If the second joint similarity of at least one to-be-detected human voice segment is greater than the second preset joint similarity, the to-be-detected voice is determined to be a real person's voice; otherwise, the to-be-detected voice is determined not to be a real person's voice.

[0144] Furthermore, in one embodiment, the speech detection device further includes a third person voice detection module, which is used to:

[0145] Acquire a real third-person voice, separate the real third-person voice from background noise, and select the longest real third-person voice segment from one or more separated segments as a third comparison segment;

[0146] Extracting the volume mean, pitch period, speech rate and spectrum distribution from the third comparison segment;

[0147] Using a probability calculation model, respectively calculating the similarity between the volume mean, pitch period, speech rate and spectrum distribution of each human voice segment to be detected and the volume mean, pitch period, speech rate and spectrum distribution of the third comparison segment, to obtain the ninth similarity, the tenth similarity, the eleventh similarity and the twelfth similarity of each human voice segment to be detected;

[0148] Calculate the product of the ninth similarity, the tenth similarity, the eleventh similarity and the twelfth similarity of each to-be-detected human voice segment to obtain a third joint similarity of each to-be-detected human voice segment;

[0149] If the third joint similarities of all the segments of the human voice to be detected are not greater than the third preset joint similarity, it is determined that the voice to be detected is a real person's voice, but there is no third person's proof, otherwise the ratio between the average volume of the third comparison segment and the average volume of the second comparison segment is calculated as the first ratio;

[0150] Calculating a ratio between an average volume of a third person's vocal segment in the vocal segment to be detected and an average volume of the target vocal segment as a second ratio;

[0151] If the difference between the first ratio and the second ratio is less than the preset difference, the voice to be detected is determined to be a real person's voice and is certified by a third party; otherwise, the voice to be detected is determined to be a real person's voice but is not certified by a third party.

[0152] Among them, the functional implementation of each module in the above-mentioned speech detection device corresponds to each step in the above-mentioned speech detection method embodiment, and its functions and implementation processes are no longer repeated here.

[0153] In a third aspect, an embodiment of the present invention provides a voice detection device, which may be a device having a data processing function, such as a personal computer (PC), a notebook computer, or a server.

[0154] Reference Figure 6 , Figure 6The hardware structure diagram of an embodiment of the speech detection device of the present invention. In the embodiment of the present invention, the speech detection device may include a processor 1001 (e.g., a central processing unit, CPU), a communication bus 1002, a user interface 1003, a network interface 1004, and a memory 1005. Among them, the communication bus 1002 is used to realize the connection and communication between these components; the user interface 1003 may include a display screen (Display), an input unit such as a microphone (Microphone) and a keyboard (Keyboard); the network interface 1004 may optionally include a standard wired interface and a wireless interface (such as a wireless fidelity WIreless-FIdelity, WI-FI interface); the memory 1005 may be a high-speed random access memory (random access memory, RAM), or a stable memory (non-volatile memory), such as a disk storage, and the memory 1005 may optionally be a storage device independent of the aforementioned processor 1001. Those skilled in the art will understand that Figure 6 The hardware structure shown in the figure does not constitute a limitation of the present invention, and may include more or less components than those shown in the figure, or combine certain components, or arrange the components differently.

[0155] Continue to refer to Figure 6 , Figure 6 The memory 1005 as a computer storage medium may include an operating system, a network communication module, a user interface module and a voice detection program. The processor 1001 may call the voice detection program stored in the memory 1005 and execute the voice detection method provided in the embodiment of the present invention.

[0156] In a fourth aspect, an embodiment of the present invention further provides a readable storage medium.

[0157] The readable storage medium of the present invention stores a voice detection program, wherein when the voice detection program is executed by a processor, the steps of the voice detection method described above are implemented.

[0158] Among them, the method implemented when the voice detection program is executed can refer to the various embodiments of the voice detection method of the present invention, and will not be repeated here.

[0159] It should be noted that, in this article, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or system including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or system. In the absence of further restrictions, an element defined by the sentence "comprises a ..." does not exclude the existence of other identical elements in the process, method, article or system including the element.

[0160] The serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages or disadvantages of the embodiments.

[0161] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) as described above, and includes a number of instructions for a terminal device to execute the methods described in each embodiment of the present invention.

[0162] The above are only preferred embodiments of the present invention, and are not intended to limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made using the contents of the present invention specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present invention.

Claims

1. A speech detection method, It is characterized in that The speech detection method comprises: Obtain the speech to be detected, the real target voice and the real background noise; Separate the human voice and background noise of the speech to be detected to obtain the human voice to be detected and the background noise to be detected; Extracting first features from the background noise to be detected and the real background noise respectively, and determining whether the background noise to be detected is real by calculating the similarity between the first feature of the background noise to be detected and the first feature of the real background noise; If the background noise to be detected is not real, it is determined that the voice to be detected is not a real person's voice; If the background noise to be detected is real, extract the second feature from the human voice to be detected and the real target human voice respectively, and determine whether the voice to be detected is a real person's voice by calculating the similarity between the second feature of the human voice to be detected and the second feature of the real target human voice; The method of separating the human voice and the background noise of the speech to be detected to obtain the human voice to be detected and the background noise to be detected comprises: Divide the speech to be detected into a plurality of continuous speech frames of equal duration; Convert each speech frame into the frequency domain using Fourier transform, and divide each speech frame into multiple frequency intervals according to the frequency domain; According to the energy value of each frequency interval and the preset energy value of each frequency interval, a probability calculation model is used to determine whether each speech frame is a human voice; Merging the speech frames that are determined to be human voices and are continuous in time, and using one or more segments obtained by merging as human voices to be detected; Merging speech frames that are determined not to be human voices and are continuous in time, and using one or more merged segments as background noise to be detected; The extracting the first feature from the background noise to be detected and the real background noise respectively, and determining whether the background noise to be detected is real by calculating the similarity between the first feature of the background noise to be detected and the first feature of the real background noise comprises: The longest segment in the background noise to be detected is used as the first comparison segment; Extracting first features from the first comparison segment, wherein the first features include standard deviation, peak value, single peak mean value and average energy; If the standard deviation of the first comparison segment is less than the preset standard deviation, the background noise to be detected is determined to be unreal, otherwise the standard deviation, peak value, single peak mean and average energy are extracted from the real background noise; Using a probability calculation model, respectively calculating the similarity between the standard deviation, peak value, single peak mean value and average energy of the real background noise and the standard deviation, peak value, single peak mean value and average energy of the first comparison segment, to obtain a first similarity, a second similarity, a third similarity and a fourth similarity; Calculate the product of the first similarity, the second similarity, the third similarity and the fourth similarity to obtain a first joint similarity; If the first joint similarity is greater than the first preset joint similarity, the background noise to be detected is determined to be real; otherwise, the background noise to be detected is determined to be unreal.

2. The speech detection method according to claim 1, It is characterized in that The extracting the second feature from the human voice to be detected and the real target human voice respectively, and determining whether the voice to be detected is a real person voice by calculating the similarity between the second feature of the human voice to be detected and the second feature of the real target human voice comprises: Separating the real target human voice from the background noise, and selecting the longest real target human voice segment from the one or more separated real target human voice segments as the second comparison segment; Extracting second features from the second comparison segment, wherein the second features include volume mean, pitch period, speech rate, and spectrum distribution; Extract the volume mean, pitch period, speech rate and spectrum distribution of each human voice segment to be detected; Using a probability calculation model, respectively calculating the volume mean, pitch period, speech rate and spectral distribution of each human voice segment to be detected and the similarity between the volume mean, pitch period, speech rate and spectral distribution of the second comparison segment, to obtain the fifth similarity, sixth similarity, seventh similarity and eighth similarity of each human voice segment to be detected; Calculate the product of the fifth similarity, the sixth similarity, the seventh similarity and the eighth similarity of each to-be-detected vocal segment to obtain a second joint similarity of each to-be-detected vocal segment; If the second joint similarity of at least one to-be-detected human voice segment is greater than the second preset joint similarity, the to-be-detected voice is determined to be a real person's voice; otherwise, the to-be-detected voice is determined not to be a real person's voice.

3. The speech detection method according to claim 2, It is characterized in that After determining that the voice to be detected is a real person's voice, the method further comprises: Acquire a real third-person voice, separate the real third-person voice from background noise, and select the longest real third-person voice segment from one or more separated segments as a third comparison segment; Extracting the volume mean, pitch period, speech rate and spectrum distribution from the third comparison segment; Using a probability calculation model, respectively calculating the similarity between the volume mean, pitch period, speech rate and spectrum distribution of each human voice segment to be detected and the volume mean, pitch period, speech rate and spectrum distribution of the third comparison segment, to obtain the ninth similarity, the tenth similarity, the eleventh similarity and the twelfth similarity of each human voice segment to be detected; Calculate the product of the ninth similarity, the tenth similarity, the eleventh similarity and the twelfth similarity of each to-be-detected human voice segment to obtain a third joint similarity of each to-be-detected human voice segment; If the third joint similarities of all the segments of the human voice to be detected are not greater than the third preset joint similarity, it is determined that the voice to be detected is a real person's voice, but there is no third person's proof, otherwise the ratio between the average volume of the third comparison segment and the average volume of the second comparison segment is calculated as the first ratio; Calculating a ratio between an average volume of a third person's vocal segment in the vocal segment to be detected and an average volume of the target vocal segment as a second ratio; If the difference between the first ratio and the second ratio is less than the preset difference, the voice to be detected is determined to be a real person's voice and is certified by a third party; otherwise, the voice to be detected is determined to be a real person's voice but is not certified by a third party.

4. A speech detection device, It is characterized in that The speech detection device comprises: An acquisition module is used to acquire the speech to be detected, the real target human voice and the real background noise; A separation module is used to separate the human voice and background noise of the speech to be detected, so as to obtain the human voice to be detected and the background noise to be detected; A determination module, used to extract first features from the background noise to be detected and the real background noise respectively, and determine whether the background noise to be detected is real by calculating the similarity between the first feature of the background noise to be detected and the first feature of the real background noise; If the background noise to be detected is not real, it is determined that the voice to be detected is not a real person's voice; If the background noise to be detected is real, extract the second feature from the human voice to be detected and the real target human voice respectively, and determine whether the voice to be detected is a real person's voice by calculating the similarity between the second feature of the human voice to be detected and the second feature of the real target human voice; The separation module is used for: Divide the speech to be detected into a plurality of continuous speech frames of equal duration; Convert each speech frame into the frequency domain using Fourier transform, and divide each speech frame into multiple frequency intervals according to the frequency domain; According to the energy value of each frequency interval and the preset energy value of each frequency interval, a probability calculation model is used to determine whether each speech frame is a human voice; Merging the speech frames that are determined to be human voices and are continuous in time, and using one or more segments obtained by merging as human voices to be detected; Merging speech frames that are determined not to be human voices and are continuous in time, and using one or more merged segments as background noise to be detected; The determination module is used to: The longest segment in the background noise to be detected is used as the first comparison segment; Extracting first features from the first comparison segment, wherein the first features include standard deviation, peak value, single peak mean value and average energy; If the standard deviation of the first comparison segment is less than the preset standard deviation, the background noise to be detected is determined to be unreal, otherwise the standard deviation, peak value, single peak mean and average energy are extracted from the real background noise; Using a probability calculation model, respectively calculating the similarity between the standard deviation, peak value, single peak mean value and average energy of the real background noise and the standard deviation, peak value, single peak mean value and average energy of the first comparison segment, to obtain a first similarity, a second similarity, a third similarity and a fourth similarity; Calculate the product of the first similarity, the second similarity, the third similarity and the fourth similarity to obtain a first joint similarity; If the first joint similarity is greater than the first preset joint similarity, the background noise to be detected is determined to be real; otherwise, the background noise to be detected is determined to be unreal.

5. A voice detection device, It is characterized in that The speech detection device comprises a processor, a memory, and a speech detection program stored in the memory and executable by the processor, wherein when the speech detection program is executed by the processor, the steps of the speech detection method as described in any one of claims 1 to 3 are implemented.

6. A readable storage medium, It is characterized in that The readable storage medium stores a speech detection program, wherein when the speech detection program is executed by a processor, the steps of the speech detection method according to any one of claims 1 to 3 are implemented.

Citation Information

Patent Citations

  • Audio detection method and device and storage medium

    CN112509598A