False voice detection method, false voice detection model acquisition method and related equipment

By building a false voice detection model that combines speaker information and using pre-trained models and speech encoders to update parameters, the problem of low accuracy of existing false voice detection solutions is solved and higher false voice detection accuracy is achieved.

CN116403603BActive Publication Date: 2025-09-05IFLYTEK CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310492726.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-28
Publication Date
2025-09-05
Estimated Expiration
2043-04-28

AI Technical Summary

Technical Problem

Existing false voice detection schemes have low accuracy and are difficult to effectively distinguish between real and false voices. In addition, existing schemes fail to fully utilize speaker information for improvement.

Method used

A false voice detection model is constructed, which includes a speaker representation module, a speech encoder, a false voice representation module, and a speech classification module. The speaker representation module is trained with the assistance of a pre-trained speech model, and the speech encoder is combined with parameter updates to introduce speaker prior information to improve detection accuracy.

Benefits of technology

By introducing speaker information and pre-training models, the accuracy and reliability of false voice detection are improved, and the ability to recognize false voices is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116403603B_ABST
    Figure CN116403603B_ABST
Patent Text Reader

Abstract

The present invention provides a false voice detection method, a false voice detection model acquisition method, and related equipment. The false voice detection method comprises: acquiring a target speech; and detecting whether the target speech is false voice based on a pre-acquired target false voice detection model. The target false voice detection model is trained using training speech labeled with speech categories to construct a false voice detection model. The constructed false voice detection model includes a speech encoder, a speaker characterization module that acquires a speaker characterization based on the output of the speech encoder, a false voice characterization module that acquires a false voice characterization based on the output of the speech encoder, and a speech classification module that performs speech classification based on the outputs of the speaker characterization module and the false voice characterization module. The speaker characterization module is acquired by combining a speaker classification task with the training of the speech encoder, and the speech encoder is a pre-trained speech model obtained through pre-training. The false voice detection method provided by the present invention can accurately detect whether a speech is false voice.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of speech detection technology, and in particular to a false voice detection method, a false voice detection model acquisition method, and related equipment. Background Art

[0002] Automatic Speaker Verification (ASV) is a technology that authenticates a speaker based on their voice. This technology authenticates a speaker based on their unique voice characteristics, providing a low-cost and flexible biometric solution.

[0003] While ASV systems are now reliable enough to support mass market adoption, concerns remain about their reliability. These concerns stem from their vulnerability to various voice spoofing attacks. These attacks involve the use of technologies like speech synthesis and voice conversion to generate false voices, which can then trick the ASV system into misidentifying the speaker.

[0004] It is understandable that if false voice detection can be performed before identity authentication based on the speaker's voice (i.e., detecting whether the voice is false voice), the security and reliability of speaker identity authentication will be greatly improved. However, how to detect false voice in voice is a problem that needs to be solved urgently. Summary of the Invention

[0005] In view of this, the present invention provides a false voice detection method, a false voice detection model acquisition method, and related devices for performing false voice detection on speech, thereby improving the security and reliability of speaker identity verification. The technical solution is as follows:

[0006] A false voice detection method, comprising:

[0007] Get the target voice;

[0008] Based on a pre-acquired target false voice detection model, detecting whether the target speech is false voice; wherein:

[0009] The target false voice detection model is trained using a false voice detection model constructed using training speech pairs labeled with speech categories, where the speech category is one of true voice and false voice;

[0010] The constructed false voice detection model includes: a speech encoder, a speaker characterization module for obtaining a speaker characterization based on an output of the speech encoder, a false voice characterization module for obtaining a false voice characterization based on an output of the speech encoder, and a speech classification module for performing speech classification based on outputs of the speaker characterization module and the false voice characterization module;

[0011] The speaker representation module is obtained by combining a speaker classification task with the training of the speech encoder, and the speech encoder is a speech pre-training model obtained through pre-training.

[0012] Optionally, detecting whether the target speech is a false voice based on a pre-acquired target false voice detection model includes:

[0013] Inputting the target speech into the speech encoder of the target false voice detection model for encoding to obtain speech features of the target speech;

[0014] Inputting the speech features of the target speech into the speaker characterization module and the false voice characterization module of the target false voice detection model respectively to obtain the speaker characterization and the false voice characterization corresponding to the target speech;

[0015] Inputting the speaker representation and the false voice representation corresponding to the target speech into the speech classification module of the target false voice detection model to obtain a speech classification result corresponding to the target speech;

[0016] Determine whether the target speech is a false voice based on a speech classification result corresponding to the target speech.

[0017] Optionally, in combination with the speaker classification task and assisted by the speech encoder, the process of training a speaker representation module includes:

[0018] Constructing a speaker classification model including the speech encoder, the speaker characterization module and the speaker classification module;

[0019] The speaker classification model is trained using training speech labeled with speaker category, wherein parameters of the speech encoder are fixed when training the speaker classification model.

[0020] Optionally, the using of training speech labeled with speaker category to train the speaker classification model includes:

[0021] Inputting the training speech into the speech encoder of the speaker classification model for encoding to obtain speech features of the training speech;

[0022] Inputting the speech features of the training speech into the speaker representation module of the speaker classification model to obtain the speaker representation corresponding to the training speech;

[0023] Inputting the speaker representation corresponding to the training speech into the speaker classification module of the speaker classification model to obtain the speaker classification result corresponding to the training speech;

[0024] According to the speaker classification result corresponding to the training speech and the speaker category annotated by the training speech, parameters of the speaker representation module and the speaker classification module in the speaker classification model are updated.

[0025] Optionally, the process of training the constructed false voice detection model using training speech labeled with speech categories includes:

[0026] The constructed false voice detection model is trained using training speech labeled with speech categories. During the first stage of training, parameters of a speech encoder and a speaker representation module in the false voice detection model are fixed.

[0027] The false voice detection model trained in the first stage is trained in the second stage using training speech labeled with speech categories. The false voice detection model trained in the second stage is used as the target false voice detection model. When the false voice detection model trained in the first stage is trained in the second stage, the parameters of each module in the false voice detection model are updated.

[0028] Optionally, the first phase of training the constructed false voice detection model using training speech labeled with speech categories includes:

[0029] Input the training speech into the speech encoder of the constructed false voice detection model to obtain the speech features of the training speech;

[0030] Inputting the speech features of the training speech into the false voice characterization module and the speaker characterization module of the constructed false voice detection model respectively, to obtain the false voice characterization and speaker characterization corresponding to the training speech;

[0031] The speech classification module of the false voice detection model constructed by inputting the false voice representation corresponding to the training speech and the speaker representation is used to classify the true voice and false voice to obtain the speech classification result corresponding to the training speech;

[0032] According to the speech classification results corresponding to the training speech and the speech categories annotated by the training speech, the parameters of the false voice characterization module and the speech classification module in the constructed false voice detection model are updated.

[0033] Optionally, the second stage training of the false voice detection model trained in the first stage using training speech labeled with speech categories includes:

[0034] Input the training speech into the speech encoder of the false voice detection model trained in the first stage to obtain the speech features of the training speech;

[0035] Inputting the speech features of the training speech into the false voice characterization module and speaker characterization module of the false voice detection model trained in the first stage, respectively, to obtain the false voice characterization and speaker characterization corresponding to the training speech;

[0036] Inputting the false voice representation and speaker representation corresponding to the training speech into the speech classification module of the false voice detection model trained in the first stage to classify the true voice and false voice, and obtaining the speech classification result corresponding to the training speech;

[0037] According to the speech classification results corresponding to the training speech and the speech categories annotated by the training speech, the parameters of all modules in the false voice detection model after the first stage of training are updated.

[0038] A method for obtaining a false voice detection model, comprising:

[0039] In combination with the speaker classification task, a speaker representation module is trained with the aid of a speech encoder, wherein the speech encoder is a speech pre-trained model obtained through pre-training, and the speaker representation module is used to obtain a speaker representation based on the output of the speech encoder;

[0040] Constructing a false voice detection model, wherein the false voice detection model includes the speech encoder, a trained speaker characterization module, a false voice characterization module that shares the output of the speech encoder with the trained speaker characterization module, and a speech classification module that performs speech classification based on the output of the speaker characterization module and the output of the false voice characterization module;

[0041] The constructed false voice detection model is trained using training voices labeled with voice categories to obtain a target false voice detection model, wherein the voice category is one of true voice and false voice.

[0042] Optionally, the step of training the constructed false voice detection model using training speech labeled with speech categories to obtain a target false voice detection model includes:

[0043] The constructed false voice detection model is trained using training speech labeled with speech categories. During the first stage of training, parameters of a speech encoder and a speaker representation module in the false voice detection model are fixed.

[0044] The false voice detection model trained in the first stage is trained in the second stage using training speech labeled with speech categories. The false voice detection model trained in the second stage is used as the target false voice detection model. When the false voice detection model trained in the first stage is trained in the second stage, the parameters of each module in the false voice detection model are updated.

[0045] A false voice detection device includes: a voice acquisition module and a false voice detection module;

[0046] The speech acquisition module is used to acquire the target speech;

[0047] The false voice detection module is configured to detect whether the target voice is false voice based on a pre-acquired target false voice detection model; wherein:

[0048] The target false voice detection model is trained using a false voice detection model constructed using training speech pairs labeled with speech categories, where the speech category is one of true voice and false voice;

[0049] The constructed false voice detection model includes: a speech encoder, a speaker characterization module for obtaining a speaker characterization based on an output of the speech encoder, a false voice characterization module for obtaining a false voice characterization based on an output of the speech encoder, and a speech classification module for performing speech classification based on outputs of the speaker characterization module and the false voice characterization module;

[0050] The speaker representation module is obtained by combining a speaker classification task with the training of the speech encoder, and the speech encoder is a speech pre-training model obtained through pre-training.

[0051] A false voice detection model acquisition device includes: a first training module, a model building module and a second training module;

[0052] The first training module is used to train the speaker characterization module in combination with the speaker classification task and supplemented by a speech encoder, wherein the speech encoder is a speech pre-training model obtained through pre-training, and the speaker characterization module is used to obtain a speaker characterization based on the output of the speech encoder;

[0053] The model construction module is configured to construct a false voice detection model, wherein the false voice detection model includes the speech encoder, a trained speaker characterization module, a false voice characterization module that shares the output of the speech encoder with the trained speaker characterization module, and a speech classification module that performs speech classification based on the output of the speaker characterization module and the output of the false voice characterization module;

[0054] The second training module is used to train the constructed false voice detection model using training voices marked with voice categories to obtain a target false voice detection model, wherein the voice category is one of true voice and false voice.

[0055] A false voice detection device, comprising: a memory and a processor;

[0056] The memory is used to store programs;

[0057] The processor is configured to execute the program to implement each step of any one of the above-mentioned false voice detection methods.

[0058] A computer-readable storage medium stores a computer program, wherein when the computer program is executed by a processor, the computer program implements each step of any one of the above-mentioned false voice detection methods.

[0059] The false voice detection method provided by the present invention can detect whether the target voice is false voice based on a pre-acquired target false voice detection model after obtaining the target voice. The target false voice detection model used in the false voice detection method provided by the present invention is trained by constructing a false voice detection model using training voice pairs marked with voice categories (true voice or false voice). It has the ability to detect whether the input voice is false voice. At the same time, in addition to including a false voice characterization module for obtaining false voice characterization, the target false voice detection model also includes a speaker characterization module that shares the output of the voice encoder with the false voice characterization module. Moreover, the speaker characterization module is obtained in advance by combining a speaker classification task and supplemented by voice encoder training. On the one hand, the speaker characterization module and the false voice characterization module share the output of the voice encoding module so that the output of the speaker characterization module is consistent with the output of the false voice characterization module. The output has a similar distribution, that is, the speaker representation output by the speaker representation module can be adapted to the false voice detection task. On the other hand, introducing the speaker representation on the basis of the false voice representation is equivalent to introducing the speaker prior information. The target false voice detection model combines the speaker prior information with the false voice representation to perform speech classification, and can obtain more accurate speech classification results. On the other hand, since the speech encoder is a pre-trained model, it is pre-trained with a large amount of data. Therefore, the speaker representation module obtained by combining the speaker classification task with the assistance of the speech encoder training can obtain a speaker representation with strong discriminability. Combining the false voice representation with the speaker representation with strong discriminability can improve the accuracy of false voice detection. In summary, the false voice detection method provided by the present invention can more accurately detect whether the speech is false voice. BRIEF DESCRIPTION OF THE DRAWINGS

[0060] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are merely embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying any creative work.

[0061] Figure 1 Schematic diagram of the structure of a false voice detection model;

[0062] Figure 2 A schematic diagram of the hardware architecture involved in the present invention;

[0063] Figure 3 A schematic diagram of a flow chart of a false voice detection method provided by an embodiment of the present invention;

[0064] Figure 4 A schematic diagram of the structure of a false voice detection model provided in an embodiment of the present invention;

[0065] Figure 5 A schematic diagram of a process for obtaining a target false voice detection model according to an embodiment of the present invention;

[0066] Figure 6 A schematic diagram of the structure of a speaker classification model provided by an embodiment of the present invention;

[0067] Figure 7 A schematic diagram of the structure of a speaker characterization module in a false voice detection model provided by an embodiment of the present invention;

[0068] Figure 8 A schematic diagram of the structure of a false voice characterization module in a false voice detection model provided in an embodiment of the present invention;

[0069] Figure 9 A schematic diagram of the structure of a false voice detection device provided by an embodiment of the present invention;

[0070] Figure 10 A schematic diagram of the structure of a device for acquiring a false voice detection model provided by an embodiment of the present invention;

[0071] Figure 11 A schematic diagram of the structure of a false voice detection device provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0072] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0073] In the process of implementing this case, the inventors of this case discovered that there are currently some false voice detection schemes. Most of these false voice detection schemes are: constructing a false voice detection model including a false voice characterization module and a speech classification module, and then using training speech labeled with speech categories to train the constructed false voice detection model to obtain a target false voice detection model, and then, based on the target false voice detection model, detecting true voice and false voice of the speech to be detected.

[0074] Although current false voice detection schemes can detect false voices, they are not very effective and have low accuracy. In view of this, the inventors of this case attempted to propose a false voice detection scheme with higher accuracy and conducted research for this purpose. During the research process, the inventors of this case thought that speaker information could be introduced into the false voice detection process to improve the false voice detection effect. Continuing along this line of thought, they came up with a false voice detection scheme that combines speaker information.

[0075] The false voice detection scheme combined with speaker information is as follows: Figure 1 As shown, a false voice detection model including a speaker representation module, a speech encoder, a false voice representation module and a speech classification module is constructed, and then the constructed false voice detection model is trained using training speech labeled with speech categories to obtain a target false voice detection model. Furthermore, based on the target false voice detection model, the true voice and false voice of the speech to be detected are detected.

[0076] Among them, the speaker representation module in the constructed false voice detection model is pre-trained by combining it with a speaker classification task, and the speech encoder adopts a speech pre-training model obtained through pre-training.

[0077] When training the constructed false voice detection model, MFCC features (Mel-Frequency Cepstral Coefficients) are extracted from the training speech to obtain the MFCC features corresponding to the training speech. The MFCC features corresponding to the training speech are then input into the speaker characterization module to obtain the speaker characterization corresponding to the training speech. Simultaneously, the training speech is input into a speech encoder for encoding to obtain the speech features corresponding to the training speech. The speech features corresponding to the training speech are then input into the false voice characterization module to obtain the false voice characterization corresponding to the training speech. Speech classification is then performed based on the speaker characterization and false voice characterization corresponding to the training speech to obtain a speech classification result corresponding to the training speech. Finally, the parameters of the false voice detection model are updated based on the speech classification result corresponding to the training speech and the annotated speech of the training speech. When updating the parameters of the false voice detection model, the parameters of the speaker characterization module and the speech encoder are fixed, and only the parameters of the false voice characterization module and the speech classification module are updated.

[0078] Compared with existing false voice detection solutions, the false voice detection solution combining speaker information introduces speaker information, and therefore improves the voice detection effect to a certain extent.

[0079] The inventors of this case studied the above-mentioned false voice detection scheme combined with speaker information and found that although the detection effect of the above-mentioned false voice detection scheme combined with speaker information has been improved, some problems still exist. For example, the input of the speaker representation acquisition part (i.e., the speaker representation module) is MFCC features, and the input features of the false voice representation acquisition part (speech encoder + false voice representation module) are WAV features (i.e., original speech). The difference in input features causes the output distribution of the speaker representation module and the false voice representation module to be dissimilar, that is, the speaker representation output by the speaker representation module is not well adapted to the false voice detection task. For example, the speaker representation output by the speaker representation module lacks sufficient distinguishability, etc. The existence of these many problems makes it difficult to significantly improve the false voice detection effect.

[0080] In view of the many problems with the above-mentioned false voice detection scheme combining speaker information, the inventors of this case continued their research and finally proposed a false voice detection method through continuous research. This false voice detection method can realize false voice detection and can perfectly overcome the defects of the above-mentioned false voice detection scheme combining speaker information.

[0081] Before introducing the false voice detection method provided by the present invention, the hardware architecture involved in the present invention is first described.

[0082] In one possible implementation, Figure 2 As shown, the hardware architecture involved in the present invention may include: an electronic device 201 and a server 202.

[0083] Exemplarily, the electronic device 201 may be any electronic product capable of performing human-computer interaction with a user, such as a PC, a laptop, a tablet computer, a PDA, a mobile phone, a learning machine, a smart TV, etc.

[0084] It should be noted that Figure 2 This is just an example. There are many types of electronic devices, not limited to Figure 2 Laptop in the.

[0085] For example, the server 202 may be a single server, a server cluster consisting of multiple servers, or a cloud computing server center. The server 202 may include a processor, a memory, a network interface, and the like.

[0086] Exemplarily, the electronic device 201 may establish a connection and communicate with the server 202 via a wireless communication network; exemplarily, the electronic device 201 may establish a connection and communicate with the server 202 via a wired communication network.

[0087] The electronic device 201 can obtain the voice to be detected and send the voice to be detected to the server 202. The server 202 performs false voice detection on the voice to be detected according to the false voice detection method provided by the present invention.

[0088] In another possible implementation, the hardware architecture involved in the present invention may include: an electronic device.

[0089] The electronic device is an electronic product with strong data processing capability. The electronic device can obtain the voice to be detected and perform false voice detection on the voice to be detected according to the false voice detection method provided by the present invention.

[0090] Those skilled in the art should understand that the above-mentioned electronic devices and servers are only examples, and other existing or future electronic devices or servers that are applicable to the present invention should also be included in the scope of protection of the present invention and are included here by reference.

[0091] Next, the false voice detection method provided by the present invention is introduced through the following embodiments.

[0092] See also Figure 3 , which shows a flow chart of a false voice detection method provided by an embodiment of the present invention. The false voice detection method may include:

[0093] Step S301: Acquire target speech.

[0094] The target speech is the speech to be detected.

[0095] Step S302: Based on the pre-acquired target false voice detection model, detect whether the target speech is false voice.

[0096] The target false voice detection model in this embodiment is trained by using a false voice detection model constructed using training speech pairs labeled with speech categories, wherein the speech category labeled with the training speech pairs is one of true voice and false voice.

[0097] like Figure 4 As shown, the constructed false voice detection model includes: a speech encoder 401 for encoding speech; a speaker characterization module 402 for obtaining a speaker characterization based on the output of speech encoder 401; a false voice characterization module 403 for obtaining a false voice characterization based on the output of speech encoder 401; and a speech classification module 404 for performing speech classification based on the outputs of speaker characterization module 402 and false voice characterization module 403. It should be noted that the speaker characterization is a feature that characterizes speaker information, and the false voice characterization is a feature that characterizes false voice information.

[0098] The speaker characterization module in the constructed false voice detection model is obtained by combining the speaker classification task with the training of the speech encoder 401. The speech encoder 401 is a speech pre-training model obtained through pre-training.

[0099] Optionally, the speech pre-training model can be a wavLM-Large model (300M pre-training model trained with 10wh). Of course, this embodiment is not limited to this. The speech pre-training model can also be other, such as wav2vec 2.0, Hubert model, wavLM model, wavlm Base+ model, etc.

[0100] Specifically, based on the pre-acquired target false voice detection model, the process of detecting whether the target speech is false voice may include:

[0101] Step S3021: Input the target speech into the speech encoder 401 of the target false voice detection model to obtain the speech features of the target speech.

[0102] The speech encoder 401 of the target false voice detection model encodes the input target speech and outputs the speech features of the target speech.

[0103] Step S3022: Input the speech features of the target speech into the speaker characterization module 402 and the false voice characterization module 403 of the target false voice detection model respectively to obtain the speaker characterization and false voice characterization corresponding to the target speech.

[0104] The speaker characterization module 402 and the false voice characterization module 403 of the target false voice detection model share the speech features output by the speech encoder 401. The speaker characterization module 402 obtains a speaker characterization corresponding to the target speech based on the speech features output by the speech encoder 401, and the false voice characterization module 403 obtains a false voice characterization corresponding to the target speech based on the speech features output by the speech encoder 401. The speaker characterization corresponding to the target speech is a feature that can characterize the speaker information of the target speech, and the false voice characterization corresponding to the target speech is a feature that can characterize the false voice information of the target speech.

[0105] Step S3023: Input the speaker representation and the false voice representation corresponding to the target speech into the speech classification module 404 of the target false voice detection model to obtain a speech classification result corresponding to the target speech.

[0106] After the speaker representation and falsetto representation corresponding to the target speech are input into the speech classification module 404 of the target falsetto detection model, the speech classification module 404 first fuses the input speaker representation with the input falsetto representation. Based on the fusion results, the module then predicts the probability of the target speech being falsetto and the probability of the target speech being true speech, thereby obtaining a predicted probability of the speech category corresponding to the target speech, i.e., the speech classification result for the target speech. The fusion of the speaker representation and the falsetto representation can be, but is not limited to, a concatenated fusion method, i.e., the speaker representation and the falsetto representation can be concatenated.

[0107] Step S3024: Determine whether the target speech is a false voice based on the speech classification result corresponding to the target speech.

[0108] Specifically, whether the target speech is false is determined based on the probability that the target speech is false and the probability that the target speech is true. In one possible implementation, it can be determined whether the probability that the target speech is false is greater than the probability that the target speech is true. If so, the target speech is determined to be false. In another possible implementation, it can be determined whether the probability that the target speech is false is greater than a preset probability threshold. If so, the target speech is determined to be false.

[0109] The false voice detection method provided by the embodiment of the present invention can detect whether the target voice is false voice based on a pre-acquired target false voice detection model after obtaining the target voice. The target false voice detection model used by the false voice detection method provided by the embodiment of the present invention is trained by a false voice detection model constructed by using training voice pairs marked with voice categories (true voice or false voice). It has the ability to detect whether the input voice is false voice. At the same time, in addition to including a false voice characterization module for obtaining false voice characterization, the target false voice detection model also includes a speaker characterization module that shares the output of the voice encoder with the false voice characterization module. Moreover, the speaker characterization module is obtained in advance by combining the speaker classification task and supplemented by voice encoder training. On the one hand, the speaker characterization module and the false voice characterization module share the output of the voice encoding module so that the output of the speaker characterization module and the false voice characterization module are consistent. The output of has a similar distribution, that is, the speaker representation output by the speaker representation module can be more suitable for the false voice detection task. On the other hand, introducing the speaker representation on the basis of the false voice representation is equivalent to introducing the speaker prior information. The target false voice detection model combines the speaker prior information on the basis of the false voice representation to perform speech classification, and can obtain more accurate speech classification results. On the other hand, since the speech encoder is a pre-trained model, it is pre-trained with a large amount of data. Therefore, combined with the speaker classification task, the speaker representation module obtained by the speech encoder training can obtain a speaker representation with strong discriminability. Combining the false voice representation with the speaker representation with strong discriminability can improve the accuracy of false voice detection. In summary, the false voice detection method provided by the embodiment of the present invention can obtain more accurate false voice detection results.

[0110] The false voice detection method provided in the above embodiment can detect whether the target speech is false voice based on a pre-acquired target false voice detection model. In another embodiment of the present invention, the process of obtaining the target false voice detection model is introduced.

[0111] See also Figure 5 , which shows a schematic flow chart of obtaining a target false voice detection model, which may include:

[0112] Step S501: combining the speaker classification task with the aid of a speech encoder, training the speaker representation module to obtain a trained speaker representation module.

[0113] The speech encoder is a speech pre-training model obtained through pre-training, and the speech pre-training model can be but is not limited to a wavLM-Large model.

[0114] The speaker representation module is used to obtain speaker representation according to the output of the speech encoder.

[0115] Specifically, in combination with the speaker classification task and supplemented by the speech encoder, the process of training the speaker representation module may include:

[0116] Step S5011: Construct a speaker classification model including a speech encoder, a speaker characterization module, and a speaker classification module.

[0117] See also Figure 6 , shows a structural diagram of a speaker classification model, which sequentially includes a speech encoder for encoding input speech, a speaker characterization module for obtaining speaker characterization according to the output of the speech encoder, and a speaker classification module for performing speaker classification according to the output of the speaker characterization module.

[0118] Step S5012: Use the training speech labeled with the speaker category to train the speaker classification model.

[0119] When training the speaker classification model, the parameters of the speech encoder are fixed.

[0120] Specifically, using training speech labeled with speaker categories, the process of training a speaker classification model may include:

[0121] Step a1: Acquire training speech from a first training speech set.

[0122] The first training speech set includes a plurality of training speech pieces labeled with speaker categories.

[0123] Step a2: input the training speech into the speech encoder of the speaker classification model to obtain the speech features of the training speech.

[0124] The speech encoder of the speaker classification model encodes the input speech and outputs speech features of the training speech. Because the speech encoder is a pre-trained speech model trained using a large amount of training data, the speech features output by the speech encoder contain rich speech information.

[0125] Step a3: Input the speech features of the training speech into the speaker representation module of the speaker classification model to obtain the speaker representation corresponding to the training speech.

[0126] Optional, such as Figure 7 As shown, the speaker representation module may include, but is not limited to, a first linear layer, an ECAPA-TDNN (Time Delay Neural Network), and a first pooling layer (such as a global average pooling layer).

[0127] The speech features of the training speech are first input into the first linear layer for dimensionality reduction. The features output by the first linear layer are input into ECAPA-TDNN. ECAPA-TDNN extracts and outputs the input features. The features output by ECAPA-TDNN are input into the first pooling layer. The first pooling layer calculates the mean and variance of the input features. The mean and variance are concatenated to obtain the speaker representation corresponding to the training speech.

[0128] It should be noted that ECAPA-TDNN generally has 1024 channels (which has a large number of parameters). To reduce the amount of computation, the speaker characterization module in this embodiment can adopt a 512-channel ECAPA-TDNN. Compared to the 1024-channel ECAPA-TDNN, the 512-channel ECAPA-TDNN has significantly fewer parameters and, accordingly, significantly reduced computation. Since the input to the speaker characterization module is the speech features output by the speech encoder, and the speech encoder is a pre-trained model trained using a large amount of training data, it can obtain relatively rich information from the input speech. Therefore, reducing the number of channels in the ECAPA-TDNN does not have a significant impact.

[0129] Step a4: input the speaker representation corresponding to the training speech into the speaker classification module of the speaker classification model to obtain the speaker classification result corresponding to the training speech.

[0130] Optionally, the speaker classification module may include a second linear layer and a softmax layer. The speaker representation corresponding to the training speech is sequentially processed by the second linear layer and the softmax layer to obtain the probability that the speaker category corresponding to the training speech is each set speaker category.

[0131] Optionally, the softmax layer in the speaker classification module may be, but is not limited to, an AAM-Softmax layer.

[0132] Step a5: update the parameters of the speaker representation module and the speaker classification module in the speaker classification model according to the speaker classification result corresponding to the training speech and the speaker category annotated by the training speech.

[0133] Specifically, the prediction loss of the speaker classification model is first determined based on the probability that the speaker category corresponding to the training speech is each predetermined speaker category and the speaker category annotated for the training speech. Then, based on the prediction loss of the speaker classification model, parameters of the speaker representation module and the speaker classification module in the speaker classification model are updated. Optionally, the prediction loss of the speaker classification model may be a cross-entropy loss. The calculation method for the cross-entropy loss is conventional and will not be detailed in this embodiment.

[0134] Repeat steps a1 to a5 until the training end conditions are met, such as model convergence or the preset number of training iterations is reached.

[0135] As mentioned above, the softmax layer in the speaker classification module can be an AAM-Softmax layer. Optionally, in the initial training phase, the distance constraint parameter margin of the AAM-Softmax layer can be set to a smaller value (for example, 0.02). After training stabilizes, this parameter can be increased (for example, to 0.2). This increases the difficulty for the speaker classification model to learn speakers, thereby improving the speaker classification model's ability to distinguish between speaker categories. It should be noted that a larger distance constraint parameter margin increases the difficulty of learning the speaker classification model.

[0136] Step S502: constructing a false voice detection model including a speech encoder, a speaker characterization module and a false voice characterization module that share the output of the speech encoder, and a speech classification module that performs speech classification according to the output of the speaker characterization module and the output of the false voice characterization module.

[0137] The speaker characterization module in the false voice detection model constructed in this step is the trained speaker characterization module obtained in step S501 , that is, the speaker characterization module in the speaker classification model trained using the training method of steps a1 to a5 .

[0138] The speaker characterization module and the false voice characterization module in the constructed false voice detection model share the output of the speech encoder. That is, the speaker characterization module and the false voice characterization module have the same input. If the speaker characterization module and the false voice characterization module have the same input, their outputs will have a similar distribution.

[0139] Step S503: using training speech labeled with speech categories to train the constructed false voice detection model to obtain a target false voice detection model.

[0140] The speech category of the training speech annotation is one of true sound and false sound.

[0141] Specifically, the process of training the constructed false voice detection model using training speech labeled with speech categories to obtain the target false voice detection model may include:

[0142] Step S5031: Use training speech labeled with speech categories to perform the first stage training on the constructed false voice detection model.

[0143] During the first phase of training of the constructed false voice detection model, the parameters of the speech encoder and speaker representation module in the false voice detection model are fixed.

[0144] Since the speaker characterization module is trained in step S501, step S5031 trains the false voice characterization module and the speech classification module with the aid of the speech encoder and the speaker characterization module trained in step S501.

[0145] Specifically, the process of performing the first phase of training on the constructed false voice detection model using training speech labeled with speech categories may include:

[0146] Step b1: Acquire training speech from the second training speech set.

[0147] The second training speech set includes a plurality of training speech pieces marked with speech categories (real voice / false voice).

[0148] Step b2: input the training speech into the speech encoder of the constructed false voice detection model to obtain the speech features of the training speech.

[0149] The speech encoder of the false voice detection model encodes the input training speech and outputs the speech features of the training speech.

[0150] Step b3: Input the speech features of the training speech into the false voice characterization module and the speaker characterization module of the constructed false voice detection model respectively to obtain the false voice characterization and the speaker characterization corresponding to the training speech.

[0151] The false voice characterization module and the speaker characterization module share the speech features output by the speech encoding module. The false voice characterization module obtains the false voice characterization corresponding to the training speech based on the input speech features (features representing the speaker information of the training speech). The speaker characterization module obtains the speaker characterization corresponding to the training speech (features representing the false voice information of the training speech) based on the input speech features.

[0152] The above content mentioned that the speaker representation module can include but is not limited to: a first linear layer, an ECAPA-TDNN, and a pooling layer (such as a global average pooling layer). The speech features of the training speech are first input into the first linear layer for dimensionality reduction processing. The features output by the first linear layer are input into the ECAPA-TDNN. ECAPA-TDNN extracts and outputs the input features. The features output by ECAPA-TDNN are input into the pooling layer. The pooling layer calculates the mean mean1 and variance var1 of the input features, and concatenates mean1 and variance var1 to obtain the speaker representation corresponding to the training speech.

[0153] Optional, such as Figure 8As shown, the pseudophonetic representation module may include a third linear layer, a BLSTM (Long Short-Term Memory) network, and a second pooling layer (e.g., a global average pooling layer). The speech features of the training speech are first input into the third linear layer for dimensionality reduction. The features output by the third linear layer are then input into the BLSTM for processing. The features output by the BLSTM are then input into the second pooling layer. The second pooling layer calculates the mean (mean2) and variance (var2) of the input features. These mean2 and variance (var2) are then concatenated to obtain the pseudophonetic representation corresponding to the training speech.

[0154] Step b4: Input the false voice representation and speaker representation corresponding to the training voice into the speech classification module of the constructed false voice detection model to classify the true voice and false voice, and obtain the speech classification result corresponding to the training voice.

[0155] After the false voice representation and speaker representation corresponding to the training speech are input into the speech classification module of the false voice detection model, the speech classification module first fuses (such as splicing) the false voice representation corresponding to the training speech with the speaker representation corresponding to the training speech, and then predicts the probability of the training speech being false voice and the probability of being true voice based on the fusion result, and obtains the speech category prediction probability corresponding to the training speech, that is, the speech classification result corresponding to the training speech.

[0156] Optionally, the speech classification module may include a feature fusion module, a fourth linear layer and a softmax layer. The false voice representation corresponding to the training speech and the speaker representation corresponding to the training speech are input into the feature fusion module for fusion, the features output by the feature fusion module are input into the fourth linear layer for mapping, the output of the fourth linear layer is input into the softmax layer, and the softmax layer outputs the probability that the training speech is a false voice and the probability that it is a true voice.

[0157] Step b5: update the parameters of the false voice characterization module and the voice classification module in the false voice detection model according to the voice classification result corresponding to the training voice and the voice category annotated by the training voice.

[0158] Specifically, the prediction loss of the false voice detection model is first determined based on the probability of the training speech being false voice and the probability of being true voice, as well as the speech category annotated with the training speech. Then, parameters of the false voice characterization module and speech classification module in the false voice detection model are updated based on the prediction loss of the false voice detection model. Optionally, the prediction loss of the false voice detection model can be a cross-entropy loss. The calculation method of the cross-entropy loss is conventional and will not be detailed in this embodiment.

[0159] Repeat steps b1 to a5 until the training end conditions are met, such as model convergence or the preset number of training iterations is reached.

[0160] Step S5032: Using the training speech labeled with the speech category, the false voice detection model trained in the first stage is trained in the second stage, and the false voice detection model trained in the second stage is used as the target false voice detection model.

[0161] When training the false voice detection model after the first stage of training, the parameters of all modules in the false voice detection model are updated.

[0162] Since step S501 trains the speaker characterization module separately, and step S5031 trains the false voice characterization module and the speech classification module separately (the separate training here means fixing the parameters of the speech encoder and the speaker characterization module, and only updating the parameters of the false voice characterization module and the speech classification module), in order to make the various modules in the false voice detection model more matched, step S5032 trains the entire model, that is, updates the parameters of all modules in the model.

[0163] Specifically, the process of performing the second-stage training on the false voice detection model trained in the first stage using training speech labeled with speech categories includes:

[0164] Step c1: Acquire training speech from the second training speech set.

[0165] Step c2: input the training speech into the speech encoder of the false voice detection model trained in the first stage to obtain the speech features of the training speech.

[0166] Step c3: Input the speech features of the training speech into the false voice characterization module and the speaker characterization module of the false voice detection model trained in the first stage, respectively, to obtain the false voice characterization and the speaker characterization corresponding to the training speech.

[0167] Step c4: input the false voice representation and speaker representation corresponding to the training speech into the speech classification module of the false voice detection model trained in the first stage to classify the true voice and false voice, and obtain the speech classification result corresponding to the training speech.

[0168] The specific implementation process and related explanations of steps c1 to c4 can be found in steps b1 to b4, which will not be described in detail in this embodiment.

[0169] Step c5: Update the parameters of all modules in the false voice detection model after the first stage training according to the speech classification results corresponding to the training speech and the speech categories annotated by the training speech.

[0170] Specifically, first, the prediction loss of the false voice detection model after the first stage of training is determined based on the speech classification results corresponding to the training speech and the speech category annotated by the training speech. Then, the parameters of the entire model are updated based on the prediction loss of the false voice detection model after the first stage of training, that is, the parameters of all modules in the false voice detection model after the first stage of training are updated.

[0171] It should be noted that the training process of the first stage training is the same as that of the second stage training. The difference is that the first stage training only updates the parameters of the false voice characterization module and the speech classification module in the false voice detection model (focusing on training the false voice characterization module and the speech classification module), while the second stage training updates the parameters of all modules in the false voice detection model to enable better matching of each module.

[0172] Furthermore, to better align the various modules of the false voice detection model while maintaining the training results of the first phase, the second phase of training can use a lower learning rate than the first phase. This prevents significant fluctuations in model parameters. This means that the model parameters are fine-tuned during the second phase of training to further improve model performance and meet practical application requirements. Furthermore, the second phase only requires a single iteration of training at a lower learning rate (one iteration of training involves feeding the model all the training speech in the second training speech set). Alternatively, multiple iterations of training at a lower learning rate can be performed.

[0173] Through the above process, a target false voice detection model with good performance can be obtained. False voice detection is performed on the speech to be detected based on the target false voice detection model, and a false voice detection result with high accuracy can be obtained.

[0174] An embodiment of the present invention further specifically provides a method for obtaining a false voice detection model, which is used to obtain a false voice detection model that can accurately detect whether speech is false voice. The process of obtaining the false voice detection model by the false voice detection model obtaining method is the same as the process of "obtaining a target false voice detection model" provided in the above embodiment, that is, first, in combination with the speaker classification task, supplemented by a speech encoder, the speaker representation module is trained, and then a false voice detection model is constructed, which includes a speech encoder, a speaker representation module that shares the output of the speech encoder, a false voice representation module, and a speech classification module that performs speech classification based on the output of the speaker representation module and the output of the false voice representation module. Finally, the constructed false voice detection model is trained using training speech labeled with speech categories to obtain the target false voice detection model. Among them, when using training speech labeled with speech categories to train the constructed false voice detection model, the training speech labeled with speech categories is first used to perform the first stage training on the constructed false voice detection model, and then the training speech labeled with speech categories is used to perform the second stage training on the false voice detection model after the first stage training. Among them, when performing the first stage training, the parameters of the speech encoder and speaker representation module in the false voice detection model are fixed, and only the parameters of the false voice representation module and the speech classification module are updated. When performing the second stage training, the parameters of all modules in the false voice detection model are updated.

[0175] For a more specific implementation process and related explanations of the method for obtaining a false voice detection model provided in an embodiment of the present invention, reference may be made to the specific implementation process and related explanations of “obtaining a target false voice detection model” in the above embodiment.

[0176] An embodiment of the present invention further provides a false voice detection device. The false voice detection device provided by the embodiment of the present invention is described below. The false voice detection device described below and the false voice detection method described above can be referenced to each other.

[0177] See also Figure 9 , shows a schematic structural diagram of a false voice detection device provided by an embodiment of the present invention, which may include: a voice acquisition module 901 and a false voice detection module 902.

[0178] The speech acquisition module 901 is used to acquire the target speech;

[0179] The false voice detection module 902 is configured to detect whether the target speech is false voice based on a pre-acquired target false voice detection model.

[0180] The target false voice detection model is trained using a false voice detection model constructed using training voice pairs labeled with voice categories, where the voice categories are either true voice or false voice. The constructed false voice detection model includes: a voice encoder, a speaker characterization module that obtains speaker characterization based on the output of the voice encoder, a false voice characterization module that obtains false voice characterization based on the output of the voice encoder, and a speech classification module that performs speech classification based on the outputs of the speaker characterization module and the false voice characterization module. The speaker characterization module is obtained by combining a speaker classification task with the assistance of the voice encoder training, and the voice encoder is a speech pre-training model obtained through pre-training.

[0181] Optionally, when detecting whether the target speech is a false voice based on a pre-acquired target false voice detection model, the false voice detection module 902 is specifically configured to:

[0182] The target speech is input into the speech encoder of the target false voice detection model for encoding to obtain the speech features of the target speech; the speech features of the target speech are respectively input into the speaker characterization module and the false voice characterization module of the target false voice detection model to obtain the speaker characterization and false voice characterization corresponding to the target speech; the speaker characterization and false voice characterization corresponding to the target speech are input into the speech classification module of the target false voice detection model to obtain the speech classification result corresponding to the target speech; and based on the speech classification result corresponding to the target speech, it is determined whether the target speech is false voice.

[0183] Optionally, the false voice detection device provided by the embodiment of the present invention may further include a false voice detection model acquisition module. The false voice detection model acquisition module includes: a first training module for training a speaker representation module in combination with the speaker classification task and the speech encoder.

[0184] The first training module is specifically used to train the speaker representation module in combination with the speaker classification task and assisted by the speech encoder:

[0185] A speaker classification model including the speech encoder, a speaker characterization module and a speaker classification module is constructed; and the speaker classification model is trained using training speech labeled with speaker categories, wherein parameters of the speech encoder are fixed when training the speaker classification model.

[0186] Optionally, when the first training module uses training speech labeled with speaker category to train the speaker classification model, it is specifically used to:

[0187] The training speech is input into the speech encoder of the speaker classification model for encoding to obtain speech features of the training speech; the speech features of the training speech are input into the speaker representation module of the speaker classification model to obtain a speaker representation corresponding to the training speech; the speaker representation corresponding to the training speech is input into the speaker classification module of the speaker classification model to obtain a speaker classification result corresponding to the training speech; and parameters of the speaker representation module and the speaker classification module in the speaker classification model are updated according to the speaker classification result corresponding to the training speech and the speaker category annotated on the training speech.

[0188] The false voice detection model acquisition module further includes: a model construction module for constructing a false voice detection model, and a second training module for training the constructed false voice detection model using training speech labeled with speech categories.

[0189] Optionally, when the second training module uses training speech labeled with speech categories to train the constructed false voice detection model, it is specifically used to:

[0190] The constructed false voice detection model is trained using training speech labeled with speech categories for the first stage of training, wherein the parameters of the speech encoder and speaker representation module in the false voice detection model are fixed during the first stage of training. The false voice detection model trained in the first stage is trained using training speech labeled with speech categories for the second stage of training, and the false voice detection model trained in the second stage is used as the target false voice detection model. wherein the parameters of each module in the false voice detection model are updated during the second stage of training of the false voice detection model trained in the first stage.

[0191] Optionally, when the second training module uses training speech labeled with speech categories to perform the first phase training on the constructed false voice detection model, it is specifically used to:

[0192] The training speech is input into the speech encoder of the constructed false voice detection model to obtain the speech features of the training speech; the speech features of the training speech are respectively input into the false voice characterization module and the speaker characterization module of the constructed false voice detection model to obtain the false voice characterization and speaker characterization corresponding to the training speech; the false voice characterization and the speaker characterization corresponding to the training speech are input into the speech classification module of the constructed false voice detection model to classify the true voice and the false voice to obtain the speech classification result corresponding to the training speech; according to the speech classification result corresponding to the training speech and the speech category annotated for the training speech, the parameters of the false voice characterization module and the speech classification module in the constructed false voice detection model are updated.

[0193] Optionally, when the second training module uses the training speech labeled with the speech category to perform the second stage training on the false voice detection model trained in the first stage, it is specifically used to:

[0194] The training speech is input into the speech encoder of the false voice detection model trained in the first stage to obtain the speech features of the training speech; the speech features of the training speech are respectively input into the false voice characterization module and the speaker characterization module of the false voice detection model trained in the first stage to obtain the false voice characterization and the speaker characterization corresponding to the training speech; the false voice characterization and the speaker characterization corresponding to the training speech are input into the speech classification module of the false voice detection model trained in the first stage to classify the true voice and the false voice, and obtain the speech classification result corresponding to the training speech; according to the speech classification result corresponding to the training speech and the speech category annotated for the training speech, the parameters of all modules in the false voice detection model trained in the first stage are updated.

[0195] The false voice detection device provided by the embodiment of the present invention can detect whether the target voice is false voice based on a pre-acquired target false voice detection model after obtaining the target voice. The target false voice detection model used by the false voice detection method provided by the embodiment of the present invention is trained by constructing a false voice detection model using training voice pairs marked with voice categories (true voice or false voice). It has the ability to detect whether the input voice is false voice. At the same time, in addition to the false voice characterization module, the target false voice detection model also includes a speaker characterization module that shares the output of the voice encoder with the false voice characterization module. Moreover, the speaker characterization module is obtained in advance by combining the speaker classification task and supplemented by the voice encoder training. On the one hand, the speaker characterization module and the false voice characterization module share the voice encoder output. The output of the speech encoding module ensures that the output of the speaker characterization module and the output of the false voice characterization module have a similar distribution, that is, the speaker characterization output by the speaker characterization module is more suitable for the false voice detection task. On the other hand, the target false voice detection model combines the false voice characterization with the speaker characterization for speech classification, and can obtain more accurate speech classification results. On the other hand, the speaker characterization module obtained by combining the speaker classification task with the assistance of speech encoder training can obtain a speaker characterization with strong discriminability. Combining the false voice characterization with the speaker characterization with strong discriminability can improve the accuracy of false voice detection. In summary, the false voice detection device provided by the embodiment of the present invention can obtain relatively accurate false voice detection results.

[0196] An embodiment of the present invention further provides a device for acquiring a false voice detection model. The device for acquiring a false voice detection model provided by an embodiment of the present invention is described below. The device for acquiring a false voice detection model described below and the method for acquiring a false voice detection model described above can refer to each other.

[0197] like Figure 10 As shown, the apparatus for acquiring a false voice detection model provided by the embodiment of the present invention may include: a first training module 1001 , a model building module 1002 and a second training module 1003 .

[0198] The first training module 1001 is used to train the speaker representation module in combination with the speaker classification task and with the aid of a speech encoder.

[0199] The speech encoder is a speech pre-training model obtained through pre-training, and the speaker representation module is used to obtain speaker representation according to the output of the speech encoder.

[0200] The model building module 1002 is used to build a false voice detection model.

[0201] Among them, the constructed false voice detection model includes a speech encoder, a trained speaker representation module, a false voice representation module that shares the output of the speech encoder with the trained speaker representation module, and a speech classification module that performs speech classification based on the output of the speaker representation module and the output of the false voice representation module.

[0202] The second training module 1003 is used to train the constructed false voice detection model using training speech labeled with speech categories to obtain a target false voice detection model.

[0203] The voice category is one of true voice and false voice.

[0204] Optionally, when the second training module 1003 uses training speech labeled with speech categories to train the constructed false voice detection model, it is specifically configured to:

[0205] The constructed false voice detection model is trained using training speech labeled with speech categories for the first stage of training, wherein the parameters of the speech encoder and speaker representation module in the false voice detection model are fixed during the first stage of training. The false voice detection model trained in the first stage is trained using training speech labeled with speech categories for the second stage of training, and the false voice detection model trained in the second stage is used as the target false voice detection model. wherein the parameters of each module in the false voice detection model are updated during the second stage of training of the false voice detection model trained in the first stage.

[0206] For a more detailed introduction of each module in this embodiment, please refer to the detailed introduction of each module in the false voice detection model acquisition module in the false voice detection device provided in the above embodiment, which will not be described in detail in this embodiment.

[0207] The false voice detection model acquisition device provided by the embodiment of the present invention can obtain a false voice detection model that can accurately detect whether a voice is false voice.

[0208] The embodiment of the present invention also provides a false voice detection device, see Figure 11, shows a schematic structural diagram of the false voice detection device, which may include: at least one processor 1101, at least one communication interface 1102, at least one memory 1103 and at least one communication bus 110.

[0209] In the embodiment of the present invention, there is at least one processor 1101 , communication interface 1102 , memory 1103 , and communication bus 1104 , and the processor 1101 , communication interface 1102 , and memory 1103 communicate with each other through the communication bus 1104 .

[0210] The processor 1101 may be a central processing unit (CPU), or an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present invention.

[0211] The memory 1103 may include a high-speed RAM memory, and may also include a non-volatile memory, such as at least one disk memory.

[0212] The memory stores a program, and the processor can call the program stored in the memory, wherein the program is used to:

[0213] Get the target voice;

[0214] Based on a pre-acquired target false voice detection model, detecting whether the target speech is false voice; wherein:

[0215] The target false voice detection model is trained using a false voice detection model constructed using training speech pairs labeled with speech categories, where the speech categories are either true voice or false voice. The constructed false voice detection model includes: a speech encoder, a speaker characterization module that obtains a speaker characterization based on the output of the speech encoder, a false voice characterization module that obtains a false voice characterization based on the output of the speech encoder, and a speech classification module that performs speech classification based on the outputs of the speaker characterization module and the false voice characterization module. The speaker characterization module is obtained by combining a speaker classification task with the assistance of training the speech encoder, and the speech encoder is a speech pre-training model obtained through pre-training.

[0216] Optionally, the detailed functions and extended functions of the program may refer to the above description.

[0217] An embodiment of the present invention further provides a computer-readable storage medium, which may store a program suitable for execution by a processor, wherein the program is used to:

[0218] Get the target voice;

[0219] Based on a pre-acquired target false voice detection model, detecting whether the target speech is false voice; wherein:

[0220] The target false voice detection model is trained using a false voice detection model constructed using training speech pairs labeled with speech categories, where the speech categories are either true voice or false voice. The constructed false voice detection model includes: a speech encoder, a speaker characterization module that obtains a speaker characterization based on the output of the speech encoder, a false voice characterization module that obtains a false voice characterization based on the output of the speech encoder, and a speech classification module that performs speech classification based on the outputs of the speaker characterization module and the false voice characterization module. The speaker characterization module is obtained by combining a speaker classification task with the assistance of training the speech encoder, and the speech encoder is a speech pre-training model obtained through pre-training.

[0221] Optionally, the detailed functions and extended functions of the program may refer to the above description.

[0222] Finally, it should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article, or device comprising the element.

[0223] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.

[0224] The above description of the disclosed embodiments is intended to enable one skilled in the art to implement or use the present invention. Various modifications to these embodiments will be readily apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention is not limited to the embodiments shown herein but is intended to conform to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A false voice detection method, characterized in that: include: Get the target voice; Based on a pre-acquired target false voice detection model, detecting whether the target speech is false voice; wherein: The target false voice detection model is trained using training speech pairs labeled with speech categories to construct a false voice detection model. The target false voice detection model classifies speech based on false voice representation and combined with prior information about the speaker. The speech category is one of true voice and false voice. The constructed false voice detection model includes: a speech encoder, a speaker characterization module for obtaining a speaker characterization based on the output of the speech encoder, a false voice characterization module for obtaining a false voice characterization based on the output of the speech encoder, and a speech classification module for fusing the speaker characterization output by the speaker characterization module and the false voice characterization output by the false voice characterization module and performing speech classification based on the fusion result; The speaker representation module is obtained by combining a speaker classification task with the training of the speech encoder, and the speech encoder is a speech pre-training model obtained through pre-training.

2. The false voice detection method according to claim 1, wherein: The detecting whether the target speech is a false voice based on a pre-acquired target false voice detection model includes: Inputting the target speech into the speech encoder of the target false voice detection model for encoding to obtain speech features of the target speech; Inputting the speech features of the target speech into the speaker characterization module and the false voice characterization module of the target false voice detection model respectively to obtain the speaker characterization and the false voice characterization corresponding to the target speech; Inputting the speaker representation and the false voice representation corresponding to the target speech into the speech classification module of the target false voice detection model to obtain a speech classification result corresponding to the target speech; Determine whether the target speech is a false voice based on a speech classification result corresponding to the target speech.

3. The false voice detection method according to claim 1, wherein: In combination with the speaker classification task and assisted by the speech encoder, the process of training the speaker representation module includes: Constructing a speaker classification model including the speech encoder, the speaker characterization module and the speaker classification module; The speaker classification model is trained using training speech labeled with speaker category, wherein parameters of the speech encoder are fixed when training the speaker classification model.

4. The false voice detection method according to claim 3, wherein: The method of training the speaker classification model using training speech labeled with speaker category includes: Inputting the training speech into the speech encoder of the speaker classification model for encoding to obtain speech features of the training speech; Inputting the speech features of the training speech into the speaker representation module of the speaker classification model to obtain the speaker representation corresponding to the training speech; Inputting the speaker representation corresponding to the training speech into the speaker classification module of the speaker classification model to obtain the speaker classification result corresponding to the training speech; According to the speaker classification result corresponding to the training speech and the speaker category annotated by the training speech, parameters of the speaker representation module and the speaker classification module in the speaker classification model are updated.

5. The false voice detection method according to claim 1, wherein: The process of training the constructed false voice detection model using training speech labeled with speech categories includes: The constructed false voice detection model is trained using training speech labeled with speech categories. During the first stage of training, parameters of a speech encoder and a speaker representation module in the false voice detection model are fixed. The false voice detection model trained in the first stage is trained in the second stage using training speech labeled with speech categories. The false voice detection model trained in the second stage is used as the target false voice detection model. When the false voice detection model trained in the first stage is trained in the second stage, the parameters of each module in the false voice detection model are updated.

6. The false voice detection method according to claim 5, characterized in that: The first stage of training the constructed false voice detection model using training speech labeled with speech categories includes: Input the training speech into the speech encoder of the constructed false voice detection model to obtain the speech features of the training speech; Inputting the speech features of the training speech into the false voice characterization module and the speaker characterization module of the constructed false voice detection model respectively, to obtain the false voice characterization and speaker characterization corresponding to the training speech; The speech classification module of the false voice detection model constructed by inputting the false voice representation corresponding to the training speech and the speaker representation is used to classify the true voice and false voice to obtain the speech classification result corresponding to the training speech; According to the speech classification results corresponding to the training speech and the speech categories annotated by the training speech, the parameters of the false voice characterization module and the speech classification module in the constructed false voice detection model are updated.

7. The false voice detection method according to claim 5, wherein: The second stage of training the false voice detection model trained in the first stage using training speech labeled with speech categories includes: Input the training speech into the speech encoder of the false voice detection model trained in the first stage to obtain the speech features of the training speech; Inputting the speech features of the training speech into the false voice characterization module and speaker characterization module of the false voice detection model trained in the first stage, respectively, to obtain the false voice characterization and speaker characterization corresponding to the training speech; Inputting the false voice representation and speaker representation corresponding to the training speech into the speech classification module of the false voice detection model trained in the first stage to classify the true voice and false voice, and obtaining the speech classification result corresponding to the training speech; According to the speech classification results corresponding to the training speech and the speech categories annotated by the training speech, the parameters of all modules in the false voice detection model after the first stage of training are updated.

8. A method for acquiring a false voice detection model, characterized in that: include: In combination with the speaker classification task, a speaker representation module is trained with the aid of a speech encoder, wherein the speech encoder is a speech pre-trained model obtained through pre-training, and the speaker representation module is used to obtain a speaker representation based on the output of the speech encoder; Constructing a false voice detection model, wherein the false voice detection model includes the speech encoder, a trained speaker representation module, a false voice representation module that shares the output of the speech encoder with the trained speaker representation module, and a speech classification module that fuses the speaker representation output by the speaker representation module and the false voice representation output by the false voice representation module and performs speech classification based on the fusion result; The constructed falsetto detection model is trained using training speech labeled with speech categories to obtain a target falsetto detection model. The target falsetto detection model performs speech classification based on falsetto representation and speaker prior information, wherein the speech category is one of true voice and false voice.

9. The method for acquiring a false voice detection model according to claim 8, wherein: The method of training the constructed false voice detection model using training speech labeled with speech categories to obtain a target false voice detection model includes: The constructed false voice detection model is trained using training speech labeled with speech categories. During the first stage of training, parameters of a speech encoder and a speaker representation module in the false voice detection model are fixed. The false voice detection model trained in the first stage is trained in the second stage using training speech labeled with speech categories. The false voice detection model trained in the second stage is used as the target false voice detection model. When the false voice detection model trained in the first stage is trained in the second stage, the parameters of each module in the false voice detection model are updated.

10. A false voice detection device, characterized in that: include: Voice acquisition module and false voice detection module; The speech acquisition module is used to acquire the target speech; The false voice detection module is configured to detect whether the target voice is a false voice based on a pre-acquired target false voice detection model; wherein: The target false voice detection model is trained using training speech pairs labeled with speech categories to construct a false voice detection model. The target false voice detection model classifies speech based on false voice representation and combined with prior information about the speaker. The speech category is one of true voice and false voice. The constructed false voice detection model includes: a speech encoder, a speaker characterization module for obtaining a speaker characterization based on the output of the speech encoder, a false voice characterization module for obtaining a false voice characterization based on the output of the speech encoder, and a speech classification module for fusing the speaker characterization output by the speaker characterization module and the false voice characterization output by the false voice characterization module and performing speech classification based on the fusion result; The speaker representation module is obtained by combining a speaker classification task with the training of the speech encoder, and the speech encoder is a speech pre-training model obtained through pre-training.

11. A device for acquiring a false voice detection model, characterized in that: include: a first training module, a model building module, and a second training module; The first training module is used to train the speaker characterization module in combination with the speaker classification task and supplemented by a speech encoder, wherein the speech encoder is a speech pre-training model obtained through pre-training, and the speaker characterization module is used to obtain a speaker characterization based on the output of the speech encoder; The model construction module is configured to construct a false voice detection model, wherein the false voice detection model includes the speech encoder, a trained speaker representation module, a false voice representation module that shares the output of the speech encoder with the trained speaker representation module, and a speech classification module that fuses the speaker representation output by the speaker representation module and the false voice representation output by the false voice representation module and performs speech classification based on the fusion result; The second training module is used to train the constructed false voice detection model using training speech labeled with speech categories to obtain a target false voice detection model. The target false voice detection model performs speech classification based on false voice representation and combined with prior information of the speaker, wherein the speech category is one of true voice and false voice.

12. A false voice detection device, characterized in that: include: memory and processor; The memory is used to store programs; The processor is configured to execute the program to implement the steps of the false voice detection method according to any one of claims 1 to 7.

13. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, each step of the false voice detection method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Voice authenticity verification method and device, electronic equipment and readable storage medium

    CN112992126A

  • Speaker representation vector distribution space creation method, speaker representation vector distribution space speech synthesis method and related equipment

    CN115762467A