Methods, apparatus, equipment, and media for decoupling voiceprints and content in synthetic audio detection.

By decoupling voiceprint and content features through deep neural networks, and combining signal separation and fully connected neural networks, the problem of insufficient recognition accuracy and robustness in synthetic audio detection in existing technologies is solved, and efficient synthetic audio detection is achieved.

CN116844552BActive Publication Date: 2026-04-03BEIJING TIMES RUILANG TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-12
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

In the detection of deep synthesized audio, conventional voiceprint identification methods are not very effective, while text-related methods have too narrow an application scenario and are difficult to effectively identify synthesized audio.

Method used

A deep neural network is used to decouple voiceprint and content features. Robust noise-resistant authenticity features are obtained through a signal separation neural network, and synthesized audio is judged by combining a fully connected neural network.

Benefits of technology

It improves the accuracy and robustness of synthetic audio recognition, completely decouples speaker identity information and text information, and enhances the accuracy and stability of detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116844552B_ABST
    Figure CN116844552B_ABST
Patent Text Reader

Abstract

This invention provides a method, apparatus, device, and medium for detecting synthesized audio by decoupling voiceprint and content, relating to the field of synthesized audio detection technology. The method includes steps S1 to S5: S1: Acquire the audio to be detected. S2: Extract voiceprint features from the audio using a deep neural network. S3: Extract content features from the audio using a content encoder. S4: Using the voiceprint and content features as noise references, obtain robust noise-resistant authenticity features stripped of voiceprint and content features through a signal separation neural network. S5: Determine whether the audio to be detected is synthesized using a fully connected neural network based on the robust noise-resistant authenticity features, and obtain the determination result. This invention's synthesized audio detection method completely decouples speaker identity information and text information from the audio, thereby performing deep synthesis detection on the remaining parts, significantly improving recognition accuracy and robustness.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of synthetic audio detection technology, and more specifically, to a method, apparatus, device, and medium for synthetic audio detection that decouples voiceprints and content. Background Technology

[0002] Voiceprints are sound wave spectra that carry speech information and are displayed using electroacoustic instruments. A person's voice can remain relatively stable over a long period after adulthood. However, each person's vocal habits differ, and each person's vocal organs are unique, thus voiceprints are specific. Deep audio synthesis technology can synthesize a voice that matches the target speaker's timbre and customized speech content information.

[0003] Deep synthesized audio detection involves analyzing audio to determine whether it is synthesized or genuinely recorded. Conventional methods of identifying deep synthesized audio based on voiceprints are largely ineffective; while text-based methods for deep synthesized audio detection have too many limitations and too narrow an application scenario.

[0004] In view of this, the applicant hereby submits this application after studying the existing technology. Summary of the Invention

[0005] The present invention provides a method, apparatus, device and medium for detecting synthetic audio by decoupling voiceprint and content, in order to improve at least one of the above-mentioned technical problems.

[0006] First aspect

[0007] This invention provides a method for detecting synthesized audio by decoupling voiceprints and content, comprising steps S1 to S5.

[0008] S1. Obtain the audio to be detected.

[0009] S2. Based on the audio to be detected, extract voiceprint features using a deep neural network.

[0010] S3. Extract content features from the audio to be detected using a content encoder.

[0011] S4. Based on the audio to be detected, using voiceprint features and content features as noise references, a robust noise-resistant authenticity feature is obtained by using a signal separation neural network to separate voiceprint features and content features.

[0012] S5. Based on the robust noise resistance authenticity features, determine whether the audio to be detected is synthetic audio through a fully connected neural network and obtain the judgment result.

[0013] The second aspect

[0014] This invention provides a synthetic audio detection device that decouples voiceprints and content, comprising:

[0015] The initial audio acquisition module is used to acquire the audio to be detected.

[0016] The voiceprint feature extraction module is used to extract voiceprint features from the audio to be detected using a deep neural network.

[0017] The content feature extraction module is used to extract content features from the audio to be detected using a content encoder.

[0018] The decoupling module is used to obtain robust noise-resistant authenticity features by separating voiceprint features and content features from the audio to be detected through a signal separation neural network, using voiceprint features and content features as noise references.

[0019] The discrimination module is used to determine whether the audio to be detected is synthetic audio based on the robust noise-resistant authenticity features through a fully connected neural network, and obtain the judgment result.

[0020] Third aspect

[0021] This invention provides a synthetic audio detection device that decouples voiceprints and content, comprising a processor, a memory, and a computer program stored in the memory. The computer program can be executed by the processor to implement the synthetic audio detection method that decouples voiceprints and content as described in any paragraph of the first aspect.

[0022] Fourth page

[0023] This invention provides a computer-readable storage medium. The computer-readable storage medium includes a stored computer program, wherein, when the computer program is executed, it controls the device containing the computer-readable storage medium to perform the synthetic audio detection method for decoupling voiceprints and content as described in any paragraph of the first aspect.

[0024] By adopting the above technical solution, the present invention can achieve the following technical effects:

[0025] The synthetic audio detection method of this invention completely decouples the speaker identity information (i.e., voiceprint) and text information (i.e., content) in the audio, thereby performing deep synthetic detection on the remaining part, which greatly improves the recognition accuracy and robustness. Attached Figure Description

[0026] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as a limitation on the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0027] Figure 1 This is a flowchart illustrating the synthetic audio detection method.

[0028] Figure 2 This is a network structure diagram of the synthetic audio detection method.

[0029] Figure 3 This is a network structure diagram for extracting voiceprint features using a deep neural network.

[0030] Figure 4 It is a network structure diagram for extracting content features through a content encoder.

[0031] Figure 5 This is a schematic diagram of the synthesized audio detection device. Implementation

[0032] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.

[0033] Example 1

[0034] Please see Figures 1 to 4 The first embodiment of the present invention provides a method for detecting synthesized audio by decoupling voiceprint and content, which can be executed by a synthesized audio detection device (hereinafter referred to as: synthesized audio detection device). In particular, it is executed by one or more processors in the synthesized audio detection device to implement steps S1 to S5.

[0035] S1. Obtain the audio to be detected.

[0036] Specifically, the audio to be detected needs to be preprocessed into a vector that can be recognized by the neural network. Preprocessing includes, but is not limited to, digitization of the speech signal, endpoint detection of the speech signal, pre-emphasis, windowing, and framing. Preprocessing the audio to be detected is a conventional technique for those skilled in the art, and will not be elaborated upon here.

[0037] It is understood that the synthesized audio detection device can be an electronic device with computing power, such as a portable laptop computer, desktop computer, server, smartphone, or tablet computer.

[0038] S2. Based on the audio to be detected, extract voiceprint features using a deep neural network.

[0039] Specifically, the voiceprint feature is a voiceprint embedding representation, that is, a sound wave spectrum that can express the characteristics of the sound source.

[0040] like Figure 3 As shown, in an optional embodiment of the present invention, based on the above embodiments, the deep neural network is an LSTM (Long Short-Term Memory) neural network. Step S2 specifically includes steps S21 and S22.

[0041] S21. Based on the audio to be detected, perform frame division and windowing according to the preset frame length and frame shift to obtain audio frames.

[0042] S22. Each audio frame is converted into an intermediate state representation using an LSTM (Long Short-Term Memory) neural network, and the intermediate state representation of each audio frame is input into the LSTM neural network of the next audio frame. The last LSTM neural network outputs the voiceprint features of the sound source corresponding to the audio to be detected.

[0043] In this embodiment, an LSTM (Long Short-Term Memory) neural network (i.e., a deep neural network) is used to extract speaker feature information (i.e., d-vector) from the input audio. The deep neural network can stack multiple layers of LSTM and can also use a neural network front-end obtained by pre-training for a speaker recognition task.

[0044] Preferably, the audio is preprocessed first, the audio to be detected W1 is divided into frames according to the set frame length and frame shift, windowed, and then the frames are passed through a Mel filter, and the logarithm is taken to obtain the vector representing this frame.

[0045] Specifically, a sliding window is constructed to select multiple frames with a fixed length. Then, an LSTM neural network is run on each window to obtain the intermediate state representation of that window. The intermediate state representation of each window is then fed into a second LSTM in sequence. The output of the last window is the d-vector representation of the entire audio segment.

[0046] In other embodiments, existing neural networks for voiceprint representation, such as pre-trained CPC, wav2vec, and wav2vec2.0, can be used instead of the LSTM long short-term memory neural network. Voiceprint features such as i-vector and x-vector can also be used instead of d-vector; this invention does not impose specific limitations on these methods.

[0047] S3. Extract content features from the audio to be detected using a content encoder.

[0048] Specifically, the content feature is a content embedding representation, that is, the natural language information contained in the audio.

[0049] like Figure 4 As shown, based on the above embodiments, in an optional embodiment of the present invention, the content encoder includes a residual block and a downsampling block. The downsampling block is a fully convolutional neural network. The fully convolutional neural network is a strided convolution.

[0050] Specifically, the content encoder consists of multiple residual blocks and downsampling blocks, used to extract content features from the audio to be detected. The content feature C has the same dimension as the speaker feature S. Preferably, the content encoder can be trained using the front end of an ASR model. The front-end encoder can be, for example, an LSTM or a transformer; this embodiment of the invention does not specifically limit its choice.

[0051] In an optional embodiment, the content encoder uses a fully convolutional neural network (WCNN) that can take audio of random length as input. The WCNN comprises four downsampling blocks, two convolutional layers, and a GELU activation function. Each downsampling block consists of four residual blocks. Each residual block consists of a dilated convolutional layer, a tanh activation function, and residual connections. The construction is similar to WaveNet.

[0052] S4. Based on the audio to be detected, using voiceprint features and content features as noise references, a robust noise-resistant authenticity feature is obtained by using a signal separation neural network to separate voiceprint features and content features.

[0053] Step S4 includes:

[0054] The voiceprint features and content features are fused using a parallel collaborative attention mechanism to obtain fused features.

[0055] Specifically, the extracted voiceprint features and content features are first encoded using an LSTM network to obtain corresponding codes s and c. Then, codes s and c are multiplied by a dot to obtain a correlation matrix T. Subsequently, the correlation matrix T is multiplied by codes s and c to obtain the corresponding attention matrix P. s With P c And multiply the code s by matrix P s Encoding c multiplied by matrix P c The results are then connected and flattened to obtain the fusion features of voiceprint features and content features.

[0056] The alternating collaborative attention mechanism is used to decouple the fused features from the audio to be detected, thereby obtaining robust noise-resistant real and fake features.

[0057] Specifically, firstly, the fused features and the audio to be detected are input into an LSTM neural network for encoding to obtain a voiceprint code sc that fuses the voiceprint and content; then, a correlation matrix A is constructed by dot product of the voiceprint code sc and the encoding w of the audio to be detected; finally, attention matrices A are obtained by dot product of the voiceprint code SC and the encoding W of the audio to be detected, respectively, based on the correlation matrix A. sc With A w Using attention matrix A w The correction vector matrix B is obtained by dot product of the voiceprint encoding sc. sc Then use B sc and A w Dot product yields the modified vector matrix B w Flattened matrix B w Obtain robust noise-resistant authenticity features by stripping away the voiceprint features and content features.

[0058] S5. Based on the robust noise reduction characteristics, determine whether the audio to be detected is a synthesized audio and obtain the judgment result.

[0059] Among these methods, it is possible to determine whether the audio is synthetic by using fully connected neural networks, convolutional neural networks with activation functions, using Siamese networks to train an auxiliary discriminator, or using triplet loss.

[0060] Taking a fully connected neural network as an example, a fully connected neural network is used to perform binary classification, and its output result is either real or synthetic.

[0061] Specifically, Neural Network 4 (i.e., a fully connected neural network) takes W2 as input and outputs a binary classification result (real or fake / synthetic). The formula for Neural Network 2 can be expressed as: .

[0062] like Figure 2 As shown, in this embodiment, neural network 1 (i.e., a deep neural network) is used to extract the voiceprint feature S from the input audio W1. Neural network 2 (i.e., a content encoder) is used to extract the content feature C from the audio W1. Then, neural network 3 (i.e., a signal separation neural network) removes S and C from W1, outputting audio W2. Subsequently, W2 is input into neural network 4 (i.e., a fully connected neural network) for deep synthesized audio detection, outputting the result Y.

[0063] In summary, in this embodiment of the invention, considering the coupling between voiceprint and content in audio, the voiceprint features and content features are extracted separately first, and then the voiceprint and content in the original audio to be detected are stripped off based on the extracted voiceprint and content features. This process removes these two parts as much as possible, so that the authenticity of the audio can be identified using pure forgery traces (i.e., robust noise-resistant authenticity features), thereby greatly improving the recognition accuracy and robustness. Furthermore, it supports adding external modules such as speech enhancement to further improve the robustness and generalization of the system.

[0064] Based on the above embodiments, in an optional embodiment of the present invention, during the training of the synthetic audio detection method, the fully connected neural network is trained using the cross-entropy loss function, and the deep neural network and content encoder are trained using adversarial training.

[0065] Specifically, during training, neural networks 1 and 2 are combined for adversarial training. The adversarial training requires the discriminator to be unable to distinguish W2's voiceprint and content information.

[0066] During adversarial training, robust noise-resistant real and fake features stripped of voiceprint and content features are input into a deep neural network. The deep neural network is then connected to a discriminator for the corresponding speaker. The training continues until the output of the discriminator for the corresponding speaker has a cosine similarity to the real speaker that is lower than a first threshold.

[0067] During adversarial training, robust noise-resistant real and fake features stripped of voiceprint and content features are input into the content encoder. The content encoder is followed by a discriminator for corresponding content recognition. Iteratively, the output of the discriminator for corresponding content recognition has a cosine similarity to the real content that is lower than the second threshold.

[0068] Preferably, the discriminator in adversarial training can use Hamming distance, Euclidean distance, etc., in addition to cosine similarity. This embodiment of the invention does not specifically limit this.

[0069] Example 2

[0070] Please see Figure 5 This invention provides a synthetic audio detection device that decouples voiceprints and content, comprising:

[0071] The initial audio acquisition module 100 is used to acquire the audio to be detected.

[0072] The voiceprint feature extraction module 200 is used to extract voiceprint features from the audio to be detected using a deep neural network.

[0073] The content feature extraction module 300 is used to extract content features from the audio to be detected by a content encoder.

[0074] The decoupling module 400 is used to obtain robust noise-resistant authenticity features by stripping voiceprint features and content features from the audio to be detected through a signal separation neural network, using voiceprint features and content features as noise references.

[0075] The discrimination module 500 is used to determine whether the audio to be detected is synthetic audio based on the robust noise-resistant authenticity features through a fully connected neural network, and obtain the judgment result.

[0076] Based on the above embodiments, in an optional embodiment of the present invention, the deep neural network is an LSTM long short-term memory neural network.

[0077] The voiceprint feature extraction module 200 specifically includes:

[0078] The audio frame acquisition unit is used to acquire audio frames by dividing the audio to be detected into frames and adding windows according to a preset frame length and frame shift.

[0079] The voiceprint feature extraction unit is used to convert each audio frame into an intermediate state representation using an LSTM (Long Short-Term Memory) neural network, and input the intermediate state representation of each audio frame into the LSTM neural network of the next audio frame. The last LSTM neural network outputs the voiceprint features of the sound source corresponding to the audio to be detected.

[0080] Example 3

[0081] This invention provides a synthetic audio detection device that decouples voiceprints and content, comprising a processor, a memory, and a computer program stored in the memory. The computer program can be executed by the processor to implement the synthetic audio detection method that decouples voiceprints and content as described in any paragraph of Embodiment 1.

[0082] Example 4

[0083] This invention provides a computer-readable storage medium. The computer-readable storage medium includes a stored computer program, wherein, when the computer program is executed, it controls the device containing the computer-readable storage medium to perform the synthetic audio detection method for decoupling voiceprints and content as described in any paragraph of Embodiment 1.

[0084] In the several embodiments provided in this invention, it should be understood that the disclosed apparatus and methods can also be implemented in other ways. The apparatus and method embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of apparatus, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than those marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram and / or flowchart, and combinations of blocks in block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0085] In addition, the functional modules in the various embodiments of the present invention can be integrated together to form an independent part, or each module can exist independently, or two or more modules can be integrated to form an independent part.

[0086] If the aforementioned functions are implemented as software functional modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, electronic device, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks. It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. In the absence of further restrictions, an element defined by the phrase "comprising a..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0087] The terminology used in the embodiments of this invention is for the purpose of describing particular embodiments only and is not intended to limit the invention. The singular forms “a,” “the,” and “the” as used in the embodiments of this invention and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise.

[0088] It should be understood that the term "and / or" used in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. Additionally, the character " / " in this article generally indicates that the preceding and following related objects have an "or" relationship.

[0089] Depending on the context, the word "if" as used here can be interpreted as "when," "when," "in response to determination," or "in response to detection." Similarly, depending on the context, the phrase "if determination" or "if detection (of the stated condition or event)" can be interpreted as "when determination," "in response to determination," "when detection (of the stated condition or event)," or "in response to detection (of the stated condition or event)."

[0090] The use of "first" and "second" in the embodiments is merely to distinguish similar objects and does not represent a specific ordering of objects. It is understood that "first" and "second" can be interchanged in a specific order or sequence where permitted. It should be understood that the objects distinguished by "first" and "second" can be interchanged where appropriate so that the embodiments described herein can be implemented in an order other than those illustrated or described herein.

[0091] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A method for detecting synthesized audio by decoupling voiceprint and content, characterized in that, Include: Obtain the audio to be detected; Based on the audio to be detected, voiceprint features are extracted using a deep neural network; Based on the audio to be detected, content features are extracted using a content encoder; Based on the audio to be detected, using the voiceprint features and content features as noise references, a robust noise-resistant authenticity feature is obtained by separating the voiceprint features and content features through a signal separation neural network; specifically: the voiceprint features and content features are fused using a parallel collaborative attention mechanism to obtain fused features; and the fused features are decoupled from the audio to be detected using an alternating collaborative attention mechanism to obtain robust noise-resistant authenticity features. Based on the robust noise resistance authenticity features, determine whether the audio to be detected is synthetic audio and obtain the determination result.

2. The method for synthesizing audio by decoupling voiceprint and content according to claim 1, characterized in that, The deep neural network is an LSTM (Long Short-Term Memory) neural network. The step of extracting voiceprint features from the audio to be detected using a deep neural network specifically includes: Based on the audio to be detected, the audio frames are obtained by segmenting and windowing according to a preset frame length and frame shift. Each audio frame is converted into an intermediate state representation using an LSTM (Long Short-Term Memory) neural network, and the intermediate state representation of each audio frame is input into the LSTM neural network of the next audio frame. The last LSTM neural network outputs the voiceprint features of the sound source corresponding to the audio to be detected.

3. The method for synthesizing audio by decoupling voiceprint and content according to claim 1, characterized in that, The content encoder includes residual blocks and downsampling blocks; wherein the downsampling blocks are fully convolutional neural networks; and the fully convolutional neural network is a strided convolution.

4. The method for detecting synthesized audio by decoupling voiceprint and content according to claim 1, characterized in that, The parallel collaborative attention mechanism is used to fuse voiceprint features and content features to obtain fused features, which specifically include: The extracted voiceprint features and content features are encoded using an LSTM network to obtain the corresponding codes s and c. Calculate the relevance matrix T based on codes s and c, and then use the relevance matrix T to calculate the attention matrix P corresponding to codes s and c. s With P c And multiply the code s by matrix P s Encoding c multiplied by matrix P c The results are then connected and flattened to obtain the fusion features of voiceprint features and content features.

5. The method for synthesizing audio by decoupling voiceprint and content according to claim 4, characterized in that, The alternating collaborative attention mechanism is used to decouple the fused features from the audio to be detected, resulting in robust noise-resistant features that can be verified by: The fusion features and the audio to be detected are respectively input into the LSTM neural network for encoding to obtain the voiceprint code sc that fuses the voiceprint and content; A correlation matrix A is constructed based on the voiceprint code sc and the code w of the audio to be detected; Calculate the attention matrix A for both the voiceprint code and the audio code to be detected based on the correlation matrix A. sc With A w ; Using attention matrix A w The correction vector matrix B is calculated using the voiceprint encoding sc. sc Then use B sc and A w Calculate the correction vector matrix B w Flattened matrix B w Obtain robust noise-resistant authenticity features by stripping away the voiceprint features and content features.

6. The method for synthesizing audio by decoupling voiceprint and content according to claim 4, characterized in that, The synthetic audio detection method uses adversarial training during training. During adversarial training, robust noise-resistant real and fake features stripped of the voiceprint features and content features are input into a deep neural network. The deep neural network is then connected to a discriminator for the corresponding speaker. The training continues until the output of the discriminator for the corresponding speaker has a cosine similarity to the real speaker that is lower than a first threshold. During adversarial training, robust noise-resistant real and fake features stripped of the voiceprint features and content features are input into the content encoder. The content encoder is followed by a discriminator for corresponding content recognition. Iteratively, the output of the discriminator for corresponding content recognition has a cosine similarity to the real content that is lower than the second threshold. During adversarial training, fully connected neural networks are trained using the cross-entropy loss function.

7. A synthesized audio detection device that decouples voiceprint and content, characterized in that, Include: The initial audio acquisition module is used to acquire the audio to be detected; The voiceprint feature extraction module is used to extract voiceprint features based on the audio to be detected using a deep neural network. The content feature extraction module is used to extract content features based on the audio to be detected through a content encoder. The decoupling module is used to obtain robust noise-resistant authenticity features by using the voiceprint features and content features as noise references, based on the audio to be detected, through a signal separation neural network; specifically, the decoupling module is used to: fuse the voiceprint features and content features using a parallel collaborative attention mechanism to obtain fused features; and decouple the fused features from the audio to be detected using an alternating collaborative attention mechanism to obtain robust noise-resistant authenticity features. The discrimination module is used to determine whether the audio to be detected is synthetic audio based on the robust noise-resistant authenticity features through a fully connected neural network, and obtain the judgment result.

8. A synthesized audio detection device that decouples voiceprint and content, characterized in that, It includes a processor, a memory, and a computer program stored in the memory; the computer program can be executed by the processor to implement the synthetic audio detection method for decoupling voiceprints and content as described in any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored computer program, wherein, when the computer program is executed, it controls the device on which the computer-readable storage medium is located to perform the synthetic audio detection method for decoupling voiceprints and content as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Counterfeit voice detection method and device, electronic equipment and storage medium

    CN115910104A