Training method for voice anti-counterfeiting detection, program product and computing equipment
By using hyperbolic spatial feature prototype training method in speech anti-counterfeiting detection, the problem of information loss in high-dimensional feature processing is solved, and higher detection accuracy and generalization ability are achieved.
Patent Information
- Application Number
- CN202510570104.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-30
- Publication Date
- 2025-08-08
AI Technical Summary
In the voice anti-counterfeiting detection, conventional dimensionality reduction methods may lose potential feature information of speech data during high-dimensional feature processing, affecting detection accuracy.
The feature prototype training method in hyperbolic space is adopted, and the speech features are mapped to the hyperbolic space through the feature extractor. The distance is calculated in the hyperbolic space by using the classifier and feature prototype, and the parameters of the classifier and feature prototype are updated to ensure the integrity and distinction of feature information.
It improves the accuracy and generalization ability of voice anti-counterfeiting detection, can better recognize the authenticity of voice, and adapt to complex high-dimensional voice data.
Smart Images

Figure CN120452456A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of this specification belong to the field of data processing technology, and more particularly, relate to a training method, program product, and computing device for voice anti-counterfeiting detection. Background Art
[0002] Generative AI is now widely used in various fields. However, this also raises security concerns. For example, using generative AI for speech synthesis can lead to the spread of false information, privacy breaches, and the destruction of social trust.
[0003] However, speech signals contain rich frequency and time domain information. Due to the diversity of the sound sources, they also possess high-dimensional and complex domain features. Therefore, for anti-counterfeiting detection, speech data is typically characterized as high-dimensional features. However, during further processing of these high-dimensional features, conventional dimensionality reduction methods can lose the underlying feature information in the speech data, thereby compromising the accuracy of anti-counterfeiting detection.
[0004] Therefore, the present invention provides a training method, program product and computing device for voice anti-counterfeiting detection. Summary of the Invention
[0005] The present invention aims to provide a training method, program product, and computing device for voice anti-counterfeiting detection, comprising:
[0006] In a first aspect, this specification provides a training method for voice anti-counterfeiting detection, the method comprising:
[0007] Determining, using a feature extractor, a first speech feature corresponding to the first sample speech;
[0008] Mapping the first speech feature to a hyperbolic space using a feature conversion module to obtain a first hyperbolic feature corresponding to the first speech feature;
[0009] Determining a first feature prototype matching the first sample speech from a plurality of feature prototypes, and determining a first prototype loss based on a first distance of the first hyperbolic feature from the first feature prototype and a second distance from the plurality of feature prototypes in a hyperbolic space;
[0010] Processing the third distance between the first hyperbolic feature and the plurality of feature prototypes using a classifier to obtain a predicted classification of the first sample speech; determining a classification loss based on the predicted classification and a first classification label indicating speech authenticity corresponding to the first sample speech;
[0011] The parameters of the classifier and the plurality of feature prototypes are updated according to a training loss, wherein the training loss includes the first prototype loss and the classification loss.
[0012] In some implementations, determining a first feature prototype matching the first sample speech from a plurality of feature prototypes specifically includes:
[0013] Determining, based on the first classification label corresponding to the first sample speech, a plurality of target prototypes from a plurality of feature prototypes, wherein the second classification labels corresponding to the plurality of target prototypes are the same as the first classification label corresponding to the first sample speech;
[0014] The feature prototype closest to the first hyperbolic feature among the several target prototypes is determined as the first feature prototype.
[0015] In some implementations, before updating the parameters of the classifier and the plurality of feature prototypes, the method further includes:
[0016] Determining, using the feature extractor, a second speech feature corresponding to a second sample speech, where the second sample speech is the enhanced speech corresponding to the first sample speech;
[0017] Mapping the second speech feature to a hyperbolic space using the feature conversion module to obtain a second hyperbolic feature corresponding to the second speech feature;
[0018] determining a second prototype loss based on a difference between the second hyperbolic feature and the first feature prototype;
[0019] The training loss also includes: the second prototype loss.
[0020] In some implementations, determining a second prototype loss based on a difference between the second hyperbolic feature and the first feature prototype specifically includes:
[0021] Determine a fourth distance between the second hyperbolic feature and the first feature prototype; and determine a second prototype loss according to a difference between the fourth distance and the first distance and a fifth distance between the second hyperbolic feature and the first hyperbolic feature.
[0022] In some implementations, wherein the training loss further includes domain feature loss, the method further includes:
[0023] Acquire a plurality of first hyperbolic features determined based on a plurality of first speech samples, and a plurality of second hyperbolic features determined based on a plurality of second speech samples corresponding to the plurality of first speech samples, where each second speech sample is an enhanced speech corresponding to each first speech sample, and each first hyperbolic feature or second hyperbolic feature is composed of a plurality of feature components;
[0024] determining a plurality of domain-sensitive feature components in each feature component according to differences between the plurality of first hyperbolic features and the plurality of second hyperbolic features;
[0025] determining a domain feature loss according to the plurality of domain-sensitive feature components;
[0026] Updating the parameters of the classifier and the plurality of feature prototypes according to the training loss specifically includes:
[0027] According to the training loss, the parameters of the classifier, the multiple feature prototypes and the feature extractor are updated.
[0028] In some implementations, the method further includes:
[0029] Determine a ternary prototype group from the plurality of feature prototypes, the ternary prototype group comprising an anchor prototype, neighboring prototypes within a neighborhood of the anchor prototype, and random prototypes outside the neighborhood of the anchor prototype;
[0030] Determine, among multiple superordinate prototypes, a first superordinate prototype that matches the anchor prototype according to the anchor prototype and the neighboring prototypes;
[0031] Determine, among the multiple superordinate prototypes, a second superordinate prototype that is unrelated to the anchor prototype based on the first superordinate prototype and the random prototype;
[0032] determining a hierarchy loss based on a sixth distance between the ternary prototype group and the first superordinate prototype and a seventh distance between the ternary prototype group and the second superordinate prototype;
[0033] The training loss also includes: layer loss.
[0034] In some implementations, the hyperbolic space is a Poincare sphere model.
[0035] In some implementations, the method further includes:
[0036] Get the voice to be tested;
[0037] Determining the features to be measured corresponding to the speech to be measured using the feature extractor;
[0038] Mapping the feature to be measured to a hyperbolic space using the feature conversion module to determine a hyperbolic feature to be measured corresponding to the feature to be measured;
[0039] The classifier is used to process the eighth distance between the hyperbolic feature to be tested and the multiple feature prototypes to obtain a detection result of the speech to be tested.
[0040] A second aspect of this specification provides a computer program product, comprising a computer program / instruction, which implements the steps of the method described in the first aspect when executed by a processor.
[0041] A third aspect of this specification provides a computing device, including a memory and a processor, wherein the memory stores executable code, and when the processor executes the executable code, the method described in the first aspect is implemented.
[0042] The present embodiment provides a training scheme for voice anti-counterfeiting detection. In view of the high data dimension and complex sample source domain of voice data, a training method using feature prototypes in hyperbolic space is adaptively determined. This training scheme can obtain classifiers and feature prototypes with strong generalization capabilities and strong fitting capabilities for high-dimensional data. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] In order to more clearly illustrate the technical solutions of the embodiments of this specification, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments recorded in this specification. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0044] Figure 1 This is a flow chart of a training method for voice anti-counterfeiting detection in an embodiment of this specification;
[0045] Figure 2 Schematic diagram of a process for determining a mask matrix in an embodiment of this specification;
[0046] Figure 3 This is a schematic diagram of a process for determining a first superordinate prototype and a second superordinate prototype in an embodiment of this specification;
[0047] Figure 4 This is a schematic diagram of the structure of the anti-counterfeiting detection model in one embodiment of this specification;
[0048] Figure 5 This is a flowchart of a method for voice anti-counterfeiting detection in one embodiment of this specification. DETAILED DESCRIPTION
[0049] To help those skilled in the art better understand the technical solutions in this specification, the following will provide a clear and complete description of the technical solutions in the embodiments of this specification, in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of this specification, not all of them. All other embodiments derived by those skilled in the art based on the embodiments in this specification without creative effort shall fall within the scope of protection of this specification.
[0050] The source domain is the sample distribution to which the sample individuals belong during training, containing the known characteristics of the sample population. The target domain is the target distribution to which the target individuals belong during prediction. Because the actual characteristics of the target population in the target domain cannot be determined before prediction, researchers typically hope to improve the model's generalization capabilities by training it with data from multiple source domains, thereby increasing the prediction accuracy of data from the target domain.
[0051] Traditional methods for voice anti-counterfeiting detection usually directly map samples from different source domains into a compact feature space. Furthermore, the feature distributions of different source domains are aligned in the feature space to improve the generalization ability of the model.
[0052] Specifically, traditional methods typically perform several rounds of dimensionality reduction on the feature vectors of individual samples during forward reasoning until an output result is obtained. This reasoning process ignores the inherent hierarchical structure of feature information, which can lead to feature information loss, especially when dealing with data types with complex hierarchical relationships. For example, in voice anti-counterfeiting tasks, different attack forms and domain information have a hierarchical structure. Direct alignment may ignore these hierarchical relationships, thereby affecting the model's generalization ability.
[0053] Figure 1 A flowchart of a training method for voice anti-counterfeiting detection in an embodiment of this specification is shown. The training method for voice anti-counterfeiting detection can be executed using a computing device with computing capabilities. For ease of description, the computing device is described below as the execution subject. The method includes:
[0054] Step S101: using a feature extractor to determine a first speech feature corresponding to a first sample speech.
[0055] First, the computing device may input a first sample speech into a feature extractor, and determine a first speech feature corresponding to the first sample speech output by the feature extractor.
[0056] The first sample speech can be obtained from a corresponding sample dataset based on the requirements of the voice anti-counterfeiting detection task. Specifically, the first sample speech can be a real, non-artificially generated speech, and thus the first classification label corresponding to the first sample speech can be the detection result - "real." The first sample speech can also be a artificially generated forged speech, and thus the first classification label corresponding to the first sample speech can be the detection result - "forged." In other words, the first classification label can represent the correct detection result of the anti-counterfeiting detection of the first sample speech.
[0057] The feature extractor can convert the original audio signal into an embedding vector in a Euclidean space. The structure of the feature extractor can refer to a convolutional neural network, for example. This specification does not limit the specific structure of the feature extractor.
[0058] In some implementations, the feature extractor may be determined based on a pre-trained coding layer, such as a coding layer of a wav2vec model, an AASIST (Audio Anti-Spoofing using Integrated Spectro-TemporalGraph Attention Networks) model, a WaveNet model, or the like. Furthermore, the feature extractor may be derived by combining several of the aforementioned coding layers.
[0059] Thus, the feature extractor is capable of extracting feature information from the original speech data. Furthermore, because the feature extractor is determined solely based on the encoder, the first speech features output by the feature extractor based on the first sample speech are not subjected to dimensionality reduction and can relatively fully include the feature information in the first sample speech.
[0060] Step S103: using a feature conversion module to map the first speech feature to a hyperbolic space to obtain a first hyperbolic feature corresponding to the first speech feature.
[0061] After determining the first speech feature, the computing device can map the first speech feature to the hyperbolic space according to the preset hyperbolic space model using a mapping function from the Euclidean space to the preset hyperbolic space model to obtain a first hyperbolic feature corresponding to the first speech feature.
[0062] Due to the topological properties of hyperbolic space - hyperbolic space does not have a finite boundary, after projecting the eigenvectors in Euclidean space to hyperbolic space, even under the condition that the spatial dimensions of hyperbolic space and Euclidean space (which also correspond to the eigendimensions of eigenvectors) are the same, the distance between two eigenvectors that are close in Euclidean space will increase exponentially after being mapped to hyperbolic space.
[0063] Therefore, for any two eigenvectors, the difference between them can be more clearly expressed by distance calculation in hyperbolic space.
[0064] It should be noted that the mapping relationship from Euclidean space to hyperbolic space and the distance measurement method in hyperbolic space will also be different depending on the corresponding hyperbolic space model. The following provides the definition function of the Poincaré Ball Model, the corresponding mapping function, and the distance calculation function when the Poincaré Ball Model is used as the preset hyperbolic space model:
[0065]
[0066] in, represents the n-dimensional Poincare sphere model, n represents n-dimensional Euclidean space, x represents a point in Euclidean space, c represents the curvature hyperparameter, ‖·‖ represents the Euclidean norm, represents the Riemannian metric, Indicates the European metric, g E =I n , represents the mapping function that maps x to the Poincare sphere model, v and u represent any two points in the Poincare sphere model, d H (u, v) is a distance calculation function that can represent the distance between any two points in the Poincare sphere model. tanh represents the hyperbolic tangent function, and arcosh represents the inverse hyperbolic cosine function.
[0067] According to the above formula, the closer the Euclidean norm of x is to The density of the Poincare sphere model is larger, that is, as the distance As we get closer, the distance in the Poincare sphere model corresponding to the same Euclidean norm grows exponentially.
[0068] For other hyperbolic space models, it is only necessary to determine the corresponding mapping function and distance calculation function, and the method shown in the figure in this embodiment can be implemented in the same way, which is not described in detail in this specification.
[0069] Step S105A: Determine a first feature prototype matching the first sample speech among multiple feature prototypes, and determine a first loss based on a first distance of the first hyperbolic feature from the first feature prototype in the hyperbolic space and a second distance from the multiple feature prototypes.
[0070] This embodiment also uses multiple feature prototypes represented in hyperbolic space, for example, there may be K (K>1). It should be noted that in the complete training process, the following can usually be executed cyclically: Figure 1 The training method shown in FIG. 1 is used to continuously update the parameters of the classifier and each feature prototype. In the first round of training, the K feature prototypes are randomly generated before executing step S105A, and the subsequent rounds of training can continue to be performed according to the feature prototypes updated in the previous round. Figure 1 The method shown does not limit the specific positions of the K feature prototypes when they are generated in the hyperbolic space.
[0071] In some implementations, each feature prototype may also correspond to a second classification label. Similar to the first classification label, the second classification label may include a detection result of "real" and a detection result of "forged." The present invention expects to use feature prototypes to represent a type of template feature in the corresponding anti-counterfeiting detection result. Therefore, when using trained feature prototypes for voice anti-counterfeiting detection, when the voice to be detected is similar to any feature prototype, it can be determined that the voice to be detected should have similar feature information to the feature prototype. Then, the detection result of the voice to be detected can be determined based on the second classification label of the feature prototype.
[0072] After obtaining the first hyperbolic feature corresponding to the first speech feature, the computing device can determine the feature prototype that matches the first sample speech among the above-mentioned K feature prototypes as the first feature prototype. Specifically, the aforementioned distance calculation function can be used to determine the distance between each of the K feature prototypes and the first hyperbolic feature, and then determine the first feature prototype that matches the first sample speech. Specifically, the feature prototype closest to the first hyperbolic feature can be directly used as the first feature prototype that matches the first sample speech; or several feature prototypes closest to the first hyperbolic feature can be determined, for example, K, and the first feature prototype that matches the first sample speech can be randomly determined among the aforementioned K feature prototypes. This specification does not impose any restrictions on this.
[0073] After determining the first feature prototype that matches the first sample speech, the computing device may determine a first loss based on the first feature prototype. Specifically, a first distance between the first hyperbolic feature and the first feature prototype in hyperbolic space and K second distances between the first hyperbolic feature and the K feature prototypes in hyperbolic space may be determined, and the first prototype loss may be determined based on the first matrix and the K second distances.
[0074] In some implementations, the first prototype loss is negatively correlated with the sum of the K second distances. On the other hand, when the second classification label corresponding to the first feature prototype is the same as the first classification label corresponding to the first hyperbolic feature, the first prototype loss is negatively correlated with the first distance; when the second classification label corresponding to the first feature prototype is different from the first classification label corresponding to the first hyperbolic feature, the first prototype loss is positively correlated with the first distance.
[0075] Therefore, the first prototype loss can reflect the similarity between the first hyperbolic feature and the first feature prototype, and simultaneously, can also reflect the similarity between the first hyperbolic feature and each feature prototype.
[0076] Specifically, this embodiment expects that each feature prototype has a certain degree of distinction, and each feature prototype can accurately represent the effective features corresponding to the second classification label. Therefore, by adjusting the feature prototype according to the above-mentioned first prototype loss, the first feature prototype can learn the feature information contained in the first sample speech with the same classification label. On the other hand, it can also avoid the effective features corresponding to each feature prototype from having too high a similarity, increase the distinction between each feature prototype, and use a limited number of feature prototypes to cover more potential features.
[0077] Step S105B: using a classifier to process the third distance between the first hyperbolic feature and the plurality of feature prototypes to obtain a predicted classification of the first sample speech; and determining a classification loss based on the predicted classification and a first classification label indicating speech authenticity corresponding to the first sample speech.
[0078] On the other hand, after obtaining the first hyperbolic feature corresponding to the first speech feature, the computing device may further calculate the distance between the first hyperbolic feature and each feature prototype, determine K third distances between the first hyperbolic feature and the K feature prototypes, and then process the distances between the first hyperbolic feature and each feature prototype using a classifier. As previously described, each feature prototype corresponds to a second classification label. Based on the third distances corresponding to the K feature prototypes and the second classification labels, the predicted classification of the first sample speech can be determined.
[0079] Furthermore, the computing device may determine a classification loss corresponding to the first sample speech according to the predicted classification and the first classification label.
[0080] In some implementations, for any feature prototype, the computing device may determine the weight corresponding to the feature prototype based on the third distance corresponding to the feature prototype. Further, based on the weights corresponding to the above-mentioned K feature prototypes and the second classification label, the computing device may determine the predicted probability of each detection result corresponding to the first sample speech as the predicted classification of the first sample speech.
[0081] It can be noted that in the process of determining the predicted classification in step S105B, this embodiment does not reduce the dimension of the first hyperbolic feature, but directly uses the distance calculation function in the hyperbolic space to determine the distance between the first hyperbolic feature and each feature prototype, thereby avoiding the process of reducing the dimension of the first speech feature in the traditional method, and can more accurately identify the corresponding relationship between the first hyperbolic feature and each feature prototype, thereby improving the accuracy of the prediction when using the trained classifier and the feature prototype for prediction.
[0082] Furthermore, a classification loss corresponding to the first sample speech can be determined based on the difference between the predicted classification and the first classification label of the first sample speech. Specifically, the classification loss can use various common loss functions, such as cross entropy loss, hinge loss, etc., and this specification does not limit this.
[0083] The following shows the classification loss determined from the binary cross entropy loss:
[0084]
[0085] in, represents classification loss, BCE represents binary cross entropy loss, W represents classifier parameters, d H (z, P) represents the distance between the first hyperbolic feature z and each feature prototype, P is the set of K feature prototypes mentioned above, and y is the first sample label corresponding to the first speech sample z.
[0086] Therefore, in the process of determining the predicted classification of the first sample speech, the feature information contained in each feature prototype is fully utilized. Accordingly, the classification loss determined thereby can be used to comprehensively adjust each feature prototype during the training process.
[0087] Step S107: updating the parameters of the classifier and the multiple feature prototypes according to the training loss, where the training loss includes the first loss and the classification loss.
[0088] After determining the classification loss and the first prototype loss, the computing device may update the parameters of the classifier according to the classification loss and the first prototype loss, and update at least one of the aforementioned K feature prototypes.
[0089] Thus, the following is performed using different first sample voice cycles: Figure 2 The method shown can enable the K feature prototypes to learn the feature information of the first sample speech. In addition, according to the above-mentioned method for determining the first prototype loss and the classification loss, each feature prototype can also have a certain degree of distinction, thereby improving the generalization ability of each feature prototype.
[0090] like Figure 1 The training method shown in the figure is used for voice anti-counterfeiting detection. In view of the high data dimension and complex sample source domain of voice data, a training method using feature prototypes in hyperbolic space is adaptively determined. The classifier and feature prototype with strong generalization ability and strong fitting ability for high-dimensional data can be trained.
[0091] In some implementations, Figure 1In step 105A shown, based on the first classification label corresponding to the first sample speech, several target prototypes are determined from multiple feature prototypes, the second classification labels corresponding to the several target prototypes are the same as the first classification label corresponding to the first sample speech, and the feature prototype among the several target prototypes that is closest to the first hyperbolic feature is determined as the first feature prototype.
[0092] As mentioned above, K feature prototypes can be preset, and the K feature prototypes can include, for example, K b The second classification label is the detection result - "real" (hereinafter the second classification label is denoted as k b )’s feature prototype and K s The second classification label is the detection result - "forgery" (hereinafter the second classification label is recorded as k s ) feature prototype, and thus K can be determined according to the first classification label b or K s The feature prototypes are used as the target prototype, and the first feature prototype is determined in the target prototype.
[0093] Furthermore, the first prototype loss can be expressed as:
[0094]
[0095] in, represents the first prototype loss, represents the first distance, represents the first feature prototype, j represents the second classification label, j∈(b,s), Indicates the second distance.
[0096] In some implementations, Figure 1 Before step 107 shown, the feature extractor is used to determine a second speech feature corresponding to a second sample speech, where the second sample speech is the enhanced speech corresponding to the first sample speech, and the feature conversion module is used to map the second speech feature to a hyperbolic space to obtain a second hyperbolic feature corresponding to the second speech feature, and a second prototype loss is determined based on the difference between the second hyperbolic feature and the first feature prototype; Figure 1 In step 107 shown, the parameters of the classifier and the multiple feature prototypes are updated according to the training loss, and the training loss includes the first prototype loss, the classification loss and the second prototype loss.
[0097] In order to further enhance the generalization ability of the trained classifier and feature prototype, in this embodiment, the second sample speech can also be processed using the same method as step S101 and step S103 to obtain a second hyperbolic feature corresponding to the second sample speech.
[0098] It should be noted that, since the second sample speech is an enhanced feature of the first sample speech, the second sample speech only includes the same feature information as the first sample speech, except for the noise caused by the enhancement process. Furthermore, the second hyperbolic feature should also include the same feature information as the first sample speech, and the noise caused by the enhancement process can simulate the domain features of data from different source domains in actual scenarios. This part of the domain features is exactly what is expected in this embodiment to be ignored by each feature prototype and the classifier. This embodiment expects the difference between the second hyperbolic feature and the first feature prototype to be the same as the difference between the first hyperbolic feature and the first feature prototype. In this way, the fourth distance between the second hyperbolic feature and the second feature prototype can be determined, and the second prototype loss can be determined based on the difference between the fourth distance and the first distance.
[0099] Furthermore, in some implementations, a fourth distance between the second hyperbolic feature and the first feature prototype is determined; and a second prototype loss is determined based on a difference between the fourth distance and the first distance, and a fifth distance between the second hyperbolic feature and the first hyperbolic feature.
[0100] Specifically, the second prototype loss can be set as follows:
[0101]
[0102] in, represents the second prototype loss, z aug represents the second hyperbolic characteristic, Represents the fourth distance, d H (z,z aug ) represents the fifth distance, ‖·‖2 represents the Euclidean norm, and the meanings of other parameters are the same as in the previous formula.
[0103] Therefore, the second prototype loss also includes a fifth distance that directly represents the difference between the second hyperbolic feature and the first hyperbolic feature. Using the second prototype loss to update the classifier and each feature prototype can make the update of each feature prototype as unaffected by the domain features as possible, thereby improving the generalization of each updated feature prototype.
[0104] In some implementations, the feature extractor can also be trained based on the training loss including the second prototype loss to reduce the content of domain features in the speech features extracted by the feature extractor, thereby improving the accuracy of the detection results when using the feature extractor for speech anti-counterfeiting detection.
[0105] It should be noted that the first sample speech is enhanced to determine the second sample speech, and various common enhancement methods may be used, such as time domain enhancement, amplitude domain enhancement, frequency domain enhancement, etc., which are not limited in this specification.
[0106] In some implementations, a plurality of first hyperbolic features determined based on a plurality of first sample voices and a plurality of second hyperbolic features determined based on a plurality of second sample voices corresponding to the plurality of first sample voices can also be obtained, each second sample voice is an enhanced voice corresponding to each first sample voice, each first hyperbolic feature or second hyperbolic feature is composed of a plurality of feature components, and a plurality of domain-sensitive feature components are determined in each feature component based on the difference between the plurality of first hyperbolic features and the plurality of second hyperbolic features, and a domain feature loss is determined based on the plurality of domain-sensitive feature components; in the example Figure 1 In step S107 shown, the parameters of the classifier, the multiple feature prototypes, and the feature extractor are updated according to the training loss, and the training loss includes the first prototype loss, the classification loss, and the domain feature loss.
[0107] Specifically, the idea of batch training can be used to obtain a number of first sample voices, for example, B, where B is the number of samples in a batch of batch training. Each first sample voice corresponds to a second sample voice. Figure 1 In steps S101-S103 shown, B first hyperbolic features corresponding to the B first sample speech and B second hyperbolic features corresponding to the B second samples are determined. Further, a first feature matrix composed of the first hyperbolic features is determined, and a second feature matrix composed of the second hyperbolic features is determined. Next, a first covariance matrix of the first feature matrix is determined, and a second covariance matrix of the second feature matrix is determined. Based on the difference between the first covariance matrix and the second covariance matrix, the domain feature loss is determined.
[0108] Under ideal conditions—that is, without interference from domain features—the first feature matrix is identical to the second feature matrix, and further, the first covariance matrix is identical to the second covariance matrix. Based on the difference between the first and second covariance matrices, the feature components that cause the difference between the first and second feature matrices can be determined as domain-sensitive feature components. Furthermore, the domain feature loss is determined based on the representation value of the domain-sensitive feature component. This domain feature loss is positively correlated with the representation value of the domain-sensitive feature vector.
[0109] The domain feature loss, for example, can be determined as follows:
[0110]
[0111]
[0112] in, It can be the first feature matrix or the second feature matrix, B is the number of samples in a batch training, D is the feature dimension, Σ represents the calculation method of covariance in hyperbolic space, org represents the first sample speech, aug represents the second sample speech, Σorg Can represent the first covariance matrix, Σ aug The second covariance matrix can be represented as It can represent the difference between the first covariance matrix and the second covariance matrix, M i,j (k c ) is a mask matrix, which is multiplied by the first feature matrix and the second feature matrix respectively to expose the top k features with the largest representation value. c feature components as domain-sensitive feature components, It is the domain feature loss.
[0113] Specifically, the process of determining the mask matrix is as follows: Figure 2 As shown in the figure, each color block in the first covariance matrix and the second covariance matrix represents each covariance component, and the darker the color block, the larger the representation value; in the difference between the first covariance matrix and the second covariance matrix, similarly, the darker the color block, the larger the representation value; each color block in the mask matrix corresponds to the position of each feature component, white represents the exposed position, and black represents the masked position.
[0114] Therefore, by updating the parameters of the feature extractor according to the domain feature loss, the interference caused by the updated feature extractor in the process of extracting features from speech samples can be reduced, thereby improving the generalization ability of the feature extractor.
[0115] In some implementations, Figure 1 Before step S107 shown, a ternary prototype group can also be determined from the multiple feature prototypes, the ternary prototype group includes an anchor prototype, a neighboring prototype in the neighborhood of the anchor prototype, and a random prototype outside the neighborhood of the anchor prototype, and among the multiple upper prototypes, a first upper prototype matching the anchor prototype is determined based on the anchor prototype and the neighboring prototype, and a second upper prototype unrelated to the anchor prototype is determined based on the first upper prototype and the random prototype, and the hierarchical loss is determined based on the sixth distance between the ternary prototype group and the first upper prototype and the seventh distance between the ternary model group and the second upper prototype; in the example shown in FIG. Figure 1 In step S107 shown, the parameters of the classifier and the multiple feature prototypes are updated according to the training loss, and the training loss includes the first prototype loss, the classification loss and the hierarchical loss.
[0116] The anchor prototype can be randomly determined from the aforementioned K feature prototypes. After the anchor prototype is determined, the neighborhood of the anchor prototype can be determined. In some implementations, a neighbor distance can be preset, and each feature prototype whose distance from the anchor prototype is less than the neighbor distance is determined as the neighborhood of the anchor prototype. Alternatively, a number of neighbors d can be preset, and the d feature prototypes closest to the anchor prototype are determined as the neighborhood of the anchor prototype. This specification does not impose any restrictions on this. Furthermore, a neighbor prototype is randomly determined within the neighborhood of the anchor prototype, and a random prototype is randomly determined outside the domain of the anchor prototype.
[0117] On the other hand, this embodiment also uses multiple superordinate prototypes represented in hyperbolic space, for example, Kt (Kt>1). Similar to the feature prototype, in the first round of training, the above Kt superordinate prototypes are Figure 1 The steps S107 shown in FIG1 are randomly generated before the training, and the subsequent rounds of training can continue to be performed according to the upper prototypes updated in the previous round. Figure 1 The method shown does not limit the specific position of the above-mentioned Kt superordinate prototype in the hyperbolic space.
[0118] Specifically, the process of determining the first superordinate prototype and the second superordinate prototype is as follows: Figure 3 As shown, where p i Anchor prototype p j is the nearest prototype p k As shown in the figure, each "x" mark represents each upper prototype. i The distance from p j The distance between each upper prototype and p is determined. The first upper prototype is determined in each upper prototype. In some implementations, each upper prototype has a distance between i The distance between j The smaller the distance, the greater the probability of being determined as the first superordinate prototype. Specifically, refer to the following formula:
[0119]
[0120] ρ ij That is the first superordinate prototype, ρ can represent any superordinate prototype, g ij It is a random disturbance, which can reduce the risk of falling into the local optimum during the determination of the first upper prototype. It can represent the probability that the upper prototype ρ is determined as the first upper prototype, according to It can be seen that the first superordinate prototype has the following characteristics compared to the other superordinate prototypes: the distance to the anchor prototype and the distance to the neighboring prototype are close, and the difference between the distance to the anchor prototype and the distance to the neighboring prototype is small. Therefore, the first superordinate prototype can better represent the hierarchical features shared by the anchor prototype and the neighboring prototype. Since the neighboring prototype is in the neighborhood of the anchor prototype, the neighboring prototype should have a certain degree of similarity with the anchor prototype. The hierarchical features represented by the first superordinate prototype are valid features. Figure 3 The solid line in the middle represents the relationship between the first superordinate prototype and the anchor prototype and the neighboring prototype, indicating that the relationship needs to be further strengthened.
[0121] It should be noted that due to the characteristics of hyperbolic space, the first superordinate prototype can represent the hierarchical features between the anchor prototype and the neighboring prototype without additional dimensions, which reduces the difficulty of fitting the hierarchical features.
[0122] Similarly, refer to ρ ij The determination process of ρ ij With p k Substitute p into the above formula i With p j The position of the second upper prototype ρ can be determined ijk However, unlike the first superordinate prototype, since the random prototype is not in the neighborhood of the anchor prototype, the hierarchical features represented by the second superordinate prototype are interference features, such as Figure 3 As shown, the dotted line represents the relationship between the second superordinate prototype and the anchor prototype and the neighboring prototype, indicating that the relationship needs to be further weakened.
[0123] The level loss can be determined as follows:
[0124]
[0125] in, This is the hierarchical loss, where δ is the margin value. This hierarchical loss is used to update both the superordinate and feature prototypes. This not only improves the expressive power of the updated superordinate prototype for hierarchical features, but also, through the relationship between the superordinate and feature prototypes, it can inversely promote each feature prototype to learn valid features and avoid interfering features, thereby improving the expressive power of each feature prototype for valid feature information.
[0126] In some implementations, Figure 1 In step S107 shown, the parameters of the feature extractor, the parameters of the classifier, and the multiple feature prototypes can be updated according to the training loss, and the training loss includes the first prototype loss, the classification loss, the second prototype loss, the domain feature loss, and the hierarchical loss.
[0127] Figure 4 The structural diagram of the anti-counterfeiting detection model in one embodiment of the present specification is shown. The anti-counterfeiting detection model includes a feature extractor, a feature conversion module, multiple feature prototypes and a classifier. The feature prototypes and classifiers of the anti-counterfeiting detection model can be processed as follows: Figure 1 In some implementations, the feature extractor of the anti-counterfeiting detection model can also be trained by the method shown in FIG. Figure 1 Training of the method shown.
[0128] Among them, the feature extractor is used to determine the feature to be tested corresponding to the speech to be tested; the feature conversion module is used to map the feature to be tested to the hyperbolic space to determine the hyperbolic feature to be tested corresponding to the feature to be tested; the classifier is used to determine the K eighth distances between the hyperbolic feature to be tested and the multiple feature prototypes (following the above example, for example, K), and based on the K eighth distances, obtain the prediction result corresponding to the speech to be tested.
[0129] Figure 5 A flow chart of a method for voice anti-counterfeiting detection in an embodiment of this specification is shown. The method for voice anti-counterfeiting detection can be used as follows: Figure 4 The anti-counterfeiting detection model shown is executed, and the method includes:
[0130] Step S501: Acquire the voice to be tested.
[0131] Because Figure 1 The method shown can improve the generalization of the classifier and each feature prototype, even for the test speech from the target domain not covered by the training sample set, the method can be used as Figure 5 The method shown is used for anti-counterfeiting detection.
[0132] Step S503: using the feature extractor to determine the features to be tested corresponding to the speech to be tested.
[0133] For details, please refer to step S101.
[0134] Step S505: using the feature conversion module to map the feature to be measured to a hyperbolic space, and determining a hyperbolic feature to be measured corresponding to the feature to be measured.
[0135] For details, please refer to step S103.
[0136] Step S507: using the classifier to process the eighth distance between the hyperbolic feature to be tested and the multiple feature prototypes to obtain a detection result of the speech to be tested.
[0137] Specifically, referring to step S105B, the prediction probability of each detection result corresponding to the speech to be tested is determined, and the detection result with the highest prediction probability is used as the detection result of the speech to be tested.
[0138] It should be noted that there is no need to determine the classification loss in step S507.
[0139] like Figure 5 The method for voice anti-counterfeiting detection shown has high detection accuracy for the test voice in an unknown target domain and can accurately identify forged voices in unknown scenarios and unknown attack methods.
[0140] In the 1990s, technological improvements could be clearly distinguished as either hardware improvements (for example, improvements to circuit structures like diodes, transistors, and switches) or software improvements (improvements to process flows). However, with the advancement of technology, many process flow improvements today can now be considered direct improvements to hardware circuit structures. Designers almost always create the corresponding hardware circuit structure by programming the improved process flow into the hardware circuit. Therefore, it cannot be said that a process flow improvement cannot be implemented using hardware modules. For example, a programmable logic device (PLD), such as a field programmable gate array (FPGA), is an integrated circuit whose logical function is determined by user programming. Designers can "integrate" a digital system on a PLD through their own programming, without having to hire a chip manufacturer to design and manufacture a dedicated integrated circuit chip. Moreover, nowadays, instead of manually fabricating integrated circuit chips, this programming is mostly done using "logic compiler" software. This is similar to the software compiler used when developing programs. Before compilation, the original code must also be written in a specific programming language, called a hardware description language (HDL). There is not just one HDL, but many, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc. The most commonly used ones are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art will also understand that by simply programming the method flow in one of these hardware description languages and then programming it into an integrated circuit, a hardware circuit that implements the logic method flow can be easily obtained.
[0141] The controller can be implemented in any suitable manner. For example, the controller can take the form of a microprocessor or processor and a computer-readable medium storing computer-readable program code (e.g., software or firmware) executable by the (micro)processor, logic gates, switches, application-specific integrated circuits (ASICs), programmable logic controllers, and embedded microcontrollers. Examples of controllers include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320. The memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art will also know that in addition to implementing the controller in a purely computer-readable program code format, the controller can be implemented in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers by logically programming the method steps. Therefore, such a controller can be considered a hardware component, and the devices included therein for implementing various functions can also be considered as structures within the hardware component. Or even, the devices for implementing various functions can be considered as both software modules that implement the method and structures within the hardware component.
[0142] The systems, devices, modules or units described in the above embodiments may be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a server system. Of course, this application does not exclude that with the future development of computer technology, the computer that implements the functions of the above embodiments may be, for example, a personal computer, a laptop computer, an in-vehicle human-computer interaction device, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.
[0143] Although one or more embodiments of this specification provide method operation steps as described in the embodiments or flow charts, more or fewer operation steps may be included based on conventional or non-creative means. The order of steps listed in the embodiments is only one way of executing the order of many steps and does not represent the only execution order. When the device or terminal product in practice is executed, it can be executed in sequence or in parallel according to the method shown in the embodiments or the drawings (for example, a parallel processor or a multi-threaded processing environment, or even a distributed data processing environment). The term "comprise", "include" or any other variant thereof is intended to cover non-exclusive inclusion, so that the process, method, product or equipment including a series of elements includes not only those elements, but also includes other elements that are not clearly listed, or also includes elements inherent to such process, method, product or equipment. In the absence of more restrictions, it is not excluded that there are other identical or equivalent elements in the process, method, product or equipment including the elements. For example, if the words first, second, etc. are used to represent the name, they do not represent any particular order.
[0144] For the convenience of description, the above devices are described in terms of functions divided into various modules. Of course, when implementing one or more of the present specifications, the functions of each module can be implemented in the same or multiple software and / or hardware, or the module that implements the same function can be implemented by a combination of multiple sub-modules or sub-units, etc. The device embodiments described above are merely schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.
[0145] The present invention is described with reference to flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0146] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0147] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0148] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.
[0149] Memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. Memory is an example of a computer-readable medium.
[0150] Computer-readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic disk storage, graphene storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory media such as modulated data signals and carrier waves.
[0151] Those skilled in the art will appreciate that one or more embodiments of this specification may be provided as a method, system, or computer program product. Thus, one or more embodiments of this specification may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, one or more embodiments of this specification may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0152] One or more embodiments of this specification may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, and the like that perform specific tasks or implement specific abstract data types. One or more embodiments of this specification may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communications network. In a distributed computing environment, program modules may be located in local and remote computer storage media, including storage devices.
[0153] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between the various embodiments can be referenced across them. Each embodiment focuses on the differences from the other embodiments. In particular, since the system embodiments are generally similar to the method embodiments, their description is relatively simple. For relevant parts, reference can be made to the description of the method embodiments. Throughout this specification, reference to the terms "one embodiment," "some embodiments," "examples," "specific examples," or "some examples" means that the specific features, structures, materials, or characteristics described in conjunction with that embodiment or example are included in at least one embodiment or example of this specification. In this specification, the schematic representations of these terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in any one or more embodiments or examples. Furthermore, those skilled in the art may combine and integrate the different embodiments or examples, and features of different embodiments or examples, described in this specification, without conflict.
[0154] The foregoing description is merely an example of one or more embodiments of this specification and is not intended to limit the one or more embodiments of this specification. Those skilled in the art will appreciate that various modifications and variations of one or more embodiments of this specification are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of this specification are intended to be included within the scope of the claims.
Claims
1. A training method for voice anti-counterfeiting detection, comprising: Determining, using a feature extractor, a first speech feature corresponding to the first sample speech; Mapping the first speech feature to a hyperbolic space using a feature conversion module to obtain a first hyperbolic feature corresponding to the first speech feature; Determining a first feature prototype matching the first sample speech from a plurality of feature prototypes, and determining a first prototype loss based on a first distance of the first hyperbolic feature from the first feature prototype and a second distance from the plurality of feature prototypes in a hyperbolic space; Processing the third distance between the first hyperbolic feature and the plurality of feature prototypes using a classifier to obtain a predicted classification of the first sample speech; determining a classification loss based on the predicted classification and a first classification label indicating speech authenticity corresponding to the first sample speech; The parameters of the classifier and the plurality of feature prototypes are updated according to a training loss, wherein the training loss includes the first prototype loss and the classification loss.
2. The method according to claim 1, wherein determining a first feature prototype matching the first sample speech from a plurality of feature prototypes comprises: Determining, based on the first classification label corresponding to the first sample speech, a plurality of target prototypes from a plurality of feature prototypes, wherein the second classification labels corresponding to the plurality of target prototypes are the same as the first classification label corresponding to the first sample speech; The feature prototype closest to the first hyperbolic feature among the several target prototypes is determined as the first feature prototype.
3. The method according to claim 1, before updating the parameters of the classifier and the plurality of feature prototypes, further comprising: Determining, using the feature extractor, a second speech feature corresponding to a second sample speech, where the second sample speech is the enhanced speech corresponding to the first sample speech; Mapping the second speech feature to a hyperbolic space using the feature conversion module to obtain a second hyperbolic feature corresponding to the second speech feature; determining a second prototype loss based on a difference between the second hyperbolic feature and the first feature prototype; The training loss also includes: the second prototype loss.
4. The method according to claim 3, wherein determining the second prototype loss based on the difference between the second hyperbolic feature and the first feature prototype comprises: Determine a fourth distance between the second hyperbolic feature and the first feature prototype; and determine a second prototype loss according to a difference between the fourth distance and the first distance and a fifth distance between the second hyperbolic feature and the first hyperbolic feature.
5. The method according to claim 1, wherein The training loss also includes domain feature loss, and the method further includes: Acquire a plurality of first hyperbolic features determined based on a plurality of first speech samples, and a plurality of second hyperbolic features determined based on a plurality of second speech samples corresponding to the plurality of first speech samples, where each second speech sample is an enhanced speech corresponding to each first speech sample, and each first hyperbolic feature or second hyperbolic feature is composed of a plurality of feature components; determining a plurality of domain-sensitive feature components in each feature component according to differences between the plurality of first hyperbolic features and the plurality of second hyperbolic features; determining a domain feature loss according to the plurality of domain-sensitive feature components; Updating the parameters of the classifier and the plurality of feature prototypes according to the training loss specifically includes: According to the training loss, the parameters of the classifier, the multiple feature prototypes and the feature extractor are updated.
6. The method of claim 1 , further comprising: Determine a ternary prototype group from the plurality of feature prototypes, the ternary prototype group comprising an anchor prototype, neighboring prototypes within a neighborhood of the anchor prototype, and random prototypes outside the neighborhood of the anchor prototype; Determine, among multiple superordinate prototypes, a first superordinate prototype that matches the anchor prototype according to the anchor prototype and the neighboring prototypes; Determine, among the multiple superordinate prototypes, a second superordinate prototype that is unrelated to the anchor prototype based on the first superordinate prototype and the random prototype; determining a hierarchy loss based on a sixth distance between the ternary prototype group and the first superordinate prototype and a seventh distance between the ternary prototype group and the second superordinate prototype; The training loss also includes: layer loss.
7. The method according to claim 1, wherein the hyperbolic space is a Poincare sphere model.
8. The method of claim 1 , further comprising: Get the voice to be tested; Determining the features to be measured corresponding to the speech to be measured using the feature extractor; Mapping the feature to be measured to a hyperbolic space using the feature conversion module to determine a hyperbolic feature to be measured corresponding to the feature to be measured; The classifier is used to process the eighth distance between the hyperbolic feature to be tested and the multiple feature prototypes to obtain a detection result of the speech to be tested.
9. A computer program product comprising a computer program / instructions, which, when executed by a processor, implement the steps of the method according to claims 1 to 8.
10. A computing device comprising a memory and a processor, wherein the memory stores executable code, and when the processor executes the executable code, the method according to any one of claims 1 to 8 is implemented.