Voice tracing model training method and device

CN122821995APending Publication Date: 2026-09-25TIANJIN UNIV +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611226114.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-13
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0003]相关技术通过语音检测的技术手段对虚假语音进行溯源的过程中,难以精准检测出虚假语音的声源身份,即,难以追溯经过语音转换的虚假语音的说话人的身份

Benefits of technology

[0016]本申请的第四方面还提供了一种计算机可读存储介质,其上存储有计算机程序或指令,上述计算机程序或指令被处理器执行时实现上述方法的步骤。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122821995A_ABST
    Figure CN122821995A_ABST
Patent Text Reader

Abstract

The application provides a voice tracing model training method and device, which can be applied to the technical field of data processing. The method comprises the following steps: processing the differences between real voice features and real voice labels, the differences between false voice features and reference sound source voice features, and the differences between false voice features and false voice labels by using a difference-guided adaptive angular interval loss function to obtain a difference-guided adaptive angular interval loss value; processing the gradient angles between the gradients of multiple voice features and the binary contrast loss values of the multiple voice features by using a gradient deviation perception contrast learning loss function to obtain a gradient deviation perception discriminant loss value; and training a voice tracing model according to the difference-guided adaptive angular interval loss value and the gradient deviation perception discriminant loss value to obtain a trained voice tracing model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing technology, and more specifically to a training method and apparatus for a speech source tracing model. Background Technology

[0002] Speech conversion technology can change the speaker's timbre to the target timbre while preserving the speaker's speech content.

[0003] While related technologies use voice detection to trace the source of fake voice messages, they struggle to accurately identify the speaker, meaning it's difficult to trace the identity of the person speaking the fake message after it has been converted. Therefore, these technologies fall short of the accuracy requirements for tracing the source in public safety scenarios and still pose security risks. Summary of the Invention

[0004] In view of the above problems, this application provides a training method and apparatus for a speech source tracing model.

[0005] According to the first aspect of this application, a training method for a speech source tracing model is provided, comprising: acquiring a source tracing speech training set, wherein the source tracing speech training set includes real speech samples and fake speech samples; extracting features from the real speech samples and fake speech samples respectively using the speech source tracing model to obtain real speech features and fake speech features; processing the differences between real speech features and real speech labels, the differences between fake speech features and reference source speech features, and the differences between fake speech features and fake speech labels using a difference-guided adaptive angle interval loss function to obtain difference-guided adaptive angle interval loss values, wherein the reference source speech features are determined by calculating the vector mean of multiple real speech samples with the same source identity; and processing the similarity between two speech features in multiple speech feature pairs using a binary contrast loss function to obtain the binary contrast loss values ​​for each of the multiple speech feature pairs. The speech feature pair includes at least one of the following: two real speech features with the same source identity; a real speech feature and a fake speech feature with the same source identity; two real speech features with different source identities; and real speech features and fake speech features with different source identities. Based on the binary contrast loss values ​​of each speech feature pair, the gradients of each speech feature pair are determined, and the gradients of the speech feature pairs represent the changing trends of the binary contrast loss values. The gradient bias-aware contrastive learning loss function is used to process the gradient angle between the gradients of each speech feature pair and the binary contrast loss values ​​of each speech feature pair to obtain the gradient bias-aware discriminative loss value. The speech source tracing model is trained according to the difference-guided adaptive angle interval loss value and the gradient bias-aware discriminative loss value to obtain the trained speech source tracing model.

[0006] According to embodiments of this application, a difference-guided adaptive angle interval loss function is used to process the differences between real speech features and real speech labels, the differences between fake speech features and reference source speech features, and the differences between fake speech features and fake speech labels, to obtain a difference-guided adaptive angle interval loss value. This includes: calculating the similarity between reference source speech features and fake speech features to obtain a cosine similarity, where the cosine similarity characterizes the degree of damage to real speech samples during the process of converting fake speech samples into fake speech samples; obtaining target angle intervals corresponding to real speech features and fake speech features respectively based on preset angle intervals and cosine similarity; obtaining a real speech loss value based on the target angle interval corresponding to the real speech features and the real similarity, where the real similarity characterizes the similarity between real speech features and real speech labels; obtaining a fake speech loss value based on the target angle interval corresponding to the fake speech features and the fake similarity, where the fake similarity characterizes the similarity between fake speech features and fake speech labels; and determining the difference-guided adaptive angle interval loss value based on the real speech loss value and the fake speech loss value.

[0007] According to an embodiment of this application, the preset angle interval includes a first preset angle interval and a second preset angle interval, wherein the second preset angle interval is greater than the first preset angle interval; wherein, obtaining the target angle interval corresponding to the real speech feature and the fake speech feature respectively based on the preset angle interval and cosine similarity includes: determining the first preset angle interval as the target angle interval corresponding to the real speech feature; processing the ratio of cosine similarity to a preset value using a truncation function to obtain a first parameter, wherein the truncation function is used to constrain the ratio of cosine similarity to the preset value within a preset numerical range; obtaining a second parameter based on the difference between the second preset angle interval and the first preset angle interval; and determining the target angle interval corresponding to the fake speech feature based on the product of the first parameter and the second parameter and the sum of the first preset angle interval.

[0008] According to embodiments of this application, a speech feature pair consisting of two real speech features with the same source identity, and a speech feature pair consisting of a real speech feature and a fake speech feature with the same source identity, have a first feature pair label; a speech feature pair consisting of two real speech features with different source identities, and a speech feature pair consisting of a real speech feature and a fake speech feature with different source identities, have a second feature pair label; wherein, a binary contrastive loss function is used to process the similarity between two speech features in multiple speech feature pairs to obtain the binary contrastive loss value for each of the multiple speech feature pairs, including: for a speech feature pair consisting of two real speech features with the same source identity, fusing the similarity between the two real speech features, the first label value corresponding to the first feature pair label, and the pre-defined feature pair label. The coefficients are set to obtain the binary contrast loss value. For speech feature pairs with the same source identity, real speech features and fake speech features are fused together with the similarity between real and fake speech features, the first label value corresponding to the first feature pair label, and the preset coefficients to obtain the binary contrast loss value. For speech feature pairs with two real speech features but different source identities, the similarity between the two real speech features, the second label value corresponding to the second feature pair label, and the preset coefficients to obtain the binary contrast loss value. For speech feature pairs with different source identities, real and fake speech features are fused together with the similarity between real and fake speech features, the second label value corresponding to the second feature pair label, and the preset coefficients to obtain the binary contrast loss value.

[0009] According to an embodiment of this application, a speech feature pair with a first feature pair label is used as a reference speech feature pair; wherein a gradient bias-aware contrastive learning loss function is used to process the gradient angle between the gradients of multiple speech feature pairs and the binary contrastive loss value of each of the multiple speech feature pairs to obtain a gradient bias-aware discriminative loss value, including: for any speech feature pair among the multiple speech feature pairs, determining the gradient bias value between the speech feature pair and the reference gradient of the reference speech feature pair based on the gradient angle between the gradient of the speech feature pair and the reference gradient of the reference speech feature pair; and determining the gradient bias-aware discriminative loss value based on the gradient bias value and the binary contrastive loss value of each of the multiple speech feature pairs.

[0010] According to embodiments of this application, determining a gradient deviation-aware discriminative loss value based on the gradient deviation values ​​and binary contrast loss values ​​of multiple speech feature pairs includes: determining a gradient deviation mean based on the mean of the multiple gradient deviation values; determining a gradient deviation coefficient for each of the multiple speech feature pairs based on the ratio of the gradient deviation values ​​to the mean gradient deviation value; fusing the gradient deviation coefficients and binary contrast loss values ​​of the multiple speech feature pairs to determine a gradient loss value for each of the multiple speech feature pairs; and determining a gradient deviation-aware discriminative loss value based on the mean of the gradient loss values ​​of the multiple speech feature pairs.

[0011] According to an embodiment of this application, determining a discrimination loss value for gradient deviation perception based on at least one of the gradient deviation values ​​and binary contrast loss values ​​of multiple speech features includes: determining multiple target gradient deviation values ​​whose gradient deviation values ​​satisfy preset screening conditions; determining target binary contrast loss values ​​corresponding to the multiple target gradient deviation values; and determining a discrimination loss value for gradient deviation perception based on the mean of the multiple target binary contrast loss values.

[0012] According to an embodiment of this application, a speech tracing model is trained based on a difference-guided adaptive angle interval loss value and a gradient deviation-aware discriminant loss value to obtain a trained speech tracing model. This includes: fusing a balance coefficient and a gradient deviation-aware discriminant loss value to obtain a target gradient deviation-aware discriminant loss value; determining a target loss value based on the sum of the difference-guided adaptive angle interval loss value and the target gradient deviation-aware discriminant loss value; and optimizing the model parameters of the speech tracing model based on the target loss value to obtain a trained speech tracing model.

[0013] According to an embodiment of this application, a speech tracing model is used to extract features from real speech samples and fake speech samples to obtain real speech features and fake speech features, including: preprocessing real speech samples and fake speech samples to obtain real speech samples and fake speech samples with the same speech duration; and extracting features from real speech samples and fake speech samples with the same speech duration to obtain real speech features and fake speech features.

[0014] The second aspect of this application provides a training apparatus for a speech source tracing model, comprising: a first acquisition module for acquiring a source tracing speech training set, wherein the source tracing speech training set includes real speech samples and fake speech samples; a first extraction module for extracting features from the real speech samples and fake speech samples respectively using the speech source tracing model to obtain real speech features and fake speech features; a first processing module for processing the differences between real speech features and real speech labels, the differences between fake speech features and reference source speech features, and the differences between fake speech features and fake speech labels using a difference-guided adaptive angle interval loss function to obtain difference-guided adaptive angle interval loss values, wherein the reference source speech features are determined by calculating the vector mean of multiple real speech samples with the same source identity; and a second processing module for processing the similarity between two speech features in multiple speech feature pairs using a binary contrast loss function to obtain the binary contrast loss values ​​for each of the multiple speech feature pairs. The system compares the loss values, where the two speech features in a speech feature pair include at least one of the following: two real speech features with the same source identity; a real speech feature and a fake speech feature with the same source identity; two real speech features with different source identities; and real speech features and fake speech features with different source identities. A first determining module is used to determine the gradients of multiple speech feature pairs based on their respective binary contrast loss values, where the gradients of the speech feature pairs represent the changing trend of the binary contrast loss values. A third processing module is used to process the gradient angles between the gradients of multiple speech feature pairs and their respective binary contrast loss values ​​using a gradient bias-aware contrastive learning loss function to obtain a gradient bias-aware discriminative loss value. A first training module is used to train the speech source tracing model based on the difference-guided adaptive angular interval loss value and the gradient bias-aware discriminative loss value to obtain the trained speech source tracing model.

[0015] A third aspect of this application provides an electronic device comprising: one or more processors; and a memory for storing one or more computer programs, wherein the one or more processors execute the one or more computer programs to implement the steps of the method described above.

[0016] A fourth aspect of this application also provides a computer-readable storage medium having a computer program or instructions stored thereon, which, when executed by a processor, implement the steps of the above-described method.

[0017] The fifth aspect of this application also provides a computer program product, including a computer program or instructions that, when executed by a processor, implement the steps of the above-described method.

[0018] According to embodiments of this application, a speech source tracing model is used to extract features from real and fake speech samples respectively, obtaining real speech features and fake speech features. This allows the influence of speech samples with different source identities on the speech source tracing model during training. Furthermore, a difference-guided adaptive angle interval loss function is used to handle the differences between real speech features and real speech labels, the differences between fake speech features and reference source speech features, and the differences between fake speech features and fake speech labels, obtaining a difference-guided adaptive angle interval loss value. This allows for dynamic constraint of the difference-guided adaptive angle interval loss value by calculating the difference between fake speech features and reference source speech features, effectively suppressing speech source tracing errors. The shift in speech features caused by speech conversion compresses the intra-class distance between genuine and fake speech features with the same source identity in the feature space. Furthermore, a binary contrastive loss function is used to process the similarity between two speech features in multiple speech feature pairs, obtaining binary contrastive loss values ​​for each pair. The gradients of these speech feature pairs are then derived based on these binary contrastive loss values. A gradient-bias-aware contrastive learning loss function is then used to process the differences between the gradients of multiple speech feature pairs, resulting in a gradient-bias-aware discriminative loss value. This allows for the discovery of highly confusing speech feature sample pairs with gradient conflicts, enhancing the model's ability to discriminate complex speech samples during training based on the gradient-bias-aware discriminative loss value. Therefore, training the speech source tracing model based on the difference-guided adaptive angular interval loss value and the gradient-bias-aware discriminative loss value improves the detection accuracy and robustness of the speech source tracing model in speech conversion scenarios, thereby enhancing the ability to identify the source identity and reducing security risks in speech-based identity authentication scenarios. Attached Figure Description

[0019] The above-mentioned contents, other objects, features and advantages of this application will become clearer from the following description of embodiments with reference to the accompanying drawings, in which:

[0020] Figure 1 The diagram illustrates application scenarios of training methods, apparatus, devices, media, and program products for speech tracing models according to embodiments of this application.

[0021] Figure 2 A flowchart illustrating a training method for a speech source tracing model according to an embodiment of this application is shown.

[0022] Figure 3 A schematic diagram of a training method for a speech source tracing model according to an embodiment of this application is shown.

[0023] Figure 4 A flowchart of a speech tracing method according to an embodiment of this application is shown.

[0024] Figure 5 A structural block diagram of a training apparatus for a speech tracing model according to an embodiment of this application is shown.

[0025] Figure 6 A block diagram of an electronic device suitable for implementing a training method for a speech source tracing model according to an embodiment of this application is shown. Detailed Implementation

[0026] The embodiments of this application will now be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of this application. In the following detailed description, numerous specific details are set forth to provide a thorough understanding of the embodiments of this application for ease of explanation. However, it will be apparent that one or more embodiments may be implemented without these specific details. Furthermore, descriptions of well-known structures and technologies are omitted in the following description to avoid unnecessarily obscuring the concepts of this application.

[0027] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the scope of this application. The terms “comprising,” “including,” etc., as used herein indicate the presence of the stated features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.

[0028] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein are to be interpreted in a manner consistent with the context of this specification, and not in an idealized or overly rigid way.

[0029] When using expressions such as "at least one of A, B and C", they should generally be interpreted in accordance with the meaning that is commonly understood by those skilled in the art (e.g., "a system having at least one of A, B and C" should include, but is not limited to, a system having A alone, a system having B alone, a system having C alone, a system having A and B, a system having A and C, a system having B and C, and / or a system having A, B and C, etc.).

[0030] In the technical solution of this application, the user information (including but not limited to user personal information, user image information, user device information, such as location information) and data (including but not limited to data used for analysis, stored data, and displayed data) involved are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, storage, use, processing, transmission, provision, disclosure, and application of related data all comply with relevant laws, regulations, and standards, take necessary confidentiality measures, do not violate public order and good morals, and provide corresponding operation entry points for users to choose to authorize or refuse.

[0031] In scenarios involving automated decision-making using personal information, the methods, devices, and systems provided in this application all offer users corresponding entry points for choosing to agree to or reject the automated decision-making results. If the user chooses to reject, the process proceeds to the expert decision-making stage. Here, "automated decision-making" refers to the activity of automatically analyzing and evaluating an individual's behavioral habits, interests, or economic, health, and credit status through computer programs, and then making a decision. Here, "expert decision-making" refers to the activity of making decisions by personnel who specialize in a particular field, possess specialized experience, knowledge, and skills, and have reached a certain level of professional expertise.

[0032] In related technologies, voice fraud detection algorithms are limited to distinguishing between real and fake speech, and cannot trace the identity of the source speaker behind fake speech. Because the speech conversion process causes significant feature shifts in the source speaker, and the features of fake speech generated by different source speakers are intertwined in the feature space, existing detection algorithms struggle to cross the boundary between real and fake speech and establish an effective mapping relationship between fake speech and the source speaker's real speech. Therefore, existing technologies cannot identify target speech from the same sound source as real speech in complex speech conversion scenarios, resulting in poor adaptability in source speaker identity tracing tasks and failing to meet the need for accurate source tracing.

[0033] This application provides a training method for a speech source tracing model. The method extracts features from real and fake speech samples using the speech source tracing model, obtaining real speech features and fake speech features respectively. During training, it considers the impact of speech samples with different source identities on the speech source tracing model. Furthermore, it uses a difference-guided adaptive angle interval loss function to handle the differences between real speech features and real speech labels, the differences between fake speech features and reference source speech features, and the differences between fake speech features and fake speech labels, obtaining a difference-guided adaptive angle interval loss value. This allows for dynamic constraints on the difference-guided adaptive angle interval loss value by calculating the difference between fake speech features and reference source speech features. This approach effectively suppresses the shift in speech features caused by speech conversion, compressing the intra-class distance between real and fake speech features with the same source identity in the feature space. Furthermore, it utilizes a binary contrastive loss function to process the similarity between two speech features in multiple speech feature pairs, obtaining the binary contrastive loss value for each pair. Based on this loss value, the gradient of the speech feature pair is obtained. Then, a gradient bias-aware contrastive learning loss function is used to process the differences between the gradients of multiple speech feature pairs, resulting in a gradient bias-aware discriminative loss value. This allows for the discovery of highly confusing speech feature sample pairs with gradient conflicts, enhancing the model's ability to discriminate complex speech samples during training based on the gradient bias-aware discriminative loss value. Therefore, training the speech source tracing model based on the difference-guided adaptive angular interval loss value and the gradient bias-aware discriminative loss value improves the detection accuracy and robustness of the speech source tracing model in speech conversion scenarios, thereby enhancing the ability to identify the source identity and reducing security risks in speech-based identity authentication scenarios.

[0034] Figure 1 The diagram illustrates application scenarios of training methods, apparatus, devices, media, and program products for speech tracing models according to embodiments of this application.

[0035] like Figure 1 As shown, application scenario 100 according to this embodiment may include a first terminal device 101, a second terminal device 102, a third terminal device 103, a network 104, and a server 105. The network 104 serves as a medium for providing a communication link between the first terminal device 101, the second terminal device 102, the third terminal device 103, and the server 105. The network 104 may include various connection types, such as wired or wireless communication links, or fiber optic cables, etc.

[0036] Users can use the first terminal device 101, the second terminal device 102, and the third terminal device 103 to interact with the server 105 via the network 104 to receive or send messages, etc. Various communication client applications can be installed on the first terminal device 101, the second terminal device 102, and the third terminal device 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social media platform software, etc. (for example only).

[0037] The first terminal device 101, the second terminal device 102, and the third terminal device 103 can be various electronic devices with displays and support web browsing, including but not limited to smartphones, tablets, laptops, and desktop computers.

[0038] Server 105 can be a server that provides various services, such as a backend management server that supports websites browsed by users using the first terminal device 101, the second terminal device 102, and the third terminal device 103 (this is just an example). The backend management server can analyze and process data such as received user requests, and feed back the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal devices.

[0039] It should be noted that the training method for the speech tracing model provided in this application embodiment can generally be executed by server 105. Correspondingly, the training device for the speech tracing model provided in this application embodiment can generally be located in server 105. The training method for the speech tracing model provided in this application embodiment can also be executed by a server or server cluster that is different from server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or server 105. Correspondingly, the training device for the speech tracing model provided in this application embodiment can also be located in a server or server cluster that is different from server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or server 105.

[0040] It should be understood that Figure 1 The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included.

[0041] The following will be based on Figure 1 The described scene, through Figures 2-3 The training method of the speech source tracing model in the embodiments of the invention is described in detail.

[0042] Figure 2 A flowchart illustrating a training method for a speech source tracing model according to an embodiment of this application is shown.

[0043] like Figure 2 As shown, the training method of the speech source tracing model in this embodiment includes operations S210 to S270.

[0044] During operation of S210, the source speech training set is obtained.

[0045] In operation S220, the speech source tracing model is used to extract features from real speech samples and fake speech samples respectively, so as to obtain real speech features and fake speech features.

[0046] In operation S230, the difference-guided adaptive angle interval loss function is used to process the differences between real speech features and real speech labels, the differences between fake speech features and reference source speech features, and the differences between fake speech features and fake speech labels, so as to obtain the difference-guided adaptive angle interval loss value.

[0047] In operation S240, the binary contrast loss function is used to process the similarity between two speech features in multiple speech feature pairs, and the binary contrast loss value of each speech feature pair is obtained.

[0048] In operation S250, based on the binary contrast loss values ​​of multiple speech feature pairs, the gradients of the speech feature pairs are determined. The gradients of the speech feature pairs represent the changing trends of the binary contrast loss values.

[0049] In operating S260, the gradient bias-aware contrastive learning loss function is used to process the gradient angle between the gradients of multiple speech features and the binary contrast loss value of multiple speech features to obtain the gradient bias-aware discriminative loss value.

[0050] In operation S270, the speech source tracing model is trained based on the difference-guided adaptive angle interval loss value and the gradient deviation-aware discriminative loss value to obtain the trained speech source tracing model.

[0051] The source speech training set includes real speech samples and fake speech samples.

[0052] Real speech samples can be the speaker's actual, unprocessed speech. Real speech training samples can include multiple real speech samples from speakers with different voice source identities.

[0053] Fake speech samples can be generated through speech conversion technology. Specifically, for any real speech sample with the same source identity, a fake speech sample with the same source identity can be obtained by processing the real speech sample using speech conversion technology (such as Adversarial Glossy Augmentation and Improved Normalization for Voice Conversion, AGAIN-VC, etc.). Fake speech samples with different source identities can be obtained by processing real speech samples with different source identities using speech conversion technology. The aforementioned speech conversion technology is used to forge the timbre of the spoken speech, that is, to convert the speaker's timbre to the target timbre while preserving the speech content.

[0054] The speech source tracing model is used to extract features from real speech samples and fake speech samples respectively. It extracts the real speech features of the source identity of real speech samples and the fake speech features of the source identity of fake speech samples.

[0055] In this model, both real and fake speech features are 80-dimensional filter acoustic features (Filter Bank, Fbank). The speech source tracing model can be a network structure with 512-dimensional hidden layer features (Emphasized ChannelAttention, Propagation and Aggregation in Time Delay Neural Network, ECAPA-TDNN). Based on this ECAPA-TDNN network structure, feature embedding is performed on real and fake speech features respectively, mapping the real and fake speech training sample features of the above-mentioned 80-dimensional Fbank acoustic features to 192 dimensions, so as to facilitate subsequent processing using difference-guided adaptive angular interval loss function, binary contrast loss function, or gradient bias-aware contrastive learning loss function.

[0056] The reference sound source speech features are determined by calculating the mean vector of multiple real speech samples with the same sound source identity, as shown in formula (1).

[0057] (1);

[0058] in, The reference speech features representing the identity of the y-th sound source. The real speech embedding information of the k-th real speech sample representing the identity of the y-th sound source. This represents the number of real speech samples for the identity of the y-th sound source. This indicates a normalization operation.

[0059] In the process of constructing reference sound source speech features, a prototype vector for each sound source identity is constructed based solely on real speech samples to describe the central embedding distribution of the speaker's voice under no forgery perturbation for that sound source identity.

[0060] During the training of the speech source tracing model, real speech labels are added to real speech samples and fake speech labels are added to fake speech samples. These labels can be the name or number of the sound source.

[0061] The difference-guided adaptive margin loss (DGAML) function can be used to analyze the degree of deviation between the spurious speech features after speech conversion and the reference source speech features, as well as the accuracy of the speech source tracing model in classifying the source identity of real and spurious speech features, thereby obtaining the difference-guided adaptive margin loss value.

[0062] The two speech features in a speech feature pair include at least one of the following: two real speech features with the same source identity; a real speech feature and a fake speech feature with the same source identity; two real speech features with different source identities; and a real speech feature and a fake speech feature with different source identities.

[0063] The binary contrastive loss function can be a contrastive learning loss function, which can calculate the similarity between two speech features in a speech feature pair, and determine the binary contrastive loss value based on the similarity calculation results of multiple speech feature pairs.

[0064] The gradient of a speech feature pair is determined based on the binary contrast loss value, and the gradient of the speech feature pair represents the changing trend of the binary contrast loss value. Specifically, it can be obtained by backpropagating the binary contrast loss value and calculating its partial derivative.

[0065] Gradient-Divergence-Aware Contrastive Loss (GDACL) can be used to measure the consistency of the gradient directions of multiple speech features. The gradient direction can be the optimization direction of the speech source tracing model, i.e., a vector containing magnitude and direction. It is used to iteratively update the model parameters of the speech source tracing model along the negative gradient direction, thereby reducing the value of the gradient-divergence-aware contrastive learning loss function.

[0066] The model parameters of the speech source tracing model are adjusted based on the difference-guided adaptive angle interval loss value and the gradient deviation-aware discriminative loss value, and the trained speech source tracing model is obtained based on the adjusted model parameters.

[0067] According to embodiments of this application, a speech source tracing model is used to extract features from real and fake speech samples respectively, obtaining real speech features and fake speech features. This allows the influence of speech samples with different source identities on the speech source tracing model during training. Furthermore, a difference-guided adaptive angle interval loss function is used to handle the differences between real speech features and real speech labels, the differences between fake speech features and reference source speech features, and the differences between fake speech features and fake speech labels, obtaining a difference-guided adaptive angle interval loss value. This allows for dynamic constraint of the difference-guided adaptive angle interval loss value by calculating the difference between fake speech features and reference source speech features, effectively suppressing speech source tracing errors. The shift in speech features caused by speech conversion compresses the intra-class distance between genuine and fake speech features with the same source identity in the feature space. Furthermore, a binary contrastive loss function is used to process the similarity between two speech features in multiple speech feature pairs, obtaining binary contrastive loss values ​​for each pair. The gradients of these speech feature pairs are then derived based on these binary contrastive loss values. A gradient-bias-aware contrastive learning loss function is then used to process the differences between the gradients of multiple speech feature pairs, resulting in a gradient-bias-aware discriminative loss value. This allows for the discovery of highly confusing speech feature sample pairs with gradient conflicts, enhancing the model's ability to discriminate complex speech samples during training based on the gradient-bias-aware discriminative loss value. Therefore, training the speech source tracing model based on the difference-guided adaptive angular interval loss value and the gradient-bias-aware discriminative loss value improves the detection accuracy and robustness of the speech source tracing model in speech conversion scenarios, thereby enhancing the ability to identify the source identity and reducing security risks in speech-based identity authentication scenarios.

[0068] According to embodiments of this application, a difference-guided adaptive angle interval loss function is used to process the differences between real speech features and real speech labels, the differences between fake speech features and reference source speech features, and the differences between fake speech features and fake speech labels, to obtain a difference-guided adaptive angle interval loss value. This includes: calculating the similarity between reference source speech features and fake speech features to obtain a cosine similarity, where the cosine similarity characterizes the degree of damage to real speech samples during the process of converting fake speech samples into fake speech samples; obtaining target angle intervals corresponding to real speech features and fake speech features respectively based on preset angle intervals and cosine similarity; obtaining a real speech loss value based on the target angle interval corresponding to the real speech features and the real similarity, where the real similarity characterizes the similarity between real speech features and real speech labels; obtaining a fake speech loss value based on the target angle interval corresponding to the fake speech features and the fake similarity, where the fake similarity characterizes the similarity between fake speech features and fake speech labels; and determining the difference-guided adaptive angle interval loss value based on the real speech loss value and the fake speech loss value.

[0069] The similarity calculation between the reference sound source speech features and the fake speech features can be achieved by calculating the cosine distance between the reference sound source speech features and the fake speech features as the cosine similarity, as shown in formula (2).

[0070] (2);

[0071] in, Let represent the cosine similarity between the i-th spurious speech feature and the reference source speech feature. This represents the i-th fake speech feature. This indicates the inner product operation.

[0072] The target angle interval is used to reflect the difference between the predicted sound source identity predicted by the speech source tracing model and the real or fake speech label, and can also suppress the feature shift caused by speech conversion.

[0073] Based on the preset angular interval and cosine similarity, the target angular interval corresponding to the real speech features and the fake speech features respectively is obtained. This can be done by using the preset angular interval as the target angular interval corresponding to the real speech features, or by fusing the preset angular interval and cosine similarity and using the fused result as the target angular interval corresponding to the fake speech features.

[0074] The real speech loss value is obtained based on the target angle interval and true similarity corresponding to the real speech features. The real speech loss value can be expressed as: , The scaling factor is denoted by e, where e is the base. The true similarity is the ratio between the i-th true speech feature and the true speech label. Similarity between them For the i-th real speech feature and excluding the real speech label Other real voice tags Similarity between them The target angle interval corresponds to the real speech features.

[0075] Similar to obtaining the real speech loss value, the fake speech loss value is obtained based on the target angle interval and fake similarity corresponding to the fake speech features. , (i.e., true similarity) is replaced with false similarity, that is, the similarity between the i-th false speech feature and the false speech label, and the... (i.e., the target angle interval corresponding to the real speech features) is replaced with the target angle interval corresponding to the fake speech features, thereby based on the above formula (i.e. By performing calculations, the value of false speech loss can be obtained.

[0076] The adaptive angle interval loss value for difference guidance can be determined based on the real speech loss value and the fake speech loss value. This can be achieved by summing the real speech loss value and the fake speech loss value and taking the mean of the summation as the adaptive angle interval loss value for difference guidance. Specifically, it can be shown in formula (3).

[0077] (3);

[0078] in, The adaptive angular interval loss value is guided by the difference. This represents the total number of real speech features and fake speech features.

[0079] According to embodiments of this application, the difference-guided adaptive angle interval loss function adaptively adjusts the angle interval by calculating the cosine distance between the reference source speech features and the fake speech features to obtain the target angle interval. This adds different interval constraints to the real speech features and the fake speech features, so as to apply a stronger angle penalty to the fake speech features with excessive differences between the reference source speech features and the fake speech features.

[0080] According to an embodiment of this application, the preset angle interval includes a first preset angle interval and a second preset angle interval, wherein the second preset angle interval is greater than the first preset angle interval; wherein, obtaining the target angle interval corresponding to the real speech feature and the fake speech feature respectively based on the preset angle interval and cosine similarity includes: determining the first preset angle interval as the target angle interval corresponding to the real speech feature; processing the ratio of cosine similarity to a preset value using a truncation function to obtain a first parameter, wherein the truncation function is used to constrain the ratio of cosine similarity to the preset value within a preset numerical range; obtaining a second parameter based on the difference between the second preset angle interval and the first preset angle interval; and determining the target angle interval corresponding to the fake speech feature based on the product of the first parameter and the second parameter and the sum of the first preset angle interval.

[0081] The preset angle interval includes a first preset angle interval and a second preset angle interval, wherein the first preset angle interval can be a preset minimum angle interval. The second preset angle interval can be the preset maximum angle interval. .

[0082] The first preset angle interval The target angular interval is determined to correspond to the real speech features.

[0083] The truncation function is used to constrain the ratio of cosine similarity to a preset value within a preset numerical range. For example, the first parameter can be expressed as... ,in, The truncation function is set to the default value. The ratio of cosine similarity to the preset value Constrained within (0,1), when the ratio of cosine similarity to a preset value... If the value is less than 0, determine The value of is 0, when the ratio of cosine similarity to the preset value is If the value is greater than 0, then determine. The value is 1.

[0084] The second parameter can be expressed as That is, the second preset angle interval and the first preset angle interval The difference.

[0085] The target angle interval corresponding to the false speech feature is determined based on the product of the first parameter and the second parameter and the sum of the first preset angle interval.

[0086] Based on the preset angular interval and cosine similarity, the target angular intervals corresponding to the real speech features and fake speech features are obtained respectively. It can be shown in formula (4).

[0087] (4);

[0088] in, Represents the characteristics of real speech. This indicates the characteristics of fake speech.

[0089] As can be seen from formula (4), when the ratio of cosine similarity to the preset value is... When the value is less than 0, that is, when the deviation between the false speech features and the reference speech features is small, the target angle interval corresponding to the false speech features is the first preset angle interval. .

[0090] According to embodiments of this application, the difference between the reference sound source speech features and the fake speech features is determined by cosine similarity. Based on the degree to which the fake speech features deviate from the reference sound source speech features, different target angle intervals are determined to dynamically apply angle penalties, so that the model compresses the intra-class distribution of the same sound source in the feature space, thereby effectively suppressing the feature shift caused by speech conversion and enhancing the robustness of the model in distinguishing the source speaker's identity.

[0091] According to embodiments of this application, a binary contrast loss function is used to process the similarity between two speech features in multiple speech feature pairs to obtain binary contrast loss values ​​for each speech feature pair. This includes: for speech feature pairs with two real speech features sharing the same source identity, fusing the similarity between the two real speech features, the first label value corresponding to the first feature pair label, and a preset coefficient to obtain a binary contrast loss value; for speech feature pairs with real and fake speech features sharing the same source identity, fusing the similarity between the real and fake speech features, the first label value corresponding to the first feature pair label, and a preset coefficient to obtain a binary contrast loss value; for speech feature pairs with two real speech features having different source identities, fusing the similarity between the two real speech features, the second label value corresponding to the second feature pair label, and a preset coefficient to obtain a binary contrast loss value; and for speech feature pairs with real and fake speech features having different source identities, fusing the similarity between the real and fake speech features, the second label value corresponding to the second feature pair label, and a preset coefficient to obtain a binary contrast loss value.

[0092] A speech feature pair consisting of two real speech features with the same source identity, and a speech feature pair consisting of a real speech feature and a fake speech feature with the same source identity, are labeled with a first feature pair label, wherein the first label value corresponding to the first feature pair label can be +1.

[0093] A speech feature pair consisting of two real speech features with different source identities, and a speech feature pair consisting of real speech features and fake speech features with different source identities, have a second feature pair label, wherein the second label value corresponding to the second feature pair label can be -1.

[0094] For two real speech feature pairs with the same sound source identity, the similarity between the two real speech features, the first label value corresponding to the first feature pair label and the preset coefficient are fused to obtain the binary contrast loss value, as shown in formula (5).

[0095] (5);

[0096] in, This represents the first label value corresponding to the first feature pair label. and This represents the preset coefficient, where, Indicates the temperature coefficient. This represents the bias coefficient. This represents the similarity between two real speech features that share the same source identity. This represents the binary contrast loss value.

[0097] For speech feature pairs with the same source identity, real speech features and fake speech features, the binary contrast loss value is obtained based on the above formula (5), where, The similarity between real and fake speech features that share the same voice source identity.

[0098] For speech feature pairs with two real speech features having different source identities, the binary contrast loss value is obtained based on the above formula (5), where, This represents the second label value corresponding to the second feature pair label. It represents the similarity between two real speech features with different source identities.

[0099] For speech feature pairs with real and fake speech features having different source identities, the binary contrast loss value is obtained based on the above formula (5), where, This represents the second label value corresponding to the second feature pair label. This represents the similarity between real speech features and fake speech features with different voice source identities.

[0100] According to embodiments of this application, by adding a first feature pair label and a second feature pair label to different speech feature pairs, it is possible to analyze the similarity between two speech features in different speech feature pairs and the gradient between two speech features, thereby facilitating the analysis of the speech tracing model's ability to recognize different speech samples.

[0101] According to an embodiment of this application, a speech feature pair with a first feature pair label is used as a reference speech feature pair; wherein, a gradient bias-aware contrastive learning loss function is used to process the gradient angle between the gradients of multiple speech feature pairs and the binary contrastive loss value of each of the multiple speech feature pairs to obtain a gradient bias-aware discriminative loss value, including: for any speech feature pair among the multiple speech feature pairs, determining the gradient bias value between the speech feature pair and the reference gradient of the reference speech feature pair based on the gradient angle between the gradient of the speech feature pair and the reference gradient of the reference speech feature pair; and determining the gradient bias-aware discriminative loss value based on the gradient bias value and the binary contrastive loss value of each of the multiple speech feature pairs.

[0102] The gradient angle between the gradient of a speech feature pair and the reference gradient of a reference speech feature pair can be the angle between the gradient direction of the gradient and the reference gradient direction of the reference gradient. The gradient direction of the gradient can be obtained by normalization, as shown in formula (6).

[0103] (6);

[0104] in, Indicates the gradient direction. Represents the gradient. It represents a tiny quantity that prevents the denominator from being zero.

[0105] The reference speech feature pair can be a speech feature pair with a first feature pair label, wherein the reference gradient direction of the reference speech feature pair indicates the desired model optimization direction.

[0106] Reference gradient direction It can be obtained based on the above formula (6).

[0107] After obtaining the gradient direction of the gradient and the reference gradient direction of the reference gradient, the gradient angle is obtained based on formula (7) and determined as the gradient deviation value between the speech feature pair and the reference speech feature pair.

[0108] (7);

[0109] in, This is the gradient angle (i.e., the gradient deviation value). The larger, the more it indicates The direction of model optimization for indicated speech feature pairs is... The more intense the conflict in the model optimization direction of the indicated reference speech feature pair, the higher the confusion of the speech feature pair (e.g., "real speech features of source B" and "fake speech features disguised as source A").

[0110] After determining the gradient bias values ​​and binary contrast loss values ​​of multiple speech feature pairs, the gradient bias values ​​and binary contrast loss values ​​of multiple speech feature pairs are fused to obtain the discriminative loss value of gradient bias perception.

[0111] According to embodiments of this application, by analyzing the consistency between the gradient direction of the gradient and the reference gradient direction of the reference gradient, it is possible to determine whether the gradient direction of the gradient deviates from the reference gradient direction, determine the degree of confusion and optimization conflict of the speech feature pair, identify complex fake speech samples that are difficult to distinguish by similar distance alone, thereby providing more targeted supervision signals for training and improving the training accuracy of the speech source tracing model during training.

[0112] According to embodiments of this application, determining a gradient deviation-aware discriminative loss value based on the gradient deviation values ​​and binary contrast loss values ​​of multiple speech feature pairs includes: determining a gradient deviation mean based on the mean of the multiple gradient deviation values; determining a gradient deviation coefficient for each of the multiple speech feature pairs based on the ratio of the gradient deviation values ​​to the mean gradient deviation value; fusing the gradient deviation coefficients and binary contrast loss values ​​of the multiple speech feature pairs to determine a gradient loss value for each of the multiple speech feature pairs; and determining a gradient deviation-aware discriminative loss value based on the mean of the gradient loss values ​​of the multiple speech feature pairs.

[0113] The mean gradient deviation is determined by the average of multiple gradient deviation values, and can be expressed as follows: ,in, This represents the total number of gradient bias values. This represents multiple gradient bias values. Used to control the amplification level of highly confusing speech feature pairs.

[0114] The gradient deviation coefficients of multiple speech features can be determined based on the ratio of their respective gradient deviation values ​​to the mean gradient deviation value, as shown in formula (8).

[0115] (8);

[0116] in, Let be the gradient bias coefficient of the ij-th speech feature pair. Let be the gradient bias value of the ij-th speech feature pair. Used to control the amplification level of highly confusing speech feature pairs. Used to adjust the contribution of speech feature pairs to the total loss, speech feature pairs The larger, The larger.

[0117] Fusing the gradient bias coefficients and binary contrast loss values ​​of multiple speech features can be achieved by multiplying the gradient bias coefficients and binary contrast loss values ​​to obtain the gradient loss values ​​for each speech feature. .

[0118] The discrimination loss value for gradient bias perception can be determined based on the average gradient loss value of multiple speech feature pairs, as shown in formula (9).

[0119] (9);

[0120] in, The discriminant loss value is for gradient bias perception.

[0121] According to embodiments of this application, by determining the gradient deviation coefficients of multiple speech features based on the ratio of their respective gradient deviation values ​​to the mean gradient deviation, the contribution of different speech features to the discrimination loss value of gradient deviation perception can be adjusted, so that speech features with larger gradient deviation values ​​can make a higher contribution to the discrimination loss value of gradient deviation perception.

[0122] According to an embodiment of this application, determining a discrimination loss value for gradient deviation perception based on at least one of the gradient deviation values ​​and binary contrast loss values ​​of multiple speech features includes: determining multiple target gradient deviation values ​​whose gradient deviation values ​​satisfy preset screening conditions; determining target binary contrast loss values ​​corresponding to the multiple target gradient deviation values; and determining a discrimination loss value for gradient deviation perception based on the mean of the multiple target binary contrast loss values.

[0123] The preset filtering condition can be to sort the gradient deviation values ​​of multiple speech feature pairs in descending order, and determine the gradient deviation values ​​of the top m speech feature pairs in the sort as the target gradient deviation values.

[0124] Determining the target binary contrast loss value corresponding to multiple target gradient deviation values ​​can be achieved by determining the target binary contrast loss value for each of the multiple target speech feature pairs based on the target speech feature pairs corresponding to the multiple target gradient deviation values ​​after determining the multiple target gradient deviation values.

[0125] The discriminant loss value for gradient bias perception is determined based on the mean of the binary contrast loss values ​​of multiple targets.

[0126] According to an embodiment of this application, by determining multiple target gradient deviation values ​​with large gradient deviation values, the mean of the target binary contrast loss value corresponding to the multiple target gradient deviation values ​​is determined, thereby obtaining a gradient deviation-aware discriminative loss value. This makes the obtained gradient deviation-aware discriminative loss value focus on highly confusing speech feature pairs with large gradient deviation values ​​that are prone to optimization conflicts in the training of the speech source tracing model, thereby solving the cross-speaker confusion problem and enhancing the model's ability to discriminate complex forgery attacks.

[0127] According to an embodiment of this application, a speech tracing model is trained based on a difference-guided adaptive angle interval loss value and a gradient deviation-aware discriminant loss value to obtain a trained speech tracing model. This includes: fusing a balance coefficient and a gradient deviation-aware discriminant loss value to obtain a target gradient deviation-aware discriminant loss value; determining a target loss value based on the sum of the difference-guided adaptive angle interval loss value and the target gradient deviation-aware discriminant loss value; and optimizing the model parameters of the speech tracing model based on the target loss value to obtain a trained speech tracing model.

[0128] The balancing coefficient is used to optimize the contribution of the gradient bias sensing discriminant loss value to obtain the target gradient bias sensing discriminant loss value.

[0129] The target loss value is determined by the sum of the difference-guided adaptive angle interval loss value and the target gradient deviation perception discrimination loss value, as shown in formula (10).

[0130] (10);

[0131] Among them, the difference-guided adaptive angle interval loss value It is responsible for adaptive shrinking and inter-class enhancement of fake speech samples through dynamic target angle intervals, and for discriminative loss values ​​that are aware of gradient bias. Responsible for identifying and enhancing the discrimination of highly confusing and difficult speech samples through gradient bias values, and balancing coefficients. The relative contributions of the difference-guided adaptive angle-interval loss and the gradient bias-aware discriminative loss during training are used to balance the differences.

[0132] If the target loss value does not converge, the parameters of the speech source tracing model are adjusted until the total loss value converges, thus completing the training of the speech source tracing model.

[0133] According to embodiments of this application, by adjusting the contribution of the discriminative loss value perceived by gradient bias to the target loss value through a balance coefficient, the speech source tracing model's ability to classify and recognize speech samples can be balanced during training, thereby improving the training accuracy of the speech source tracing model.

[0134] According to an embodiment of this application, a speech tracing model is used to extract features from real speech samples and fake speech samples to obtain real speech features and fake speech features, including: preprocessing real speech samples and fake speech samples to obtain real speech samples and fake speech samples with the same speech duration; and extracting features from real speech samples and fake speech samples with the same speech duration to obtain real speech features and fake speech features.

[0135] Preprocessing real and fake speech samples separately can be performed by cropping the real and fake speech samples to obtain real and fake speech samples with speech duration.

[0136] Feature extraction can be performed on real speech samples and fake speech samples with the same speech duration, respectively. This can be done based on an attention mechanism to extract features from real speech samples and fake speech samples with the same speech duration, thus obtaining real speech features and fake speech features.

[0137] According to the embodiments of this application, real speech samples and fake speech samples are cropped respectively to make the speech lengths of real speech samples and fake speech samples consistent, and to align real speech samples and fake speech samples to facilitate subsequent feature extraction.

[0138] Figure 3 A schematic diagram of a training method for a speech source tracing model according to an embodiment of this application is shown.

[0139] like Figure 3 As shown, real speech sample 301 and fake speech sample 302 are input into the speech source tracing model. Real speech sample 301 and fake speech sample 302 are respectively cropped to obtain real speech sample segment 303 and fake speech sample segment 304. Further, an extractor is used to extract features from real speech sample segment 303 and fake speech sample segment 304 respectively, obtaining real speech features 305 and fake speech features 306. Then, an adaptive feature clustering module is used to process real speech features 305 and fake speech features 306 to obtain a difference-guided adaptive angle interval loss value 307. A confusion sample discrimination module is then used to process real speech features 305 and fake speech features 306 to obtain a gradient bias-aware discrimination loss value 308. Finally, a target loss value 309 is obtained based on the difference-guided adaptive angle interval loss value 307 and the gradient bias-aware discrimination loss value 308. The speech source tracing model is then tuned based on the target loss value 309.

[0140] According to embodiments of this application, in order to evaluate the effectiveness of this application, a 7945-hour fake speech corpus was synthesized based on three speech conversion algorithms (AGAIN-VC, VQMIVC, and FreeVC) and three corpora (Voxceleb1, Voxceleb2, and LibriSpeech). Speakers in Voxceleb1 and Voxceleb2 were considered real speakers, while speakers in LibriSpeech were considered fake speakers.

[0141] In the training dataset, VoxCeleb2 was used as the speech data of real speakers, and speakers from LibriSpeech were selected as fake speakers. Considering the large size of the VoxCeleb2 dataset, one-tenth of the samples were randomly selected, containing 109,200 speech data points from 630 real speakers. Simultaneously, 2,400 fake speakers were selected from the LibriSpeech dataset, totaling 282,610 speech data points. In the test dataset, this embodiment selected 1,251 real speakers from VoxCeleb1, totaling 153,516 speech data points. The LibriSpeech dataset contains 2,484 speakers; after allocating the 2,400 speakers to the training set, the remaining 84 speakers were used to construct the test dataset, containing 9,757 speech data points.

[0142] For each real speaker's real speech sample, three fake speaker speech samples are randomly selected, and a speech conversion algorithm is used to convert these three speech samples into the real speaker's voice, thus generating three fake speech samples. Based on this method, for each speech conversion algorithm, this embodiment generated 327,600 fake speech samples (109,200 × 3) in the training dataset and 460,548 fake speech samples (153,516 × 3) in the test dataset. Since this embodiment uses three different speech conversion algorithms, the final training dataset contains 982,800 speech samples (327,600 × 3), and the test dataset contains 1,381,644 speech samples (460,548 × 3).

[0143] Furthermore, to evaluate the new model performance, this application designs three test lists of different difficulties—S2B-O, S2B-E, and S2B-H—for speech tracing tasks, based on the three test lists Vox-O, Vox-E, and Vox-H provided in VoxCeleb1. S2B stands for "Spoofed to Bonafide," representing the matching relationship between fake and real speech samples. Specifically, this application uses fake speech samples as the registration set and real speech samples as the test set. For each speech sample in the registration set, eight sample pairs are repeatedly generated, including four pairs of positive samples and four pairs of negative samples. Positive sample pairs are constructed by randomly selecting a real speech sample from the test set that comes from the same speaker as the registered fake speech sample; negative sample pairs are constructed by randomly selecting a real speech sample from the test set that comes from a different speaker than the registered fake speech sample.

[0144] The model performance was evaluated using equal error rate (EER) and minimum detection cost function (minDCF). It was then compared with an automatic speaker verification model trained solely on real speech (baseline-1) and a fake speaker attribution algorithm employing a standard loss function (without adaptive margin and gradient-aware mechanisms) (baseline-2).

[0145] Table 1 shows the performance of different methods in different speech conversion scenarios.

[0146] Table 1

[0147]

[0148] The results show that, compared with the two baseline models, the proposed dual-branch collaborative training method can effectively reduce the error rate and minimize the detection cost of the speech source tracing model in different speech conversion scenarios. Specifically, the results of baseline-1 indicate that traditional automatic speaker verification models are not suitable for fake speaker source tracing tasks. Due to the lack of training on forged speech, the model struggles to overcome the feature shift caused by speech conversion and cannot establish an effective mapping between real and fake speech. After introducing an embedding alignment strategy (baseline-2), the model learns the commonalities between real and fake speech, gaining a certain matching ability, and its performance is significantly improved compared to baseline-1. However, due to the use of a fixed classification boundary loss function, its discrimination ability remains limited when facing severely deviated forged samples or highly confused samples. This application introduces a difference-guided adaptive angular interval loss value... The classification boundary is dynamically adjusted based on the degree of deviation of the forged sample from the speaker prototype, effectively suppressing feature shift caused by speech conversion and significantly reducing the intra-class distance of homologous samples. Meanwhile, the gradient bias-aware discriminative loss value... By analyzing gradient direction conflicts, this application accurately identifies and enhances the learning of highly confusing and difficult samples, thus solving the cross-speaker confusion problem. This enables the application to achieve state-of-the-art performance across all metrics.

[0149] According to an embodiment of this application, in order to further verify the effectiveness of this application, fake speech generated by different speech conversion algorithms is mixed, and the performance under different conditions is evaluated.

[0150] Table 2 shows the performance of different methods after mixing the fake speech generated by FreeVC with the fake speech generated by AGAIN-VC.

[0151] Table 2

[0152]

[0153] Table 3 shows the performance of different methods after mixing the fake speech generated by FreeVC with the fake speech generated by VQMIVC.

[0154] Table 3

[0155]

[0156] Experimental results demonstrate that even under complex hybrid attack scenarios, this application still exhibits state-of-the-art performance, showcasing its robustness to complex environments. In summary, the difference-guided adaptive angle interval loss function and gradient bias-aware discriminative loss function proposed in this application, through the collaborative optimization of feature clustering and sample pair discrimination, can effectively improve the model's ability to trace the true identity of the source speaker.

[0157] Figure 4 A flowchart of a speech tracing method according to an embodiment of this application is shown.

[0158] This application provides a method for tracing the source of speech, such as... Figure 4 As shown, the voice tracing method includes operations S410-S420.

[0159] When operating the S410, acquire the actual voice and the voice to be traced.

[0160] In operation S420, the speech source tracing model is used to process the real speech and the speech to be traced to determine the target speech in the speech to be traced that comes from the same sound source as the real speech. The speech source tracing model is trained based on the training method of the above speech source tracing model.

[0161] According to an embodiment of this application, a speech tracing model is used to process real speech and speech to be traced to obtain a similarity value between real speech and speech to be traced.

[0162] According to an embodiment of this application, a similarity threshold is set. If the similarity value between the real speech and the speech to be traced meets the similarity threshold, it proves that the real speech and the speech to be traced come from the same sound source. If the similarity threshold is not met, it proves that the real speech and the speech to be traced are different sound sources.

[0163] According to embodiments of this application, based on the above comparison of similarity values ​​and similarity thresholds, it is possible to determine the target speech from the same sound source as the real speech in the speech to be traced.

[0164] According to the embodiments of this application, the speech source tracing model of this application can accurately identify target speech that comes from the same sound source as real speech, avoid the problem that fake speech cannot be identified, and prevent the impact of fake speech.

[0165] Based on the training method of the above-mentioned speech source tracing model, this application also provides a training device for the speech source tracing model. The following will combine... Figure 5 The device is described in detail.

[0166] Figure 5 A structural block diagram of a training apparatus for a speech tracing model according to an embodiment of this application is shown.

[0167] like Figure 5 As shown, the training device 500 for the speech source tracing model in this embodiment includes a first acquisition module 510, a first extraction module 520, a first processing module 530, a second processing module 540, a first determination module 550, a third processing module 560, and a first training module 570.

[0168] The first acquisition module 510 is used to acquire a source-tracing speech training set, wherein the source-tracing speech training set includes real speech samples and fake speech samples. In one embodiment, the first acquisition module 510 can be used to perform the operation S210 described above, which will not be repeated here.

[0169] The first extraction module 520 is used to extract features from real speech samples and fake speech samples respectively using a speech source tracing model to obtain real speech features and fake speech features. In one embodiment, the first extraction module 520 can be used to perform the operation S220 described above, which will not be repeated here.

[0170] The first processing module 530 is used to process the differences between real speech features and real speech labels, the differences between fake speech features and reference source speech features, and the differences between fake speech features and fake speech labels using a difference-guided adaptive angle interval loss function, to obtain the difference-guided adaptive angle interval loss value. The reference source speech features are determined by calculating the vector mean of multiple real speech samples with the same source identity. In one embodiment, the first processing module 530 can be used to perform the operation S230 described above, which will not be repeated here.

[0171] The second processing module 540 is used to process the similarity between two speech features in multiple speech feature pairs using a binary contrastive loss function, obtaining the binary contrastive loss value for each of the multiple speech feature pairs. The two speech features in a speech feature pair include at least one of the following: two real speech features with the same source identity; a real speech feature and a fake speech feature with the same source identity; two real speech features with different source identities; or a real speech feature and a fake speech feature with different source identities. In one embodiment, the second processing module 540 can be used to perform the operation S240 described above, which will not be repeated here.

[0172] The first determining module 550 is used to determine the gradient of each speech feature pair based on the respective binary contrast loss value of multiple speech feature pairs. The gradient of the speech feature pair represents the changing trend of the binary contrast loss value. In one embodiment, the first determining module 550 is used to perform the operation S250 described above, which will not be repeated here.

[0173] The third processing module 560 is used to process the gradient angle between the gradients of multiple speech features and the binary contrast loss values ​​of multiple speech features using a gradient bias-aware contrastive learning loss function, to obtain a gradient bias-aware discriminative loss value. In one embodiment, the third processing module 560 can be used to perform the operation S260 described above, which will not be repeated here.

[0174] The first training module 570 is used to train the speech source tracing model based on the difference-guided adaptive angle interval loss value and the gradient deviation-aware discriminative loss value, thereby obtaining the trained speech source tracing model. In one embodiment, the first training module 570 can be used to perform the operation S270 described above, which will not be repeated here.

[0175] According to embodiments of this application, a speech source tracing model is used to extract features from real and fake speech samples respectively, obtaining real speech features and fake speech features. This allows the influence of speech samples with different source identities on the speech source tracing model during training. Furthermore, a difference-guided adaptive angle interval loss function is used to handle the differences between real speech features and real speech labels, the differences between fake speech features and reference source speech features, and the differences between fake speech features and fake speech labels, obtaining a difference-guided adaptive angle interval loss value. This allows for dynamic constraint of the difference-guided adaptive angle interval loss value by calculating the difference between fake speech features and reference source speech features, effectively suppressing speech source tracing errors. The shift in speech features caused by speech conversion compresses the intra-class distance between genuine and fake speech features with the same source identity in the feature space. Furthermore, a binary contrastive loss function is used to process the similarity between two speech features in multiple speech feature pairs, obtaining binary contrastive loss values ​​for each pair. The gradients of these speech feature pairs are then derived based on these binary contrastive loss values. A gradient-bias-aware contrastive learning loss function is then used to process the differences between the gradients of multiple speech feature pairs, resulting in a gradient-bias-aware discriminative loss value. This allows for the discovery of highly confusing speech feature sample pairs with gradient conflicts, enhancing the model's ability to discriminate complex speech samples during training based on the gradient-bias-aware discriminative loss value. Therefore, training the speech source tracing model based on the difference-guided adaptive angular interval loss value and the gradient-bias-aware discriminative loss value improves the detection accuracy and robustness of the speech source tracing model in speech conversion scenarios, thereby enhancing the ability to identify the source identity and reducing security risks in speech-based identity authentication scenarios.

[0176] According to an embodiment of this application, the first processing module 530 includes a first calculation submodule, a first obtaining submodule, a second obtaining submodule, a third obtaining submodule, and a first determining submodule.

[0177] The first calculation submodule is used to calculate the similarity between the speech features of the reference sound source and the fake speech features to obtain the cosine similarity. The cosine similarity represents the degree of damage to the real speech sample during the process of converting the fake speech sample into a fake speech sample.

[0178] The first submodule is used to obtain the target angle intervals corresponding to the real speech features and the fake speech features respectively, based on the preset angle intervals and cosine similarity.

[0179] The second submodule is used to obtain the real speech loss value based on the target angle interval and real similarity corresponding to the real speech features. The real similarity represents the similarity between the real speech features and the real speech labels.

[0180] The third submodule is used to obtain the false speech loss value based on the target angle interval and false similarity corresponding to the false speech features. The false similarity represents the similarity between the false speech features and the false speech labels.

[0181] The first determination submodule is used to determine the difference-guided adaptive angle interval loss value based on the real speech loss value and the spurious speech loss value.

[0182] According to an embodiment of this application, the preset angle interval includes a first preset angle interval and a second preset angle interval, wherein the second preset angle interval is greater than the first preset angle interval.

[0183] The first obtaining submodule includes a first determining unit, a first processing unit, a first obtaining unit, and a second determining unit.

[0184] The first determining unit is used to determine the first preset angle interval as the target angle interval corresponding to the real speech features.

[0185] The first processing unit is used to process the ratio of cosine similarity to a preset value using a truncation function to obtain a first parameter. The truncation function is used to constrain the ratio of cosine similarity to a preset value within a preset range.

[0186] The first obtaining unit is used to obtain the second parameter based on the difference between the second preset angle interval and the first preset angle interval.

[0187] The second determining unit is used to determine the target angle interval corresponding to the false speech feature based on the product of the first parameter and the second parameter and the sum of the first preset angle interval.

[0188] According to embodiments of this application, a voice feature pair consisting of two real voice features with the same voice source identity, and a voice feature pair consisting of a real voice feature and a fake voice feature with the same voice source identity, are labeled with a first feature pair label; a voice feature pair consisting of two real voice features with different voice source identities, and a voice feature pair consisting of a real voice feature and a fake voice feature with different voice source identities, are labeled with a second feature pair label.

[0189] The second processing module 540 includes a first fusion submodule, a second fusion submodule, a third fusion submodule, and a fourth fusion submodule.

[0190] The first fusion submodule is used to fuse the similarity between two real speech features, the first label value corresponding to the first feature pair label, and the preset coefficient for two real speech feature pairs with the same sound source identity, to obtain a binary contrast loss value.

[0191] The second fusion submodule is used to fuse real speech features and fake speech features with the same sound source identity into a binary contrast loss value by fusing the similarity between real speech features and fake speech features, the first label value corresponding to the first feature pair label, and a preset coefficient.

[0192] The third fusion submodule is used to fuse the similarity between two real speech features, the second label value corresponding to the second feature pair label, and the preset coefficient for two real speech feature pairs with different sound source identities, to obtain a binary contrast loss value.

[0193] The fourth fusion submodule is used to fuse real speech features and fake speech features with different sound source identities into a binary contrast loss value by fusing the similarity between real speech features and fake speech features, the second label value corresponding to the second feature pair label, and the preset coefficient.

[0194] According to an embodiment of this application, a speech feature pair having a first feature pair label is used as a reference speech feature pair.

[0195] The third processing module 560 includes a second determining submodule and a third determining submodule.

[0196] The second determining submodule is used to determine the gradient deviation value between a speech feature pair and a reference speech feature pair for any speech feature pair among multiple speech feature pairs, based on the gradient angle between the gradient of the speech feature pair and the reference gradient of the reference speech feature pair.

[0197] The third determination submodule is used to determine the discrimination loss value for gradient deviation perception based on the gradient deviation values ​​and binary contrast loss values ​​of multiple speech features.

[0198] According to an embodiment of this application, the third determining submodule includes a third determining unit, a fourth determining unit, a fifth determining unit, and a sixth determining unit.

[0199] The third determining unit is used to determine the average gradient deviation based on the average of multiple gradient deviation values.

[0200] The fourth determining unit is used to determine the gradient deviation coefficients of multiple speech features based on the ratio of their respective gradient deviation values ​​to the mean gradient deviation value.

[0201] The fifth determining unit is used to fuse the gradient bias coefficients and binary contrast loss values ​​of multiple speech features to determine the gradient loss values ​​of multiple speech features.

[0202] The sixth determining unit is used to determine the discrimination loss value for gradient bias perception based on the mean of the gradient loss values ​​of multiple speech feature pairs.

[0203] According to an embodiment of this application, the third determining submodule includes a seventh determining unit, an eighth determining unit, and a ninth determining unit.

[0204] The seventh determining unit is used to determine multiple target gradient deviation values ​​that satisfy preset screening conditions.

[0205] The eighth determination unit is used to determine the target binary contrast loss value corresponding to the gradient deviation values ​​of multiple targets.

[0206] The ninth determination unit is used to determine the discrimination loss value for gradient bias sensing based on the mean of multiple target binary comparison loss values.

[0207] According to an embodiment of this application, the first training module 570 includes a fifth fusion submodule, a fourth determination submodule, and a first optimization submodule.

[0208] The fifth fusion submodule is used to fuse the balance coefficient and the discriminant loss value of gradient bias perception to obtain the discriminant loss value of target gradient bias perception.

[0209] The fourth determination submodule is used to determine the target loss value based on the sum of the difference-guided adaptive angle interval loss value and the target gradient deviation-aware discrimination loss value.

[0210] The first optimization submodule is used to optimize the model parameters of the speech source tracing model based on the target loss value, so as to obtain the trained speech source tracing model.

[0211] According to an embodiment of this application, the first extraction module 520 includes a first processing submodule and a first extraction submodule.

[0212] The first processing submodule is used to preprocess real speech samples and fake speech samples respectively to obtain real speech samples and fake speech samples with the same speech duration.

[0213] The first extraction submodule is used to extract features from real speech samples and fake speech samples with the same speech duration, respectively, to obtain real speech features and fake speech features.

[0214] According to embodiments of this application, any multiple modules among the first acquisition module 510, first extraction module 520, first processing module 530, second processing module 540, third processing module 560, and first training module 570 can be combined into one module, or any one of these modules can be split into multiple modules. Alternatively, at least some of the functions of one or more of these modules can be combined with at least some of the functions of other modules and implemented in one module. According to embodiments of this application, at least one of the first acquisition module 510, first extraction module 520, first processing module 530, second processing module 540, third processing module 560, and first training module 570 can be at least partially implemented as hardware circuitry, such as a field-programmable gate array (FPGA), a programmable logic array (PLA), a system-on-a-chip, a system-on-a-substrate, a system-on-package, an application-specific integrated circuit (ASIC), or implemented in hardware or firmware by any other reasonable means of integrating or packaging the circuitry, or implemented in software, hardware, or firmware, or in any suitable combination of any of these three implementation methods. Alternatively, at least one of the first acquisition module 510, the first extraction module 520, the first processing module 530, the second processing module 540, the third processing module 560, and the first training module 570 may be implemented at least partially as a computer program module, which can perform corresponding functions when the computer program module is run.

[0215] Figure 6 A block diagram of an electronic device suitable for implementing a training method for a speech source tracing model according to an embodiment of this application is shown.

[0216] like Figure 6 As shown, an electronic device 600 according to an embodiment of this application includes a processor 601, which can perform various appropriate actions and processes according to a program stored in a ROM 602 (Read-Only Memory) or a program loaded from a storage portion 608 into a RAM 603 (Random Access Memory). The processor 601 may include, for example, a general-purpose microprocessor (e.g., a CPU), an instruction set processor and / or an associated chipset and / or a special-purpose microprocessor (e.g., an application-specific integrated circuit (ASIC)), etc. The processor 601 may also include onboard memory for caching purposes. The processor 601 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of this application.

[0217] RAM 603 stores various programs and data required for the operation of electronic device 600. Processor 601, ROM 602, and RAM 603 are interconnected via bus 604. Processor 601 executes various operations of the method flow according to embodiments of this application by executing programs in ROM 602 and / or RAM 603. It should be noted that the programs may also be stored in one or more memories other than ROM 602 and RAM 603. Processor 601 may also execute various operations of the method flow according to embodiments of this application by executing programs stored in said one or more memories.

[0218] According to embodiments of this application, the electronic device 600 may further include an input / output (I / O) interface 605, which is also connected to a bus 604. The electronic device 600 may also include one or more of the following components connected to the I / O interface 605: an input section 606 including a keyboard, mouse, etc.; an output section 607 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 608 including a hard disk, etc.; and a communication section 609 including a network interface card such as a LAN card, modem, etc. The communication section 609 performs communication processing via a network such as the Internet. A drive 610 is also connected to the I / O interface 605 as needed. A removable medium 611, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed on the drive 610 as needed so that computer programs read from it can be installed into the storage section 608 as needed.

[0219] This application also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments; or it may exist independently and not assembled into the device / apparatus / system. The computer-readable storage medium carries one or more programs, which, when executed, implement the method according to the embodiments of this application.

[0220] According to embodiments of this application, the computer-readable storage medium can be a non-volatile computer-readable storage medium, such as including but not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this application, the computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. For example, according to embodiments of this application, the computer-readable storage medium may include ROM 602 and / or RAM 603 and / or one or more memories other than ROM 602 and RAM 603 described above.

[0221] Embodiments of this application also include a computer program product comprising a computer program containing program code for performing the methods shown in the flowchart. When the computer program product is run on a computer system, the program code is used to enable the computer system to implement the training method for the speech tracing model provided in the embodiments of this application.

[0222] When the computer program is executed by the processor 601, it performs the functions defined in the system / apparatus of this application embodiment. According to the embodiments of this application, the systems, apparatuses, modules, units, etc., described above can be implemented by computer program modules.

[0223] In one embodiment, the computer program may rely on a tangible storage medium such as an optical storage device or a magnetic storage device. In another embodiment, the computer program may also be transmitted and distributed in the form of signals over a network medium, and downloaded and installed via the communication section 609, and / or installed from the removable medium 611. The program code contained in the computer program can be transmitted using any suitable network medium, including but not limited to: wireless, wired, etc., or any suitable combination thereof.

[0224] In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 609, and / or installed from the removable medium 611. When the computer program is executed by the processor 601, it performs the functions defined in the system of this application embodiment. According to the embodiments of this application, the systems, devices, apparatuses, modules, units, etc., described above can be implemented by computer program modules.

[0225] According to embodiments of this application, program code for executing the computer programs provided in the embodiments of this application can be written in any combination of one or more programming languages. Specifically, these computational programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages ​​include, but are not limited to, languages ​​such as Java, C++, Python, "C", or similar programming languages. The program code can be executed entirely on the user's computing device, partially on the user's device, partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computing devices, the remote computing device can be connected to the user's computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., via the Internet using an Internet service provider).

[0226] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0227] Those skilled in the art will understand that the features described in the various embodiments of this application can be combined or combined in various ways, even if such combinations or combinations are not explicitly described in this application. In particular, the features described in the various embodiments of this application can be combined and / or combined in various ways without departing from the spirit and teachings of this application. All such combinations and / or combinations fall within the scope of this application.

[0228] The embodiments of this application have been described above. However, these embodiments are merely illustrative and not intended to limit the scope of this application. Although various embodiments have been described above, this does not mean that the measures in the various embodiments cannot be used advantageously in combination. Without departing from the scope of this application, those skilled in the art can make various substitutions and modifications, all of which should fall within the scope of this application.

Claims

1. A training method for a speech source tracing model, characterized in that, The method includes: Obtain a source-tracing speech training set, wherein the source-tracing speech training set includes real speech samples and fake speech samples; The speech source tracing model is used to extract features from the real speech samples and the fake speech samples respectively to obtain real speech features and fake speech features; The difference-guided adaptive angle interval loss function is used to process the difference between the real speech features and the real speech labels, the difference between the fake speech features and the reference source speech features, and the difference between the fake speech features and the fake speech labels, to obtain the difference-guided adaptive angle interval loss value. The reference source speech features are determined by calculating the mean vector of multiple real speech samples with the same source identity. The similarity between two speech features in multiple speech feature pairs is processed using a binary contrastive loss function to obtain the binary contrastive loss value for each of the multiple speech feature pairs, wherein the two speech features in the speech feature pairs include at least one of the following: Two real speech features that share the same source identity; Real and fake speech features with the same voice source identity; Two real speech features with different source identities; Real speech features and fake speech features with different voice source identities; Based on the binary contrast loss values ​​of each of the multiple speech feature pairs, the gradients of each of the multiple speech feature pairs are determined, and the gradients of the speech feature pairs represent the changing trends of the binary contrast loss values. The gradient bias-aware contrastive learning loss function is used to process the gradient angle between the gradients of multiple speech features and the binary contrast loss value of multiple speech features to obtain the gradient bias-aware discriminative loss value. The speech source tracing model is trained based on the difference-guided adaptive angle interval loss value and the gradient deviation-aware discriminative loss value to obtain the trained speech source tracing model.

2. The method according to claim 1, characterized in that, The method utilizes a difference-guided adaptive angle-interval loss function to process the differences between real speech features and real speech labels, the differences between fake speech features and reference speech features, and the differences between fake speech features and fake speech labels, to obtain a difference-guided adaptive angle-interval loss value, including: The similarity between the reference sound source speech features and the fake speech features is calculated to obtain a cosine similarity. The cosine similarity represents the degree of damage to the real speech sample during the process of converting the fake speech sample into the fake speech sample. Based on the preset angle interval and the cosine similarity, the target angle interval corresponding to the real speech feature and the fake speech feature respectively is obtained; The real speech loss value is obtained based on the target angle interval and real similarity corresponding to the real speech features, wherein the real similarity characterizes the similarity between the real speech features and the real speech labels; The false speech loss value is obtained based on the target angle interval and false similarity corresponding to the false speech feature, wherein the false similarity characterizes the similarity between the false speech feature and the false speech label; The difference-guided adaptive angle interval loss value is determined based on the real speech loss value and the fake speech loss value.

3. The method according to claim 2, characterized in that, The preset angle interval includes a first preset angle interval and a second preset angle interval, wherein the second preset angle interval is greater than the first preset angle interval; The step of obtaining the target angle interval corresponding to the real speech feature and the fake speech feature respectively based on the preset angle interval and the cosine similarity includes: The first preset angle interval is determined as the target angle interval corresponding to the real speech features; The ratio of the cosine similarity to the preset value is processed using a truncation function to obtain a first parameter. The truncation function is used to constrain the ratio of the cosine similarity to the preset value within a preset numerical range. The second parameter is obtained based on the difference between the second preset angle interval and the first preset angle interval; The target angle interval corresponding to the false speech feature is determined based on the sum of the product of the first parameter and the second parameter and the first preset angle interval.

4. The method according to claim 1, characterized in that, A speech feature pair consisting of two real speech features with the same source identity, and a speech feature pair consisting of a real speech feature and a fake speech feature with the same source identity, are labeled with a first feature pair label; a speech feature pair consisting of two real speech features with different source identities, and a speech feature pair consisting of a real speech feature and a fake speech feature with different source identities, are labeled with a second feature pair label. The step of using a binary contrast loss function to process the similarity between two speech features in multiple speech feature pairs to obtain binary contrast loss values ​​for each of the multiple speech feature pairs includes: For two real speech feature pairs with the same sound source identity, the similarity between the two real speech features, the first label value corresponding to the first feature pair label, and the preset coefficient are fused to obtain the binary contrast loss value. For speech feature pairs with the same source identity, real speech features and fake speech features, the similarity between real speech features and fake speech features, the first label value corresponding to the first feature pair label and the preset coefficient are fused to obtain the binary contrast loss value. For two real speech feature pairs with different sound source identities, the similarity between the two real speech features, the second label value corresponding to the label of the second feature pair, and the preset coefficient are fused to obtain the binary contrast loss value. For speech feature pairs with real and fake speech features having different sound source identities, the similarity between real and fake speech features, the second label value corresponding to the second feature pair label, and the preset coefficient are fused to obtain the binary contrast loss value.

5. The method according to claim 4, characterized in that, The speech feature pair with the first feature pair label is used as the reference speech feature pair; The step of using a gradient bias-aware contrastive learning loss function to process the gradient angles between the gradients of multiple speech features and the binary contrastive loss values ​​of multiple speech features to obtain a gradient bias-aware discriminative loss value includes: For any one of the multiple speech feature pairs, the gradient deviation value between the speech feature pair and the reference speech feature pair is determined based on the gradient angle between the gradient of the speech feature pair and the reference gradient of the reference speech feature pair. The discrimination loss value for gradient deviation perception is determined based on the gradient deviation values ​​of each of the multiple speech feature pairs and the binary contrast loss value.

6. The method according to claim 5, characterized in that, The step of determining the discrimination loss value for gradient deviation perception based on the gradient deviation values ​​of the respective speech features and the binary contrast loss value includes: The average gradient deviation is determined based on the average of multiple gradient deviation values. The gradient deviation coefficients of the multiple speech features are determined based on the ratio of their respective gradient deviation values ​​to the mean gradient deviation value. By fusing the gradient bias coefficients of multiple speech feature pairs with their respective gradient loss values ​​and the binary contrast loss values, the gradient loss values ​​of multiple speech feature pairs with their respective gradient loss values ​​are determined. The discrimination loss value for gradient bias perception is determined based on the mean of the gradient loss values ​​of multiple speech feature pairs.

7. The method according to claim 6, characterized in that, Determining the discriminative loss value for gradient deviation perception based on at least one of the gradient deviation values ​​and the binary contrast loss value for each of the multiple speech features includes: Determine multiple target gradient deviation values ​​that satisfy preset screening conditions; Determine the target binary contrast loss value corresponding to the multiple target gradient deviation values; The discrimination loss value for gradient bias perception is determined based on the mean of multiple target binary contrast loss values.

8. The method according to claim 1, characterized in that, The difference-guided adaptive angle interval loss value and the gradient deviation-aware discriminative loss value are used to train the speech source tracing model, resulting in a trained speech source tracing model, including: By fusing the balance coefficient and the discriminant loss value sensed by the gradient deviation, the discriminant loss value sensed by the target gradient deviation is obtained; The target loss value is determined by summing the difference-guided adaptive angle interval loss value and the target gradient deviation-aware discriminative loss value. Based on the target loss value, the model parameters of the speech source tracing model are optimized to obtain the trained speech source tracing model.

9. The method according to claim 1, characterized in that, The step of using the speech source tracing model to extract features from the real speech samples and the fake speech samples respectively, to obtain real speech features and fake speech features, includes: The real speech samples and the fake speech samples are preprocessed respectively to obtain real speech samples and fake speech samples with the same speech duration; Feature extraction is performed on the real speech samples and the fake speech samples with the same speech duration to obtain the real speech features and the fake speech features.

10. A training device for a speech source tracing model, characterized in that, The device includes: The first acquisition module is used to acquire the source-tracing speech training set, wherein the source-tracing speech training set includes real speech samples and fake speech samples; The first extraction module is used to extract features from the real speech sample and the fake speech sample using the speech tracing model, so as to obtain real speech features and fake speech features. The first processing module is used to process the differences between the real speech features and the real speech labels, the differences between the fake speech features and the reference sound source speech features, and the differences between the fake speech features and the fake speech labels using a difference-guided adaptive angle interval loss function, to obtain the difference-guided adaptive angle interval loss value. The reference sound source speech features are determined by performing vector mean calculation on multiple real speech samples with the same sound source identity. The second processing module is used to process the similarity between two speech features in multiple speech feature pairs using a binary contrastive loss function, to obtain the binary contrastive loss value for each of the multiple speech feature pairs, wherein the two speech features in the speech feature pairs include at least one of the following: Two real speech features that share the same source identity; Real and fake speech features with the same voice source identity; Two real speech features with different source identities; Real speech features and fake speech features with different voice source identities; The first determining module is used to determine the gradient of each of the multiple speech feature pairs based on their respective binary contrast loss values, wherein the gradient of the speech feature pair represents the changing trend of the binary contrast loss value. The third processing module is used to process the gradient angles between the gradients of multiple speech features and their respective binary contrast loss values ​​using a gradient bias-aware contrastive learning loss function, thereby obtaining a gradient bias-aware discriminative loss value; and The first training module is used to train the speech source tracing model based on the difference-guided adaptive angle interval loss value and the gradient deviation-aware discriminative loss value, so as to obtain the trained speech source tracing model.