Model training method, voiceprint detection method, related device and medium
By using the transfer learning model training method in the voiceprint verification system, the distribution difference between training data and actual application data is reduced, which solves the problem of the voiceprint verification system's performance degradation in new environments and achieves efficient and accurate physical attack identification.
Patent Information
- Application Number
- CN202410529102.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-04-28
- Publication Date
- 2025-10-28
AI Technical Summary
Existing voiceprint verification systems experience reduced detection performance when facing physical attacks, especially when deployed in new environments or using unknown devices, making it difficult to effectively identify physical attacks on recording and playback devices.
Through the transfer learning model training method, the voiceprint data and known label data collected by the terminal device are used to train the transfer learning model and classifier model, reducing the distribution difference between the training data and the actual application data, and enhancing the robustness of the system in different environments.
It improves the accuracy of the voiceprint verification system in practical applications and the recognition rate of physical attacks, reduces the error rate, optimizes the use of storage and computing resources, and improves the overall efficiency of the system.
Smart Images

Figure CN120853583A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of electronic technology, and in particular to a model training method, a voiceprint detection method, related devices and media. Background Technology
[0002] Automatic speaker verification (ASV) technology is an important branch of biometrics, primarily used to confirm whether a person's identity matches a voice sample. However, current ASV systems are vulnerable to physical attacks, which refer to attempts to deceive voiceprint verification systems by using the physical properties of sound. This typically involves attackers attempting to simulate or replicate the target user's voice characteristics to bypass security measures.
[0003] To detect physical attacks, a method based on channel noise feature enhancement and single-class classification is currently commonly used. In this method, physical attacks leave identifiable traces in the audio track, which are summarized as "channel noise." It removes components not directly related to physical attack detection, such as speaker characteristics and speech content, by performing a difference operation between the original audio and the resynthesized audio, focusing on preserving and analyzing the "channel noise."
[0004] However, while it can effectively identify known recording and playback devices, its detection performance often declines once the system is deployed to a new environment or when attackers use unknown devices. Summary of the Invention
[0005] This application provides a model training method, a voiceprint detection method, related devices, and a storage medium to improve the accuracy of voiceprint physical attack identification.
[0006] The first aspect of this application provides a model training method. In this method, a model training device acquires a first voiceprint audio sample to be tested, extracts feature vectors from the first voiceprint audio sample to obtain first voiceprint data. The first voiceprint audio sample is audio generated in the deployment environment of the terminal device. The model training device makes a judgment based on the first voiceprint data. If the first voiceprint data meets preset conditions, the model training device trains a transfer learning model based on the first voiceprint data and first training data with known labels to obtain a trained transfer learning model. The model training device inputs the first voiceprint data and the first training data into the trained transfer learning model for mapping to obtain second training data. The model training device trains an initialized classifier model based on the second training data to obtain a trained classifier model.
[0007] In this embodiment, the model training device trains a transfer learning model to transfer the sample distribution in the first training data to the first voiceprint data, thereby reducing the distribution difference between the first voiceprint data and the first training data. This enables effective feature transfer and adaptation, enhances the robustness of the system in different environments, and maintains high performance unaffected by environmental changes.
[0008] In some possible implementations, the model training device determines whether the trained transfer learning model and the trained classifier model meet the acceptance criteria. If they do, the model training device updates the transfer learning model and the classifier with the trained transfer learning model and the trained classifier model, and then uses the trained transfer learning model and the trained classifier model to perform voiceprint detection.
[0009] In some possible implementations, acceptance criteria include that the error rates of the trained transfer learning model and the trained classifier model are lower than a preset threshold.
[0010] In this embodiment, by judging whether the error rates of the trained transfer learning model and the trained classifier model are lower than a preset threshold, the performance of the model can be comprehensively evaluated, resulting in a better evaluation effect.
[0011] In some possible implementations, if the trained transfer learning model and the trained classifier model do not meet the acceptance criteria, the model training device will retrain the model or adjust the model parameters.
[0012] In some possible implementations, the model training model responds to input operations from the user's graphical interface by performing the step of acquiring first voiceprint data.
[0013] In this embodiment, the model training device can adapt to the data collection strategy, thus reducing unnecessary data processing, optimizing the use of storage and computing resources, and improving the overall efficiency of the system.
[0014] In some possible implementations, the preset conditions include: the number of first voiceprint data reaches a preset value, or the collection time of the first voiceprint data reaches a preset time.
[0015] In some possible implementations, if the first voiceprint data meets a preset condition, the collection of the first voiceprint data is stopped.
[0016] In this embodiment, the model training device reduces unnecessary data processing, optimizes the use of storage and computing resources, and improves the overall efficiency of the system by setting preset conditions.
[0017] A second aspect of this application provides a voiceprint detection method. Optionally, the execution subject of this method can be a terminal device, a component or device applied to the terminal device (e.g., a processor, chip, or chip system), or a logic module or software capable of implementing all or part of the terminal device's functions. Taking a terminal device as an example, in this method, the terminal device acquires second voiceprint data, wherein the second voiceprint data is a feature vector of a second voiceprint audio to be tested, and the second voiceprint audio to be tested is audio generated in the deployment environment of the terminal device. The terminal device processes the second voiceprint data according to a trained transfer learning model to obtain voiceprint features. The distribution difference between the voiceprint features and the first training data is smaller than the distribution difference between the second voiceprint data and the first training data. The terminal device detects the voiceprint features according to a trained classifier model to obtain a detection result.
[0018] In this embodiment, by using a transfer learning model to process the second voiceprint data, the distribution difference between the obtained voiceprint features and the first training data is reduced, thereby improving the accuracy of the system, reducing the error rate, and enabling the system to have a higher recognition rate against physical attacks in practical applications.
[0019] In some possible implementations, the trained transfer learning model is obtained by training as described in the first aspect above.
[0020] In some possible implementations, the trained classifier model is obtained by training as described in the first aspect above.
[0021] A third aspect of this application provides a model training apparatus, the model training apparatus comprising:
[0022] An interface unit is used to acquire first voiceprint data, which is the feature vector of the first voiceprint audio to be tested, and the first voiceprint audio to be tested is the audio generated in the deployment environment of the terminal device.
[0023] The processing unit is used to train the transfer learning model based on the first voiceprint data and the first training data if the first voiceprint data meets the preset conditions, so as to obtain the trained transfer learning model. The first training data is feature data of known labels.
[0024] The processing unit is also used to map the first voiceprint data and the first training data according to the trained transfer learning model to obtain the second training data;
[0025] The processing unit is also used to train the initialized classifier model based on the second training data to obtain the trained classifier model.
[0026] In some possible implementations, the processing unit is further configured to:
[0027] If the trained transfer learning model and the trained classifier model meet the acceptance criteria, then the trained transfer learning model and the trained classifier model are used for voiceprint detection.
[0028] In some possible implementations, acceptance criteria include: the error rates of the trained transfer learning model and the trained classifier model are lower than a preset threshold.
[0029] In some possible implementations, the processing unit is further configured to:
[0030] If the trained transfer learning model and the trained classifier model do not meet the acceptance criteria, then retrain or adjust the parameters.
[0031] In some possible implementations, the interface unit is specifically used for:
[0032] In response to user input via the graphical interface, the step of acquiring the first voiceprint data is performed.
[0033] In some possible implementations, the preset conditions include: the number of first voiceprint data reaches a preset value, or the collection time of the first voiceprint data reaches a preset time.
[0034] In some possible implementations, the interface unit is also used for:
[0035] If the first voiceprint data meets the preset conditions, then stop collecting the first voiceprint data.
[0036] A fourth aspect of this application provides a voiceprint detection device. Optionally, the voiceprint detection device can be a terminal device, a component or device applied to the terminal device (e.g., a processor, chip, or chip system), or a logic module or software capable of implementing all or part of the functions of the terminal device. The voiceprint detection device includes:
[0037] The interface unit is used to acquire the second voiceprint data, which is the feature vector of the second voiceprint audio to be tested, and the second voiceprint audio to be tested is the audio generated in the deployment environment of the terminal device.
[0038] The processing unit is used to process the second voiceprint data according to the trained transfer learning model to obtain voiceprint features;
[0039] The processing unit is also used to detect voiceprint features based on the trained classifier model and obtain detection results.
[0040] The fifth aspect of this application provides a model training apparatus, which includes: a processor, a memory, an input / output device, and a bus; the memory stores computer instructions; when the processor executes the computer instructions in the memory, the memory stores computer instructions; when the processor executes the computer instructions in the memory, it is used to implement any of the implementation methods described in the first aspect.
[0041] The sixth aspect of this application provides a voiceprint detection device, the model training device comprising: a processor, a memory, an input / output device, and a bus; the memory stores computer instructions; when the processor executes the computer instructions in the memory, the memory stores computer instructions; when the processor executes the computer instructions in the memory, it is used to implement any of the implementation methods of the second aspect described above.
[0042] A seventh aspect of this application provides a chip system comprising a processor and input / output ports, wherein the processor is used to implement the processing functions involved in the methods described in the first, second, or third aspects above, and the input / output ports are used to implement the transceiver functions involved in the methods described in the first, second, or third aspects above.
[0043] In one possible design, the chip system also includes a memory for storing program instructions and data for implementing the functions involved in the methods described in the first or second aspect above.
[0044] This chip system can consist of chips or include chips and other discrete components.
[0045] An eighth aspect of this application provides a computer-readable storage medium including instructions that, when executed on a computer, cause the computer to perform the method described in the first aspect above, or the method described in the second aspect above.
[0046] The ninth aspect of this application provides a computer program product containing instructions that, when run on a computer, cause the computer to perform the method described in the first aspect above, or the method described in the second aspect above.
[0047] The beneficial effects from the third to the ninth aspects can be understood by referring to the beneficial effects of the first or second aspects and their corresponding implementation methods; details will not be elaborated here. Attached Figure Description
[0048] Figure 1 This is a schematic diagram of one embodiment of the system architecture described in this application.
[0049] Figure 2This is a schematic diagram of one embodiment of the model training method in this application;
[0050] Figure 3 This is a schematic diagram of another embodiment of the model training method in this application;
[0051] Figure 4 This is a schematic diagram of one embodiment of the training transfer learning model in this application;
[0052] Figure 5 This is a schematic diagram of one embodiment of the voiceprint detection method in this application;
[0053] Figure 6 This is a schematic diagram of one embodiment of the model training device in this application;
[0054] Figure 7 This is a schematic diagram of one embodiment of the voiceprint detection device in this application;
[0055] Figure 8 This is a schematic diagram of another embodiment of the model training device in this application;
[0056] Figure 9 This is a schematic diagram of another embodiment of the voiceprint detection device in this application. Detailed Implementation
[0057] This application provides a model training method, a voiceprint detection method, related devices, and a storage medium, which can improve the accuracy of the ASV system and enable the ASV system to have a higher recognition rate of physical attacks in practical applications.
[0058] The embodiments of this application will now be described with reference to the accompanying drawings. Those skilled in the art will recognize that, with technological advancements and the emergence of new scenarios, the technical solutions provided in the embodiments of this application are equally applicable to similar technical problems.
[0059] The terms "first," "second," etc., used in the specification, claims, and drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such terms are interchangeable where appropriate; this is merely a way of distinguishing objects with the same attributes in the description of embodiments of this application. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion, so that a process, method, system, product, or apparatus that comprises a series of elements is not necessarily limited to those elements, but may include other elements not expressly listed or inherent to those processes, methods, products, or apparatuses.
[0060] In this application, "for indicating" can include both direct and indirect indication. When describing an indication message as indicating A, it can include whether the indication message directly indicates A or indirectly indicates A, but does not necessarily mean that the indication message carries A.
[0061] Furthermore, the specific indication method can also be any existing indication method, such as, but not limited to, the above-mentioned indication methods and their various combinations. Specific details of various indication methods can be found in existing technologies, and will not be repeated here. As described above, for example, when multiple pieces of information of the same type need to be indicated, the indication methods for different pieces of information may differ. In specific implementation, the required indication method can be selected according to specific needs. This application embodiment does not limit the selected indication method; therefore, the indication methods involved in this application embodiment should be understood to cover various methods that enable the party to be indicated to obtain the information to be indicated.
[0062] In the embodiments of this application, descriptions such as "when," "under the circumstances," "if," and "if" all refer to the device making corresponding processing under certain objective circumstances, and are not limited to a specific time. They do not require the device to make a judgment action during implementation, nor do they imply any other limitations.
[0063] First, some technical terms involved in the embodiments of this application will be introduced.
[0064] 1) Automatic Speaker Verification (ASV) is a computer-based method for verifying the identity of a speaker. This technology analyzes a speaker's voice characteristics to determine if they are a pre-registered target speaker. It utilizes individual voice characteristics for identification, as each person's voiceprint is unique, similar to a fingerprint or retinal pattern. In multi-person conversations, ASV can automatically identify each speaker. ASV has wide applications in speech recognition and voice security; for example, in call center customer authentication and contactless facility access, speaker verification simplifies the process of verifying the identity of registered speakers.
[0065] 2) Voice physical attacks are a type of attack targeting voiceprint detection or voice recognition systems. They involve directly manipulating the sound signal or sound propagation path at the physical level to interfere with, disrupt, or deceive the system. This type of attack differs from traditional software or network attacks, as it directly targets the sound signal itself or the sound acquisition device. One of the most common physical attack methods is the recording and playback attack. This type of attack is relatively simple to execute but extremely destructive: the attacker first secretly records the target speaker's voice, then plays these recordings to the ASV system in situations requiring authentication, attempting to deceive the system into passing identity verification. The danger of this strategy lies in its ability to bypass voiceprint-based security verification without requiring complex technical means.
[0066] 3) Transfer learning is an important method in machine learning. Its core idea is to apply knowledge or models learned on one task to another related but different task. The advantage of this method is that when data for the target task is scarce or labeling is costly, transfer learning can leverage the rich knowledge and experience already present in the source task to assist in learning the target task, thereby improving learning efficiency and reducing reliance on new data.
[0067] Please see Figure 1 The system architecture on which the model training method in this embodiment is based is briefly described below:
[0068] The voiceprint detection method is applied to a terminal device 101, which includes a feature extractor 102, an online transfer learning module 102, and a classifier 104. The feature extractor 102 extracts feature vectors from the input voiceprint audio to represent the voiceprint features of the entire audio segment. The online transfer learning module 103 acts as a bridge, utilizing data collected in different environments to enhance the model's generalization ability. The classifier 104 outputs detection results based on the feature vectors, identifying whether the audio is a genuine voiceprint or a physical attack. A model training device 105 trains the model on the online transfer learning module 103 and the classifier 104. The model training device 105 can also be a component or device applied to the terminal device (e.g., a processor, chip, or chip system), or a logic module or software that implements all or part of the terminal device's functions; specific limitations are not specified here. In one possible implementation, the model training device 105 can be a server, a component or device applied to the server (e.g., a processor, chip, or chip system), or a logic module or software capable of implementing all or part of the server's functions; no specific limitation is made here. The terminal device can send the extracted feature vectors to the server, which will then train the model and return the trained model.
[0069] Figure 1 The terminal device 101 can also be called user equipment (UE), mobile station (MS), or mobile terminal (MT), etc. Specifically, Figure 1 The terminal device 101 can be a mobile phone, tablet computer, or computer with wireless transceiver capabilities. It can also be a virtual reality (VR) terminal, an augmented reality (AR) terminal, a wireless terminal in industrial control, a wireless terminal in autonomous driving, a wireless terminal in telemedicine, a wireless terminal in a smart grid, a wireless terminal in a smart city, a wireless terminal in a smart home, an in-vehicle terminal, a vehicle with vehicle-to-vehicle (V2V) communication capabilities, an intelligent connected vehicle, a drone with drone-to-UAV (U2U) communication capabilities, etc., and is not specifically limited here.
[0070] Currently, since physical attacks leave identifiable traces in the audio track, these traces are categorized as "channel noise." Therefore, detecting physical attacks on voiceprints involves reconstructing the audio from the voiceprint audio using a vocoder, and then performing a difference operation between the reconstructed audio and the original voiceprint audio to analyze the "channel noise."
[0071] However, since both the resynthesized audio and the original voiceprint audio are manually labeled data, meaning that both belong to the development domain, and there are certain differences between the development domain and the deployment domain, such as changes in data distribution or missing features, these differences may have a negative impact on the system's performance. Therefore, this method cannot effectively identify audio data generated in the actual deployment environment of the terminal device.
[0072] Based on this, an embodiment of this application provides a method. Please refer to... Figure 2 An embodiment of this application includes a model training method comprising:
[0073] 201. Obtain the first voiceprint data;
[0074] The feature extractor receives the first audio sample of the voiceprint to be tested and extracts its feature vector to obtain the first voiceprint data. The model training device collects the first voiceprint data generated by the feature extractor and stores it in a local database, such as... Figure 3 As shown. In one possible implementation, the model training device needs to be configured to collect the first voiceprint data. Alternatively, the model training device responds to user input via the graphical interface by performing the step of acquiring the first voiceprint data.
[0075] In one possible implementation, the model training device responds to user input via the graphical interface by setting a start time for collecting the first voiceprint data, enabling the model training device to collect the first voiceprint data at a predetermined time. Alternatively, the model training device may trigger the collection of the first voiceprint data based on changes in the coordinates of the terminal device; however, this is not specifically limited here.
[0076] In this embodiment, the model training device can adapt to the data collection strategy, thereby reducing unnecessary data processing, optimizing the use of storage and computing resources, and improving the overall efficiency of the system.
[0077] 202. Determine whether the preset conditions are met;
[0078] like Figure 3 As shown, the model training device determines whether the collected first voiceprint data meets preset conditions. In one possible implementation, the model training device stops collecting the first voiceprint data when the number of first voiceprint data reaches a preset value. For example, the model training device stops collecting the first voiceprint data when the number of first voiceprint data reaches 50, 500, or 1000. In another possible implementation, the model training device stops collecting the first voiceprint data when the collection time reaches a preset duration; the specific implementation is not limited here.
[0079] It should be understood that the aforementioned preset values and preset durations can be set by the user or pre-configured on the model training device; no specific limitations are made here.
[0080] 203. Train the transfer learning model and the classifier model;
[0081] The model training device uses the collected first voiceprint data and a small amount of first training data with known labels to train a transfer learning model. This transfer learning model can also be a mapping matrix; the specific type is not limited here.
[0082] Specifically, in one possible implementation, the model training device uses transfer component analysis (TCA) to train a transfer learning model using the first voiceprint data from the actual application data domain and the first training data from the training data domain, thereby reducing the distributional differences between the training data domain and the actual application data domain. For example... Figure 4As shown, the training data domain includes real training data and attack training data. Real training data includes training data manually labeled as real voiceprints, while attack training data includes training data manually labeled as physical attacks. The practical application data domain includes real voiceprint data and attack voiceprint data, where both are voiceprint features extracted by the feature extractor from the audio of the voiceprint to be tested.
[0083] TCA (Transfer Learning Computation) is a method for dimensionality reduction in transfer learning. Specifically, TCA learns the common transfer components across all domains (i.e., components that do not cause changes in inter-domain distribution and maintain the inherent structure of the original data), thus reducing the distributional differences between different domains in the mapped subspace. It achieves this by finding a feature map such that the probability density and conditional probability density of the mapped data distribution are equal in both the source and target domains.
[0084] The model training device inputs the training data domain and the real-world application data domain into the transfer learning model, and transfers the knowledge from the training data domain to the real-world application data domain, obtaining the trained model parameters. The transfer learning model M is then updated based on these model parameters. i The trained transfer learning model M is obtained. i+1 , where i is an integer greater than or equal to 0, such that after transfer learning model M i+1 The difference in data distribution between the mapped training data domain and the real-world application data domain is reduced. In one possible implementation, the transfer learning model M... i+1 The output is the online migration public domain, where the data is divided into real data and attack data.
[0085] It should be understood that real-world application data also includes voiceprints generated by the deployment environment. For example, when a terminal device is in a noisy environment and uses ASV technology to unlock, the audio of the voiceprint to be tested collected by the terminal device includes ambient noise. Training a transfer learning model using TCA technology can find common features between the training data domain and the real-world application data domain by mapping them to the same feature space, thereby reducing the impact of ambient noise on the real-world application data domain while preserving knowledge from the training data domain.
[0086] It should be noted that the training data domain is also called the source domain, and the data domain used in actual applications is also called the target domain. Manually labeled samples in the source domain are used to train the model, and the distribution of samples in the source domain is used to improve the model performance in the target domain, thus achieving transfer learning.
[0087] In this embodiment, the model training device trains a transfer learning model to transfer the sample distribution in the training data domain to the actual application data domain, thereby reducing the distribution difference between the training data domain and the actual application data domain. This enables effective feature transfer and adaptation, enhances the robustness of the system in different environments, and maintains high performance unaffected by environmental changes.
[0088] The model training device is based on the transfer learning model M i+1 The first voiceprint data and the first training data are mapped to the same feature space to obtain the mapped second training data. The model training device randomly initializes a new classifier model C. i The classifier model C is then used with the mapped second training data. i Perform iterative training until the optimization objective is achieved, resulting in the classifier model C. i+1 .
[0089] In one possible implementation, the model training device is deployed on a server. Specifically, the terminal device sends the collected voiceprint data to the server, where the model training device trains the transfer learning model M. i+1 and classifier model C i+1 And send it to the terminal device, the specifics of which are not specified here.
[0090] 204. Determine whether the acceptance conditions are met;
[0091] The terminal device uses the trained transfer learning model M i+1 and classifier model C i+1 The evaluation was conducted to determine the effectiveness of the transfer learning model M. i+1 and classifier model C i+1 Does it meet the acceptance criteria?
[0092] In one possible implementation, the acceptance criterion can be one or more of the following: equal error rate (EER), accuracy, recall, or specific performance metrics. EER is a commonly used performance metric in audio domains and speech recognition evaluation tasks, particularly for evaluating the performance of speaker recognition and voiceprint recognition systems. EER refers to the error rate in binary classification tasks when the false positive rate (FAR) equals the false negative rate (FRR). The false positive rate is the probability that a speech that is actually negative (not the target speaker) is incorrectly classified as positive (the target speaker); the false negative rate is the probability that a speech that is actually positive is incorrectly classified as negative. EER is usually expressed as a percentage, with lower EER values indicating better system performance.
[0093] In one possible implementation, the terminal device responds to user input via the graphical interface by setting acceptance criteria. Specifically, the user can set acceptance criteria through the graphical interface to one or more of EER, precision, recall, or specific performance metrics; the specific criteria are not limited here.
[0094] If the transfer learning model M i+1 and classifier model C i+1 If the acceptance criteria are met, the terminal device proceeds to step 205. If the transfer learning model M... i+1 and classifier model C i+1 If the acceptance criteria are not met, the terminal device returns to step 201. Specifically, the terminal device reacquires the voiceprint data and retrains the transfer learning model and classifier model based on the new voiceprint data until the transfer learning model and classifier model meet the acceptance criteria. In one possible implementation, if the transfer learning model M... i+1 and classifier model C i+1 If the acceptance criteria are not met, the terminal device sends feedback to the user's graphical interface, prompting the user to check the transfer learning model M. i+1 and classifier model C i+1 If the acceptance criteria are not met, the user decides whether to recollect data to train the transfer learning model and classifier model. The terminal device responds to the user's graphical interface and executes step 201, the specifics of which are not limited here.
[0095] 205. Use the trained transfer learning model and the trained classifier model for voiceprint detection.
[0096] The terminal device will transfer learning model M i+1 and classifier model C i+1 This is applied to the online transfer learning module and classifier, updating the models on both the online transfer learning module and the classifier, such as... Figure 4 As shown. The subsequent inference process will use the transfer learning model M. i+1 and classifier model C i+1 To reason.
[0097] In this embodiment, the terminal device uses an online transfer learning module to train the model based on the voiceprint data generated by the user in the deployment environment, which improves the accuracy of the system and reduces the equal error rate (EER), enabling the system to have a higher recognition rate of physical attacks in practical applications.
[0098] Please see Figure 5 An embodiment of this application includes a voiceprint detection method comprising:
[0099] 501. Obtain the second voiceprint data;
[0100] The voiceprint detection device receives the input second voiceprint audio to be tested and extracts the feature vector of the second voiceprint audio to obtain the second voiceprint data. In one possible implementation, the first voiceprint audio and the second voiceprint audio can be the same audio or different audio; the first voiceprint data and the second voiceprint data can be the same voiceprint data or different voiceprint data, which is not limited here.
[0101] 502. Process the second voiceprint data according to the trained transfer learning model;
[0102] The voiceprint detection device inputs the second voiceprint data into the transfer learning model M. i+1 The second voiceprint data is mapped to the feature space to obtain voiceprint features. It should be noted that the distribution difference between the voiceprint features and the first training data is smaller than the distribution difference between the second voiceprint data and the first training data.
[0103] 503. Detect voiceprint features based on the trained classifier model;
[0104] The voiceprint detection device inputs voiceprint features into a classifier model, which outputs detection results based on the voiceprint features to identify whether the audio is a real voiceprint or a physical attack.
[0105] For example, the voiceprint detection method provided in this application was tested using a publicly available dataset, and the test results are shown in Table 1 below:
[0106] Table 1
[0107] EER (%) n1 = 0 (baseline) n1=50 n1=500 n1=5000 n2=50 24.02 26.89 27.42 29.4 n2=500 24.05 22.7 23.82 27.41 n2=1000 24.02 22.01 22.93 25.29 n2=5000 24.04 21.63 22.11 22.48 n2=10000 24.07 21.41 21.66 21.79
[0108] As shown in Table 1, n1 represents the number of audio samples in the training data, and n2 represents the number of audio samples collected by the terminal device in the deployment environment. Compared to the existing solution's EER of 24.0%, this invention achieves a significant performance improvement, reducing the EER to 21.4%. Experiments were conducted to explore the impact of data volume on detection efficiency by adjusting n1 and n2. Experimental data show that as the amount of user data collected increases, the system's EER continuously decreases, highlighting the importance of large datasets in improving the system's recognition capabilities.
[0109] Specifically, in real-world testing, this embodiment of the application reduced the EER to 3.59% when using 50 initial training audios and 1000 collected audios, a decrease of 17.09% compared to the baseline level, as shown in Table 2.
[0110] Table 2
[0111]
[0112] After initial verification of system performance, the experiment also conducted detailed tests on the number of parameters and computational load of each module in the embodiments of this application to evaluate its feasibility on actual devices. As shown in Table 3, the number of parameters and computational load required by each module are within a reasonable range.
[0113] Table 3
[0114] Module Specific implementation method Number of parameters (K) Inference computation Training computation Feature extractor Vocoder+PCA 154 79M 1.3G Online transfer learning module TCA 15 4M 20G Classifier VAE 40 0.25M 0.0125G
[0115] The experiment also referenced the operating frequency and memory configuration of common mobile phone devices. For example, Table 4 below shows four different terminal devices with varying configurations, thus comparing the potential operating efficiency of the present invention on mobile devices.
[0116] Table 4
[0117] Terminal equipment model Operating frequency (GHz) Memory (GB) Terminal device 1 3.7 6 Terminal device 2 2.8 12 Terminal device 3 3.7 8 Terminal device 4 3.3 16
[0118] Test results show that the resource overhead required to update the online transfer learning module and classifier is relatively low. Considering the processing power and memory capacity of modern mobile devices, the embodiments of this application theoretically have the potential to be deployed on these mobile devices.
[0119] The embodiments of this application outperform existing solutions in detection performance and also demonstrate superior resource utilization efficiency, making it a solution suitable for widespread deployment on various hardware platforms, especially on resource-constrained mobile devices.
[0120] The model training device and voiceprint detection device in the embodiments of this application are described below with reference to the accompanying drawings.
[0121] Please see Figure 6 One embodiment of the model training device in this application includes:
[0122] Interface unit 601 is used to acquire first voiceprint data, the first voiceprint data being the feature vector of the first voiceprint audio to be tested, and the first voiceprint audio to be tested being the audio generated in the deployment environment of the terminal device.
[0123] The processing unit 602 is used to train the transfer learning model based on the first voiceprint data and the first training data if the first voiceprint data meets the preset conditions, so as to obtain the trained transfer learning model. The first training data is feature data of known labels.
[0124] Processing unit 602 is also used to map the first voiceprint data and the first training data according to the trained transfer learning model to obtain the second training data;
[0125] The processing unit 602 is also used to train the initialized classifier model based on the second training data to obtain the trained classifier model.
[0126] Optionally, the processing unit 602 is also used for:
[0127] If the trained transfer learning model and the trained classifier model meet the acceptance criteria, then the trained transfer learning model and the trained classifier model are used for voiceprint detection.
[0128] Optionally, acceptance criteria include: the error rates of the trained transfer learning model and the trained classifier model are lower than a preset threshold.
[0129] Optionally, the processing unit 602 is also used for:
[0130] If the trained transfer learning model and the trained classifier model do not meet the acceptance criteria, then retrain or adjust the parameters.
[0131] Optionally, interface unit 601 is specifically used for:
[0132] In response to user input via the graphical interface, the step of acquiring the first voiceprint data is performed.
[0133] Optionally, the preset conditions include: the number of first voiceprint data reaches a preset value; or, the collection time of the first voiceprint data reaches a preset time.
[0134] Optionally, interface unit 601 is also used for:
[0135] If the first voiceprint data meets the preset conditions, then stop collecting the first voiceprint data.
[0136] Please see Figure 7 One embodiment of the voiceprint detection device in this application includes:
[0137] Interface unit 701 is used to acquire second voiceprint data, the second voiceprint data being the feature vector of the second voiceprint audio to be tested, and the second voiceprint audio to be tested being the audio generated in the deployment environment of the terminal device.
[0138] The processing unit 702 is used to process the second voiceprint data according to the trained transfer learning model to obtain voiceprint features;
[0139] The processing unit 702 is also used to detect voiceprint features based on the trained classifier model and obtain detection results.
[0140] The following describes a model training apparatus provided in an embodiment of this application. Please refer to [link to relevant documentation]. Figure 8 , Figure 8 This is a schematic diagram of a model training device provided in an embodiment of this application.
[0141] The model training device specifically includes:
[0142] Processor 801, memory 802, input / output unit 803, bus 804;
[0143] The processor 801 is connected to the memory 802, the input / output unit 803, and the bus 804;
[0144] The program is stored in memory 802;
[0145] The processor 801 executes the program in the memory 802, causing the model training device to perform the method described in the foregoing embodiments.
[0146] The following describes a voiceprint detection device provided in an embodiment of this application. Please refer to [link to relevant documentation]. Figure 9 , Figure 9 This is a schematic diagram of a voiceprint detection device provided in an embodiment of this application. The voiceprint detection device can be a terminal device as described in the above method embodiments, or it can be a chip, chip system, or processor that supports the terminal device in implementing the above methods. This voiceprint detection device can be used to implement the methods described in the above method embodiments; for details, please refer to the descriptions in the above method embodiments.
[0147] The voiceprint detection device may include one or more processors 901, which are connected to a memory 902, an input / output unit 903, and a bus 904. The processor 901 may be a general-purpose processor or a dedicated processor, such as a baseband processor or a central processing unit. The baseband processor can be used to process communication protocols and communication data, while the central processing unit can be used to control the voiceprint detection device (e.g., a base station, baseband chip, terminal, terminal chip, DU or CU, etc.), execute software programs, and process the data from the software programs.
[0148] Optionally, the voiceprint detection device may include one or more memories 902, which may store instructions that can be executed on the processor 901, causing the voiceprint detection device to perform the methods described in the above method embodiments. Optionally, the memory 902 may also store data. The processor 901 and the memory 902 may be configured separately or integrated together.
[0149] Optionally, the voiceprint detection device may further include a transceiver and an antenna. The transceiver, also known as a transceiver unit, transceiver, or transceiver circuit, is used to implement transceiver functions. The transceiver may include a receiver and a transmitter; the receiver, also known as a receiver circuit, is used to implement a receiving function; the transmitter, also known as a transmitter or transmitting circuit, is used to implement a transmitting function.
[0150] In another possible design, the processor 901 may include a transceiver for implementing receive and transmit functions. For example, the transceiver may be a transceiver circuit, an interface, or an interface circuit. The transceiver circuit, interface, or interface circuit for implementing receive and transmit functions may be separate or integrated. The aforementioned transceiver circuit, interface, or interface circuit may be used for reading and writing code / data, or for transmitting or relaying signals.
[0151] In another possible design, the processor 901 may optionally store instructions that, when executed, cause the voiceprint detection device to perform the methods described in the above method embodiments. The instructions may be stored in the processor 901; in this case, the processor 901 may be implemented in hardware.
[0152] In another possible design, the voiceprint detection device may include circuitry that enables the terminal device in the aforementioned method embodiments to transmit, receive, or communicate. The processor and transceiver described in this application embodiment can be implemented on integrated circuits (ICs), analog ICs, radio frequency integrated circuits (RFICs), mixed-signal ICs, application-specific integrated circuits (ASICs), printed circuit boards (PCBs), electronic devices, etc. The processor and transceiver can also be manufactured using various IC process technologies, such as complementary metal-oxide-semiconductor (CKOS), M-type metal-oxide-semiconductor (MKOS), p-type metal-oxide-semiconductor (PKOS), bipolar junction transistors (BJTs), bipolar CKOS (BiCKOS), silicon germanium (SiGe), gallium arsenide (GaAs), etc.
[0153] The voiceprint detection device described in the above embodiments can be a terminal device, but the scope of the voiceprint detection device described in the embodiments of this application is not limited to this, and the structure of the voiceprint detection device can be unrestricted. Figure 9 The limitations. The voiceprint detection device can be a standalone device or part of a larger device. For example, the voiceprint detection device can be:
[0154] (1) Independent integrated circuit IC, or chip, or chip system or subsystem;
[0155] (2) A collection of one or more ICs, optionally including a storage component for storing data and instructions;
[0156] (3) ASIC, such as modems (KSK);
[0157] (4) Modules that can be embedded in other devices;
[0158] (5) Receivers, terminals, smart terminals, cellular phones, wireless devices, handheld devices, mobile units, vehicle-mounted devices, network devices, cloud devices, artificial intelligence devices, etc.
[0159] (6) Others, etc.
[0160] It is understood that some optional features in the embodiments of this application can be implemented independently in certain scenarios without relying on other features, such as the current solution on which they are based, to solve the corresponding technical problems and achieve the corresponding effects. Alternatively, they can be combined with other features as needed in certain scenarios. Correspondingly, the communication device given in the embodiments of this application can also implement these features or functions, which will not be elaborated here.
[0161] It should be understood that the processor in the embodiments of this application can be an integrated circuit chip with signal processing capabilities. In implementation, each step of the above method embodiments can be completed by the integrated logic circuitry in the processor's hardware or by instructions in software form. The processor described above can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.
[0162] It is understood that the memory in the embodiments of this application can be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory can be read-only memory (ROK), programmable ROK (PROK), erasable PROK (EPROK), electrically erasable programmable ROK (EEPROK), or flash memory. The volatile memory can be random access memory (RAK), which serves as an external cache. By way of example, but not limitation, many forms of RAK are available, such as static random access memory (SRAK), dynamic random access memory (DRAK), synchronous dynamic random access memory (SDRAK), double data rate synchronous dynamic random access memory (DDR SDRAK), enhanced synchronous dynamic random access memory (ESDRAK), synchronous linked dynamic random access memory (SLDRAK), and direct memory bus RAK (DR RAK). It should be noted that the memory used in the systems and methods described herein is intended to include, but is not limited to, these and any other suitable types of memory.
[0163] This application also provides a computer-readable storage medium including instructions that, when executed on a computer, cause the computer to perform the methods described in the foregoing embodiments.
[0164] This application also provides a computer program product containing instructions that, when run on a computer, cause the computer to perform the methods described in the foregoing embodiments.
[0165] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0166] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.
[0167] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0168] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0169] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0170] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer instructions. When the computer instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium accessible to a computer or a data storage device such as a server or data center that integrates one or more available media. The available media may be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., high-density digital video discs (DVDs)), or semiconductor media (e.g., solid-state disks (SSDs)).
Claims
1. A model training method, characterized in that, include: Acquire first voiceprint data, which is the feature vector of the first voiceprint audio to be tested, and the first voiceprint audio to be tested is the audio generated in the deployment environment of the terminal device. If the first voiceprint data meets the preset conditions, then the transfer learning model is trained based on the first voiceprint data and the first training data to obtain the trained transfer learning model. The first training data is feature data of known labels. The first voiceprint data and the first training data are mapped to the first training data according to the trained transfer learning model to obtain the second training data; The initialized classifier model is trained based on the second training data to obtain the trained classifier model.
2. The method according to claim 1, characterized in that, The method further includes: If the trained transfer learning model and the trained classifier model meet the acceptance criteria, then the trained transfer learning model and the trained classifier model are used for voiceprint detection.
3. The method according to claim 2, characterized in that, The acceptance criteria include: the error rates of the trained transfer learning model and the trained classifier model are lower than a preset threshold.
4. The method according to claim 2 or 3, characterized in that, The method further includes: If the trained transfer learning model and the trained classifier model do not meet the acceptance criteria, then retraining or parameter adjustment is required.
5. The method according to any one of claims 1 to 4, characterized in that, The acquisition of the first voiceprint data includes: In response to input operations from the user's graphical interface, the step of acquiring the first voiceprint data is performed.
6. The method according to any one of claims 1 to 5, characterized in that, The preset conditions include: The number of the first voiceprint data has reached a preset value; or, The collection time for the first voiceprint data reaches the preset duration.
7. The method according to claim 6, characterized in that, The method further includes: If the first voiceprint data meets the preset conditions, then the collection of the first voiceprint data will stop.
8. A voiceprint detection method, characterized in that, include: Acquire second voiceprint data, which is the feature vector of the second voiceprint audio to be tested, and the second voiceprint audio to be tested is the audio generated in the deployment environment of the terminal device. The second voiceprint data is processed using the trained transfer learning model to obtain voiceprint features; The voiceprint features are detected using the trained classifier model to obtain the detection results.
9. The method according to claim 8, characterized in that, The trained transfer learning model is obtained by training using the method described in any one of claims 1 to 7.
10. The method according to claim 8 or 9, characterized in that, The trained classifier model is obtained by training using the method described in any one of claims 1 to 7.
11. A model training device, characterized in that, include: An interface unit is used to acquire first voiceprint data, wherein the first voiceprint data is a feature vector of a first voiceprint audio to be tested, and the first voiceprint audio to be tested is audio generated in the deployment environment of the terminal device. The processing unit is configured to train the transfer learning model based on the first voiceprint data and the first training data if the first voiceprint data meets the preset conditions, thereby obtaining the trained transfer learning model, wherein the first training data is feature data of known labels. The processing unit is further configured to map the first voiceprint data and the first training data according to the trained transfer learning model to obtain the second training data; The processing unit is further configured to train the initialized classifier model based on the second training data to obtain the trained classifier model.
12. The apparatus according to claim 11, characterized in that, The processing unit is also used for: If the trained transfer learning model and the trained classifier model meet the acceptance criteria, then the trained transfer learning model and the trained classifier model are used for voiceprint detection.
13. The apparatus according to claim 12, characterized in that, The acceptance criteria include: the error rates of the trained transfer learning model and the trained classifier model are lower than a preset threshold.
14. The apparatus according to claim 12 or 13, characterized in that, The processing unit is also used for: If the trained transfer learning model and the trained classifier model do not meet the acceptance criteria, then retraining or parameter adjustment is required.
15. The apparatus according to any one of claims 11 to 14, characterized in that, The interface unit is specifically used for: In response to input operations from the user's graphical interface, the step of acquiring the first voiceprint data is performed.
16. The apparatus according to any one of claims 11 to 15, characterized in that, The preset conditions include: The number of the first voiceprint data has reached a preset value; or, The collection time for the first voiceprint data reaches the preset duration.
17. The apparatus according to claim 16, characterized in that, The interface unit is also used for: If the first voiceprint data meets the preset conditions, then the collection of the first voiceprint data will stop.
18. A voiceprint detection device, characterized in that, include: An interface unit is used to acquire second voiceprint data, which is a feature vector of a second voiceprint audio to be tested, and the second voiceprint audio to be tested is audio generated in the deployment environment of the terminal device. The processing unit is used to process the second voiceprint data according to the trained transfer learning model to obtain voiceprint features; The processing unit is also used to detect the voiceprint features based on the trained classifier model to obtain the detection results.
19. A model training device, characterized in that, include: A processor and a memory, wherein the processor is coupled to the memory; The memory is used to store programs; The processor is configured to execute a program in the memory, causing the model training apparatus to perform the method as described in any one of claims 1 to 7.
20. A voiceprint detection device, characterized in that, include: A processor and a memory, wherein the processor is coupled to the memory; The memory is used to store programs; The processor is configured to execute a program in the memory, causing the model training apparatus to perform the method as described in any one of claims 8 to 10.
21. A computer-readable storage medium comprising instructions that, when executed on a computer, cause the computer to perform the method as claimed in any one of claims 1 to 7, or cause the computer to perform the method as claimed in any one of claims 8 to 10.
22. A computer program product comprising instructions that, when run on a computer, cause the computer to perform the method as claimed in any one of claims 1 to 7, or cause the computer to perform the method as claimed in any one of claims 8 to 10.