Method for constructing voiceprint recognition model and related products
During the construction of the voiceprint recognition model, the sample center is determined based on the similarity of the speech samples, which solves the problem of wrong user labels, improves the accuracy and recognition effect of the model, and reduces the construction cost.
Patent Information
- Application Number
- CN202310011024.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-01-05
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2043-01-05
AI Technical Summary
During the construction of the voiceprint recognition model, it is impossible to ensure that the obtained user voice samples belong to the user's voice completely, resulting in wrong user tags and affecting the model training effect.
By obtaining multiple voice samples from at least two users, the voiceprint feature vector is extracted, and the same loss function and non-similar loss function are constructed to determine the sample center and improve the accuracy of the voiceprint recognition model.
It improves the accuracy of the voiceprint recognition model, improves the recognition effect, saves model construction time and cost, and improves user experience.
Smart Images

Figure CN116312555B_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the general data processing technology field of the Internet industry, and specifically relates to a method for constructing a voiceprint recognition model and related products. Background Art
[0002] When constructing a voiceprint recognition model, when using a user's telephone number or user ID as a label to construct a voice sample training set, it is impossible to ensure that multiple voices of the user obtained are all completely the voice of this user. Voice samples of non-this-user voices will be labeled as the voice of this user, resulting in voice samples with incorrect user labels in the training set.
[0003] Currently, by using a trained voiceprint recognition model to compare different voice segments under a certain user label and removing samples with a similarity less than a certain threshold, existing voiceprint recognition models often cannot distinguish such incorrect labels, resulting in the model being affected by incorrect label samples during training construction, leading to poor recognition effects. Or when there are incorrect labels in the samples, additional means are required, such as model filtering, manual user annotation, etc., to propose voice samples with incorrect labels, resulting in an increase in the time and cost of model construction. Summary of the Invention
[0004] This application provides a method for constructing a voiceprint recognition model and related products, aiming to determine a sample center according to the similarity between voice samples during the construction of the voiceprint recognition model, improve the accuracy of the voiceprint recognition model, enhance the recognition effect, save the time and cost of model construction, and enhance the user experience.
[0005] In a first aspect, an embodiment of this application provides a method for constructing a voiceprint recognition model, and the method includes:
[0006] Obtain multiple voice samples of at least two users, where the multiple voice samples include at least three voice samples of each user among the at least two users;
[0007] Extract multiple voiceprint feature vectors corresponding to the multiple voice samples;
[0008] Construct a within-class loss function based on the multiple voiceprint feature vectors. The within-class loss function is used to characterize that the similarity between any target voice sample of each user among the at least two users and its own first center point is the largest. The first center point is used to characterize the sum of the products of the first similarity weights of each first candidate voice sample of the target voice sample and its own voiceprint feature vector. The first candidate voice sample refers to the voice samples other than the target voice sample among all the voice samples of the user corresponding to the target voice sample. The first similarity weight refers to the proportion of the similarity between the target voice sample and a single first candidate voice sample in the sum of the similarities of all the first candidate voice samples of the target voice sample. The first similarity refers to the similarity between the target voice sample and a single first candidate voice sample;
[0009] Construct a between-class loss function based on the multiple voiceprint feature vectors. The between-class loss function is used to characterize that the similarity between any target voice sample of each user among the at least two users and its own second center point is the smallest. The second center point is used to characterize the sum of the products of the second similarity weights of each second candidate voice sample of the target voice sample and its own voiceprint feature vector. The second candidate voice sample refers to the voice samples of other users except the user corresponding to the target voice sample. The second similarity weight refers to the proportion of the similarity between the target voice sample and a single second candidate voice sample in the sum of the similarities of all the second candidate voice samples of the target voice sample. The second similarity refers to the similarity between the target voice sample and a single second candidate voice sample;
[0010] Obtain a target loss function according to the within-class loss function and the between-class loss function. The target loss function is used to characterize the correct voice samples and incorrect voice samples among at least three voice samples of each user among the at least two users. The correct voice sample is used to characterize that the correspondence between the current voice sample and the user is correct, and the incorrect voice sample is used to characterize that the correspondence between the current voice sample and the user is incorrect;
[0011] Update the parameters of the voiceprint recognition model according to the target loss function to obtain a trained voiceprint recognition model.
[0012] In a second aspect, an embodiment of the present application provides a device for constructing a voiceprint recognition model. The device includes:
[0013] An acquisition unit, configured to acquire multiple voice samples of at least two users, where the multiple voice samples include at least three voice samples of each user among the at least two users;
[0014] An extraction unit, configured to extract multiple voiceprint feature vectors corresponding to the multiple voice samples;
[0015] A first construction unit, configured to construct a same-class loss function according to the multiple voiceprint feature vectors. The same-class loss function is used to characterize that the similarity between any target voice sample of each user among the at least two users and its own first center point is the largest. The first center point is used to characterize the sum of the products of the first similarity weights of each first candidate voice sample of the target voice sample and its own voiceprint feature vector. The first candidate voice sample refers to the voice samples other than the target voice sample among all the voice samples of the user corresponding to the target voice sample. The first similarity weight refers to the proportion of the similarity between the target voice sample and a single first candidate voice sample in the sum of the similarities of all the first candidate voice samples of the target voice sample. The first similarity refers to the similarity between the target voice sample and a single first candidate voice sample;
[0016] A second construction unit, configured to construct a different-class loss function according to the multiple voiceprint feature vectors. The different-class loss function is used to characterize that the similarity between any target voice sample of each user among the at least two users and its own second center point is the smallest. The second center point is used to characterize the sum of the products of the second similarity weights of each second candidate voice sample of the target voice sample and its own voiceprint feature vector. The second candidate voice sample refers to the voice samples of other users except the user corresponding to the target voice sample. The second similarity weight refers to the proportion of the similarity between the target voice sample and a single second candidate voice sample in the sum of the similarities of all the second candidate voice samples of the target voice sample. The second similarity refers to the similarity between the target voice sample and a single second candidate voice sample;
[0017] A loss function generation unit, configured to obtain a target loss function according to the same-class loss function and the different-class loss function. The target loss function is used to characterize the correct voice samples and incorrect voice samples among at least three voice samples of each user among the at least two users. The correct voice sample is used to characterize that the correspondence between the current voice sample and the user is correct, and the incorrect voice sample is used to characterize that the correspondence between the current voice sample and the user is incorrect;
[0018] A model training unit, configured to update the parameters of the voiceprint recognition model according to the target loss function to obtain a trained voiceprint recognition model.
[0019] In a third aspect, an embodiment of the present application provides an electronic device, including an application processor, a communication module, a memory, and one or more programs. The application processor is communicatively connected to the memory and the communication module through an internal communication bus. The one or more programs are stored in the memory and are configured to be executed by the application processor. The one or more programs include instructions for performing the steps in the method described in the first aspect of the embodiment of the present application.
[0020] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, on which a computer program / instructions are stored. When the computer program / instructions are executed by an application processor, the steps of the method described in the first aspect of the embodiment of the present application are implemented.
[0021] In a fifth aspect, an embodiment of the present application provides a computer program product, including a computer program / instructions. When the computer program / instructions are executed by an application processor, the steps of the method described in the first aspect of the embodiment of the present application are implemented.
[0022] It can be seen that in the embodiments of the present application, first, a plurality of voice samples of at least two users are obtained. The plurality of voice samples include at least three voice samples of each user among the at least two users. Then, a plurality of voiceprint feature vectors corresponding to the plurality of voice samples are extracted. After that, a same-class loss function is constructed according to the plurality of voiceprint feature vectors. The same-class loss function is used to characterize that the similarity between any target voice sample of each user among the at least two users and its own first center point is the largest. The first center point is used to characterize the sum of the products of the first similarity weights of each first candidate voice sample of the target voice sample and its own voiceprint feature vector. The first candidate voice sample refers to the voice samples other than the target voice sample among all the voice samples of the user corresponding to the target voice sample. The first similarity weight refers to the proportion of the similarity between the target voice sample and a single first candidate voice sample in the sum of the similarities of all the first candidate voice samples of the target voice sample. The first similarity refers to the similarity between the target voice sample and a single first candidate voice sample. Then, a different-class loss function is constructed according to the plurality of voiceprint feature vectors. The different-class loss function is used to characterize that the similarity between any target voice sample of each user among the at least two users and its own second center point is the smallest. The second center point is used to characterize the sum of the products of the second similarity weights of each second candidate voice sample of the target voice sample and its own voiceprint feature vector. The second candidate voice sample refers to the voice samples of other users except the user corresponding to the target voice sample. The second similarity weight refers to the proportion of the similarity between the target voice sample and a single second candidate voice sample in the sum of the similarities of all the second candidate voice samples of the target voice sample. The second similarity refers to the similarity between the target voice sample and a single second candidate voice sample. Next, a target loss function is obtained according to the same-class loss function and the different-class loss function. The target loss function is used to characterize the correct voice samples and incorrect voice samples among at least three voice samples of each user among the at least two users. The correct voice sample is used to characterize that the corresponding relationship between the current voice sample and the user is correct, and the incorrect voice sample is used to characterize that the corresponding relationship between the current voice sample and the user is incorrect. Finally, the parameters of the voiceprint recognition model are updated according to the target loss function to obtain the trained voiceprint recognition model. In this way, it is possible to determine the sample center according to the similarity between voice samples during the construction of the voiceprint recognition model, improve the accuracy of the voiceprint recognition model, enhance the recognition effect, save the time and cost of model construction, and improve the user experience. Description of the Drawings
[0023] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0024] Figure 1 is a schematic structural diagram of an electronic device provided by an embodiment of the present application;
[0025] Figure 2 is a schematic flowchart of a method for constructing a voiceprint recognition model provided by an embodiment of the present application;
[0026] Figure 3 is a schematic flowchart of a method for constructing a homogeneous loss function provided by an embodiment of the present application;
[0027] Figure 4 is a schematic flowchart of a method for constructing a non - homogeneous loss function provided by an embodiment of the present application;
[0028] Figure 5a is a block diagram of the functional units of a voiceprint recognition model construction device provided by an embodiment of the present application;
[0029] Figure 5b is a block diagram of the functional units of another voiceprint recognition model construction device provided by an embodiment of the present application. Detailed implementation manners
[0030] In order to enable those skilled in the art of the present technology to better understand the solution of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without making creative efforts belong to the scope of protection of the present application.
[0031] The terms "first", "second", etc. in the specification and claims of the present application and the above - mentioned drawings are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "include" and "have" and any variations thereof are intended to cover non - exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products, or devices.
[0032] Referring to "embodiment" in this article means that a specific feature, structure, or characteristic described in connection with the embodiment can be included in at least one embodiment of the present application. The phrase appears in various places in the specification and does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.
[0033] First, the system architecture involved in the embodiments of the present application will be introduced.
[0034] To better understand the technical solutions of the embodiments of the present application, first, the electronic devices that the embodiments of the present application may involve will be introduced.
[0035] Please refer to Figure 1 , Figure 1 which is a schematic structural diagram of an electronic device provided by the embodiments of the present application. As Figure 1 shown, the electronic device includes one or more application processors 120, a memory 130, a communication module 140, and one or more programs 131. The application processor 120 is communicatively connected to the memory 130 and the communication module 140 through an internal communication bus.
[0036] Among them, the one or more programs 131 are stored in the memory 130 and are configured to be executed by the application processor 120. The one or more programs 131 include instructions for executing any step in the above method embodiments.
[0037] Among them, the application processor 120 can be, for example, a central application processor (Central Processing Unit, CPU), a general-purpose application processor, a digital signal application processor (Digital Signal Processor, DSP), an application-specific integrated circuit (Application-Specific Integrated Circuit, ASIC), a field-programmable gate array (Field Programmable Gate Array, FPGA) or other programmable logic devices, transistor logic devices, hardware components or any combination thereof. It can implement or execute various exemplary logic blocks, units, and circuits described in connection with the disclosure of the present application. The application processor 120 can also be a combination that implements computing functions, such as a combination of one or more micro-application processors, a combination of a DSP and a micro-application processor, and so on. The communication unit can be the communication module 140, a transceiver, a transceiver circuit, etc., and the storage unit can be the memory 130.
[0038] The memory 130 can be a volatile memory, a non-volatile memory, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), or a flash memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of random access memory (RAM) are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchlink DRAM (SLDRAM), and directrambus RAM (DR RAM).
[0039] Among them, the electronic device can be a server, specifically, it can be a single server, or a server cluster composed of several servers, or a cloud computing service center. The electronic device can also be a terminal device, specifically, it can be a mobile phone terminal, a tablet computer, a laptop computer, etc.
[0040] Based on this, an embodiment of the present application provides a method for constructing a voiceprint recognition model. The following will describe the embodiment of the present application in detail with reference to the accompanying drawings.
[0041] Please refer to Figure 2 , Figure 2 which is a schematic flowchart of a method for constructing a voiceprint recognition model provided by an embodiment of the present application. The method is applied to an electronic device as shown in Figure 1 and, as shown in Figure 2 , the method includes steps 201-step 206:
[0042] Step 201, obtain multiple voice samples of at least two users, where the multiple voice samples include at least three voice samples of each user among the at least two users.
[0043] Among them, obtaining multiple voice samples of at least two users can be: randomly obtaining multiple voice samples of at least two users from a voice data set. This voice data set can be a set constructed using online recording data. At this time, in this voice data set, the voice samples of users may be incorrect voice samples. An incorrect voice sample means that the corresponding relationship between the voice sample and the user is incorrect, that is, a voice sample that is not the voice sample of the user himself is classified as the voice sample of this user. The voice samples of users may be correct voice samples. A correct voice sample means that the corresponding relationship between the voice sample and the user is correct. For example, in the project voice of Project A, there are situations such as call forwarding, call transfer, and background user voices in the same user voice communication. Therefore, in the obtained voice data set, the same user voice may have incorrect voice samples that are not the voice samples of the user himself.
[0044] In specific implementation, the voiceprint recognition model can be trained multiple times, and the number of voice samples for each training can be the same or different.
[0045] Step 202: Extract multiple voiceprint feature vectors corresponding to the multiple voice samples.
[0046] First, obtain the voiceprint feature H of each voice sample through existing voice feature engineering. For example, the voiceprint feature can be a feature based on a filter bank (Filter bank, Fbank), Mel-frequency cepstral coefficients feature (Mel-Frequency Ceptral Coefficients, MFCC), logarithmic Mel-spectrum feature, etc.
[0047] Then, use a deep learning voice framework, such as the Resnet network, ECAPA-TDNN, etc. to construct an encoder of the voiceprint recognition model, and obtain the voiceprint feature vector of each voiceprint feature.
[0048] For example, when the number of at least two users is n and the number of at least three voice samples is m, the size of the training data for one time is n*m. For example, n = 32 and m = 5, that is, 5 voice samples of each of the 32 users are trained at one time, and the multiple voiceprint feature vectors are x i, , where i ∈ [0, n - 1], j ∈ [0, m - 1]. i is the user serial number, j is the serial number of the voice sample of a single user, and x i, refers to the voiceprint feature vector of the target voice sample with the voice sample serial number j of the user with the user serial number i.
[0049] Step 203: Construct a homogeneous loss function according to the multiple voiceprint feature vectors.
[0050] Among them, the same-class loss function is used to characterize that the similarity between any target voice sample of each user among the at least two users and its own first center point is the largest. The first center point is used to characterize the sum of the products of the first similarity weights of each first candidate voice sample of the target voice sample and its own voiceprint feature vector. The first candidate voice sample refers to all voice samples of the user corresponding to the target voice sample except the target voice sample. The first similarity weight refers to the proportion of the similarity between the target voice sample and a single first candidate voice sample in the sum of the similarities of all first candidate voice samples of the target voice sample. The first similarity refers to the similarity between the target voice sample and a single first candidate voice sample.
[0051] Next, the process of constructing the same-class loss function in the embodiments of the present application will be described.
[0052] Please refer to Figure 3 , Figure 3 which is a schematic diagram of a process for constructing a same-class loss function provided by an embodiment of the present application. As Figure 3 shown, the process includes steps A1 - A4:
[0053] Step A1: Calculate the similarities between different voiceprint feature vectors of the same user among the at least two users respectively to obtain the first similarity matrix of the at least two users.
[0054] Specifically, the step of calculating the similarities between different voiceprint feature vectors of the same user among the at least two users respectively to obtain the first similarity matrix of the at least two users includes performing the following operations for each voice sample of each user among the at least two users: determining the current voice sample of the current user as the target voice sample; determining the voiceprint feature vector of the target voice sample as the target vector; determining the other voice samples of the current user except the current voice sample as the first candidate voice samples; determining the voiceprint feature vectors of the first candidate voice samples as the first candidate vectors; calculating multiple first similarities between the target vector and the first candidate vectors respectively.
[0055] For n users in each training batch, each user has m voice samples, n is greater than or equal to 2, and m is greater than or equal to 3. For the same user, one voice sample is sequentially selected as the target voice sample, the voiceprint feature vector of the target voice sample is the target vector, the remaining m - 1 voices are used as the first candidate voice samples, the voiceprint feature vectors of the first candidate voice samples are the first candidate vectors, and the first similarity matrix is calculated.
[0056] For example, when m = 5, five voice samples [y0, y1, y2, y3, y4] of a user can obtain voiceprint feature vectors [x0, x1, x2, x3, x4]. When y0 is selected as the target voice sample, then [y1, y2, y3, y4] are the first candidate voice samples, x0 is the target vector, and [x1, x2, x3, x4] are the first candidate vectors.
[0057] The first similarity is: sim(i, j, k) = cos(x i, , x i,, ), where i ∈ [0, n - 1], j ∈ [0, m - 1], k ∈ [0, m - 1] and j ≠ k. Here, x i,, refers to the voiceprint feature vector corresponding to the first candidate voice with the voice sample number k corresponding to the target voice sample with the voice sample number j of the user with the user number i.
[0058] Therefore, for one user, a first similarity matrix of size m * (m - 1) can be obtained. For n users in the same batch, a first similarity matrix of size n * m * (m - 1) can be obtained.
[0059] Step A2: Calculate the first similarity weights of each voice sample of each user among the at least two users with each first candidate voice sample in all the first candidate voice samples of itself, to obtain the first similarity weight matrix of the at least two users.
[0060] It can be seen that according to the first similarity matrix in step A1, a weight size measuring the relationship between the first candidate voice sample and the target voice sample can be calculated, thus creating conditions for constructing a central vector based on the first candidate voice sample.
[0061] Based on the above m = 5, five voiceprint feature vectors [y0, y1, y2, y3, y4] can be obtained from five voice samples [x0, x1, x2, x3, x4] of a user. When x0 is selected as the target voice sample, then [x1, x2, x3, x4] are the first candidate voices, y0 is the target vector, and [y1, y2, y3, y4] are the first candidate vectors.
[0062] The first similarity weight a i,, is calculated as follows:
[0063] where a i,, refers to the similarity weight between the first candidate voice with the voice sample number k corresponding to the target voice sample with the voice sample number j of the user with the user number i and all the first candidate voices corresponding to the target voice sample with the voice sample number j of the user with the user number i.
[0064] At this time, γ is a hyperparameter, for example, γ = 0.05.
[0065] Therefore, for one user, a first similarity weight matrix of size m*(m - 1) can be obtained. For N users in the same batch, a first similarity weight matrix of size n*m*(m - 1) can be obtained.
[0066] Step A3: Obtain the first center point matrix of the at least two users according to the first similarity weight matrix and the first voiceprint feature vector matrix.
[0067] Based on the foregoing when m = 5, 5 voice samples [x0, x1, x2, x3, x4] of one user can obtain voiceprint feature vectors [y0, y1, y2, y3, y4]. According to the first similarity weights in step A2, the first center point c of n users can be constructed using the m - 1 first candidate vectors of the m - 1 first candidate voice samples of each target voice sample of each user among the n users. i,j . Among them, At this time k ≠ j, c i, refers to the first center point of the target voice sample with voice sample serial number j of the user with user serial number i, and a i,, *x i,, refers to the product of the first similarity weight of the first candidate voice with voice sample serial number k of the target voice sample with voice sample serial number j of the user with user serial number i and the voiceprint feature vector.
[0068] Therefore, for one user, m first center points can be obtained. For n users, n*m first center points can be obtained in one training.
[0069] Step A4: Construct the same - class loss function according to the first voiceprint feature vector matrix and the first center point matrix.
[0070] Based on the foregoing when m = 5, the same - class loss function L1 is constructed as follows.
[0071] L1 = -(w * cos(x i, , c i, ) + b).
[0072] Among them, w and b are hyperparameters. For example, w = 10, b = - 5, i is the user serial number, j is the voice sample serial number, x i, is the voiceprint feature vector of the target voice sample with voice sample serial number j of the user with user serial number i. c i, is the first center point of the target voice sample with voice sample serial number j of the user with user serial number i.
[0073] It can be seen that for the m voice samples belonging to the same user, when each voice sample is used as the target voice sample, the similarity with the center point constructed based on its corresponding first candidate voice is the largest, that is, the cosine angle is the smallest.
[0074] Preferably, step A4 includes: determining whether there is a historical global center vector matrix of the at least two users; if so, updating the historical global center vector matrix according to the first center point matrix to obtain an updated global center vector matrix; obtaining the updated first center point matrix of the at least two users according to the first similarity weight matrix and the updated global center vector matrix; obtaining the same-class loss function according to the first voiceprint feature vector matrix and the updated first center point matrix.
[0075] For example, y0 is the target voice sample of a user, and [y1, y2, y3, y4] are the first candidate voice samples.
[0076] Suppose represents the global center vector of the i-th user in the current training batch, represents the global center vector of the i-th user in the previous training batch of the current training batch, The calculation method of is as follows:
[0077]
[0078] where t is the training batch, h is a hyperparameter, for example h = 0.99, c i is the center point determined for the m voice samples of this user, and the calculation method is as follows:
[0079]
[0080] At this time, a i, The calculation method is as follows:
[0081] u i,j = s T *tanh(x i,j *v + d)
[0082] In this formula, s T is the transpose of s, s, v, and d are matrices, hyperparameters, and weights trained.
[0083] Since there may be errors in the target voice samples, the similarity weights between the target voice samples of each user and their global center vectors are calculated at this time.
[0084]
[0085] At this time, γ is a hyperparameter, for example γ = 0.05.
[0086] It can be seen that if x0 is an incorrect voice sample, there is a large difference between x0 and the global central voiceprint vector of the user himself. A global central vector is constructed for each user for filtering, so that the voiceprint recognition model can avoid the influence of incorrect voice samples.
[0087] Optionally, after constructing the same-class loss function according to the first voiceprint feature vector matrix and the first center point matrix, if not, obtain a preset first hyperparameter and a second hyperparameter; construct the same-class loss function according to the first voiceprint feature vector matrix, the first center point matrix, the first hyperparameter and the second hyperparameter.
[0088] Step 204, construct a non-same-class loss function according to the multiple voiceprint feature vectors.
[0089] Wherein, the non-same-class loss function is used to characterize that the similarity between any target voice sample of each user among at least two users and its own second center point is the smallest. The second center point is used to characterize the sum of the products of the second similarity weights of each second candidate voice sample of the target voice sample and its own voiceprint feature vector. The second candidate voice sample refers to the voice samples of other users except the user corresponding to the target voice sample. The second similarity weight refers to the proportion of the similarity between the target voice sample and a single second candidate voice sample in the sum of the similarities of all second candidate voice samples of the target voice sample. The second similarity refers to the similarity between the target voice sample and a single second candidate voice sample.
[0090] Next, the process of constructing the non-same-class loss function in the embodiments of the present application will be described.
[0091] Please refer to Figure 4 , Figure 4 is a schematic flowchart of a process for constructing a non-same-class loss function provided by an embodiment of the present application. As Figure 4 shown, the process includes steps B1-step B4:
[0092] Step B1, calculate the similarity between the voiceprint feature vectors of different users among the at least two users respectively, and obtain the second similarity matrix of the at least two users.
[0093] Specifically, calculating the similarity between different voiceprint feature vectors of the same user among the at least two users to obtain the second similarity matrix of the at least two users includes performing the following operations on each voice sample of each user among the at least two users: determining the current voice sample of the current user as the target voice sample; determining the voiceprint feature vector of the target voice sample as the target vector; determining the voice samples of other users except the current user as the second candidate voice samples; determining the voiceprint feature vectors of the second candidate voice samples as the second candidate vectors; and respectively calculating multiple second similarities between the target vector and the second candidate vectors.
[0094] For n users in each training batch, each user has m voice samples, n is greater than or equal to 2, and m is greater than or equal to 3. For the same user, one voice sample is sequentially selected as the target voice sample, the voiceprint feature vector of the target voice sample is the target vector, and the (n - 1)*m voices of other users are used as the second candidate voice samples. The voiceprint feature vectors of the second candidate voice samples are the second candidate vectors, and the second similarity matrix is calculated.
[0095] For example, when m = 5, the 5 voice samples [y0, y1, y2, y3, y4] of a user can obtain voiceprint feature vectors [x0, x1, x2, x3, x4]. The other 31 users have a total of 155 voice samples. When x0 is selected as the target voice sample, the 155 voice samples of the other 31 users are the second candidate voices, x0 is the target vector, and the voiceprint feature vectors of the 155 voice samples of the other 31 users are the second candidate vectors.
[0096] The second similarity is:
[0097] sim(i,j,p,u) = cos(x i, ,x p, ), x [, refers to the voiceprint feature vector corresponding to the second candidate voice with the voice sample number u of other users with the user number p.
[0098] Therefore, for one user, a second similarity matrix of size m*(n - 1)*m can be obtained. For n users in the same batch, a second similarity matrix of size n*m*(n - 1)*m can be obtained.
[0099] Step B2: Calculate the second similarity weights between the target voice sample of each user among the at least two users and each second candidate voice sample in all its second candidate voice samples according to the second similarity matrix to obtain the second similarity matrix of the at least two users.
[0100] It can be seen that, according to the second similarity matrix in step B1, a weight measuring the similarity between the second candidate voice sample and the target voice sample can be calculated, thus creating conditions for constructing a central vector based on the second candidate voice sample.
[0101] Based on the above-mentioned case where m = 5, 5 voice samples [x0, x1, x2, x3, x4] of a user can obtain voiceprint feature vectors [y0, y1, y2, y3, y4]. When x0 is selected as the target voice sample, the (n - 1) * m voice samples of other users are used as the second candidate voice samples.
[0102] The calculation of the second similarity weight is as follows:
[0103] a i,,, refers to the similarity weight between the second candidate voice corresponding to the target voice sample with voice sample serial number j of the target user with user serial number i and all the second candidate voices of the second candidate voice with voice sample serial number u of other users with user serial number p.
[0104] At this time, γ is a hyperparameter, for example, γ = 0.05.
[0105] Therefore, for one user, a second similarity weight matrix of size m * (n - 1) * m can be obtained. For n users in the same batch, a second similarity weight matrix of size n * m * (n - 1) * m can be obtained.
[0106] Step B3: Obtain the second center point matrix of the at least two users according to the second similarity weight matrix and the second voiceprint feature vector matrix.
[0107] Based on the above-mentioned case where m = 5, 5 voice samples [x0, x1, x2, x3, x4] of a user can obtain voiceprint feature vectors [y0, y1, y2, y3, y4]. According to the second similarity weight in step B2, the second center point c can be constructed using the (n - 1) * m second candidate vectors of the (n - 1) * m second candidate voice samples of each of the n users k, :
[0108] At this time, l ≠ q, where i is the serial number of the user, and c k, refers to the second center point of the target voice sample with voice sample serial number j of the user with user serial number i, and a k, *x k, refers to the product of the second similarity weight of the second candidate voice with voice sample serial number q of other users with user serial number k and the voiceprint feature vector.
[0109] Therefore, for one user, m*(n - 1)*m second central points can be obtained. For n users, n*m*(n - 1)*m second central points can be obtained in one training.
[0110] Step B4: Construct the non - same - class loss function according to the second voiceprint feature vector matrix and the second central point matrix.
[0111] Based on the foregoing when m = 5, the non - same - class loss function L2 is constructed as:
[0112]
[0113] where, sim i,,, = w*cos(x i, , c k, ), w and b are hyperparameters. For example, w = 10, b = - 5, i is the user number, j is the voice sample number, where x i, is the j - th voice sample of the i - th user, c k, is the l - th second candidate central point of the k - th user, i≠k.
[0114] It can be seen that for the (n - 1)*m samples that do not belong to one user, when each voice sample of the same user is used as the target voice sample, the similarity with the central point constructed based on its corresponding second candidate voice is the smallest, that is, the cosine angle is the largest.
[0115] Step 205: Obtain the target loss function according to the same - class loss function and the non - same - class loss function.
[0116] Among them, the target loss function is used to represent the correct voice samples and incorrect voice samples in at least three voice samples of each user among the at least two users. The correct voice samples are used to represent that the corresponding relationship between the current voice sample and the user is correct, and the incorrect voice samples are used to represent that the corresponding relationship between the current voice sample and the user is incorrect.
[0117] The target loss function Loss is:
[0118]
[0119] Step 206: Update the parameters of the voiceprint recognition model according to the target loss function to obtain the trained voiceprint recognition model.
[0120] Training the voiceprint recognition model based on the above objective loss function Loss, the voiceprint recognition model will tend to correct voice samples. It can be seen that in the embodiments of the present application, first, multiple voice samples of at least two users are obtained, and the multiple voice samples include at least three voice samples of each user among the at least two users. Then, multiple voiceprint feature vectors corresponding to the multiple voice samples are extracted. After that, a within-class loss function is constructed according to the multiple voiceprint feature vectors. The within-class loss function is used to characterize that the similarity between any target voice sample of each user among the at least two users and its own first center point is the largest. The first center point is used to characterize the sum of the products of the first similarity weights of each first candidate voice sample of the target voice sample and its own voiceprint feature vector. The first candidate voice sample refers to the voice samples other than the target voice sample among all the voice samples of the user corresponding to the target voice sample. The first similarity weight refers to the proportion of the similarity between the target voice sample and a single first candidate voice sample in the sum of the similarities of all the first candidate voice samples of the target voice sample. The first similarity refers to the similarity between the target voice sample and a single first candidate voice sample. Then, a between-class loss function is constructed according to the multiple voiceprint feature vectors. The between-class loss function is used to characterize that the similarity between any target voice sample of each user among the at least two users and its own second center point is the smallest. The second center point is used to characterize the sum of the products of the second similarity weights of each second candidate voice sample of the target voice sample and its own voiceprint feature vector. The second candidate voice sample refers to the voice samples of other users except the user corresponding to the target voice sample. The second similarity weight refers to the proportion of the similarity between the target voice sample and a single second candidate voice sample in the sum of the similarities of all the second candidate voice samples of the target voice sample. The second similarity refers to the similarity between the target voice sample and a single second candidate voice sample. Next, an objective loss function is obtained according to the within-class loss function and the between-class loss function. The objective loss function is used to characterize the correct voice samples and incorrect voice samples among at least three voice samples of each user among the at least two users. The correct voice sample is used to characterize that the corresponding relationship between the current voice sample and the user is correct, and the incorrect voice sample is used to characterize that the corresponding relationship between the current voice sample and the user is incorrect. Finally, the parameters of the voiceprint recognition model are updated according to the objective loss function to obtain the trained voiceprint recognition model. In this way, it is possible to determine the sample center according to the similarity between voice samples during the construction of the voiceprint recognition model, improve the accuracy of the voiceprint recognition model, enhance the recognition effect, save the time and cost of model construction, and improve the user experience.
[0121] It can be understood that since the method embodiment and the device embodiment are different presentation forms of the same technical concept, the content of the method embodiment part in the present application should be synchronously adapted to the device embodiment part, and will not be elaborated here.
[0122] Consistent with the embodiments shown above, as Figure 5a shown, Figure 5a is a functional unit composition block diagram of a voiceprint recognition model construction device provided by an embodiment of the present application. In Figure 5a the voiceprint recognition model construction device 500 is applied to an electronic device, and the voiceprint recognition model construction device 500 includes:
[0123] An acquisition unit 501, configured to acquire a plurality of voice samples of at least two users, where the plurality of voice samples include at least three voice samples of each user among the at least two users;
[0124] An extraction unit 502, configured to extract a plurality of voiceprint feature vectors corresponding to the plurality of voice samples;
[0125] A first construction unit 503, configured to construct a same-class loss function according to the plurality of voiceprint feature vectors, where the same-class loss function is used to characterize that the similarity between any target voice sample of each user among the at least two users and its own first center point is the largest, and the first center point is used to characterize the sum of the products of the first similarity weights of each first candidate voice sample of the target voice sample and its own voiceprint feature vector, the first candidate voice sample refers to the voice samples other than the target voice sample among all the voice samples of the user corresponding to the target voice sample, the first similarity weight refers to the proportion of the similarity between the target voice sample and a single first candidate voice sample in the sum of the similarities between the target voice sample and all its first candidate voice samples, and the first similarity refers to the similarity between the target voice sample and a single first candidate voice sample;
[0126] A second construction unit 504, configured to construct a non-same-class loss function according to the plurality of voiceprint feature vectors, where the non-same-class loss function is used to characterize that the similarity between any target voice sample of each user among the at least two users and its own second center point is the smallest, and the second center point is used to characterize the sum of the products of the second similarity weights of each second candidate voice sample of the target voice sample and its own voiceprint feature vector, the second candidate voice sample refers to the voice samples of other users except the user corresponding to the target voice sample, the second similarity weight refers to the proportion of the similarity between the target voice sample and a single second candidate voice sample in the sum of the similarities between the target voice sample and all its second candidate voice samples, and the second similarity refers to the similarity between the target voice sample and a single second candidate voice sample;
[0127] A loss function generation unit 505 is configured to obtain a target loss function according to the same-class loss function and the different-class loss function, where the target loss function is used to characterize correct speech samples and incorrect speech samples in at least three speech samples of each of the at least two users. The correct speech samples are used to characterize that the corresponding relationship between the current speech sample and the user is correct, and the incorrect speech samples are used to characterize that the corresponding relationship between the current speech sample and the user is incorrect.
[0128] A model training unit 506 is configured to update parameters of the voiceprint recognition model according to the target loss function to obtain a trained voiceprint recognition model.
[0129] It can be understood that since the method embodiment and the device embodiment are different presentation forms of the same technical concept, the content of the method embodiment part in this application should be synchronously adapted to the device embodiment part, and will not be elaborated here.
[0130] In the case of adopting an integrated unit, as Figure 5b shown, Figure 5b FIG. is a block diagram of the functional units of another voiceprint recognition model construction device provided by an embodiment of the present application. In Figure 5b it, the voiceprint recognition model construction device 510 includes: a processing module 512 and a communication module 511.
[0131] The processing module 512 is configured to obtain multiple voice samples of at least two users through the communication module 511, where the multiple voice samples include at least three voice samples of each of the at least two users; and extract multiple voiceprint feature vectors corresponding to the multiple voice samples; and construct a same-class loss function according to the multiple voiceprint feature vectors, where the same-class loss function is used to characterize that the similarity between any target voice sample of each of the at least two users and its own first center point is the largest, and the first center point is used to characterize the sum of the products of the first similarity weights of each first candidate voice sample of the target voice sample and its own voiceprint feature vector, the first candidate voice sample refers to the voice samples other than the target voice sample among all the voice samples of the user corresponding to the target voice sample, the first similarity weight refers to the proportion of the similarity between the target voice sample and a single first candidate voice sample in the sum of the similarities of all the first candidate voice samples of the target voice sample, and the first similarity refers to the similarity between the target voice sample and a single first candidate voice sample; and construct a different-class loss function according to the multiple voiceprint feature vectors, where the different-class loss function is used to characterize that the similarity between any target voice sample of each of the at least two users and its own second center point is the smallest, and the second center point is used to characterize the sum of the products of the second similarity weights of each second candidate voice sample of the target voice sample and its own voiceprint feature vector, the second candidate voice sample refers to the voice samples of other users except the user corresponding to the target voice sample, the second similarity weight refers to the proportion of the similarity between the target voice sample and a single second candidate voice sample in the sum of the similarities of all the second candidate voice samples of the target voice sample, and the second similarity refers to the similarity between the target voice sample and a single second candidate voice sample; and obtain a target loss function according to the same-class loss function and the different-class loss function, where the target loss function is used to characterize the correct voice samples and incorrect voice samples among at least three voice samples of each of the at least two users, the correct voice sample is used to characterize that the corresponding relationship between the current voice sample and the user is correct, and the incorrect voice sample is used to characterize that the corresponding relationship between the current voice sample and the user is incorrect; and update the parameters of the voiceprint recognition model according to the target loss function to obtain a trained voiceprint recognition model.
[0132] For example, the processing module 512 executes some steps of the obtaining unit 501, the extracting unit 502, the first constructing unit 503, the second constructing unit 504, the loss function generating unit 505, and the model training unit 506, and / or is used to execute other processes of the technologies described herein. The communication module 511 is used to support the interaction between the voiceprint recognition model construction device 510 and other devices. AsFigure 5b As shown in the figure, the voiceprint recognition model construction device 510 may further include a storage module 513, and the storage module 513 is used to store the program code and data of the voiceprint recognition model construction device 510.
[0133] Among them, the processing module 512 may be an application processor or a controller. For example, it may be a Central Processing Unit (CPU), a general application processor, a Digital Signal Processor (DSP), an ASIC, an FPGA, or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of this application. The application processor may also be a combination that implements computing functions, such as a combination of one or more micro-application processors, a combination of a DSP and a micro-application processor, and so on. The communication module 511 may be a transceiver, an RF circuit, or a communication interface, etc. The storage module 513 may be a memory.
[0134] Among them, all relevant contents of each scenario involved in the above method embodiments can be cited in the function descriptions of the corresponding functional modules, and will not be elaborated here. The above voiceprint recognition model construction device 510 can all execute the above Figure 2 shown voiceprint recognition model construction method.
[0135] The above embodiments can be implemented in whole or in part by software, hardware, firmware, or any other arbitrary combination. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions according to the embodiments of this application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired or wireless manner. The computer-readable storage medium can be any available medium that the computer can access or a data storage device such as a server or a data center that includes one or more collections of available media. The available medium can be a magnetic medium (such as a floppy disk, a hard disk, a magnetic tape), an optical medium (such as a DVD), or a semiconductor medium. The semiconductor medium can be a solid-state drive.
[0136] Embodiments of the present application also provide a computer storage medium, on which computer programs / instructions are stored. When the computer programs / instructions are executed by an application processor, some or all of the steps of any of the methods described in the foregoing method embodiments are implemented.
[0137] Embodiments of the present application also provide a computer program product, including computer programs / instructions. When the computer programs / instructions are executed by an application processor, the steps of the method described in the first aspect of the embodiments of the present application are implemented.
[0138] It should be understood that in various embodiments of the present application, the sequence numbers of the foregoing processes do not mean the order of execution. The execution order of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.
[0139] In several embodiments provided by the present application, it should be understood that the disclosed methods, apparatuses, and systems can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for example, the division of the units is only a logical function division, and there may be other division methods in actual implementation; for example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings, direct couplings, or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of the apparatuses or units can be in electrical, mechanical, or other forms.
[0140] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0141] In addition, in each embodiment of the present invention, the functional units can be integrated in one processing unit, or each unit can be physically included separately, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware, or in the form of a combination of hardware and software functional units.
[0142] The integrated unit implemented in the form of software functional units can be stored in a computer-readable storage medium. The above-mentioned software functional units are stored in a storage medium and include several instructions for causing a computer device (which may be a user computer, a server, or a network device, etc.) to execute some steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: USB flash drive, mobile hard disk, magnetic disk, optical disk, volatile memory, or non-volatile memory. Among them, the non-volatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. The volatile memory may be random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of random access memory (RAM) are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct rambus random access memory (DR RAM), etc., all of which are media that can store program code.
[0143] Although the present invention is disclosed as above, the present invention is not limited thereto. Any person skilled in the art can easily think of changes or substitutions without departing from the spirit and scope of the present invention, and can make various changes and modifications, including combinations of the above different functions and implementation steps, including software and hardware implementation manners, all within the protection scope of the present invention.
Claims
1. A method for constructing a voiceprint recognition model, characterized in that The method includes: Obtaining a plurality of voice samples of at least two users, where the plurality of voice samples include at least three voice samples of each of the at least two users; Extracting a plurality of voiceprint feature vectors corresponding to the plurality of voice samples; Constructing a within-class loss function according to the plurality of voiceprint feature vectors, where the within-class loss function is used to characterize that the similarity between any target voice sample of each of the at least two users and its own first center point is the largest. The first center point is used to characterize the sum of the products of the first similarity weights of each first candidate voice sample of the target voice sample and its own voiceprint feature vector. The first candidate voice sample refers to the voice samples other than the target voice sample among all the voice samples of the user corresponding to the target voice sample. The first similarity weight refers to the proportion of the similarity between the target voice sample and a single first candidate voice sample in the sum of the similarities of all the first candidate voice samples of the target voice sample. The first similarity refers to the similarity between the target voice sample and a single first candidate voice sample; Constructing a between-class loss function according to the plurality of voiceprint feature vectors, where the between-class loss function is used to characterize that the similarity between any target voice sample of each of the at least two users and its own second center point is the smallest. The second center point is used to characterize the sum of the products of the second similarity weights of each second candidate voice sample of the target voice sample and its own voiceprint feature vector. The second candidate voice sample refers to the voice samples of other users except the user corresponding to the target voice sample. The second similarity weight refers to the proportion of the similarity between the target voice sample and a single second candidate voice sample in the sum of the similarities of all the second candidate voice samples of the target voice sample. The second similarity refers to the similarity between the target voice sample and a single second candidate voice sample; Obtaining a target loss function according to the within-class loss function and the between-class loss function, where the target loss function is used to characterize the correct voice samples and incorrect voice samples among at least three voice samples of each of the at least two users. The correct voice sample is used to characterize that the corresponding relationship between the current voice sample and the user is correct, and the incorrect voice sample is used to characterize that the corresponding relationship between the current voice sample and the user is incorrect; Updating the parameters of the voiceprint recognition model according to the target loss function to obtain a trained voiceprint recognition model.
2. The method according to claim 1, characterized in that The constructing a within-class loss function according to the plurality of voiceprint feature vectors includes: Calculating the similarities between different voiceprint feature vectors of the same user among the at least two users respectively to obtain a first similarity matrix of the at least two users; Calculating the first similarity weights between each voice sample of each of the at least two users and each first candidate voice sample among all its own first candidate voice samples according to the first similarity matrix to obtain a first similarity weight matrix of the at least two users; Based on the first similarity weight matrix and the first voiceprint feature vector matrix, obtain the first center point matrix of the at least two users; Based on the first voiceprint feature vector matrix and the first center point matrix, construct the within-class loss function.
3. The method according to claim 2, characterized in that, The constructing the within-class loss function based on the first voiceprint feature vector matrix and the first center point matrix includes: Determine whether there is a historical global center vector matrix of the at least two users; If so, update the historical global center vector matrix according to the first center point matrix to obtain an updated global center vector matrix; Based on the first similarity weight matrix and the updated global center vector matrix, obtain the updated first center point matrix of the at least two users; Based on the first voiceprint feature vector matrix and the updated first center point matrix, obtain the within-class loss function.
4. The method according to claim 3, wherein After determining whether there is a historical global center vector matrix of the at least two users, the method further includes: If not, obtain a pre-set first hyperparameter and a second hyperparameter; Based on the first voiceprint feature vector matrix, the first center point matrix, the first hyperparameter, and the second hyperparameter, construct the within-class loss function.
5. The method according to claim 2, wherein The calculating the similarity between different voiceprint feature vectors of the same user among the at least two users to obtain the first similarity matrix of the at least two users includes performing the following operations for each voice sample of each user among the at least two users: Determine the current voice sample of the current user as the target voice sample; Determine the voiceprint feature vector of the target voice sample as the target vector; Determine the other voice samples of the current user except the current voice sample as the first candidate voice samples; Determine the voiceprint feature vectors of the first candidate voice samples as candidate vectors; Calculate multiple similarities between the target vector and the candidate vectors respectively.
6. The method according to claim 5, characterized in that, The constructing the between-class loss function based on the multiple voiceprint feature vectors includes: Calculate the similarity between the voiceprint feature vectors of different users among the at least two users respectively to obtain the second similarity matrix of the at least two users; Calculate the second similarity weights of each target voice sample of each user among the at least two users and each second candidate voice sample in all the second candidate voice samples of itself according to the second similarity matrix to obtain the second similarity matrix of the at least two users; Based on the second similarity weight matrix and the second voiceprint feature vector matrix, obtain the second center point matrix of the at least two users; Based on the second voiceprint feature vector matrix and the second center point matrix, construct the between-class loss function.
7. An apparatus for constructing a voiceprint recognition model, characterized in that The device includes: An acquisition unit, configured to acquire multiple voice samples of at least two users, where the multiple voice samples include at least three voice samples of each user among the at least two users; An extraction unit, configured to extract multiple voiceprint feature vectors corresponding to the multiple voice samples; A first construction unit for constructing a same-class loss function according to the multiple voiceprint feature vectors, where the same-class loss function is used to characterize that the similarity between any target voice sample of each user among the at least two users and its own first center point is the largest, and the first center point is used to characterize the sum of the products of the first similarity weights of each first candidate voice sample of the target voice sample and its own voiceprint feature vector, the first candidate voice sample refers to the voice samples other than the target voice sample among all the voice samples of the user corresponding to the target voice sample, the first similarity weight refers to the proportion of the similarity between the target voice sample and a single first candidate voice sample in the sum of the similarities of all the first candidate voice samples of the target voice sample, and the first similarity refers to the similarity between the target voice sample and a single first candidate voice sample; A second construction unit for constructing a different-class loss function according to the multiple voiceprint feature vectors, where the different-class loss function is used to characterize that the similarity between any target voice sample of each user among the at least two users and its own second center point is the smallest, and the second center point is used to characterize the sum of the products of the second similarity weights of each second candidate voice sample of the target voice sample and its own voiceprint feature vector, the second candidate voice sample refers to the voice samples of other users except the user corresponding to the target voice sample, the second similarity weight refers to the proportion of the similarity between the target voice sample and a single second candidate voice sample in the sum of the similarities of all the second candidate voice samples of the target voice sample, and the second similarity refers to the similarity between the target voice sample and a single second candidate voice sample; A loss function generation unit for obtaining a target loss function according to the same-class loss function and the different-class loss function, where the target loss function is used to characterize the correct voice samples and incorrect voice samples among at least three voice samples of each user among the at least two users, the correct voice sample is used to characterize that the corresponding relationship between the current voice sample and the user is correct, and the incorrect voice sample is used to characterize that the corresponding relationship between the current voice sample and the user is incorrect; A model training unit for updating the parameters of the voiceprint recognition model according to the target loss function to obtain a trained voiceprint recognition model.
8. An electronic device, characterized in that, It includes an application processor, a communication module, a memory, and one or more programs. The application processor is communicatively connected to the memory and the communication module through an internal communication bus. The one or more programs are stored in the memory and are configured to be executed by the application processor. The one or more programs include instructions for performing the steps in the method according to any one of claims 1-6.
9. A computer-readable storage medium having computer programs / instructions stored thereon, characterized in that, When the computer program / instructions are executed by the application processor, the steps of the method according to any one of claims 1-6 are implemented.
10. A computer program product comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by the application processor, the steps of the method according to any one of claims 1-6 are implemented.
Citation Information
Patent Citations
Speaker recognition method and device and computer equipment and computer readable media
CN106683680A
Voiceprint recognition method, device and apparatus for original voice, and storage medium
CN111524525A