Speaker recognition method based on trivial pronunciation and related device
By dividing the training set into multiple training tasks and using meta-learning techniques to train the speaker embedding layer and classification model, the problem of low accuracy in ordinary pronunciation recognition is solved, achieving higher recognition accuracy and robustness.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- NANJING INST OF INTELLIGENT TECH INST OF MICROELECTRONICS OF THE CHINESE ACAD OF
- Filing Date
- 2023-01-09
- Publication Date
- 2026-04-17
AI Technical Summary
Existing speaker recognition systems lack accurate matching methods when dealing with ordinary pronunciations, resulting in low recognition accuracy.
The training set is divided into multiple training tasks, each containing multiple speakers and audio. Speaker embedding layers and classification models are trained using cross-entropy loss and backpropagation. Meta-learning techniques are used to improve the model's generalization ability, making it applicable to all vocalization methods.
It improves the speaker recognition system's accuracy under ordinary pronunciation conditions and enhances its robustness to the randomness of pronunciation.
Smart Images

Figure CN116052644B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent recognition, and in particular to a speaker recognition method and related equipment based on ordinary pronunciation. Background Technology
[0002] Current speaker recognition systems are mostly based on "normal pronunciation," that is, pronunciation produced by human consciousness with clear audio content. These pronunciations record the process of vocal cord vibration and vocal tract modulation, containing rich speaker information, and are therefore well-suited for speaker recognition. Speaker recognition is a biometric identification technology that identifies a speaker based on the speaker's personality information in the audio signal. With technological advancements, speaker recognition systems have achieved impressive performance.
[0003] However, some pronunciations, limited by physiological characteristics or pronunciation habits, are subject to less speaker control. This makes speaker recognition based on these pronunciations potentially more effective in overcoming the randomness of pronunciation. Examples include coughs and laughter during speech, the "hello" on the phone, the "tsk-tsk" sound made with the tongue to express dissatisfaction, and the "uh-huh" sound to indicate doubt or uncertainty. These pronunciations vary depending on individual habits, and although they generally contain no content information, they are rich in speaker information. We call these pronunciations, which frequently occur in spoken dialogue and are subject to less subjective speaker control, "ordinary pronunciations."
[0004] Using ordinary pronunciations in speaker recognition may enhance the system's robustness to pronunciation randomness. Ordinary pronunciations have several characteristics that distinguish them from normal pronunciations, the most important of which are short pronunciation duration and less audio content. Therefore, there is still a lack of a more accurate method to match ordinary pronunciations with their corresponding speakers. Summary of the Invention
[0005] In view of the above problems, the present invention provides a speaker recognition method and related equipment based on ordinary pronunciation, the main purpose of which is to solve the problem of the lack of a more accurate method for matching ordinary pronunciation with its corresponding speaker.
[0006] To address at least one of the aforementioned technical problems, in a first aspect, the present invention provides a speaker recognition method based on ordinary pronunciation, the method comprising:
[0007] The training set is divided into at least two training tasks. Each training task includes at least two speakers, each speaker includes at least two audio sentences, and each audio sentence is equipped with frame-level phoneme labels, speaker labels and corresponding target spectral features. Each training task includes a support set and six query sets.
[0008] Based on all the target spectral features of the aforementioned support set, the initial speaker embedding layer model, and the initial speaker classification model, the cross-entropy loss of the aforementioned support set is determined through the first operation;
[0009] Based on all the aforementioned support sets, the cross-entropy loss and backpropagation method determine the first speaker classification model and the first speaker embedding layer model through the second operation.
[0010] The target speaker embedding model is determined by the average loss of the six query sets of all training tasks determined by the first speaker classification model and the first speaker embedding model described above, through a third operation.
[0011] Optionally, the initial speaker embedding layer model described above is composed of at least two convolutional layers with BN and ReLU layers stacked with a fully connected layer.
[0012] The initial speaker classification model described above consists of a fully connected layer. The number of input nodes of the fully connected layer is determined based on the number of output nodes of the fully connected layer of the speaker embedding layer model described above. The number of output nodes is the number of speakers in the training set described above.
[0013] The above methods also include:
[0014] All initial spectral features of the support set for the target training task are segmented based on the step size to determine all the aforementioned target spectral features of the support set.
[0015] Optionally, the cross-entropy loss of the support set is determined through a first operation based on all the target spectral features, the initial speaker embedding layer model, and the initial speaker classification model of the support set, including:
[0016] Input all the target spectral features mentioned above into the initial speaker embedding layer model to obtain the speaker embedding layer;
[0017] The speaker embedding layer is input into the initial speaker classification model and the cross-entropy loss of the support set is determined based on the speaker labels.
[0018] Optionally, the cross-entropy loss and backpropagation method based on all the aforementioned support sets determine the first speaker classification model and the first speaker embedding layer model through a second operation, including:
[0019] Based on the aforementioned cross-entropy loss, the gradients of the initial speaker classification model and the initial speaker embedding layer model are calculated sequentially using the backpropagation method.
[0020] The first parameters of the initial speaker classification model and the initial speaker embedding layer model are obtained based on the gradients of the initial speaker classification model and the initial speaker embedding layer model.
[0021] The first speaker classification model and the first speaker embedding layer model are determined based on the first parameters of the above-mentioned initial speaker classification model and the above-mentioned initial speaker embedding layer model.
[0022] Optionally, the average loss of the six query sets for all training tasks determined by the first speaker classification model and the first speaker embedding layer model is used to determine the target speaker embedding layer model through a third operation, including:
[0023] Based on the above first speaker classification model and the above first speaker embedding layer model, calculate the average loss of the six query sets for all the above training tasks;
[0024] The average loss of the six query sets based on all the above training tasks is used to update and obtain the second parameters of the above initial speaker classification model and the above initial speaker embedding layer model through backpropagation, wherein the above second parameters are used to determine the target speaker embedding layer model.
[0025] Optionally, the above methods also include:
[0026] Repeat the first, second, and third operations described above;
[0027] If there exists a second parameter that can make the average loss of the six query sets of all training tasks converge, then the second parameter is determined as the target parameter of the initial speaker classification model and the initial speaker embedding layer model.
[0028] Based on the target parameters of the above-mentioned initial speaker classification model and the above-mentioned initial speaker embedding layer model, the above-mentioned target speaker classification model and the above-mentioned target speaker embedding layer model are determined.
[0029] Optionally, the above methods also include:
[0030] Real-time spectral feature extraction of real-time audio;
[0031] Based on the above real-time spectral features, ordinary pronunciations in the above real-time audio are detected.
[0032] In the case where the target real-time audio contains trivial pronunciations, the target real-time spectral features of the target real-time audio are input into the target speaker embedding layer model to obtain the speaker embedding layer of the target real-time audio.
[0033] Based on the spectral features of the registrant's ordinary pronunciation, the preset ordinary pronunciation embedding layer of the registrant is determined through the aforementioned target speaker embedding layer model;
[0034] Calculate the cosine similarity between the speaker embedding layer of the aforementioned target real-time audio and the preset trivial pronunciation embedding layer of the aforementioned registrant;
[0035] Based on the aforementioned cosine similarity, the matching status between the speaker of the aforementioned target real-time audio and the aforementioned registrant is determined.
[0036] Secondly, embodiments of the present invention also provide a speaker recognition device based on ordinary pronunciation, comprising:
[0037] A partitioning unit is used to divide the training set into at least two training tasks. Each training task includes at least two speakers, each speaker includes at least two audio sentences, each audio sentence is equipped with frame-level phoneme labels and speaker labels and corresponding target spectral features, and each training task includes a support set and six query sets.
[0038] The first determining unit is used to determine the cross-entropy loss of the support set based on all the target spectral features of the support set, the initial speaker embedding layer model, and the initial speaker classification model through a first operation.
[0039] The acquisition unit is used to determine the first speaker classification model and the first speaker embedding layer model through a second operation based on the cross-entropy loss and backpropagation method of all the above-mentioned support sets.
[0040] The second determining unit is used to determine the target speaker embedding model through a third operation based on the average loss of the six query sets of all training tasks determined by the first speaker classification model and the first speaker embedding layer model.
[0041] To achieve the above objectives, according to a third aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium comprising a stored program, wherein, when the program is executed by a processor, the steps of the above-described speaker recognition method based on ordinary pronunciation are implemented.
[0042] To achieve the above objectives, according to a fourth aspect of the present invention, an electronic device is provided, comprising at least one processor and at least one memory connected to the processor; wherein the processor is configured to invoke program instructions in the memory to execute the steps of the above-described speaker recognition method based on ordinary pronunciation.
[0043] By employing the above technical solution, the speaker recognition method and related equipment based on ordinary pronunciation provided by this invention address the current lack of a more accurate method for matching ordinary pronunciations with their corresponding speakers. This invention divides the training set into at least two training tasks, where each training task includes at least two speakers, each speaker includes at least two audio sentences, and each audio sentence has frame-level phoneme labels, speaker labels, and corresponding target spectral features. Each training task includes a support set and six query sets. Based on all the target spectral features of the support set, the initial speaker embedding layer model, and the initial speaker classification model, a first operation determines the cross-entropy loss of the support set. Based on the cross-entropy loss of all the support sets and the backpropagation method, a second operation determines the first speaker classification model and the first speaker embedding layer model. Based on the average loss of the six query sets of all training tasks determined by the first speaker classification model and the first speaker embedding layer model, a third operation determines the target speaker embedding layer model. In the above scheme, compared with joint training that directly uses ordinary audio and adds audio corresponding to phonemes, this scheme uses ordinary audio in the training set as the support set and phoneme information as different query sets for meta-learning training. This allows the speaker embedding layer network to find parameter update directions applicable to all phonemes during the training process, so that the trained model can generalize to all speech patterns of the speaker, including ordinary speech patterns, thereby improving the accuracy of the speaker recognition system in recognizing speakers using ordinary speech patterns.
[0044] Accordingly, the speaker recognition device, apparatus, and computer-readable storage medium based on ordinary pronunciation provided in the embodiments of the present invention also have the above-mentioned technical effects.
[0045] The above description is merely an overview of the technical solution of the present invention. In order to better understand the technical means of the present invention and to implement it in accordance with the contents of the specification, and in order to make the above and other objects, features and advantages of the present invention more apparent and understandable, specific embodiments of the present invention are described below. Attached Figure Description
[0046] Various other advantages and benefits will become apparent to those skilled in the art upon reading the following detailed description of preferred embodiments. The accompanying drawings are for illustrative purposes only and are not intended to limit the invention. Furthermore, the same reference numerals denote the same parts throughout the drawings. In the drawings:
[0047] Figure 1 The diagram illustrates a flowchart of a speaker recognition method based on common pronunciation provided by an embodiment of the present invention.
[0048] Figure 2This invention illustrates a process for training a model based on the MAML method to update the parameters of the speaker embedding layer model and the classification model in a certain training task, as provided in an embodiment of the present invention.
[0049] Figure 3 The diagram illustrates a class diagram of the training task partitioning for a speaker recognition method based on ordinary pronunciation provided in an embodiment of the present invention.
[0050] Figure 4 This invention provides a flowchart illustrating how to update parameters using MAML methods based on the PyTorch framework.
[0051] Figure 5 This diagram illustrates the composition of a speaker recognition device based on ordinary pronunciation provided in an embodiment of the present invention.
[0052] Figure 6 This diagram illustrates the composition of an electronic device for speaker recognition based on ordinary pronunciation, as provided in an embodiment of the present invention. Detailed Implementation
[0053] Exemplary embodiments of the invention will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the invention are shown in the drawings, it should be understood that the invention may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that this invention will be thorough and complete, and will fully convey the scope of the invention to those skilled in the art.
[0054] To address the current lack of a more accurate method for matching ordinary pronunciations with their corresponding speakers, this invention provides a speaker recognition method based on ordinary pronunciations, such as... Figure 1 As shown, the method includes:
[0055] S101. Divide the training set into at least two training tasks, wherein each training task includes at least two speakers, each speaker includes at least two audio sentences, each audio sentence is equipped with frame-level phoneme labels and speaker labels and corresponding target spectral features, and each training task includes a support set and six query sets.
[0056] For example, a speaker recognition dataset D, i.e., the training set mentioned above, is determined where each audio sentence has both phoneme and speaker labels. The spectral features of each audio sentence in D are extracted, and the phoneme labels of each audio sentence are converted into one-hot encoded frame-level phoneme labels. The number of audio sentences in D is U. 句The size of the spectral features is [spectral dimension, number of time frames]; the size of the frame-level phoneme labels is [6, number of time frames], because the number of phoneme categories is 6, which divides all phonemes into 6 categories: plosives, affricates, fricatives, nasals, glissandos, and vowels. The dataset D can be the TIMIT dataset.
[0057] For example, the training set D is divided into several training tasks. Each training task has T speakers. 人 Each speaker has a randomly selected T. 句 Sentence spectral features and their corresponding frame-level phoneme labels and speaker labels are used. Each task has one support set and six query sets. Optionally, in one support set for each task in the training set, each speaker has one sentence; the six query sets in the training set each speaker (T)... 句 -1) Sentences. Each sentence from each person in the query set can be repeated; sentences in both the support set and the query set cannot be repeated. Where T 人 Optional: 100, T 句 The number of training tasks N can be set to 3. 任务 for:
[0058]
[0059] For example, where Floor represents floor rounding down. The support set has T. 人 Each speaker has a spectral feature of 33 frames over a period of time, which is randomly extracted from the spectral features of a single audio sentence. The remaining (T) for each speaker... 句 —1) The spectral features of the sentence audio are used for 6 query sets, each containing only one distinct phoneme category. In each query set, T 人 Each speaker has two audio spectrum segments, each with a length of 9 frames, obtained using the frame-level phoneme tags of the audio. When the length of selectable phonemes in a sentence of audio in the query set is less than 9 frames, the length is expanded to 9 frames from the current phoneme frame to the left and right sides of the original audio; when the phoneme length exceeds 9 frames, only the middle 9 frames are taken.
[0060] S102. Based on all the target spectral features of the above support set, the initial speaker embedding layer model and the initial speaker classification model, the cross-entropy loss of the above support set is determined through the first operation;
[0061] For example, step S102 above further includes S1021, as... Figure 2 As shown in (1), all the target spectral features are input into the initial speaker embedding layer model to obtain the speaker embedding layer; the speaker embedding layer is input into the initial speaker classification model and the cross-entropy loss of the support set is determined based on the speaker labels.
[0062] For example, taking the VGG model, the target spectral features are input into the speaker embedding layer model. In the average pooling layer, the intermediate representations of the S segments are aggregated into segment-level representations. Then, the segment-level representations are passed through a fully connected layer to obtain the embedding layer representations. The batch size of the embedding layer is [T]. 人 [Embedding layer dimension]. The embedding layer is input into the classification model, and the cross-entropy loss L of the support set is obtained using speaker labels. 支持 .
[0063] S103. Based on the cross-entropy loss and backpropagation method of all the above support sets, the first speaker classification model and the first speaker embedding layer model are determined through the second operation.
[0064] For example, such as Figure 2 As shown in (2), the steps of S103 above further include S1031: calculating the gradients of the initial speaker classification model and the initial speaker embedding layer model in sequence by backpropagation based on the cross-entropy loss; obtaining the first parameters of the initial speaker classification model and the initial speaker embedding layer model based on the gradients of the initial speaker classification model and the initial speaker embedding layer model; and determining the first speaker classification model and the first speaker embedding layer model based on the first parameters of the initial speaker classification model and the initial speaker embedding layer model.
[0065] For example, the parameter θ of the speaker embedding layer model is θ ′ Copy the parameters θ of the classification layer model fc For θ ′ fc Using L 支持 Backpropagation calculates the gradients of each layer and updates θ. ′ and θ ′ fc The value of θ can be understood as... ′ and θ ′ fc The value is the parameter of the first speaker embedding layer model and the first speaker classification model.
[0066] S104. The target speaker embedding model is determined by the average loss of the six query sets of all training tasks determined by the first speaker classification model and the first speaker embedding layer model.
[0067] For example, such as Figure 2 As shown in (3), step S104 above further includes S1041, calculating the average loss of the six query sets for all the above training tasks based on the above first speaker classification model and the above first speaker embedding layer model; as Figure 2As shown in (4), the average loss of the six query sets based on all the above training tasks is updated by backpropagation and the second parameters of the above initial speaker classification model and the above initial speaker embedding layer model are obtained, wherein the above second parameters are used to determine the target speaker embedding layer model.
[0068] For example, using the updated θ ′ and θ ′ fc Calculate the average loss of 6 query sets in the same task in
[0069]
[0070] This represents the prototype loss for each query set, where the prototype of each person in the query set is the updated θ in the support set. ′ The calculated speaker embedding layer, This represents the cross-entropy loss for each query set.
[0071] For example, step S104 above further includes S1042, repeating the first operation, the second operation and the third operation above; if there is a second parameter that can make the average loss of the six query sets of all training tasks converge, determining the second parameter as the target parameter of the initial speaker classification model and the initial speaker embedding layer model; determining the target speaker classification model and the target speaker embedding layer model based on the target parameter of the initial speaker classification model and the initial speaker embedding layer model.
[0072] For example, for N 任务 Repeat the above steps for each training task until L... 查询 Convergence completes the training of the speaker embedding layer model.
[0073] By employing the above technical solution, the speaker recognition method based on ordinary pronunciation provided by this invention addresses the current lack of a more accurate method for matching ordinary pronunciations with their corresponding speakers. This invention divides the training set into at least two training tasks, where each training task includes at least two speakers, each speaker includes at least two audio sentences, and each audio sentence has frame-level phoneme labels, speaker labels, and corresponding target spectral features. Each training task includes a support set and six query sets. Based on all the target spectral features of the support set, the initial speaker embedding layer model, and the initial speaker classification model, a first operation determines the cross-entropy loss of the support set. Based on the cross-entropy loss of all the support sets and the backpropagation method, a second operation determines the first speaker classification model and the first speaker embedding layer model. Based on the average loss of the six query sets of all training tasks determined by the first speaker classification model and the first speaker embedding layer model, a third operation determines the target speaker embedding layer model. In the above scheme, compared with joint training that directly uses ordinary audio and adds audio corresponding to phonemes, this scheme uses ordinary audio in the training set as the support set and phoneme information as different query sets for meta-learning training. This allows the speaker embedding layer network to find parameter update directions applicable to all phonemes during the training process, so that the trained model can generalize to all speech patterns of the speaker, including ordinary speech patterns, thereby improving the accuracy of the speaker recognition system in recognizing speakers using ordinary speech patterns.
[0074] In one embodiment, the initial speaker embedding layer model described above is composed of at least two convolutional layers with BN and ReLU layers stacked together with a fully connected layer.
[0075] The initial speaker classification model described above consists of a fully connected layer. The number of input nodes of the fully connected layer is determined based on the number of output nodes of the fully connected layer of the speaker embedding layer model described above. The number of output nodes is the number of speakers in the training set described above.
[0076] The above methods also include:
[0077] All initial spectral features of the support set for the target training task are segmented based on the step size to determine all the aforementioned target spectral features of the support set.
[0078] For example, an initial speaker embedding layer model is constructed and trained using a meta-learning method to obtain a speaker embedding layer model that generalizes well to all phonemes. The parameter of the embedding layer model is θ. The minimum number of input frames in the time dimension of the speaker embedding layer model should be 9 frames or less, as the duration of phonemes is usually short. The speaker embedding layer model can be a VGG model, and the parameters of each layer are shown below from top to bottom, where the BN layer is a batch normalization layer, and the ReLU layer is a rectified linear unit layer. Table 1 shows one method for constructing an initial speaker embedding layer model:
[0079] Table 1
[0080]
[0081]
[0082] For example, an initial speaker classification model is constructed, which is a fully connected layer with parameters denoted as θ. fc It is only applied during the training process. Its input node count is the output node count of the last fully connected layer of the speaker embedding layer, and the output node count is the number of speakers in the dataset D.
[0083] For example, obtain the cross-entropy loss L of the support set in a task. 支持 Supports batching of spectral features from all speakers in the set, with a size of [T]. 人 [Spectral dimension, 33 frames]. To reduce the variability of longer audio segments, the spectral features are first segmented according to a certain step size, as shown in the following formula:
[0084]
[0085] Where S is the number of segments, and Frames is the number of time frames for spectral features, i.e., 33; Frames win The minimum number of input frames for the speaker embedding layer model can be set to 9 when the speaker embedding layer model is the VGG model from step 1; stride is the segmentation step size, which can be set to 4. The spectral feature size then becomes [T] after segmentation. 人 Spectrum dimension, Frames win S]
[0086] In one embodiment, the above method further includes:
[0087] Real-time spectral feature extraction of real-time audio;
[0088] Based on the above real-time spectral features, ordinary pronunciations in the above real-time audio are detected.
[0089] In the case where the target real-time audio contains trivial pronunciations, the target real-time spectral features of the target real-time audio are input into the target speaker embedding layer model to obtain the speaker embedding layer of the target real-time audio.
[0090] Based on the spectral features of the registrant's ordinary pronunciation, the preset ordinary pronunciation embedding layer of the registrant is determined through the aforementioned target speaker embedding layer model;
[0091] Calculate the cosine similarity between the speaker embedding layer of the aforementioned target real-time audio and the preset trivial pronunciation embedding layer of the aforementioned registrant;
[0092] Based on the aforementioned cosine similarity, the matching status between the speaker of the aforementioned target real-time audio and the aforementioned registrant is determined.
[0093] For example, the ordinary pronunciation embedding layer of the registered speaker is obtained. A common pronunciation is selected, such as the "um" sound, and the spectral features of the common pronunciation of the registered speaker are passed through the trained speaker embedding layer model to obtain the ordinary pronunciation embedding layer of the registered speaker.
[0094] For example, real-time audio data is acquired and spectral features are extracted. The spectral features are input to a trivial pronunciation detector to detect the presence of trivial pronunciations in the audio. The trivial pronunciation detector is a neural network model trained using positive samples of a certain type of trivial pronunciation and other audio samples as negative samples. The neural network model consists of several convolutional layers and one binary classification fully connected layer stacked together. The binary classification fully connected layer outputs two nodes: one representing a trivial pronunciation and the other representing a negative sample. When the audio contains a trivial pronunciation, the value of the output node belonging to the trivial pronunciation is higher than the value of the negative sample.
[0095] Understandably, when there are no trivial pronunciations in the audio, the system will continue to acquire real-time audio; when there are trivial pronunciations in the audio, the spectral features will be input into the trained speaker embedding layer network to obtain the speaker embedding layer of the real-time audio.
[0096] For example, the speaker embedding layer of the real-time audio and the trivial pronunciation embedding layer of the registered speaker are compared using cosine similarity calculation to determine if they are the same speaker. If the cosine similarity exceeds a set threshold, the speaker in the real-time audio is considered to be the same speaker as the registered speaker; otherwise, they are not.
[0097] Furthermore, the specific steps for implementing this solution can be:
[0098] 1) The model consists of at least two convolutional layers with BN and ReLU layers stacked together with a fully connected layer;
[0099] 2) Divide the training set into at least two training tasks, where each training task includes at least two speakers, each speaker includes at least two audio sentences, each audio sentence has frame-level phoneme labels and speaker labels and corresponding initial spectral features, and each training task includes a support set and six query sets.
[0100] 3) Construct an initial speaker classification model, wherein the initial speaker classification model consists of a fully connected layer, the number of input nodes of the fully connected layer is determined based on the number of output nodes of the fully connected layer of the speaker embedding layer model, and the number of output nodes is the number of speakers in the training set;
[0101] 4) Segment all initial spectral features of the support set for the target training task based on the step size to determine all target spectral features of the support set;
[0102] 5) Input all the target spectral features into the initial speaker embedding layer model to obtain the speaker embedding layer;
[0103] 6) Input the speaker embedding layer into the initial speaker classification model and determine the cross-entropy loss of the support set based on the speaker labels;
[0104] 7) Based on the cross-entropy loss, the gradients of the initial speaker classification model and the initial speaker embedding layer model are calculated sequentially using the backpropagation method;
[0105] 8) Update the gradients of the initial speaker classification model and the initial speaker embedding layer model, and obtain the first parameters of the initial speaker classification model and the initial speaker embedding layer model;
[0106] 9) Determine the first speaker classification model and the first speaker embedding layer model based on the first parameters of the initial speaker classification model and the initial speaker embedding layer model.
[0107] 10) Calculate the average loss of the six query sets for all the training tasks based on the first speaker classification model and the first speaker embedding layer model;
[0108] 11) The second parameters of the initial speaker classification model and the initial speaker embedding layer model are updated and obtained by backpropagation based on the average loss of the six query sets of all the training tasks, wherein the second parameters are used to determine the target speaker embedding layer model.
[0109] 12) Repeat steps 5)-11);
[0110] 13) If there exists a second parameter that can make the average loss of the six query sets of all training tasks converge, determine the second parameter as the target parameter of the initial speaker classification model and the initial speaker embedding layer model;
[0111] 14) Determine the target speaker classification model and the target speaker embedding layer model based on the target parameters of the initial speaker classification model and the initial speaker embedding layer model.
[0112] 15) Perform real-time spectral feature extraction on real-time audio;
[0113] 16) Detect ordinary pronunciations in the real-time audio based on the real-time spectral features;
[0114] 17) If the target real-time audio contains trivial pronunciations, input the target real-time spectral features of the target real-time audio into the target speaker embedding layer model to obtain the speaker embedding layer of the target real-time audio;
[0115] 18) Based on the spectral features of the registrant's ordinary pronunciation, the preset ordinary pronunciation embedding layer of the registrant is determined through the target speaker embedding layer model;
[0116] 19) Calculate the cosine similarity between the speaker embedding layer of the target real-time audio and the preset trivial pronunciation embedding layer of the registrant;
[0117] 20) Determine the matching status between the speaker of the target real-time audio and the registrant based on the cosine similarity.
[0118] For example, this solution mainly applies the concept of meta-learning and adapts MAML (Model-Agnostic Meta-Learning) to speaker recognition of ordinary pronunciations, utilizing phoneme information to improve the system's recognition performance for ordinary pronunciations. Therefore, there are two main considerations in implementation: 1) how to divide the dataset into training tasks and acquire training tasks during the training process; 2) how to implement the MAML method to train the model and update the model parameters.
[0119] 1. Use the Dataset and Dataloader classes from torch.utils.data to implement training task partitioning:
[0120] like Figure 3 As shown, PhoneDataset is first defined. It is a subclass of the Dataset class in torch.utils.data and needs to implement the three functions __init__(), __len__(), and __getitem__().
[0121] In the `__init__()` function, a dictionary member variable needs to be implemented. The key of this dictionary is the speaker label in the form of an index. When there are N... 人 When a person is present, the value of the key can be in the range [0, N]. 人 [―1], the dictionary's value is in list form, and each element in the list is a tuple. The tuple consists of the speaker's speech spectrum features and the corresponding frame-level phoneme label. The list contains all the speaker's speech, as shown in Table 2:
[0122] Table 2. List of member variables in PhoneDataset that record speaker tags and their corresponding spectral characteristics
[0123] Key (Speaker Tag) Value 0 [(statement, frame-level phoneme label),...,(statement, frame-level phoneme label)] 1 [(statement, frame-level phoneme label),...,(statement, frame-level phoneme label)] … … <![CDATA[N 人 ―1]]> [(statement, frame-level phoneme label),...,(statement, frame-level phoneme label)]
[0124] The `__len__()` function needs to determine how many people are needed to iterate through all the tasks in a dataset, based on...
[0125]
[0126] Where N 任务 U represents the number of tasks. 句 T represents the number of statements in the dataset. 人 The number of speakers for each task, T 句 T is randomly selected for each task speaker 句 Sentence spectrum characteristics, because the dataloader object treats each batch as a task, and each batch contains T 人 Then the value returned by the __len__() function should be
[0127] The `__getitem__()` function returns the corresponding speaker's tag and the corresponding dictionary value based on the index generated by the dataloader. The dataloader generates the index randomly each time. One index, T per batch 人 Then return N 任务 There are batches, and the index range is [0, -1]. Due to With N 人 The number may vary, therefore the speaker tag returned by the __getitem__() function should be index%N. 人 .
[0128] Because the return values need to be batched, the labels, spectral features, etc., must have the same size, and the spectral feature sizes of each speaker's support set and query set in each task should be consistent. Therefore, in the `collate_fn()` function, the T values obtained in each task are... 人 The tags and spectral features are organized. The return values and their dimensions are shown in Table 3. Additionally, it supports non-overlapping statements between the set and the query set, allows overlap between the statements in the six query sets, and randomly selects the speaker's statement for each person.
[0129] Table 3 illustrates the data returned by the collate_fn() function.
[0130]
[0131] During training, if a for data in dataloader statement is used on the torch.utils.data.Dataloader class object dataloader, then the data obtained in each iteration corresponds to the various return values in the data.
[0132] 2. Train the model using MAML and update the model parameters:
[0133] Since obtaining the loss function only requires processing the return data from the example in Table 3, it will be omitted. The focus is on the implementation of parameter updates. The Python and PyTorch functions used are as follows: Figure 4 As shown. First, the copied speaker embedding layer model parameters θ are updated using data from the support set. ′ and classification model parameters θ′ fc The temp_weights are composed of these weights, and then the query set loss is calculated using the temp_weights and functions in torch.nn.functional. Last use .backward() directly updates the speaker embedding layer model parameters θ and the classification model parameters θ. fc .
[0134] It is understandable that the above two design ideas are based on the PyTorch framework for project implementation.
[0135] Furthermore, as a response to the above Figure 1 In addition to the implementation of the method shown, this embodiment of the invention also provides a speaker recognition device based on ordinary pronunciation, used for the above-mentioned... Figure 1The method shown is implemented accordingly. This device embodiment corresponds to the foregoing method embodiment. For ease of reading, this device embodiment will not repeat the details of the foregoing method embodiment, but it should be clear that the device in this embodiment can implement all the contents of the foregoing method embodiment. Figure 5 As shown, the device includes: a first acquisition unit 21, a determination unit 22, a second acquisition unit 23, and a generation unit 24, wherein,
[0136] The partitioning unit 21 is used to divide the training set into at least two training tasks. Each training task includes at least two speakers, each speaker includes at least two audio sentences, each audio sentence is equipped with frame-level phoneme labels and speaker labels and corresponding target spectral features, and each training task includes a support set and six query sets.
[0137] The first determining unit 22 is used to determine the cross-entropy loss of the support set based on all the target spectral features of the support set, the initial speaker embedding layer model, and the initial speaker classification model through a first operation.
[0138] The acquisition unit 23 is used to determine the first speaker classification model and the first speaker embedding layer model through a second operation based on the cross-entropy loss and backpropagation method of all the above-mentioned support sets.
[0139] The second determining unit 24 is used to determine the target speaker embedding model through a third operation based on the average loss of the six query sets of all training tasks determined by the first speaker classification model and the first speaker embedding layer model.
[0140] For example, the initial speaker embedding layer model described above is composed of at least two convolutional layers with BN and ReLU layers stacked with a fully connected layer.
[0141] The initial speaker classification model described above consists of a fully connected layer. The number of input nodes of the fully connected layer is determined based on the number of output nodes of the fully connected layer of the speaker embedding layer model described above. The number of output nodes is the number of speakers in the training set described above.
[0142] The above methods also include:
[0143] All initial spectral features of the support set for the target training task are segmented based on the step size to determine all the aforementioned target spectral features of the support set.
[0144] For example, the cross-entropy loss of the support set is determined through a first operation based on all the target spectral features of the support set, the initial speaker embedding layer model, and the initial speaker classification model, including:
[0145] Input all the target spectral features mentioned above into the initial speaker embedding layer model to obtain the speaker embedding layer;
[0146] The speaker embedding layer is input into the initial speaker classification model and the cross-entropy loss of the support set is determined based on the speaker labels.
[0147] For example, the cross-entropy loss and backpropagation method based on all the aforementioned support sets determines the first speaker classification model and the first speaker embedding layer model through a second operation, including:
[0148] Based on the aforementioned cross-entropy loss, the gradients of the initial speaker classification model and the initial speaker embedding layer model are calculated sequentially using the backpropagation method.
[0149] The first parameters of the initial speaker classification model and the initial speaker embedding layer model are obtained based on the gradients of the initial speaker classification model and the initial speaker embedding layer model.
[0150] The first speaker classification model and the first speaker embedding layer model are determined based on the first parameters of the above-mentioned initial speaker classification model and the above-mentioned initial speaker embedding layer model.
[0151] For example, the average loss of the six query sets of all training tasks determined by the first speaker classification model and the first speaker embedding layer model determines the target speaker embedding layer model through a third operation, including:
[0152] Based on the above first speaker classification model and the above first speaker embedding layer model, calculate the average loss of the six query sets for all the above training tasks;
[0153] The average loss of the six query sets based on all the above training tasks is used to update and obtain the second parameters of the above initial speaker classification model and the above initial speaker embedding layer model through backpropagation, wherein the above second parameters are used to determine the target speaker embedding layer model.
[0154] For example, the above-mentioned unit further includes:
[0155] Repeat the first, second, and third operations described above;
[0156] If there exists a second parameter that can make the average loss of the six query sets of all training tasks converge, then the second parameter is determined as the target parameter of the initial speaker classification model and the initial speaker embedding layer model.
[0157] Based on the target parameters of the above-mentioned initial speaker classification model and the above-mentioned initial speaker embedding layer model, the above-mentioned target speaker classification model and the above-mentioned target speaker embedding layer model are determined.
[0158] For example, the above-mentioned unit further includes:
[0159] Real-time spectral feature extraction of real-time audio;
[0160] Based on the above real-time spectral features, ordinary pronunciations in the above real-time audio are detected.
[0161] In the case where the target real-time audio contains trivial pronunciations, the target real-time spectral features of the target real-time audio are input into the target speaker embedding layer model to obtain the speaker embedding layer of the target real-time audio.
[0162] Based on the spectral features of the registrant's ordinary pronunciation, the preset ordinary pronunciation embedding layer of the registrant is determined through the aforementioned target speaker embedding layer model;
[0163] Calculate the cosine similarity between the speaker embedding layer of the aforementioned target real-time audio and the preset trivial pronunciation embedding layer of the aforementioned registrant;
[0164] Based on the aforementioned cosine similarity, the matching status between the speaker of the aforementioned target real-time audio and the aforementioned registrant is determined.
[0165] By employing the above technical solution, the speaker recognition device based on ordinary pronunciation provided by this invention addresses the current lack of a more accurate method for matching ordinary pronunciations with their corresponding speakers. This invention divides the training set into at least two training tasks, where each training task includes at least two speakers, each speaker includes at least two audio sentences, and each audio sentence has frame-level phoneme labels, speaker labels, and corresponding target spectral features. Each training task includes a support set and six query sets. Based on all the target spectral features of the support set, the initial speaker embedding layer model, and the initial speaker classification model, a first operation determines the cross-entropy loss of the support set. Based on the cross-entropy loss of all the support sets and the backpropagation method, a second operation determines the first speaker classification model and the first speaker embedding layer model. Based on the average loss of the six query sets of all training tasks determined by the first speaker classification model and the first speaker embedding layer model, a third operation determines the target speaker embedding layer model. In the above scheme, compared with joint training that directly uses ordinary audio and adds audio corresponding to phonemes, this scheme uses ordinary audio in the training set as the support set and phoneme information as different query sets for meta-learning training. This allows the speaker embedding layer network to find parameter update directions applicable to all phonemes during the training process, so that the trained model can generalize to all speech patterns of the speaker, including ordinary speech patterns, thereby improving the accuracy of the speaker recognition system in recognizing speakers using ordinary speech patterns.
[0166] The processor contains a kernel, which retrieves the corresponding program units from memory. One or more kernels can be configured, and by adjusting kernel parameters, a speaker recognition method based on ordinary pronunciation can be implemented. This method can address the current lack of a more accurate way to match ordinary pronunciations with their corresponding speakers.
[0167] This invention provides a computer-readable storage medium including a stored program that, when executed by a processor, implements the aforementioned speaker recognition method based on ordinary pronunciation.
[0168] This invention provides a processor for running a program, wherein the program executes the speaker recognition method based on ordinary pronunciation.
[0169] This invention provides an electronic device, which includes at least one processor and at least one memory connected to the processor; wherein the processor is configured to call program instructions in the memory to execute the speaker recognition method based on ordinary pronunciation as described above.
[0170] This invention provides an electronic device 30, such as... Figure 6 As shown, the electronic device includes at least one processor 301, and at least one memory 302 and bus 303 connected to the processor; wherein, the processor 301 and the memory 302 communicate with each other through the bus 303; the processor 301 is used to call program instructions in the memory to execute the above-mentioned speaker recognition method based on ordinary pronunciation.
[0171] The smart electronic devices mentioned in this article can be PCs, tablets, mobile phones, etc.
[0172] This application also provides a computer program product, which, when executed on a process management electronic device, is suitable for executing a program that initializes the following method steps:
[0173] The training set is divided into at least two training tasks. Each training task includes at least two speakers, each speaker includes at least two audio sentences, and each audio sentence is equipped with frame-level phoneme labels, speaker labels and corresponding target spectral features. Each training task includes a support set and six query sets.
[0174] Based on all the target spectral features of the aforementioned support set, the initial speaker embedding layer model, and the initial speaker classification model, the cross-entropy loss of the aforementioned support set is determined through the first operation;
[0175] Based on all the aforementioned support sets, the cross-entropy loss and backpropagation method determine the first speaker classification model and the first speaker embedding layer model through the second operation.
[0176] The target speaker embedding model is determined by the average loss of the six query sets of all training tasks determined by the first speaker classification model and the first speaker embedding model described above, through a third operation.
[0177] Furthermore, the aforementioned initial speaker embedding layer model is composed of at least two convolutional layers with BN and ReLU layers stacked with a fully connected layer.
[0178] The initial speaker classification model described above consists of a fully connected layer. The number of input nodes of the fully connected layer is determined based on the number of output nodes of the fully connected layer of the speaker embedding layer model described above. The number of output nodes is the number of speakers in the training set described above.
[0179] The above methods also include:
[0180] All initial spectral features of the support set for the target training task are segmented based on the step size to determine all the aforementioned target spectral features of the support set.
[0181] Furthermore, the cross-entropy loss of the support set is determined through a first operation using all the target spectral features, the initial speaker embedding layer model, and the initial speaker classification model based on the support set, including:
[0182] Input all the target spectral features mentioned above into the initial speaker embedding layer model to obtain the speaker embedding layer;
[0183] The speaker embedding layer is input into the initial speaker classification model and the cross-entropy loss of the support set is determined based on the speaker labels.
[0184] Furthermore, the aforementioned cross-entropy loss and backpropagation method based on all the aforementioned support sets determines the first speaker classification model and the first speaker embedding layer model through a second operation, including:
[0185] Based on the aforementioned cross-entropy loss, the gradients of the initial speaker classification model and the initial speaker embedding layer model are calculated sequentially using the backpropagation method.
[0186] The first parameters of the initial speaker classification model and the initial speaker embedding layer model are obtained based on the gradients of the initial speaker classification model and the initial speaker embedding layer model.
[0187] The first speaker classification model and the first speaker embedding layer model are determined based on the first parameters of the above-mentioned initial speaker classification model and the above-mentioned initial speaker embedding layer model.
[0188] Furthermore, the average loss of the six query sets for all training tasks determined by the aforementioned first speaker classification model and the aforementioned first speaker embedding layer model is used to determine the target speaker embedding layer model through a third operation, including:
[0189] Based on the above first speaker classification model and the above first speaker embedding layer model, calculate the average loss of the six query sets for all the above training tasks;
[0190] The average loss of the six query sets based on all the above training tasks is used to update and obtain the second parameters of the above initial speaker classification model and the above initial speaker embedding layer model through backpropagation, wherein the above second parameters are used to determine the target speaker embedding layer model.
[0191] Furthermore, the above methods also include:
[0192] Repeat the first, second, and third operations described above;
[0193] If there exists a second parameter that can make the average loss of the six query sets of all training tasks converge, then the second parameter is determined as the target parameter of the initial speaker classification model and the initial speaker embedding layer model.
[0194] Based on the target parameters of the above-mentioned initial speaker classification model and the above-mentioned initial speaker embedding layer model, the above-mentioned target speaker classification model and the above-mentioned target speaker embedding layer model are determined.
[0195] Furthermore, the above methods also include:
[0196] Real-time spectral feature extraction of real-time audio;
[0197] Based on the above real-time spectral features, ordinary pronunciations in the above real-time audio are detected.
[0198] In the case where the target real-time audio contains trivial pronunciations, the target real-time spectral features of the target real-time audio are input into the target speaker embedding layer model to obtain the speaker embedding layer of the target real-time audio.
[0199] Based on the spectral features of the registrant's ordinary pronunciation, the preset ordinary pronunciation embedding layer of the registrant is determined through the aforementioned target speaker embedding layer model;
[0200] Calculate the cosine similarity between the speaker embedding layer of the aforementioned target real-time audio and the preset trivial pronunciation embedding layer of the aforementioned registrant;
[0201] Based on the aforementioned cosine similarity, the matching status between the speaker of the aforementioned target real-time audio and the aforementioned registrant is determined.
[0202] It should be noted that the descriptions of each embodiment in the above embodiments have different focuses. For parts that are not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.
[0203] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0204] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create a machine for implementing the flowchart illustrations. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0205] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0206] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0207] This application also provides a computer program product, which includes computer software instructions that, when executed on a processing device, cause the processing device to perform actions such as... Figure 1 The control flow of the memory in the corresponding embodiment.
[0208] A computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the flow or function according to the embodiments of this application is generated. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium may be any available medium that a computer can store or a data storage device such as a server or data center that integrates one or more available media. The available medium may be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state disk (SSD)).
[0209] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0210] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces, or indirect coupling or communication connection between apparatuses or units, and may be electrical, mechanical, or other forms.
[0211] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0212] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0213] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0214] The above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit it. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.
Claims
1. A speaker recognition method based on ordinary pronunciation, characterized in that, include: The training set is divided into at least two training tasks. Each training task includes at least two speakers, each speaker includes at least two audio sentences, and each audio sentence is equipped with frame-level phoneme labels, speaker labels and corresponding target spectral features. Each training task includes a support set and six query sets. Based on all the target spectral features of the support set, the initial speaker embedding layer model, and the initial speaker classification model, the cross-entropy loss of the support set is determined through a first operation; The cross-entropy loss and backpropagation method based on all the said support sets determine the first speaker classification model and the first speaker embedding layer model through the second operation; The target speaker embedding model is determined by the average loss of six query sets for all training tasks determined by the first speaker classification model and the first speaker embedding layer model through a third operation. Real-time spectral feature extraction of real-time audio; Based on the real-time spectral features, ordinary pronunciations in the real-time audio are detected, wherein ordinary pronunciations are pronunciations that occur in spoken dialogue and are less subject to the speaker's subjective control; In the case where the target real-time audio contains trivial pronunciations, the target real-time spectral features of the target real-time audio are input into the target speaker embedding layer model to obtain the speaker embedding layer of the target real-time audio; Based on the spectral characteristics of the registrant's ordinary pronunciation, the preset ordinary pronunciation embedding layer of the registrant is determined through the target speaker embedding layer model; Calculate the cosine similarity between the speaker embedding layer of the target real-time audio and the preset trivial pronunciation embedding layer of the registrant; The matching status between the speaker and the registrant in the target real-time audio is determined based on the cosine similarity.
2. The method according to claim 1, characterized in that, The initial speaker embedding layer model is composed of at least two convolutional layers with BN and ReLU layers stacked together with a fully connected layer. The initial speaker classification model consists of a fully connected layer. The number of input nodes of the fully connected layer is determined based on the number of output nodes of the fully connected layer of the speaker embedding layer model. The number of output nodes is the number of speakers in the training set. The method further includes: All initial spectral features of the support set for the target training task are segmented based on a step size to determine all the target spectral features of the support set.
3. The method according to claim 1, characterized in that, The first operation determines the cross-entropy loss of the support set based on all the target spectral features, the initial speaker embedding layer model, and the initial speaker classification model of the support set, including: All target spectral features are input into the initial speaker embedding layer model to obtain the speaker embedding layer; The speaker embedding layer is input into the initial speaker classification model, and the cross-entropy loss of the support set is determined based on the speaker labels.
4. The method according to claim 1, characterized in that, The cross-entropy loss and backpropagation method based on all the support sets determines the first speaker classification model and the first speaker embedding layer model through a second operation, including: The gradients of the initial speaker classification model and the initial speaker embedding layer model are calculated sequentially using the backpropagation method based on the cross-entropy loss. The first parameters of the initial speaker classification model and the initial speaker embedding layer model are obtained based on the gradients of the initial speaker classification model and the initial speaker embedding layer model. The first speaker classification model and the first speaker embedding layer model are determined based on the first parameters of the initial speaker classification model and the initial speaker embedding layer model.
5. The method according to claim 1, characterized in that, The average loss of the six query sets for all training tasks determined based on the first speaker classification model and the first speaker embedding layer model is used to determine the target speaker embedding layer model through a third operation, including: The average loss of the six query sets for all the training tasks is calculated based on the first speaker classification model and the first speaker embedding layer model. The second parameters of the initial speaker classification model and the initial speaker embedding layer model are updated and obtained by backpropagation based on the average loss of the six query sets of all the training tasks, wherein the second parameters are used to determine the target speaker embedding layer model.
6. The method according to claim 5, characterized in that, Also includes: Repeat the first, second, and third operations; If there exists a second parameter that can make the average loss of the six query sets of all training tasks converge, then the second parameter is determined as the target parameter of the initial speaker classification model and the initial speaker embedding layer model. The target speaker classification model and the target speaker embedding layer model are determined based on the target parameters of the initial speaker classification model and the initial speaker embedding layer model.
7. A speaker recognition device based on ordinary pronunciation, characterized in that, include: A partitioning unit is used to divide the training set into at least two training tasks. Each training task includes at least two speakers, each speaker includes at least two audio sentences, each audio sentence is equipped with frame-level phoneme labels and speaker labels and corresponding target spectral features, and each training task includes a support set and six query sets. The first determining unit is configured to determine the cross-entropy loss of the support set based on all the target spectral features of the support set, the initial speaker embedding layer model, and the initial speaker classification model through a first operation. An acquisition unit is used to determine a first speaker classification model and a first speaker embedding layer model through a second operation based on the cross-entropy loss and backpropagation method of all the support sets. The second determining unit is used to determine the target speaker embedding model through a third operation based on the average loss of six query sets of all training tasks determined by the first speaker classification model and the first speaker embedding layer model. Real-time spectral feature extraction of real-time audio; Based on the real-time spectral features, ordinary pronunciations in the real-time audio are detected, wherein ordinary pronunciations are pronunciations that occur in spoken dialogue and are less subject to the speaker's subjective control; In the case where the target real-time audio contains trivial pronunciations, the target real-time spectral features of the target real-time audio are input into the target speaker embedding layer model to obtain the speaker embedding layer of the target real-time audio; Based on the spectral characteristics of the registrant's ordinary pronunciation, the preset ordinary pronunciation embedding layer of the registrant is determined through the target speaker embedding layer model; Calculate the cosine similarity between the speaker embedding layer of the target real-time audio and the preset trivial pronunciation embedding layer of the registrant; The matching status between the speaker and the registrant in the target real-time audio is determined based on the cosine similarity.
8. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored program, wherein, when the program is executed by a processor, it implements the speaker recognition method based on ordinary pronunciation as described in any one of claims 1 to 6.
9. An electronic device, characterized in that, The electronic device includes at least one processor and at least one memory connected to the processor; wherein the processor is configured to call program instructions in the memory to execute the speaker recognition method based on ordinary pronunciation as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Speaker embedded layer model training method based on data set difficulty, medium and equipment
CN117423333A
Voiceprint recognition and voice deception detection integration method
CN119832917A