Voiceprint extraction method, identity recognition method and related devices
By extracting and statistics based on the spectral map in the voiceprint extraction method, obtaining phoneme features and forming voiceprint features, the problem of poor performance of voiceprint features in the prior art is solved, and higher robustness and accuracy are achieved, and the processing effect of applications such as identity recognition is improved.
Patent Information
- Application Number
- CN202210239481.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-11
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2042-03-11
AI Technical Summary
The voiceprint feature performance obtained by the existing voiceprint extraction methods is not good enough, which affects the processing effect of application scenarios such as identity recognition.
Feature extraction is performed based on the spectral map of the target object, feature sequences of phoneme fragments are obtained, feature statistics are performed to obtain phoneme characteristics, and voiceprint characteristics are obtained based on phoneme characteristics. This method weakens the differences between phoneme-level text information through feature statistics, retains information related to pronunciation characteristics, and improves the robustness and accuracy of voiceprint characteristics.
Effectively utilize phoneme-level text information to reduce its interference to voiceprint features, improve the robustness and accuracy of voiceprint features, and thus improve the processing effect of application scenarios such as identity recognition.
Smart Images

Figure CN114783415B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of speech processing, and in particular, to a voiceprint extraction method, an identity recognition method, a voiceprint extraction device, an identity recognition device, an electronic device, and a computer-readable storage medium. Background Art
[0002] Voiceprint features play an important role in application scenarios such as identity recognition and big data analysis. Taking identity recognition as an example, identity recognition can be further divided into identity recognition in the financial field, identity recognition in the security field, identity recognition in the smart home field, and so on. Applying voiceprint features to identity recognition can achieve identity recognition without the knowledge of the identity recognition object, and has a high acceptance rate.
[0003] The performance of voiceprint features affects the processing effects in various application scenarios. However, the performance of the voiceprint features obtained by the current voiceprint extraction methods is not good enough. Summary of the Invention
[0004] The present application provides a voiceprint extraction method, an identity recognition method, a voiceprint extraction device, an identity recognition device, an electronic device, and a computer-readable storage medium, which can solve the problem that the performance of the voiceprint features obtained by the current voiceprint extraction methods is not good enough.
[0005] To solve the above technical problems, one technical solution adopted by the present application is: to provide a voiceprint extraction method. The method includes: performing feature extraction on a first spectrogram of a target object to obtain a feature sequence of a plurality of phoneme segments; wherein the feature sequence includes at least one frame-level feature; performing feature statistics on the feature sequence of the phoneme segments to obtain the phoneme features of the phoneme segments; and obtaining the voiceprint features of the target object based on the phoneme features of the plurality of phoneme segments. To solve the above technical problems, another technical solution adopted by the present application is: to provide an identity recognition method. The method includes: obtaining a first voiceprint feature of an object to be recognized, and obtaining a voiceprint feature library; wherein the voiceprint feature library contains a plurality of second voiceprint features, and each second voiceprint feature is labeled with the identity information of the belonging object, and the first voiceprint feature and / or the second voiceprint feature are extracted based on the aforementioned voiceprint extraction method; and analyzing based on the first voiceprint feature and the voiceprint feature library to obtain the identity information of the object to be recognized.
[0006] To solve the above technical problems, another technical solution adopted by this application is: to provide a voiceprint extraction device, which includes: a feature extraction module, configured to perform feature extraction based on the first spectrogram of a target object to obtain a feature sequence of several phoneme segments; wherein, the feature sequence includes at least one frame-level feature; a feature statistics module, configured to perform feature statistics based on the feature sequence of the phoneme segments to obtain the phoneme features of the phoneme segments; a voiceprint acquisition module, configured to obtain the voiceprint features of the target object based on the phoneme features of several phoneme segments.
[0007] To solve the above technical problems, another technical solution adopted by this application is: to provide an identity recognition device, which includes: a feature acquisition module, configured to acquire the first voiceprint features of an object to be recognized and acquire a voiceprint feature library; wherein, the voiceprint feature library contains several second voiceprint features, and each second voiceprint feature is labeled with the identity information of the belonging object, and the first voiceprint features and / or the second voiceprint features are extracted based on the voiceprint extraction device as described above; a voiceprint analysis module, configured to perform analysis based on the first voiceprint features and the voiceprint feature library to obtain the identity information of the object to be recognized.
[0008] To solve the above technical problems, another technical solution adopted by this application is: to provide an electronic device, which includes a processor and a memory connected to the processor, wherein the memory stores program instructions; the processor is configured to execute the program instructions stored in the memory to implement the above method.
[0009] To solve the above technical problems, another technical solution adopted by this application is: to provide a computer-readable storage medium, storing program instructions, which can implement the above method when the program instructions are executed.
[0010] In the above manner, this application first obtains the feature sequence of phoneme segments from the first spectrogram of the target object, then converts the feature sequence of phoneme segments into the phoneme features of phoneme segments through feature statistics, and then obtains the voiceprint features based on the phoneme features. Since feature statistics weakens the differences between different phoneme-level text information covered in the feature sequence, the phoneme features and the voiceprint features obtained based on the phoneme features can cover as little phoneme-level text information of the target object as possible and retain as much information related to the pronunciation characteristics of the target object itself as possible, that is, decouple from the phoneme-level text information as much as possible, effectively utilize the phoneme-level text information and reduce the interference of the phoneme-level text information on the voiceprint features, and improve the robustness and accuracy of the voiceprint features. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] Figure 1 is a schematic flowchart of an embodiment of the voiceprint extraction method of this application;
[0012] Figure 2It is a schematic flowchart of another embodiment of the voiceprint extraction method of the present application;
[0013] Figure 3 It is a schematic flowchart of another embodiment of the voiceprint extraction method of the present application;
[0014] Figure 4 It is a schematic flowchart of the attention statistical pooling of the present application;
[0015] Figure 5 It is a schematic flowchart of a specific example of the voiceprint extraction of the present application;
[0016] Figure 6 It is a schematic flowchart of another embodiment of the voiceprint extraction method of the present application;
[0017] Figure 7 It is a schematic flowchart of another embodiment of the voiceprint recognition method of the present application;
[0018] Figure 8 It is a schematic structural diagram of an embodiment of the voiceprint extraction device of the present application;
[0019] Figure 9 It is a schematic structural diagram of an embodiment of the identity recognition device of the present application;
[0020] Figure 10 It is a schematic structural diagram of an embodiment of the electronic device of the present application;
[0021] Figure 11 It is a schematic structural diagram of an embodiment of the computer-readable storage medium of the present application. Detailed implementation manners
[0022] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0023] The terms "first", "second", and "third" in the present application are only used for descriptive purposes, and cannot be understood as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first", "second", and "third" may explicitly or implicitly include at least one of the features. In the description of the present application, the meaning of "a plurality" is at least two, such as two, three, etc., unless otherwise specifically and clearly defined.
[0024] References to "embodiments" in this specification mean that the particular features, structures, or characteristics described in connection with the embodiments can be included in at least one embodiment of the present application. The phrase appears in various places in the specification and does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that, without conflict, the embodiments described herein can be combined with other embodiments.
[0025] Figure 1 is a schematic flowchart of an embodiment of the voiceprint extraction method of the present application. It should be noted that if there are substantially the same results, this embodiment is not limited to Figure 1 the process sequence shown. As Figure 1 shown, this embodiment may include:
[0026] S11: Extract features from the first spectrogram of the target object to obtain a feature sequence of a plurality of phoneme segments.
[0027] Among them, the feature sequence includes at least one frame-level feature.
[0028] The target object can be a person, an animal, a robot, or other entities that can emit sounds.
[0029] The first spectrogram can be obtained based on the voice data of the target object. The voice data of the target object can be real-time or non-real-time, depending on the specific application scenario. For example, in scenarios with high real-time requirements, the voice data of the target object is real-time; in scenarios with low real-time requirements, the semantic data of the target object is non-real-time.
[0030] In some embodiments, the first spectrogram can be constructed from the voice data of the target object. Specifically, a window function and Fourier transform can be applied to the voice data to obtain frequency-domain features (filterbank features) of dimension d, and the spectrogram composed of the frequency-domain features is used as the first spectrogram. The window function can be a rectangular window function, a Hanning window function, a Hamming window function, etc.
[0031] In some embodiments, to avoid the situation of too long voice data, the spectrogram composed of frequency-domain features can be segmented to obtain the first spectrogram. Specifically, a second spectrogram can be constructed based on the voice data of the target object, and the second spectrogram is segmented to obtain a plurality of spectrogram segments; at least one spectrogram segment is selected as the first spectrogram respectively.
[0032] For example, the first spectrogram is segmented according to the window length l to obtain N spectrogram segments {Seg 1 、Seg 2 、...、Seg N}, and the size of each spectrogram segment is l×d (if the length is less than l, the voice data is copied several times and the excess is discarded).
[0033] The size of the window length l can be set arbitrarily, or set to a specified value. For example, the specified value is 1 / 2 of the average effective duration of the voice data. It can be understood that if the window length l is set too small, the first spectrogram may be fragmented, and the coherent spectrogram is cut into several small spectrogram segments, resulting in excessive loss of information between the spectrogram segments and being unable to model the long-term correlation of the voice. If the window length l is set too large, it will affect the processing efficiency and consume too much computing resources.
[0034] A phoneme segment has a feature sequence. In the feature sequence of the phoneme segment, each frame-level feature corresponds to a time frame.
[0035] In some embodiments, a trained phoneme feature sequence extraction model can be used to extract features from the first spectrogram of the target object to obtain the feature sequences of several phoneme segments.
[0036] In some embodiments, the audio feature sequence related to the phoneme-level text information can be extracted from the first spectrogram first, and then the audio feature sequence is segmented into the feature sequences of several phoneme segments.
[0037] S12: Perform feature statistics based on the feature sequences of the phoneme segments to obtain the phoneme features of the phoneme segments.
[0038] The feature statistics result can be any feature that can characterize all the frame-level features in the feature sequence. For example, feature variance (variance of all frame-level features), feature mean (mean of all frame-level features), feature standard deviation (standard deviation of all frame-level features), feature range (range of all features), etc.
[0039] In some embodiments, the feature statistics result can be directly used as the phoneme feature of the phoneme segment.
[0040] In some embodiments, the processing results such as the concatenation result and fusion result of different types of feature statistics results can be used as the phoneme features of the phoneme segment.
[0041] S13: Obtain the voiceprint feature of the target object based on the phoneme features of several phoneme segments.
[0042] If there is only one first spectrogram, the voiceprint feature of the target object can be obtained by processing the phoneme features of several phoneme segments through a neural network.
[0043] If there are multiple first spectrograms, for each first spectrogram, the voiceprint feature corresponding to the first spectrogram can be obtained based on the phoneme features of several phoneme segments; the voiceprint features corresponding to each first spectrogram are fused to obtain the voiceprint feature of the target object. Among them, the voiceprint feature corresponding to the first spectrogram can be obtained by processing the phoneme features of several phoneme segments through a neural network. The neural network for processing phoneme features can be, but is not limited to, a DNN (Deep Neural Network).
[0044] It can be understood that in the application task of voiceprint features, the text information in the source speech data (the speech data of the target object) of the voiceprint features is not concerned. In other words, the application of voiceprint features has nothing to do with the text information in the source speech data. Therefore, one of the criteria for judging the quality of voiceprint features is whether the text information they cover is sufficiently small.
[0045] Through the implementation of this embodiment, first, the feature sequence of phoneme segments is obtained from the first spectrogram of the target object, then the feature sequence of phoneme segments is converted into the phoneme features of phoneme segments through feature statistics, and then the voiceprint feature is obtained based on the phoneme features. Since feature statistics will weaken the differences between different phoneme-level text information covered in the feature sequence, the phoneme features and the voiceprint features obtained based on the phoneme features can cover as little phoneme-level text information of the target object as possible and retain as much information related to the pronunciation characteristics of the target object itself as possible, that is, be decoupled from the phoneme-level text information as much as possible, effectively utilize the phoneme-level text information and reduce the interference of the phoneme-level text information on the voiceprint features, and improve the robustness and accuracy of the voiceprint features.
[0046] Figure 2 It is a schematic flowchart of another embodiment of the voiceprint extraction method of the present application. It should be noted that if there are substantially the same results, this embodiment is not limited to Figure 2 the shown process sequence. This embodiment is a further expansion of S11. As Figure 2 shown, this embodiment may include:
[0047] S21: Perform feature extraction based on the first spectrogram to obtain frame-level features.
[0048] The frame-level features are the audio feature sequences mentioned above. In some embodiments, a trained neural network for frame-level feature extraction can be used to process the first spectrogram to obtain frame-level features.
[0049] In some embodiments, multiple different spectrogram features of the first spectrogram can be obtained by performing multiple different spectrogram feature extractions based on the first spectrogram; the multiple different spectrogram features are integrated to obtain frame-level features.
[0050] Among them, the spectrogram feature extraction can be implemented by a neural network. The neural network can be, but is not limited to, a CNN (Convolutional Neural Network). Taking the CNN as an example, M different convolutional kernels can be used to process the first spectrogram to obtain M spectrogram features (local features). It can be understood that when the CNN performs spectrogram feature extraction, it can jointly analyze the first spectrogram from the time domain and the frequency domain, deeply excavate the information in the first spectrogram, or express more detailed features, so as to obtain more accurate spectrogram features.
[0051] The global information of multiple spectrogram features can be integrated through a Transformer (machine translation model), an LSTM (Long Short Term Memory), etc. to obtain frame-level features.
[0052] The frame-level feature can be expressed as c t (t = 1, 2,..., T), where T is the time frame serial number. The frame-level feature has good text representation ability and strong correlation with phonemes. A phoneme segment corresponds to several consecutive features in the frame-level feature. Therefore, the feature sequences of each phoneme segment can be obtained through clustering subsequently.
[0053] S22: Perform feature clustering on the frame-level features to obtain the feature sequences of each phoneme segment.
[0054] The ways of feature clustering can be K-Means clustering, density-based clustering (DBSCAN), mean shift clustering, etc. For example, through K-Means clustering, the frame-level feature c t (t = 1, 2,..., T) can be segmented into K categories, and each category k (k = 1, 2,..., K) represents the feature sequence of a phoneme segment where, t k represents the starting time frame serial number, and p represents the frame length of the phoneme segment.
[0055] Furthermore, in the above S12, the ways of feature statistics can be statistical pooling (Static Pooling), attention statistical pooling (Attention Static Pooling), etc. In the statistical pooling method, the weights of each frame-level feature in the feature sequence of the phoneme segment are the same. In the attention statistical pooling method, the weights of each frame-level feature in the feature sequence of the phoneme segment are attention weights. The other processing processes of the statistical pooling method and the attention statistical pooling method are similar and will not be elaborated here. The attention statistical pooling method is introduced in detail as follows:
[0056] Figure 3 is a schematic flowchart of another embodiment of the voiceprint extraction method of the present application. It should be noted that if there are substantially the same results, this embodiment does not Figure 3Limited to the shown process sequence. This embodiment is a further expansion of S12. For example, Figure 3 As shown, this embodiment may include:
[0057] S31: Obtain the attention weights of each frame-level feature in the feature sequence.
[0058] The linear transformation can be performed on each frame-level feature to obtain the linear transformation result. The linear transformation result of the k-th frame-level feature can be calculated according to the following formula:
[0059]
[0060] where W represents the linear transformation projection matrix, B represents the bias, represents the k-th frame-level feature and represents the linear transformation result.
[0061] In some embodiments, the linear transformation results of each frame-level feature can be sequentially used as the linear transformation result of the current frame-level feature. Divide the current linear transformation result by all linear transformation results to obtain the initial attention weight of the current frame-level feature. Normalize the initial attention weights of each frame-level feature to obtain the attention weights of each frame-level feature.
[0062] In some embodiments, the softmax function can be directly used to process each linear transformation result to obtain the attention weights of each frame-level feature. The formula for calculating the attention weight of the k-th frame-level feature can be as follows:
[0063]
[0064] where, represents the attention weight of the k-th frame-level feature.
[0065] S32: Based on each frame-level feature and its attention weight in the feature sequence, obtain the statistical data of the feature sequence.
[0066] Among them, the statistical data includes: feature mean and / or feature variance.
[0067] Regarding the feature mean: The feature sequence can be weighted according to the attention weights to obtain the feature mean. The formula can be as follows:
[0068]
[0069] where μ k represents the feature mean.
[0070] For the feature variance: the feature differences between each frame-level feature in the feature sequence and the feature mean can be obtained; the product of the transposed result of each feature difference and the feature difference can be obtained; based on each product and the attention weights, the feature variance can be obtained. The formula can be as follows:
[0071]
[0072] where diag represents taking the diagonal elements, and (.) T represents taking the transpose, and σ k represents the feature variance.
[0073] S33: Based on the statistical data of the feature sequence, obtain the phoneme features.
[0074] If the statistical data only includes the feature mean, the feature mean can be directly used as the phoneme feature; if the statistical data only includes the feature variance, the feature variance can be directly used as the phoneme feature; if the statistical data includes the feature mean and the feature variance, the concatenated result of the feature mean and the feature variance can be used as the phoneme feature.
[0075] In this embodiment, the method of first extracting the feature sequence related to the phoneme-level text content and then converting the feature sequence into phoneme features can effectively utilize the phoneme-level text content, and since the phoneme features are the statistical data of the feature sequence, the potential phoneme-level text content is reduced compared to the feature sequence.
[0076] The following combines Figure 4 , in the form of an example, to elaborate on S31 - S33 in detail:
[0077] 1) Perform a linear transformation on the nth first spectrogram and the kth phoneme segment based on W and B to obtain the linear transformation result
[0078] 2) Use the softmax function to process to obtain the attention weights
[0079] 3) Calculate the product of and to obtain the mean μ k and the variance σ k ;
[0080] 4) Concatenate the mean μ k and the variance σ k to obtain the phoneme feature of
[0081] The phoneme features of K phoneme segments of the nth first spectrogram form a phoneme feature sequence h corresponding to the nth first spectrogram k (k = 1, 2, ..., K).
[0082] As follows, in combination with Figure 5 , a detailed description of voiceprint extraction is given in the form of an example:
[0083] 1) Use CNN, Transformer, K-Means clustering, and Attention Static Pooling to process the nth first spectrogram in sequence to obtain the phoneme feature sequence h corresponding to the nth first spectrogram k (k = 1, 2, ..., K).
[0084] 2) Use DNN to compress the dimension of h k to obtain the voiceprint feature w corresponding to the nth first spectrogram n .
[0085] 3) Perform weighted averaging on the voiceprint features corresponding to N first spectrograms to obtain the voiceprint feature of the target object
[0086] Furthermore, the voiceprint extraction method provided by this application is implemented based on a voiceprint extraction model, that is, the voiceprint feature is extracted based on the voiceprint extraction model. The voiceprint extraction model is trained based on sample spectrograms
[0087] In some embodiments, the sample spectrogram is labeled with the true results of the phoneme features of several phoneme segments. During the training process, use the voiceprint extraction model to process to obtain the predicted results of the phoneme features of several phoneme segments corresponding to the sample spectrogram, and based on the difference between the true results and the predicted results of the phoneme features of several phoneme segments, adjust the parameters of the voiceprint extraction model
[0088] In some embodiments, the sample spectrogram is labeled with the true results of the sample voiceprint features. During the training process, use the voiceprint extraction model to extract the predicted results of the sample voiceprint features, and based on the difference between the true results and the predicted results of the sample voiceprint features, adjust the parameters of the voiceprint extraction model
[0089] In some embodiments, the sample spectrogram is labeled with the sample object to which it belongs. The following details this training method:
[0090] Figure 6 is a schematic flowchart of another embodiment of the voiceprint extraction method of this application. It should be noted that if there are substantially the same results, this embodiment is not limited to Figure 6 the process sequence shown. In this embodiment. As Figure 6 shown, this embodiment may include:
[0091] S41: Extract the voiceprint from the sample spectrogram based on the voiceprint extraction model to obtain the sample voiceprint feature.
[0092] The extraction method of the sample voiceprint feature is similar to that of the voiceprint feature of the aforementioned target object, and will not be elaborated here.
[0093] S42: Make a prediction based on the sample voiceprint feature to obtain the predicted object to which the sample spectrogram belongs.
[0094] The softmax function can be used to predict the probabilities that the sample spectrogram belongs to each candidate object. The sum of the probabilities belonging to each candidate object is 1, and the candidate object with the largest corresponding probability is taken as the predicted object.
[0095] S43: Adjust the network parameters of the voiceprint extraction model based on the difference between the sample object and the predicted object.
[0096] Based on the difference between the sample object and the predicted object, a loss function (such as the cross-entropy loss function) can be constructed, and the parameters of the voiceprint extraction model are adjusted based on the loss function. The conditions for the end of training can include that the number of training times reaches the expectation, the training effect reaches the expectation, the training time reaches the expectation, and so on.
[0097] It can be understood that when predicting the predicted object to which the sample spectrogram belongs, it is based on the information of the pronunciation characteristics of the pronunciation object covered by the sample voiceprint feature (such as phonemes, pauses, etc.), and has nothing to do with the text information of the pronunciation. Therefore, the prediction result can measure the expression ability of the sample voiceprint feature for the information of the pronunciation characteristics of the pronunciation object. In this way, by adjusting the network parameters based on the difference between the sample object and the predicted object, the voiceprint extraction model can learn the information of the pronunciation characteristics and reduce the attention to the text information, so that the expression ability of the extracted sample voiceprint feature for the content of the pronunciation characteristics of the pronunciation object is getting stronger and stronger.
[0098] By implementing this embodiment, the voiceprint extraction effect of the voiceprint extraction model can be measured through the prediction result based on the sample voiceprint feature, and the network parameters of the voiceprint extraction model can be adjusted reversely according to the voiceprint recognition effect, so as to realize the training of the voiceprint extraction model.
[0099] The voiceprint feature extracted by the aforementioned voiceprint extraction method can be used for storage and can be applied to identity recognition scenarios such as determining the same object from multiple objects to be recognized and determining whether the object to be recognized is a specified object. The following are several application scenarios of the voiceprint feature in identity recognition:
[0100] Identity recognition application scenario 1: To improve security, the financial management APP has set up multiple identity verification links such as account password verification and voiceprint feature verification. Party A needs to log in to his financial management APP. After the account password verification, the financial management APP will prompt Party A to input voice data. The financial management APP obtains the first spectrogram from the voice data, extracts the voiceprint feature from the first spectrogram, obtains the voiceprint feature of Party A, compares the voiceprint feature of Party A with the pre-stored voiceprint feature under the account password, judges whether the similarity meets the requirements. If it meets, the voiceprint feature verification passes and logging in is allowed; otherwise, it is not allowed, etc.
[0101] Identity recognition application scenario 2: Party B needs to enter a certain community. The community checkpoint management personnel do not know whether Party B belongs to the community. The community access control device is used to verify the voiceprint feature of Party B. The access control device will prompt Party B to input voice data. The financial management APP obtains the first spectrogram from the voice data, extracts the voiceprint feature from the first spectrogram, obtains the voiceprint feature of Party A, compares the voiceprint feature of Party A with the voiceprint feature in the personnel management database of the community, judges whether the highest similarity between the voiceprint feature of Party A and the voiceprint feature in the management database meets the requirements. If it meets, it is determined that Party B is a community member; otherwise, it is determined that Party B is not a community member.
[0102] Identity recognition application scenario 3: For two people who appear at different times and in different regions, by collecting their voice data, their first spectrograms are determined, the voiceprint features are respectively extracted from their first spectrograms, and based on the similarity of their voiceprint features, it is determined whether they belong to the same person.
[0103] Identity recognition application scenario 4: At a scene where it is necessary to distinguish how many people are speaking (such as a conference scene), by collecting segmented or clause-based voice data at the scene, the corresponding voice data for one segment or one clause is regarded as being emitted by one person, the voiceprint feature corresponding to the voice data is extracted with one segment or one clause of corresponding voice data as a unit, the different voiceprint features are compared pairwise, and the number of speakers at the scene is determined according to the comparison result.
[0104] Figure 7 It is a schematic flowchart of another embodiment of the voiceprint recognition method of the present application. It should be noted that if there are substantially the same results, this embodiment is not limited to Figure 7 the shown process sequence. In this embodiment, the voiceprint extraction model is trained based on sample spectrograms, and the sample spectrograms are labeled with the corresponding sample objects. As Figure 7 shown, this embodiment may include:
[0105] S51: Obtain the first voiceprint feature of the object to be recognized, and obtain the voiceprint feature library.
[0106] Among them, the voiceprint feature library contains a number of second voiceprint features, each of which is labeled with the identity information of the object to which it belongs, and the first voiceprint feature and / or the second voiceprint feature are extracted based on the aforementioned voiceprint extraction method.
[0107] S52: Analyze based on the first voiceprint feature and the voiceprint feature library to obtain the identity information of the object to be recognized.
[0108] The first voiceprint feature can be respectively matched with a number of second voiceprint features in the voiceprint feature library, and the identity information labeled by the second voiceprint feature that meets the matching condition is used as the identity information of the object to be recognized. The matching condition can include at least one of the highest similarity with the first voiceprint feature and a similarity greater than a threshold.
[0109] Through the implementation of this embodiment, the first voiceprint feature is applied to voiceprint recognition to determine the identity information of the object to be recognized. Since the first voiceprint feature is extracted based on the aforementioned voiceprint extraction method, the first voiceprint feature has high robustness and high accuracy, so the accuracy of the determined identity information of the object to be recognized is high.
[0110] Figure 8 It is a schematic structural diagram of an embodiment of the voiceprint extraction device of the present application. As Figure 8 shown, the voiceprint extraction device 10 may include a feature extraction module 11, a feature statistics module 12, and a voiceprint acquisition module 13.
[0111] The feature extraction module 11 can be used to extract features based on the first spectrogram of the target object to obtain a feature sequence of a number of phoneme segments. Among them, the feature sequence may include at least one frame-level feature. The feature statistics module 12 can be used to perform feature statistics based on the feature sequence of the phoneme segments to obtain the phoneme features of the phoneme segments. The voiceprint acquisition module 13 can be used to obtain the voiceprint features of the target object based on the phoneme features of a number of phoneme segments.
[0112] Through the implementation of this embodiment, the voiceprint extraction device first uses the feature extraction module to obtain the feature sequence of the phoneme segments from the first spectrogram of the target object, then uses the feature statistics module to convert the feature sequence of the phoneme segments into the phoneme features of the phoneme segments through feature statistics, and then uses the voiceprint acquisition module to obtain the voiceprint features based on the phoneme features. Since feature statistics weakens the differences between different phoneme-level text information covered in the feature sequence, the phoneme features and the voiceprint features obtained based on the phoneme features can cover as little phoneme-level text information of the target object as possible and retain as much information related to the pronunciation characteristics of the target object itself as possible, that is, decouple from the phoneme-level text information as much as possible, effectively utilize the phoneme-level text information and reduce the interference of the phoneme-level text information on the voiceprint features, and improve the robustness and accuracy of the voiceprint features.
[0113] Further, feature extraction is performed on the first spectrogram of the target object to obtain the feature sequences of several phoneme segments, which specifically may include: performing feature extraction on the first spectrogram to obtain frame-level features; performing feature clustering on the frame-level features to obtain the feature sequences of each phoneme segment.
[0114] Therefore, frame-level features can be obtained based on the first spectrogram, and then the feature sequences of each phoneme segment can be obtained based on the method of feature clustering.
[0115] Further, performing feature extraction on the first spectrogram to obtain frame-level features specifically includes: performing multiple different spectrogram feature extractions on the first spectrogram to obtain multiple different spectrogram features of the first spectrogram; integrating the multiple different spectrogram features to obtain frame-level features.
[0116] Therefore, by performing multiple different spectrogram feature extractions on the first spectrogram, the first spectrogram can be jointly analyzed from the time domain and frequency domain perspectives, so that the obtained spectrogram features have higher accuracy, and further the frame-level features obtained by integrating the spectrogram features have higher accuracy.
[0117] Further, performing feature statistics on the feature sequences of the phoneme segments to obtain the phoneme features of the phoneme segments specifically may include: obtaining the attention weights of each frame-level feature in the feature sequence; obtaining the statistical data of the feature sequence based on each frame-level feature and its attention weight in the feature sequence; wherein, the statistical data includes: feature mean and / or feature variance; obtaining the phoneme features based on the statistical data of the feature sequence.
[0118] Therefore, the weights of each frame-level feature are attention weights, rather than being consistent. The greater the attention weight of the frame-level feature, the higher the importance of the frame-level feature represents. Therefore, the obtained statistical data has stronger expression ability for the frame-level features with high importance.
[0119] Further, the step of obtaining the feature mean may include: weighting each frame-level feature in the feature sequence according to the attention weight to obtain the feature mean.
[0120] Further, the step of obtaining the feature variance may include: obtaining the feature differences between each frame-level feature in the feature sequence and the feature mean; obtaining the product of the transposed result of each feature difference and the feature difference; obtaining the feature variance based on each product and the attention weight.
[0121] Further, obtaining the voiceprint feature of the target object based on the phoneme features of several phoneme segments specifically may include: for each first spectrogram, obtaining the voiceprint feature corresponding to the first spectrogram based on the phoneme features of several phoneme segments; fusing the voiceprint features respectively corresponding to each first spectrogram to obtain the voiceprint feature of the target object.
[0122] Therefore, in the case of the first spectrogram with multiple target objects, the voiceprint features corresponding to different first spectrograms can be fused to obtain the voiceprint features of the target object.
[0123] Furthermore, the voiceprint extraction device 10 may further include a spectrogram acquisition module. The spectrogram acquisition module can be used to construct a second spectrogram based on the voice data of the target object, segment the second spectrogram to obtain a plurality of spectrogram segments, and select at least one spectrogram segment as the first spectrogram respectively.
[0124] Therefore, in the case where the voice data / second spectrogram is long, the second spectrogram can be segmented into a plurality of spectrogram segments, and the spectrogram segments are used as the first spectrogram and applied to voiceprint extraction. The voiceprint extraction method based on the first spectrogram can improve the processing efficiency.
[0125] Furthermore, the voiceprint extraction device 10 may further include a training module. The training module can be used to train a voiceprint extraction model for extracting voiceprint features.
[0126] Furthermore, the voiceprint extraction model is trained based on sample spectrograms, and the sample spectrograms are labeled with the sample objects to which they belong.
[0127] Furthermore, the training steps of the voiceprint extraction model specifically include: performing voiceprint extraction on the sample spectrogram based on the voiceprint extraction model to obtain sample voiceprint features; making a prediction based on the sample voiceprint features to obtain the predicted object to which the sample spectrogram belongs; and adjusting the network parameters of the voiceprint extraction model based on the difference between the sample object and the predicted object.
[0128] Since the predicted object to which the sample spectrogram belongs is based on the information of the pronunciation characteristics of the pronunciation object covered by the sample voiceprint features (such as phonemes, pauses, etc.) and has nothing to do with the text information of the pronunciation, the prediction result can measure the expression ability of the sample voiceprint features for the information of the pronunciation characteristics of the pronunciation object. Thus, by adjusting the network parameters based on the difference between the sample object and the predicted object, the voiceprint extraction model can learn the information of the pronunciation characteristics and reduce the attention to the text information, so that the sample voiceprint features extracted can have an increasingly strong expression ability for the content of the pronunciation characteristics of the pronunciation object.
[0129] For other detailed descriptions of the voiceprint extraction device, please refer to the previous embodiments and will not be elaborated here.
[0130] Figure 9 is a schematic structural diagram of an embodiment of the identity recognition device of the present application. As Figure 9 shown, the identity recognition device includes a feature acquisition module and a voiceprint analysis module.
[0131] The feature acquisition module is used to acquire the first voiceprint feature of the object to be recognized and acquire the voiceprint feature library; wherein, the voiceprint feature library contains a number of second voiceprint features, each of the second voiceprint features is labeled with the identity information of the object to which it belongs, and the first voiceprint feature and / or the second voiceprint feature are extracted based on the aforementioned voiceprint extraction device.
[0132] The voiceprint analysis module is used to perform analysis based on the first voiceprint feature and the voiceprint feature library to obtain the identity information of the object to be recognized.
[0133] Through the implementation of this embodiment, the identity recognition device uses the feature acquisition module to acquire the first voiceprint feature, and performs analysis based on the first voiceprint feature and the second voiceprint features in the voiceprint feature library to obtain the identity information of the object to be recognized. Since the first voiceprint feature and / or the second voiceprint feature are extracted based on the aforementioned voiceprint extraction device, they have high accuracy and robustness, and have strong expressive ability for the pronunciation characteristics of the object to which the first voiceprint feature and / or the second voiceprint feature belong. Therefore, the accuracy of the identity information of the object to be recognized obtained is high.
[0134] For other detailed descriptions of the identity recognition device, please refer to the previous embodiments and will not be elaborated here.
[0135] Figure 10 It is a schematic structural diagram of an embodiment of an electronic device according to the present application. As Figure 10 shown, the electronic device includes a processor 21 and a memory 22 coupled to the processor 21.
[0136] Among them, the memory 22 stores program instructions for implementing the method of any of the above embodiments; the processor 21 is used to execute the program instructions stored in the memory 22 to implement the steps of the above method embodiment. Among them, the processor 21 can also be called a CPU (Central Processing Unit, central processing unit). The processor 21 may be an integrated circuit chip with signal processing capabilities. The processor 21 may also be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0137] Figure 11 It is a schematic structural diagram of an embodiment of a computer-readable storage medium according to the present application. As Figure 11As shown, the computer-readable storage medium 30 of the embodiment of the present application stores program instructions 31, and when the program instructions 31 are executed, the method provided in the above embodiments of the present application is implemented. Among them, the program instructions 31 can form a program file and be stored in the above computer-readable storage medium 30 in the form of a software product, so that a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor can execute all or part of the steps of the methods in various embodiments of the present application. The aforementioned computer-readable storage medium 30 includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs, or terminal devices such as computers, servers, mobile phones, and tablets.
[0138] In several embodiments provided by the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of devices or units can be in electrical, mechanical, or other forms.
[0139] In addition, each functional unit in various embodiments of the present application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit. The above are only the embodiments of the present application, and do not limit the patent scope of the present application. Any equivalent structural or equivalent process transformation made by using the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present application.
Claims
1. A voiceprint extraction method, characterized in that, Including: Performing feature extraction on a first spectrogram of a target object to obtain a feature sequence of a plurality of phoneme segments; wherein, the feature sequence includes at least one frame-level feature, and each frame-level feature corresponds to a time frame; Performing feature statistics on all the frame-level features in the feature sequence of the phoneme segments to obtain the phoneme features of the phoneme segments; Obtaining the voiceprint feature of the target object based on the phoneme features of the plurality of phoneme segments; The performing feature extraction on a first spectrogram of a target object to obtain a feature sequence of a plurality of phoneme segments includes: Performing feature extraction on the first spectrogram to obtain the frame-level features; Performing feature clustering on the frame-level features to obtain the feature sequences of the respective phoneme segments.
2. The method according to claim 1, wherein The performing feature extraction on the first spectrogram to obtain the frame-level features includes: Performing multiple different spectrogram feature extractions on the first spectrogram to obtain multiple different spectrogram features of the first spectrogram; Integrating the multiple different spectrogram features to obtain the frame-level features.
3. The method according to claim 1, characterized in that, The performing feature statistics on all the frame-level features in the feature sequence of the phoneme segments to obtain the phoneme features of the phoneme segments includes: Obtaining the attention weights of the respective frame-level features in the feature sequence; Based on the respective frame-level features in the feature sequence and their attention weights, obtaining the statistical data of the feature sequence; wherein, the statistical data includes: feature mean and / or feature variance; Obtaining the phoneme features based on the statistical data of the feature sequence.
4. The method according to claim 3, characterized in that The step of obtaining the feature mean includes: Weighting the respective frame-level features in the feature sequence according to the attention weights to obtain the feature mean.
5. The method according to claim 3, characterized in that The step of obtaining the feature variance includes: Obtaining the feature differences between the respective frame-level features in the feature sequence and the feature mean; Obtaining the product of the transposed result of each feature difference and the feature difference; Based on each product and the attention weights, obtaining the feature variance.
6. The method according to claim 1, characterized in that, Before the performing feature extraction on a first spectrogram of a target object to obtain a feature sequence of a plurality of phoneme segments, the method further includes: Based on the speech data of the target object, constructing a second spectrogram, and segmenting the second spectrogram to obtain a plurality of spectrogram segments; Selecting at least one of the spectrogram segments as the first spectrogram respectively; The obtaining the voiceprint feature of the target object based on the phoneme features of the plurality of phoneme segments includes: For each of the first spectrograms, obtaining the voiceprint feature corresponding to the first spectrogram based on the phoneme features of the plurality of phoneme segments; Fusing the voiceprint features respectively corresponding to the respective first spectrograms to obtain the voiceprint feature of the target object.
7. The method according to claim 1, characterized in that The voiceprint feature is extracted based on a voiceprint extraction model, and the voiceprint extraction model is trained based on sample spectrograms, and the sample spectrograms are labeled with the sample objects to which they belong.
8. The method according to claim 7, wherein The training step of the voiceprint extraction model includes: Performing voiceprint extraction on the sample spectrograms based on the voiceprint extraction model to obtain sample voiceprint features; Perform prediction based on the sample voiceprint features to obtain the predicted object to which the sample spectrogram belongs; Adjust the network parameters of the voiceprint extraction model based on the difference between the sample object and the predicted object.
9. An identity recognition method, characterized in that, Comprising: Obtain the first voiceprint features of the object to be recognized, and obtain a voiceprint feature library; wherein, the voiceprint feature library contains a number of second voiceprint features, each of the second voiceprint features is labeled with the identity information of the object to which it belongs, and the first voiceprint feature and / or the second voiceprint feature are extracted based on the voiceprint extraction method according to any one of claims 1 to 8; Analyze based on the first voiceprint features and the voiceprint feature library to obtain the identity information of the object to be recognized.
10. A voiceprint extraction device, characterized in that, Comprising: A feature extraction module, configured to perform feature extraction based on the first spectrogram of the target object to obtain a feature sequence of a number of phoneme segments; wherein, the feature sequence includes at least one frame-level feature, and each of the frame-level features corresponds to a time frame; A feature statistics module, configured to perform feature statistics based on the feature sequences of all the frame-level features in the phoneme segments to obtain the phoneme features of the phoneme segments; A voiceprint acquisition module, configured to obtain the voiceprint features of the target object based on the phoneme features of the number of phoneme segments.
11. An identity recognition device, characterized in that, Comprising: A feature acquisition module, configured to obtain the first voiceprint features of the object to be recognized, and obtain a voiceprint feature library; wherein, the voiceprint feature library contains a number of second voiceprint features, each of the second voiceprint features is labeled with the identity information of the object to which it belongs, and the first voiceprint feature and / or the second voiceprint feature are extracted based on the voiceprint extraction device according to claim 10; A voiceprint analysis module, configured to analyze based on the first voiceprint features and the voiceprint feature library to obtain the identity information of the object to be recognized.
12. An electronic device, characterized in that, Comprising a processor and a memory coupled to each other, the memory stores program instructions, and the processor is configured to execute the program instructions to implement the voiceprint extraction method according to any one of claims 1 to 8, or to implement the identity recognition method according to claim 9.
13. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores program instructions, and when the program instructions are executed, the voiceprint extraction method according to any one of claims 1 to 8 is implemented, or the identity recognition method according to claim 9 is implemented.
Citation Information
Patent Citations
Electronic device, authentication method and storage medium
CN108154371A
Voice checking method and device, electronic equipment and readable storage medium
CN110689895A
Deep neural network-based multi-class acoustic feature integration method and system
CN111276131A