A method and device for generating multi-level phonemes for speaker recognition
Through a multi-level phoneme generation method, frequent and highly distinguishing phonemes are screened, which solves the problem of short duration of linguistic phoneme units and improves the accuracy of voiceprint recognition.
Patent Information
- Application Number
- CN202211424536.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-15
- Publication Date
- 2025-09-30
- Estimated Expiration
- 2042-11-15
AI Technical Summary
In existing voiceprint recognition technology, the duration of linguistically defined phoneme units is short, making it difficult to provide sufficient speaker identity information, and phoneme transition information may be damaged or lost, resulting in poor recognition performance.
A multi-level phoneme generation method is adopted to screen out frequent and highly discriminative phonemes by calculating the frequency and distinctiveness of phonemes, and to construct the optimal phoneme set for speaker recognition.
The accuracy of speaker recognition has been improved. By mining phoneme combinations at multiple levels, more speaker-related information can be obtained, the optimal phoneme combination can be selected, and system performance can be improved.
Smart Images

Figure CN115731936B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of speech signal processing and voiceprint recognition, and in particular to a method and device for generating multi-level phonemes for speaker recognition. Background Art
[0002] Voiceprint information, as an important biometric, is an effective means of user authentication. Voiceprint recognition, which identifies a speaker based on a given voice signal, has a wide range of applications, particularly in security and smart devices. Compared to text-dependent voiceprint recognition, text-independent speaker recognition, as it is not restricted by the textual content of the voice signal, is more susceptible to text variations, resulting in reduced recognition performance. Therefore, phoneme / syllable-based voiceprint recognition systems can effectively mitigate the negative impact of text variations on recognition performance by modeling individual phonemes / syllables. However, selecting the appropriate speech units for modeling directly impacts the performance of the voiceprint recognition system. First, the modeled speech units must occur frequently so that the system can utilize them for modeling and recognition. Furthermore, the modeled speech units must exhibit good speaker discrimination to enhance the voiceprint recognition system. However, currently, no method exists to construct an optimal set of speech units for speaker recognition tasks.
[0003] Currently, voiceprint recognition systems that use phoneme units for modeling typically use linguistically defined phonemes as units and extract the speaker identity information contained therein. However, these methods often have the following problems:
[0004] 1) For voiceprint recognition tasks, the phoneme unit defined by linguistics may not be the optimal speech unit for identifying the speaker's identity;
[0005] 2) For linguistically defined phonemes, most phonemes have a very short duration, making it difficult to provide rich and sufficient speaker identity information for subsequent modeling;
[0006] 3) Modeling only a single phoneme may miss and damage the speaker-related information contained in the conversion between phonemes, resulting in poor performance of the speaker recognition system. Summary of the Invention
[0007] The present invention proposes a method and apparatus for multi-level phoneme generation for speaker recognition, aiming to address the shortcomings of existing voiceprint recognition technology, including the problem of poor recognition rate of speaker recognition systems based on phoneme modeling, caused by the fact that the duration of a single phoneme is too short to provide sufficient speaker identity information, and the speaker information existing between phoneme transitions may be damaged or lost.
[0008] To achieve the above object, the present invention provides the following technical solutions:
[0009] A method for multi-level phoneme generation for speaker recognition, comprising:
[0010] Determine the set of primary phonemes;
[0011] Obtaining a speech database and a primary phoneme sequence corresponding to each piece of speech data;
[0012] Starting from the first-level phoneme, by calculating the frequency of occurrence of the phoneme and the preset threshold, the frequent phonemes at each level are screened and a higher-level phoneme candidate set is generated until the stopping condition is met;
[0013] Starting from the first-level phoneme, by calculating the speaker discrimination of the phoneme and setting the discrimination requirement, the highly discriminative phonemes at each level are screened until the stopping condition is met to obtain the final multi-level phoneme set.
[0014] Further technical solution: The method for determining the primary phoneme is: using the phoneme category defined by linguistics, or using the minimum speech unit defined by the unsupervised learning method as the primary phoneme.
[0015] Further technical solution: The acquisition of the speech database and the primary phoneme sequence corresponding to each speech data includes: obtaining the phoneme sequence by manual annotation, or obtaining the phoneme sequence by using a speech recognition or phoneme recognition model.
[0016] Further technical solution: The obtained phoneme sequence is a phoneme category marked according to the order in which the phonemes appear in the speech signal.
[0017] Further technical solution: The method for screening frequent phonemes at each level and generating a phoneme candidate set at a higher level is specifically as follows:
[0018] Starting from the first-level phonemes, the set containing all the first-level phonemes is taken as the first-level phoneme candidate set, and the first-level phonemes that meet the frequent conditions constitute the first-level phoneme frequent set, and the second-level phoneme candidate set is generated from the first-level phoneme frequent set, and the frequent second-level phonemes are selected to constitute the second-level phoneme frequent set. Similarly, the k-1-level phoneme frequent set is used to construct the k-level phoneme candidate set, and the k-level phonemes that meet the frequent conditions are selected from the k-level phoneme candidate set to constitute the k-level phoneme frequent set, until a higher-level candidate set cannot be generated or there are no frequent phones that meet the conditions to construct a frequent set, where the k-level phoneme refers to an ordered combination formed by the merger of k first-level phonemes.
[0019] Further technical solution: The method for screening frequent phonemes at each level and generating a phoneme candidate set at a higher level is specifically as follows:
[0020] When k is greater than or equal to 2, the method for constructing a k-level phoneme candidate set from a k-1-level phoneme frequent set is: merging two k-1-level phonemes having k-2 intersections in the k-1-level phoneme frequent set.
[0021] Further technical solution: The k-level phonemes that meet the frequent conditions are selected from the k-level phoneme candidate set to form the k-level phoneme frequent set. The frequent conditions are: the frequency of the phoneme appearing in the data set is greater than a preset value, or the ratio of the number of sentences in which the phoneme appears to the total number of sentences in the database is greater than a preset value.
[0022] Further technical solution: The method for screening the highly distinguishing phonemes at each level is:
[0023] Starting from the first-level phoneme, the obtained k-level phoneme frequent set is used as a new candidate set, and the final k-level phoneme set is composed of k-level phonemes that meet the strong distinguishability condition, and so on, until there is no higher-level candidate set.
[0024] Further technical solution: The strong distinguishing condition in the final k-level phoneme set formed by the k-level phonemes satisfying the strong distinguishing condition includes:
[0025] A universal speaker recognition model is used to perform speaker recognition on data belonging to a phoneme category, so that the recognition accuracy is higher than a preset value.
[0026] At the same time, the present invention also provides the following technical solutions:
[0027] A device for multi-level phoneme generation for speaker recognition, comprising:
[0028] A data unit, which acquires and stores voice data and the primary phoneme sequence corresponding to each voice data;
[0029] The frequent candidate set generation unit, based on the determined first-level phonemes, takes the set containing all first-level phonemes as the first-level phoneme candidate set. For second-level and above phonemes, a k-level phoneme candidate set is generated from the k-1-level phoneme frequent set according to the constraint conditions.
[0030] The frequent phoneme screening unit calculates the frequency of occurrence of the k-level phonemes using the phoneme sequence tags in the speech data for the generated k-level phoneme candidate set, and screens the phonemes that meet the frequent conditions from the k-level phonemes according to the set frequent conditions to form a k-level phoneme frequent set;
[0031] The strongly discriminative phoneme screening unit calculates the discriminability of each k-level phoneme based on the obtained k-level phoneme frequent set as the candidate set, and screens out the phonemes that meet the conditions according to the set strong discriminative conditions to form a k-level strongly discriminative phoneme set.
[0032] Compared with the prior art, the present invention has the following beneficial effects:
[0033] 1) The multi-level phoneme generation method provided by the present invention can simultaneously consider the universality of phonemes and the distinguishability of speaker identities, which helps to comprehensively evaluate the role of phoneme units in speaker recognition and select the optimal phoneme combination to improve speaker recognition performance;
[0034] 2) The multi-level phoneme generation method provided by the present invention mines valuable phoneme combinations at multiple levels, fully obtains speaker-related information in the speech signal, and improves the accuracy of speaker recognition. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] Figure 1 Schematic diagram of a method flow for multi-level phoneme generation for speaker recognition according to an embodiment of the present invention;
[0036] Figure 2 This is a structural block diagram of a device for multi-level phoneme generation for speaker recognition according to an embodiment of the present invention. DETAILED DESCRIPTION
[0037] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0038] Example 1
[0039] like Figure 1 As shown, a flowchart of a method for multi-level phoneme generation for speaker recognition according to the present invention includes:
[0040] Step 1: Determine the set of primary phonemes.
[0041] The determined primary phoneme can be a phoneme category defined by linguistics, with the phoneme category serving as the primary phoneme; or a minimum speech unit defined by methods such as unsupervised learning, with the minimum speech unit serving as the primary phoneme.
[0042] In a specific embodiment, when the first-level phonemes adopt the phoneme categories defined in linguistics, for Chinese data, there are a total of 65 phoneme categories. All the phoneme categories are regarded as first-level phonemes, and a set of first-level phonemes including 65 phonemes is obtained.
[0043] Step 2: Obtain the speech database and the primary phoneme sequence corresponding to each piece of speech data.
[0044] Obtaining a speech database and the primary phoneme sequence corresponding to each piece of speech data, including obtaining the phoneme sequence through manual annotation, or obtaining the phoneme sequence through speech recognition, phoneme recognition, etc. The extracted phoneme sequence is specifically a phoneme category labeled according to the order in which the phonemes appear in the speech signal.
[0045] In a specific embodiment, preferably, a method of speech recognition and forced alignment is used for each sentence in the speech database to obtain a phoneme sequence, and the phoneme sequence corresponds to the order of appearance of the phonemes in the speech data.
[0046] Step 3: Starting from the first-level phoneme, by calculating the frequency of occurrence of the phoneme and the preset threshold, the frequent phonemes of each level are screened and a higher-level phoneme candidate set is generated until the stopping condition is met.
[0047] Starting from the first-level phonemes, the first-level phoneme set is used as the candidate set; the frequent condition is set to the frequent support of the phoneme is not less than R, where the specific calculation method of frequent support is: frequent support = the number of times the phoneme appears in the speech signal / the total number of speech signals in the data set; the first-level phonemes that meet the frequent condition constitute the first-level phoneme frequent set; the second-level phoneme candidate set is generated from the first-level phoneme frequent set, and the specific generation method is to merge the first-level frequent phonemes into two to form a second-level phoneme, and the set of all generated second-level phonemes is the second-level phoneme candidate set; similarly, the second-level phonemes that meet the frequent condition constitute the second-level phoneme frequent set; and so on, the k-1 level phoneme frequent set is used to construct the k-level candidate set, and the k-level phonemes that meet the frequent condition are selected from the k-level candidate set to form the k-level frequent set, until a higher-level candidate set cannot be generated or there are no frequent phones that meet the condition to construct a frequent set, where the k-level phoneme refers to an ordered combination of k first-level phonemes. A method for constructing a k-level candidate set from a k-1-level phoneme frequent set may be to merge two k-1-level phonemes that have k-2 intersections in the k-1-level frequent set.
[0048] In a specific embodiment, the frequent support of each of the 65 generated first-level phonemes is calculated, and the first-level phonemes with a frequent support of no less than R are retained to form a first-level phoneme frequent set. The phonemes in the first-level phoneme frequent set are merged in pairs to form second-level phonemes. The set of all second-level phonemes is then a second-level phoneme candidate set. The support of each second-level phoneme is calculated, and the second-level phonemes with a frequent support of no less than R are retained to form a second-level phoneme frequent set. Similarly, a k-level candidate set is constructed using the k-1-level phoneme frequent set. Specifically, two k-1-level phonemes X and Y with k-2 intersections are taken, where X and Y each contain k-1 first-level phonemes and are internally ordered. If X is completely identical from the 2nd to the k-1th position and Y from the 1st to the k-2th position, all the first-level phonemes of X and the last first-level phoneme of Y are merged to form a higher-level k-level phoneme, which is then placed in the k-level phoneme candidate set. After traversing all combinations that meet the construction conditions and completing the construction of the k-level phoneme candidate set, k-level phonemes that meet the frequent conditions are selected from the k-level candidate set to form a k-level frequent set. This process stops when no higher-level candidate sets can be generated or no frequent phonemes that meet the conditions can be used to construct a frequent set. The frequent sets of all levels constitute the final multi-level frequent set.
[0049] Step 4: Starting from the first-level phoneme, by calculating the speaker discrimination of the phoneme and setting the discrimination requirement, the highly discriminative phonemes at each level are screened until the stopping condition is met to obtain the final multi-level phoneme set.
[0050] Starting from the first-level phoneme, the first-level phoneme frequent set is used as the candidate set; the strong discriminability condition is set as the discriminability support of the phoneme is not less than S, where the specific evaluation method of the discriminability support of the phoneme can be to train a general model to perform speaker recognition on the data of the phoneme category in the training set, and use the recognition rate as the discriminability support of the phoneme; the final first-level phoneme set is composed of the first-level phonemes with discriminability support greater than S; similarly, the k-level phoneme frequent set is used as the candidate set, and the final k-level phoneme set is composed of the k-level phonemes that meet the strong discriminability condition; and so on, stop when there is no candidate set at a higher level, and output the final phoneme set of all levels.
[0051] In a specific embodiment, a general speaker recognition model is first trained using the phoneme data corresponding to the frequent phoneme classes at each level. For the phonemes in the first-level frequent phoneme set, speaker recognition is performed on the validation set, and the recognition rate corresponding to each phoneme is calculated. The first-level frequent phonemes with a recognition rate not lower than S are placed in the final optimal first-level phoneme set. Similarly, the k-level frequent phoneme set is used as a candidate set, and the universal model is used to estimate the discriminability of each k-level frequent phoneme. All k-level phonemes with a speaker recognition rate not lower than S constitute the final optimal k-level phoneme set. The process stops when there is no candidate set at a higher level, and the optimal phoneme sets of all levels are output.
[0052] The method provided by this invention simultaneously considers the universality of phonemes and their ability to distinguish speaker identities, helping to comprehensively evaluate the impact of phoneme units on speaker recognition systems and select optimal phoneme combinations to enhance the effectiveness of speaker recognition systems. Furthermore, by mining valuable phoneme combinations at multiple levels, it fully captures speaker-related information from speech signals, improving the accuracy of speaker recognition.
[0053] Example 2
[0054] like Figure 2 FIG. 1 is a device for generating multi-level phonemes for speaker recognition according to the present invention, comprising:
[0055] A data unit, which acquires and stores voice data and the primary phoneme sequence corresponding to each voice data;
[0056] The frequent candidate set generation unit, based on the determined first-level phonemes, takes the set containing all first-level phonemes as the first-level phoneme candidate set. For second-level and above phonemes, a k-level phoneme candidate set is generated from the k-1-level phoneme frequent set according to the constraint conditions.
[0057] The frequent phoneme screening unit calculates the frequency of occurrence of the k-level phonemes using the phoneme sequence tags in the speech data for the generated k-level phoneme candidate set, and screens the phonemes that meet the frequent conditions from the k-level phonemes according to the set frequent conditions to form a k-level phoneme frequent set;
[0058] The strongly discriminative phoneme screening unit calculates the discriminability of each k-level phoneme based on the obtained k-level phoneme frequent set as the candidate set, and screens out the phonemes that meet the conditions according to the set strong discriminative conditions to form a k-level strongly discriminative phoneme set.
[0059] It should be noted that the various units in this embodiment are logical in nature. In a specific implementation process, one unit can be split into multiple units, and multiple units can also be combined into one unit.
[0060] According to the second embodiment of the present invention, a device for multi-level phoneme generation for speaker recognition is provided. The device can simultaneously consider the universality of phonemes and the distinguishability of speaker identities, which helps to comprehensively evaluate the role of phoneme units in speaker recognition systems, and mine valuable phoneme combinations at multiple levels to fully obtain speaker-related information in speech signals, thereby improving the accuracy of speaker recognition.
[0061] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A method for multi-level phoneme generation for speaker recognition, characterized in that: include: Determine the set of primary phonemes; Obtaining a speech database and a primary phoneme sequence corresponding to each piece of speech data; Starting from the first-level phoneme, by calculating the frequency of occurrence of the phoneme and the preset threshold, the frequent phonemes at each level are screened and a higher-level phoneme candidate set is generated until the stopping condition is met; Starting from the first-level phoneme, by calculating the speaker discrimination of the phoneme and setting the discrimination requirement, the highly discriminative phonemes at each level are screened until the stopping condition is met, and the final multi-level phoneme set is obtained; The method for screening frequent phonemes at each level and generating a phoneme candidate set at a higher level is specifically as follows: starting from the first-level phoneme, a set containing all the first-level phonemes is used as the first-level phoneme candidate set, and the first-level phonemes that meet the frequent condition constitute the first-level phoneme frequent set; the phonemes in the first-level phoneme frequent set are merged in pairs to form a second-level phoneme, a set containing all the second-level phonemes is used as the second-level phoneme candidate set, and the second-level phonemes that meet the frequent condition constitute the second-level phoneme frequent set; and so on, when k is greater than or equal to 2, A k-level phoneme candidate set is constructed from a k-1-level phoneme frequent set, and k-level phonemes that meet the frequent conditions are selected from the k-level phoneme candidate set to form a k-level phoneme frequent set, until a higher-level candidate set cannot be generated or no frequent phonemes that meet the conditions can be used to construct a frequent set. The method for constructing a k-level phoneme candidate set from a k-1-level phoneme frequent set is: two k-1-level phonemes with k-2 intersections in the k-1-level phoneme frequent set are merged; wherein, a k-level phoneme refers to an ordered combination formed by merging k first-level phonemes.
2. The method for multi-level phoneme generation for speaker recognition according to claim 1, characterized in that: The method for determining the primary phoneme is: using the phoneme category defined in linguistics, or using the minimum speech unit defined by an unsupervised learning method as the primary phoneme.
3. The method for multi-level phoneme generation for speaker recognition according to claim 1, characterized in that: The obtaining of the speech database and the primary phoneme sequence corresponding to each piece of speech data includes: obtaining the phoneme sequence by manual annotation, or obtaining the phoneme sequence by using a speech recognition or phoneme recognition model.
4. The method for multi-level phoneme generation for speaker recognition according to claim 3, characterized in that: The obtained phoneme sequence is a phoneme category marked according to the order in which the phonemes appear in the speech signal.
5. The method for multi-level phoneme generation for speaker recognition according to claim 1, characterized in that: The frequent condition for selecting k-level phonemes that meet the frequent condition from the k-level phoneme candidate set to form the k-level phoneme frequent set is: the frequency of the phoneme appearing in the data set is greater than a preset value, or the ratio of the number of sentences in which the phoneme appears to the total number of sentences in the database is greater than a preset value.
6. The method for multi-level phoneme generation for speaker recognition according to claim 1, characterized in that: The method for screening the highly distinguishing phonemes at each level is as follows: Starting from the first-level phoneme, the obtained k-level phoneme frequent set is used as a new candidate set, and the final k-level phoneme set is composed of k-level phonemes that meet the strong distinguishability condition, and so on, until there is no higher-level candidate set.
7. The method for multi-level phoneme generation for speaker recognition according to claim 6, characterized in that: The strong distinguishability condition in the final k-level phoneme set formed by the k-level phonemes satisfying the strong distinguishability condition includes: A universal speaker recognition model is used to perform speaker recognition on data belonging to a phoneme category, so that the recognition accuracy is higher than a preset value.
8. A device for generating multi-level phonemes for speaker recognition, used to implement the method for generating multi-level phonemes for speaker recognition as claimed in any one of claims 1 to 7, characterized in that: include: A data unit, which acquires and stores voice data and the primary phoneme sequence corresponding to each voice data; The frequent candidate set generating unit, based on the determined first-level phonemes, takes a set containing all the first-level phonemes as a first-level phoneme candidate set, and forms a first-level phoneme frequent set from the first-level phonemes that meet the frequent condition; merges the phonemes in the first-level phoneme frequent set into second-level phonemes, takes a set containing all the second-level phonemes as a second-level phoneme candidate set, and forms a second-level phoneme frequent set from the second-level phonemes that meet the frequent condition; and for second-level and above phonemes, generates a k-level phoneme candidate set from the k-1-level phoneme frequent set according to the constraint condition; The frequent phoneme screening unit calculates the frequency of occurrence of the k-level phonemes using the phoneme sequence tags in the speech data for the generated k-level phoneme candidate set, and screens the phonemes that meet the frequent conditions from the k-level phonemes according to the set frequent conditions to form a k-level phoneme frequent set; The strongly discriminative phoneme screening unit calculates the discriminability of each k-level phoneme based on the obtained k-level phoneme frequent set as the candidate set, and screens out the phonemes that meet the conditions according to the set strong discriminative conditions to form a k-level strongly discriminative phoneme set.
Citation Information
Patent Citations
Phoneme selection method and device for voiceprint recognition
CN115966210A