A method and device for screening phonemes, an electronic device and a readable storage medium

By performing phoneme stability analysis and screening on the speech stream, target phonemes within the stable phoneme range are identified for voiceprint modeling, which solves the problem of low accuracy in voiceprint recognition and matching in complex scenarios and achieves higher recognition and matching accuracy.

CN119851694BActive Publication Date: 2025-12-05BEIJING YUANJIAN INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510010962.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-03
Publication Date
2025-12-05
Estimated Expiration
2045-01-03

AI Technical Summary

Technical Problem

Existing technologies have poor accuracy in voiceprint recognition and matching in complex scenarios and fail to effectively take into account the changes in a speaker's voice in different scenarios.

Method used

By performing phoneme stability analysis on the speech stream, several relatively stable target phonemes located within the phoneme stability interval are selected for speakerprint modeling analysis, including phoneme classification, acoustic information extraction, dimensionality reduction and clustering, to determine the phoneme stability interval, and speakerprint comparison is performed based on this.

Benefits of technology

It improves the accuracy of voiceprint recognition and matching, and increases the accuracy of recognition and matching in complex scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119851694B_ABST
    Figure CN119851694B_ABST
Patent Text Reader

Abstract

The application provides a phoneme screening method and device, electronic equipment and readable storage medium. A to-be-detected speech stream is acquired, and the phoneme category of each phoneme contained in each syllable in the to-be-detected speech stream is determined according to phoneme classification information corresponding to a target language to which the to-be-detected speech stream belongs. Acoustic information of each phoneme is extracted, and the acoustic information of each phoneme is subjected to dimension reduction processing and clustering processing to determine a plurality of clustering clusters corresponding to each phoneme category. For each phoneme category, a phoneme stable interval corresponding to the phoneme category is determined based on the plurality of clustering clusters corresponding to the phoneme category. The to-be-detected speech stream is screened based on the phoneme stable interval corresponding to each phoneme category, and a plurality of target phonemes located in the phoneme stable interval in the to-be-detected speech stream are determined. The to-be-detected speech stream is subjected to voiceprint comparison and analysis based on the plurality of determined target phonemes. In this way, the accuracy of voiceprint recognition and matching can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of speech processing technology, and in particular to a method, apparatus, electronic device, and readable storage medium for selecting phonemes. Background Technology

[0002] Before detecting speech, it is generally necessary to acquire the speech stream to be detected and perform voiceprint modeling or voiceprint detection on the speech stream.

[0003] In related technologies, voiceprint recognition methods or voiceprint identification techniques that directly extract acoustic signals for modeling focus primarily on the acoustic characteristics of the extracted speech. These methods typically perform voiceprint modeling and recognition at the speech flow and syllable level. However, the voiceprint recognition process does not consider the speaker's speech variations in different scenarios. Consequently, the accuracy of voiceprint recognition and matching is relatively poor in complex scenarios. Summary of the Invention

[0004] In view of this, the purpose of this application is to provide a method, apparatus, electronic device and readable storage medium for phoneme screening. Before detecting a speech stream, the method performs phoneme stability analysis on the phonemes in the speech stream to determine the phoneme stability range, so as to screen out multiple relatively stable target phonemes located within the phoneme stability range for subsequent voiceprint modeling analysis, thereby improving the accuracy of voiceprint recognition and matching.

[0005] In a first aspect, embodiments of this application provide a method for screening phonemes, the screening method comprising:

[0006] Acquire the speech stream to be detected and segment the speech stream to be detected into multiple syllables;

[0007] Based on the phoneme classification information corresponding to the target language to which the speech stream to be detected belongs, determine the multiple phonemes contained in each syllable and the phoneme category of each phoneme;

[0008] The acoustic information of each phoneme is extracted, and the acoustic information of each phoneme is subjected to dimensionality reduction and clustering to determine multiple clusters corresponding to each phoneme category.

[0009] For each phoneme category, based on the multiple clusters corresponding to that phoneme category, the stable range of the phoneme is determined.

[0010] Based on the stable phoneme intervals corresponding to each phoneme category, the speech stream to be detected is filtered to identify multiple target phonemes located in the stable phoneme intervals of the speech stream to be detected, and voiceprint comparison analysis is performed on the speech stream to be detected based on the identified multiple target phonemes.

[0011] In one possible implementation, the phoneme category of the phoneme is determined by the following steps:

[0012] Based on the phoneme classification information, determine the phoneme classification category to which each phoneme in each syllable belongs;

[0013] For each phoneme category, based on the speech environment information of the speech stream to be detected, the environment category to which each phoneme belongs is determined; wherein, the environment category to which the phoneme belongs is a sub-category of the phoneme category to which it belongs.

[0014] In one possible implementation, the phonemes include vowel types and consonant types; the extraction of acoustic information for each phoneme includes:

[0015] When the phoneme is a vowel, extract the formant information of the phoneme;

[0016] When the phoneme is a consonant, the phoneme is analyzed to extract its spectral characteristics, energy level information, duration information, spectral envelope information, spectral tilt information, bandwidth, and glottal wave characteristics.

[0017] In one possible implementation, for each phoneme, the acoustic information of that phoneme is reduced in dimensionality through the following steps:

[0018] For the multiple acoustic features included in the acoustic information of the phoneme, each acoustic feature is centered, and the centered covariance matrix is ​​calculated.

[0019] Based on the eigenvalues ​​and eigenvectors of the covariance matrix, multiple target eigenvectors are determined;

[0020] Each acoustic feature is projected onto its corresponding target feature vector to obtain the acoustic information of the phoneme after dimensionality reduction.

[0021] In one possible implementation, for each phoneme category, the acoustic information of that phoneme is clustered through the following steps:

[0022] For each phoneme category, multiple cluster centers for that phoneme category are determined based on the mean of that phoneme category;

[0023] For each phoneme category, based on the distance between each phoneme in that phoneme category and the determined cluster centers, the phonemes in that phoneme category are clustered to obtain multiple clusters.

[0024] In one possible implementation, determining the phoneme stability interval corresponding to each phoneme category based on multiple clusters corresponding to that phoneme category includes:

[0025] For each phoneme category, the clusters containing more than a preset phoneme number threshold are identified as the classification stability intervals for that phoneme category.

[0026] For each phoneme category, dimensionality reduction clustering analysis is performed on the environmental classification corresponding to each phoneme in that phoneme category to determine the stable interval of the environmental phonemes corresponding to that phoneme category.

[0027] In one possible implementation, the screening method further includes:

[0028] For each phoneme category, based on the acoustic information of each phoneme included in the phoneme category after dimensionality reduction and clustering, a phoneme profile corresponding to the phoneme category is determined; wherein, the phoneme profile includes the phoneme stability range and phoneme fluctuation amplitude of the phoneme category.

[0029] Voiceprint comparison and analysis are performed based on the phoneme profile corresponding to each phoneme category.

[0030] Secondly, embodiments of this application also provide a phoneme screening device, the screening device comprising:

[0031] The speech stream segmentation module is used to acquire the speech stream to be detected and segment the speech stream to be detected into multiple syllables;

[0032] The phoneme classification and extraction module is used to determine the multiple phonemes contained in each syllable and the phoneme category of each phoneme according to the phoneme classification information corresponding to the target language to which the speech stream to be detected belongs.

[0033] The phoneme dimensionality reduction and clustering module is used to extract the acoustic information of each phoneme, and to perform dimensionality reduction and clustering on the acoustic information of each phoneme to determine multiple clusters corresponding to each phoneme category.

[0034] The phoneme stability interval determination module is used to determine the phoneme stability interval corresponding to each phoneme category based on multiple clusters under the phoneme category.

[0035] The phoneme filtering module is used to filter the speech stream to be detected based on the stable interval of each phoneme category, and determine multiple target phonemes in the speech stream to be detected that are located in the stable interval of the phoneme, so as to perform voiceprint comparison analysis on the speech stream to be detected based on the determined multiple target phonemes.

[0036] Thirdly, embodiments of this disclosure also provide an electronic device, including: a processor, a storage medium, and a bus, wherein the storage medium stores machine-readable instructions executable by the processor, and when the electronic device is running, the processor communicates with the storage medium via the bus, and the processor executes the machine-readable instructions to perform the virtual scene editing method as described in any of the first aspects.

[0037] Fourthly, embodiments of this disclosure also provide a computer-readable storage medium storing a computer program, which, when executed by a processor, performs the virtual scene editing method as described in any of the first aspects.

[0038] This application provides a method, apparatus, electronic device, and readable storage medium for phoneme screening. The method involves acquiring a speech stream to be detected and segmenting it into multiple syllables. Based on the phoneme classification information corresponding to the target language of the speech stream, the method determines the multiple phonemes contained in each syllable and the phoneme category of each phoneme. It extracts the acoustic information of each phoneme and performs dimensionality reduction and clustering processing on the acoustic information of each phoneme to determine multiple clusters corresponding to each phoneme category. For each phoneme category, based on the multiple clusters corresponding to the phoneme category, it determines the phoneme stability interval corresponding to that phoneme category. Based on the phoneme stability interval corresponding to each phoneme category, the method screens the speech stream to be detected, identifying multiple target phonemes located within the phoneme stability interval in the speech stream, and then performing voiceprint comparison analysis on the speech stream based on the identified multiple target phonemes. In this way, before detecting the speech stream, phoneme stability analysis is performed on the phonemes in the speech stream to determine the phoneme stability range. This allows for the selection of several relatively stable target phonemes located within the phoneme stability range for subsequent voiceprint modeling analysis, thereby improving the accuracy of voiceprint recognition and matching.

[0039] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description

[0040] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0041] Figure 1 A flowchart illustrating a phoneme selection method provided in an embodiment of this application;

[0042] Figure 2 This is one of the structural schematic diagrams of a phoneme screening device provided in an embodiment of this application;

[0043] Figure 3 A second schematic diagram of a phoneme screening device provided in an embodiment of this application;

[0044] Figure 4 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0045] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. The components of the embodiments of this application described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely represents selected embodiments of this application. Based on the embodiments of this application, every other embodiment obtained by those skilled in the art without inventive effort falls within the scope of protection of this application.

[0046] First, the applicable scenarios for this application will be introduced. This application can be applied to the field of speech processing technology.

[0047] Before detecting speech, it is generally necessary to acquire the speech stream to be detected and perform voiceprint modeling or voiceprint detection on the speech stream.

[0048] Common voiceprint recognition methods or voiceprint identification techniques that directly extract acoustic signals for modeling focus on the acoustic characteristics of the extracted speech. The extracted speech is mostly in units of speech flow and syllables, lacking preprocessing and analysis of speech data at the phoneme level.

[0049] In isolated word speech recognition, the simplest and most effective method is the DTW (Dynamic Time Warping) algorithm. This algorithm, based on dynamic programming (DP), solves the template matching problem for speech sounds of varying lengths and is one of the earliest and most classic algorithms in speech recognition. However, this method does not consider the inherent variability and instability of the speaker's pronunciation; it directly inputs speech data for comparison, effectively ignoring the speaker's speech variations in different scenarios. Therefore, its accuracy in voiceprint recognition and matching is relatively poor in complex scenarios.

[0050] Based on this, embodiments of this application provide a phoneme screening method to improve the accuracy of voiceprint recognition and matching.

[0051] Please see Figure 1 , Figure 1 This is a flowchart illustrating a phoneme selection method provided in an embodiment of this application. Figure 1 As shown in the embodiments of this application, the phoneme screening method includes:

[0052] S101. Obtain the speech stream to be detected and divide the speech stream to be detected into multiple syllables.

[0053] S102. Based on the phoneme classification information corresponding to the target language to which the speech stream to be detected belongs, determine the multiple phonemes contained in each syllable and the phoneme category of each phoneme.

[0054] S103. Extract the acoustic information of each phoneme, and perform dimensionality reduction and clustering processing on the acoustic information of each phoneme to determine multiple clusters corresponding to each phoneme category.

[0055] S104. For each phoneme category, based on the multiple clusters corresponding to the phoneme category, determine the stable range of the phoneme corresponding to that phoneme category.

[0056] S105. Based on the stable interval of each phoneme category, the speech stream to be detected is filtered to determine each target phoneme in the speech stream to be detected that is located in the stable interval of the phoneme, and the speech stream to be detected is subjected to voiceprint comparison analysis based on the determined multiple target phonemes.

[0057] The phoneme screening method provided in this application performs phoneme stability analysis on the phonemes in the speech stream before detecting the speech stream, determines the phoneme stability range, and screens out multiple relatively stable target phonemes located within the phoneme stability range for subsequent voiceprint modeling analysis, so as to improve the accuracy of voiceprint recognition and matching.

[0058] The exemplary steps of the embodiments of this disclosure are described below:

[0059] S101. Obtain the speech stream to be detected and divide the speech stream to be detected into multiple syllables.

[0060] In this embodiment, before detecting speech, it is generally necessary to acquire the speech stream to be detected and perform voiceprint modeling or voiceprint detection on the speech stream. In related technologies, voiceprint recognition methods or voiceprint identification techniques that directly extract acoustic signals for modeling focus primarily on the acoustic characteristics of the extracted speech, typically performing voiceprint modeling and recognition on a per-speech or per-syllable basis. However, these related technologies do not consider the speaker's speech variations in different scenarios during the voiceprint recognition process. In complex scenarios, the accuracy of voiceprint recognition and matching is relatively poor.

[0061] In this embodiment of the application, after obtaining the speech stream to be detected, it is necessary to segment the speech stream to be detected to obtain multiple discrete syllables. For syllables, the syllables corresponding to different languages ​​are different. For example, if the language is Chinese, one syllable is one Chinese character; if the language is English, one syllable is one word.

[0062] Furthermore, after the speech stream to be detected is segmented into multiple syllables, the phonemes in the syllables can be scanned and analyzed to determine the multiple phonemes contained in each syllable and the phoneme category of each phoneme.

[0063] S102. Based on the phoneme classification information corresponding to the target language to which the speech stream to be detected belongs, determine the multiple phonemes contained in each syllable and the phoneme category of each phoneme.

[0064] Among them, a phoneme is a basic concept in phonetics. It refers to the smallest unit of speech based on the natural properties of speech. It can be divided into two categories: vowels and consonants. Vowels and consonants can be further combined to form syllables.

[0065] In one possible implementation, for each syllable, it is necessary to determine the multiple phonemes contained in each syllable and the phoneme category corresponding to each phoneme. Furthermore, for the entire speech stream to be detected, the set of phonemes corresponding to the same phoneme category is determined to include the phonemes contained in each phoneme category.

[0066] For example, if the audio stream to be detected is "hahahaha, ah? Auntie", the syllable "ha" contains the phoneme vowel a and the consonant h; "ah" contains the phoneme vowel a; "ah" contains the phoneme vowel a. When counting the vowel a, the vowel a in "ha", the vowel a in "ah" and the vowel a in "ah" are included.

[0067] Specifically, the phoneme category of the phoneme is determined through the following steps:

[0068] a1: Based on the phoneme classification information, determine the phoneme classification category to which each phoneme in each syllable belongs.

[0069] a2: For each phoneme category, based on the speech environment information of the speech stream to be detected, determine the environment category to which each phoneme in the phoneme category belongs; wherein, the environment category to which the phoneme belongs is a sub-category of the phoneme category to which it belongs.

[0070] In one possible implementation, the phoneme classification category represents the category to which the phoneme belongs. For example, the phoneme classification category may include vowel a, consonant u, etc., and then the phoneme classification category to which each phoneme belongs is determined according to the corresponding phoneme category.

[0071] Furthermore, after classifying the speech stream to be detected, phonemes can be extracted by category. Even phonemes of the same category may exhibit differences in different speech environments. Therefore, the extracted phonemes need to be further subdivided within their phoneme categories according to speech environment information to determine the environment category to which each phoneme belongs. The environment category to which a phoneme belongs is a sub-category of its original phoneme category.

[0072] For example, if the vowel 'a' can appear after different categories of consonants, and different consonants can all affect the vowel 'a' through co-sound change effects, then within the major category of vowel 'a', 'a' should be further subdivided into secondary subcategories according to consonant category.

[0073] Furthermore, after determining the phoneme categories for the speech stream to be detected, acoustic information can be extracted for each phoneme. In order to reduce the excessive computational load on subsequent phoneme analysis due to excessive acoustic information of phonemes and reduce the efficiency of phoneme processing, dimensionality reduction processing can be performed on the acoustic information of each phoneme after extracting the acoustic features of each phoneme.

[0074] S103. Extract the acoustic information of each phoneme, and perform dimensionality reduction and clustering processing on the acoustic information of each phoneme to determine multiple clusters corresponding to each phoneme category.

[0075] In the embodiments of this application, phonemes are divided into vowel types and consonant types. When extracting acoustic information from different types of phonemes, the acoustic information to be extracted also differs to some extent.

[0076] Specifically, the step "extracting acoustic information for each phoneme" includes:

[0077] b1: When the phoneme is a vowel, extract the formant information of the phoneme.

[0078] b2: When the phoneme is a consonant, the phoneme is analyzed to extract its spectral characteristics, energy level information, duration information, spectral envelope information, spectral tilt information, bandwidth, and glottal wave characteristics.

[0079] In one possible implementation, if the phoneme is a vowel type, the acoustic information of the phoneme generally mainly includes formant information, where the formant is a specific frequency that is amplified when the sound wave passes through the vocal tract and is the main frequency component in the speech signal.

[0080] Specifically, algorithms for extracting formants can include spectral peak search, cepstral method, linear predictive coding (LPC), autocorrelation and cross-correlation based methods, broadband spectrum analysis, multi-window spectrum estimation, and maximum likelihood estimation. These algorithms estimate the frequency and bandwidth of formants and other useful signals by analyzing the spectrum, cepstral, autocorrelation or cross-correlation of speech signals, and by using linear prediction models and probabilistic models.

[0081] In another possible implementation, if the phoneme is a consonant type, the acoustic information of the phoneme generally includes spectral characteristics, energy level, duration, spectral envelope, spectral tilt, bandwidth, and glottal wave characteristics.

[0082] Specifically, algorithms for extracting acoustic information from consonant-type phonemes can include Short-Time Fourier Transform (STFT) to obtain spectral features, Linear Predictive Coding (LPC) and cepstral analysis to estimate the spectral envelope, spectral centroid and spectral entropy to describe spectral characteristics and complexity, Mel-frequency cepstral coefficients (MFCC) and perceptual linear prediction (PLP) to combine spectral analysis and auditory perception characteristics, and deep learning methods such as Convolutional Neural Networks (CNN) and Recurrent Neural Networks (RNN) to extract features directly from the original speech waveform.

[0083] In one possible implementation, after extracting acoustic information for each phoneme, the acoustic information of the phoneme can be dimensionality reduced to reduce the excessive computational load on subsequent phoneme analysis caused by too much acoustic information of the phoneme, thereby improving the efficiency of phoneme processing.

[0084] Specifically, for each phoneme, the acoustic information of that phoneme is reduced in dimensionality through the following steps:

[0085] c1: For the multiple acoustic features included in the acoustic information of the phoneme, each acoustic feature is centered, and the centered covariance matrix is ​​calculated.

[0086] c2: Based on the eigenvalues ​​and eigenvectors of the covariance matrix, determine multiple target eigenvectors.

[0087] c3: Project each acoustic feature according to the corresponding target feature vector to obtain the acoustic information of the phoneme after dimensionality reduction.

[0088] In this embodiment of the application, the acoustic features of the extracted phonemes can be reduced in dimensionality using the statistical method of principal component analysis (PCA). Specifically, the multiple acoustic features contained in the obtained acoustic information of the phonemes are centered. The specific processing method is as follows: for each acoustic feature, the feature mean of the acoustic feature is determined. Further, for each acoustic feature, the feature value of the acoustic feature is subtracted from the corresponding feature mean to complete the centering process.

[0089] Furthermore, the covariance matrix of multiple acoustic features after centering is calculated. The covariance matrix can be used to represent the probability density of multidimensional random variables, and thus the correlation between different acoustic features can be captured through the covariance matrix.

[0090] Furthermore, the eigenvalues ​​and eigenvectors of the covariance matrix are solved; the eigenvectors are sorted according to the eigenvalues ​​in descending order, and the eigenvectors corresponding to the first few largest eigenvalues ​​are selected as the target eigenvectors. The acoustic features are then projected onto the determined target eigenvectors to achieve dimensionality reduction.

[0091] In this way, by reducing the dimensionality of the data, the various acoustic information in the phonemes can be reduced to two dimensionless dimensions, which reduces the number of variables in the dataset while preserving as much important information as possible in the original acoustic information.

[0092] In one possible implementation, linear discriminant analysis, nonnegative matrix factorization (NMF), and t-SNE (t-Distributed Stochastic Neighbor Embedding) can also be used to reduce the dimensionality of the extracted phoneme acoustic features.

[0093] Furthermore, after dimensionality reduction of the acoustic information of phonemes, it is necessary to cluster the multiple phonemes included in each phoneme category, and then determine the stable range of phonemes in each phoneme category based on the multiple clusters obtained from the clustering.

[0094] Specifically, for each phoneme category, the acoustic information of that phoneme is clustered through the following steps:

[0095] d1: For each phoneme category, determine multiple cluster centers for that phoneme category based on the mean of that phoneme category.

[0096] d2: For each phoneme category, based on the distance between each phoneme in that phoneme category and the determined cluster centers, the phonemes in that phoneme category are clustered to obtain multiple clusters.

[0097] In this embodiment of the application, when clustering phonemes under each phoneme category, K-means clustering can be used. K-means clustering iteratively assigns data points to K clusters and continuously updates the centers of these clusters to achieve the purpose of grouping similar data points into one category and separating data points of different categories.

[0098] In one possible implementation, algorithms such as hierarchical clustering, DBSCAN, and spectral clustering can also be used to cluster multiple phonemes contained in each phoneme category.

[0099] In one possible implementation, for each phoneme category, after clustering, the individual phonemes under that phoneme category are clustered into different clusters, and then the stable range of each phoneme category can be determined by analyzing each cluster.

[0100] S104. For each phoneme category, based on the multiple clusters corresponding to the phoneme category, determine the stable range of the phoneme corresponding to that phoneme category.

[0101] In the embodiments of this application, after performing clustering processing on each phoneme category, it is necessary to determine multiple clusters under each phoneme category, and determine the clusters containing a large number of phonemes as the phoneme stable intervals.

[0102] Specifically, the step "for each phoneme category, based on the multiple clusters corresponding to that phoneme category, determine the phoneme stability interval corresponding to that phoneme category" includes:

[0103] e1: For each phoneme category, the clusters containing more than a preset phoneme number threshold are identified as the classification stability intervals for that phoneme category.

[0104] e2: For each phoneme category, perform dimensionality reduction clustering analysis on the environmental classification corresponding to each phoneme in that phoneme category to determine the stable interval of the environmental phonemes corresponding to that phoneme category.

[0105] In this embodiment of the application, the preset phoneme quantity threshold can be set according to the clustering results and phoneme filtering requirements, and no specific limitation is made here.

[0106] In one possible implementation, if a cluster contains a number of phonemes that is greater than a preset phoneme number threshold, the phonemes in that cluster are identified as relatively stable phonemes.

[0107] For example, for the phoneme category of vowel 'a', which includes several phonemes such as vowel a1, vowel a2, and vowel a3, when clustering the phoneme category of vowel 'a', cluster 1 and cluster 2 are generated. Cluster 1 contains more phonemes than a preset phoneme number threshold, while cluster 2 contains less phonemes than the preset phoneme number threshold. That is, cluster 1 is a cluster with a larger number of phonemes, and cluster 2 is a cluster with a larger number of phonemes. Therefore, vowels a1 and a3 contained in cluster 1 are stable phonemes, while vowel a2 contained in cluster 2 is an unstable phoneme.

[0108] In another possible implementation, for each phoneme category, the occurrence environment is divided into different environment categories, and stable clusters are captured from each environment category for analysis, thereby performing a second dimensionality reduction on the acoustic information of the phoneme.

[0109] For example, the dimensionality reduction method is a Gaussian Mixture Model (GMM). Each Gaussian component has its mean, covariance matrix, and mixing coefficients. The mean represents the location of the data centers, the covariance matrix describes the shape and diffusion of the data, and the mixing coefficients represent the weight of each component in the overall distribution. GMM optimizes the model parameters through an iterative expectation-maximization (EM) algorithm, including initializing the parameters, calculating the posterior probability in the E-step, and updating the parameters in the M-step, until the parameters converge.

[0110] Gaussian mixture models can be used to determine the stable range of specific phonemes and how this stability is affected by different speech environments. The speech stream to be tested can then be filtered based on the stable range of the phonemes, and analysis can be performed using relatively stable phonemes.

[0111] S105. Based on the stable interval of each phoneme category, the speech stream to be detected is filtered to determine multiple target phonemes located in the stable interval of the speech stream to be detected, and the speech stream to be detected is analyzed by voiceprint comparison based on the determined multiple target phonemes.

[0112] In this embodiment, for each phoneme category, the frames corresponding to the target phoneme located in the stable phoneme interval of the speech stream to be detected are retained, while frames of unstable phonemes that are not located in the stable phoneme interval are removed, resulting in a speech stream to be detected containing multiple stable phonemes for voiceprint comparison analysis, thereby improving the accuracy of voiceprint recognition and voiceprint identification.

[0113] For the example above, vowels a1 and a3 are determined to be stable phonemes in the speech stream to be detected, while vowel a2 is an unstable phoneme. Therefore, the frame containing vowel a2 needs to be removed from the speech stream to be detected, and vowels a1 and a3 in the speech stream to be detected are retained for subsequent speaker modeling analysis and speaker identification.

[0114] In another possible implementation, by analyzing the stable range of phonemes, a phoneme profile corresponding to each phoneme category can be generated, and then voiceprint analysis and identification can be performed based on the phoneme profile of each phoneme category.

[0115] Specifically, the screening method further includes:

[0116] f1: For each phoneme category, based on the acoustic information of each phoneme included in the phoneme category after dimensionality reduction and clustering, determine the phoneme profile corresponding to the phoneme category; wherein, the phoneme profile includes the phoneme stability range and phoneme fluctuation amplitude of the phoneme category.

[0117] f2: Perform voiceprint comparison analysis based on the phoneme profile corresponding to each phoneme category.

[0118] In the embodiments of this application, for each phoneme category, the stable range and fluctuation range of the phonemes can be determined by analyzing the phonemes. Here, the fluctuation range of the phonemes can be the range of distance fluctuation between each phoneme in each stable cluster and the cluster center of the cluster.

[0119] In one possible implementation, the phoneme profiles corresponding to different speakers vary greatly. Therefore, by constructing a set of speaker phoneme profiles (containing a phoneme profile for each phoneme), preliminary voiceprint identification can be performed to determine whether the voice to be detected is the voice of the corresponding speaker, which can improve the accuracy of voiceprint identification.

[0120] This application provides a method for phoneme screening, which involves acquiring a speech stream to be detected and segmenting it into multiple syllables; determining the multiple phonemes contained in each syllable and the phoneme category of each phoneme according to the phoneme classification information corresponding to the target language to which the speech stream belongs; extracting the acoustic information of each phoneme and performing dimensionality reduction and clustering processing on the acoustic information of each phoneme to determine multiple clusters corresponding to each phoneme category; for each phoneme category, determining the phoneme stability interval corresponding to the phoneme category based on the multiple clusters corresponding to the phoneme category; screening the speech stream to be detected based on the phoneme stability interval corresponding to each phoneme category to determine multiple target phonemes in the speech stream to be detected that are located in the phoneme stability interval, and performing voiceprint comparison analysis on the speech stream to be detected based on the determined multiple target phonemes. In this way, before detecting the speech stream, phoneme stability analysis is performed on the phonemes in the speech stream to determine the phoneme stability range. This allows for the selection of several relatively stable target phonemes located within the phoneme stability range for subsequent voiceprint modeling analysis, thereby improving the accuracy of voiceprint recognition and matching.

[0121] Based on the same inventive concept, this application also provides a phoneme screening device corresponding to the phoneme screening method. Since the principle of the device in this application is similar to the phoneme screening method described above in this application, the implementation of the device can refer to the implementation of the method, and the repeated parts will not be described again.

[0122] Please see Figure 2 , Figure 3 , Figure 2 This is one of the structural schematic diagrams of a phoneme screening device provided in an embodiment of this application. Figure 3 This is a second schematic diagram of a phoneme screening device provided in an embodiment of this application. Figure 2 As shown, the screening device 200 includes:

[0123] The speech stream segmentation module 210 is used to acquire the speech stream to be detected and segment the speech stream to be detected into multiple syllables;

[0124] The phoneme classification and extraction module 220 is used to determine the multiple phonemes contained in each syllable and the phoneme category of each phoneme according to the phoneme classification information corresponding to the target language to which the speech stream to be detected belongs.

[0125] The phoneme dimensionality reduction and clustering module 230 is used to extract the acoustic information of each phoneme, and to perform dimensionality reduction and clustering processing on the acoustic information of each phoneme to determine multiple clusters corresponding to each phoneme category.

[0126] The phoneme stability interval determination module 240 is used to determine the phoneme stability interval corresponding to each phoneme category based on multiple clusters corresponding to the phoneme category.

[0127] The phoneme filtering module 250 is used to filter the speech stream to be detected based on the stable interval of each phoneme category, determine multiple target phonemes in the speech stream to be detected that are located in the stable interval, and perform voiceprint comparison analysis on the speech stream to be detected based on the determined multiple target phonemes.

[0128] In one possible implementation, the phoneme classification extraction module 220 is used to determine the phoneme category through the following steps:

[0129] Based on the phoneme classification information, determine the phoneme classification category to which each phoneme in each syllable belongs;

[0130] For each phoneme category, based on the speech environment information of the speech stream to be detected, the environment category to which each phoneme belongs is determined; wherein, the environment category to which the phoneme belongs is a sub-category of the phoneme category to which it belongs.

[0131] In one possible implementation, the phonemes include vowel types and consonant types; when the phoneme dimensionality reduction clustering module 230 is used to extract the acoustic information of each phoneme, the phoneme dimensionality reduction clustering module 230 is used to:

[0132] When the phoneme is a vowel, extract the formant information of the phoneme;

[0133] When the phoneme is a consonant, the phoneme is analyzed to extract its spectral characteristics, energy level information, duration information, spectral envelope information, spectral tilt information, bandwidth, and glottal wave characteristics.

[0134] In one possible implementation, for each phoneme, the phoneme dimensionality reduction clustering module 230 is used to perform dimensionality reduction processing on the acoustic information of the phoneme through the following steps:

[0135] For the multiple acoustic features included in the acoustic information of the phoneme, each acoustic feature is centered, and the centered covariance matrix is ​​calculated.

[0136] Based on the eigenvalues ​​and eigenvectors of the covariance matrix, multiple target eigenvectors are determined;

[0137] Each acoustic feature is projected onto its corresponding target feature vector to obtain the acoustic information of the phoneme after dimensionality reduction.

[0138] In one possible implementation, for each phoneme category, the phoneme dimensionality reduction clustering module 230 is used to cluster the acoustic information of the phoneme through the following steps:

[0139] For each phoneme category, multiple cluster centers for that phoneme category are determined based on the mean of that phoneme category;

[0140] For each phoneme category, based on the distance between each phoneme in that phoneme category and the determined cluster centers, the phonemes in that phoneme category are clustered to obtain multiple clusters.

[0141] In one possible implementation, when the phoneme stability interval determination module 240 determines the phoneme stability interval corresponding to each phoneme category based on multiple clusters corresponding to that phoneme category, the phoneme stability interval determination module 240 is used to:

[0142] For each phoneme category, the clusters containing more than a preset phoneme number threshold are identified as the classification stability intervals for that phoneme category.

[0143] For each phoneme category, dimensionality reduction clustering analysis is performed on the environmental classification corresponding to each phoneme in that phoneme category to determine the stable interval of the environmental phonemes corresponding to that phoneme category.

[0144] In one possible implementation, such as Figure 3 As shown, the screening device 200 further includes a voiceprint analysis module 260, which is used for:

[0145] For each phoneme category, based on the acoustic information of each phoneme included in the phoneme category after dimensionality reduction and clustering, a phoneme profile corresponding to the phoneme category is determined; wherein, the phoneme profile includes the phoneme stability range and phoneme fluctuation amplitude of the phoneme category.

[0146] Voiceprint comparison and analysis are performed based on the phoneme profile corresponding to each phoneme category.

[0147] This application provides a phoneme screening device that acquires a speech stream to be detected and segments it into multiple syllables. Based on the phoneme classification information corresponding to the target language to which the speech stream belongs, it determines the multiple phonemes contained in each syllable and the phoneme category of each phoneme. It extracts the acoustic information of each phoneme and performs dimensionality reduction and clustering processing on the acoustic information of each phoneme to determine multiple clusters corresponding to each phoneme category. For each phoneme category, based on the multiple clusters corresponding to the phoneme category, it determines the phoneme stability interval corresponding to that phoneme category. Based on the phoneme stability interval corresponding to each phoneme category, it filters the speech stream to be detected to determine multiple target phonemes located in the phoneme stability interval of the speech stream, and performs voiceprint comparison analysis on the speech stream based on the determined multiple target phonemes. In this way, before detecting the speech stream, phoneme stability analysis is performed on the phonemes in the speech stream to determine the phoneme stability range. This allows for the selection of several relatively stable target phonemes located within the phoneme stability range for subsequent voiceprint modeling analysis, thereby improving the accuracy of voiceprint recognition and matching.

[0148] Please see Figure 4 , Figure 4 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Figure 4 As shown, the electronic device 400 includes a processor 410, a memory 420, and a bus 430.

[0149] The memory 420 stores machine-readable instructions executable by the processor 410. When the electronic device 400 is running, the processor 410 communicates with the memory 420 via the bus 430. When the machine-readable instructions are executed by the processor 410, they can perform the operations described above. Figure 1 The steps of the phoneme selection method in the illustrated method embodiment can be found in the method embodiment for specific implementation, and will not be repeated here.

[0150] This application also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, can perform the above-described actions. Figure 1 The steps of the phoneme selection method in the illustrated method embodiment can be found in the method embodiment for specific implementation, and will not be repeated here.

[0151] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0152] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. The apparatus embodiments described above are merely illustrative. For example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. Furthermore, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Additionally, the shown or discussed mutual couplings, direct couplings, or communication connections may be through some communication interfaces; indirect couplings or communication connections between devices or units may be electrical, mechanical, or other forms.

[0153] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0154] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0155] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a processor-executable, non-volatile, computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0156] Finally, it should be noted that the above-described embodiments are merely specific implementations of this application, used to illustrate the technical solutions of this application, and not to limit them. The scope of protection of this application is not limited thereto. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can still modify or easily conceive of changes to the technical solutions described in the foregoing embodiments, or make equivalent substitutions for some of the technical features, within the scope of the technology disclosed in this application. Such modifications, changes, or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be covered within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A method for selecting phonemes, characterized in that, The screening method includes: Acquire the speech stream to be detected and segment the speech stream to be detected into multiple syllables; Based on the phoneme classification information corresponding to the target language to which the speech stream to be detected belongs, determine the multiple phonemes contained in each syllable and the phoneme category of each phoneme; The acoustic information of each phoneme is extracted, and the acoustic information of each phoneme is subjected to dimensionality reduction and clustering to determine multiple clusters corresponding to each phoneme category. For each phoneme category, based on the multiple clusters corresponding to that phoneme category, the stable range of the phoneme is determined. Based on the stable phoneme intervals corresponding to each phoneme category, the speech stream to be detected is filtered to identify multiple target phonemes located in the stable phoneme intervals of the speech stream to be detected, and voiceprint comparison analysis is performed on the speech stream to be detected based on the identified multiple target phonemes.

2. The screening method according to claim 1, characterized in that, The phoneme category of the phoneme is determined by the following steps: Based on the phoneme classification information, determine the phoneme classification category to which each phoneme in each syllable belongs; For each phoneme category, based on the speech environment information of the speech stream to be detected, the environment category to which each phoneme belongs is determined; wherein, the environment category to which the phoneme belongs is a sub-category of the phoneme category to which it belongs.

3. The screening method according to claim 1, characterized in that, The phonemes include vowel types and consonant types; the extraction of acoustic information for each phoneme includes: When the phoneme is a vowel, extract the formant information of the phoneme; When the phoneme is a consonant, the phoneme is analyzed to extract its spectral characteristics, energy level information, duration information, spectral envelope information, spectral tilt information, bandwidth, and glottal wave characteristics.

4. The screening method according to claim 1, characterized in that, For each phoneme, the acoustic information of that phoneme is reduced in dimensionality using the following steps: For the multiple acoustic features included in the acoustic information of the phoneme, each acoustic feature is centered, and the centered covariance matrix is ​​calculated. Based on the eigenvalues ​​and eigenvectors of the covariance matrix, multiple target eigenvectors are determined; Each acoustic feature is projected onto its corresponding target feature vector to obtain the acoustic information of the phoneme after dimensionality reduction.

5. The screening method according to claim 1, characterized in that, For each phoneme category, the acoustic information of that phoneme is clustered using the following steps: For each phoneme category, multiple cluster centers for that phoneme category are determined based on the mean of that phoneme category; For each phoneme category, based on the distance between each phoneme in that phoneme category and the determined cluster centers, the phonemes in that phoneme category are clustered to obtain multiple clusters.

6. The screening method according to claim 1, characterized in that, For each phoneme category, based on multiple clusters corresponding to that phoneme category, the method for determining the phoneme stability interval for that phoneme category includes: For each phoneme category, the clusters containing more than a preset phoneme number threshold are identified as the classification stability intervals for that phoneme category. For each phoneme category, dimensionality reduction clustering analysis is performed on the environmental classification corresponding to each phoneme in that phoneme category to determine the stable interval of the environmental phonemes corresponding to that phoneme category.

7. The screening method according to claim 1, characterized in that, The screening method further includes: For each phoneme category, based on the acoustic information of each phoneme included in the phoneme category after dimensionality reduction and clustering, a phoneme profile corresponding to the phoneme category is determined; wherein, the phoneme profile includes the phoneme stability range and phoneme fluctuation amplitude of the phoneme category. Voiceprint comparison and analysis are performed based on the phoneme profile corresponding to each phoneme category.

8. A phoneme screening device, characterized in that, The screening device includes: The speech stream segmentation module is used to acquire the speech stream to be detected and segment the speech stream to be detected into multiple syllables; The phoneme classification and extraction module is used to determine the multiple phonemes contained in each syllable and the phoneme category of each phoneme according to the phoneme classification information corresponding to the target language to which the speech stream to be detected belongs. The phoneme dimensionality reduction and clustering module is used to extract the acoustic information of each phoneme, and to perform dimensionality reduction and clustering on the acoustic information of each phoneme to determine multiple clusters corresponding to each phoneme category. The phoneme stability interval determination module is used to determine the phoneme stability interval corresponding to each phoneme category based on multiple clusters under the phoneme category. The phoneme filtering module is used to filter the speech stream to be detected based on the stable interval of each phoneme category, and determine multiple target phonemes in the speech stream to be detected that are located in the stable interval of the phoneme, so as to perform voiceprint comparison analysis on the speech stream to be detected based on the determined multiple target phonemes.

9. An electronic device, characterized in that, include: The device includes a processor, a memory, and a bus, wherein the memory stores machine-readable instructions executable by the processor, and when the electronic device is in operation, the processor communicates with the memory via the bus, and the processor executes the machine-readable instructions to perform the steps of the phoneme screening method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, performs the steps of the phoneme selection method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Phoneme selection method and device for voiceprint recognition

    CN115966210A

  • Voiceprint recognition method based on phoneme information and electronic equipment

    CN116403587A