A dialect generation method and system based on voiceprint features
By collecting individual dialect recordings, using deep learning models to extract voiceprint feature vectors and training a dialect classifier, and combining individual voiceprints with regional prosodic features, dialect speech that conforms to the target region is generated. This solves the problem of insufficient personalization and prosodic feature fusion in existing technologies, and improves the accuracy and efficiency of dialect generation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SIMAI INTELLIGENT TECHNOLOGY (SHENZHEN) CO LTD
- Filing Date
- 2026-01-30
- Publication Date
- 2026-05-29
Smart Images

Figure CN122116872A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a dialect generation method and system based on voiceprint features, belonging to the field of artificial intelligence. Background Technology
[0002] Dialect generation technology refers to the process of simulating or reproducing specific dialect speech through technical means. It has important application value in fields such as dialect culture protection, intelligent voice interaction, and regional services. In order to meet the demand for personalized dialect speech in different scenarios, it is necessary to generate speech content that not only conforms to dialect characteristics but also fits the individual's speech style. Therefore, the dialect generation process needs to be precisely processed.
[0003] Current dialect generation methods mainly employ manual recording and conversion, as well as general speech synthesis technology. These methods involve having professionals record dialect speech samples, combining them with speech conversion algorithms to process general speech into dialects, or adjusting the speech based on preset dialect acoustic parameters to obtain the target dialect speech. However, these methods do not fully consider individual voiceprint features during processing, resulting in a lack of personalized recognition in the generated dialect speech. Furthermore, the generated results do not accurately integrate the prosodic features of the regional dialect, thus reducing the accuracy of dialect generation. Summary of the Invention
[0004] This invention provides a dialect generation method and system based on voiceprint features, the main purpose of which is to improve the accuracy of dialect generation.
[0005] To achieve the above objectives, the present invention provides a dialect generation method based on voiceprint features, comprising: Collect individual dialect recordings in the target area, and extract individual voiceprint identifiers and recording content information of the target area based on the individual dialect recordings; The voiceprint acoustic feature vector of the recorded content information is extracted using a pre-trained deep learning model to construct the dialect feature vector corresponding to the individual in the target area; Using the dialect feature vector, a dialect classifier corresponding to the target region is trained. The target dialect data to be processed is received, and the voiceprint baseline features corresponding to the target dialect data are analyzed. Based on the voiceprint baseline features, the dialect classifier is used to analyze the dialect type corresponding to the target dialect data. By combining the individual voiceprint identifier with the dialect type, the target dialect data is subjected to voiceprint feature fusion processing to obtain the adapted initial speech, and the regional prosodic features of the target region are analyzed. Based on the regional prosodic features, the initial adapted speech is subjected to prosodic optimization processing to generate the dialect generation result corresponding to the target region.
[0006] Optionally, the step of extracting individual voiceprint identifiers and recording content information of the target area based on the individual dialect recording includes: The individual dialect recordings are analyzed to obtain the original audio stream and vocal tract feature set; The original audio stream is preprocessed to obtain a clean audio stream; Based on preset dialect feature requirements, the pure audio stream is subjected to acoustic unit segmentation to obtain an acoustic sequence; Extract the spectral feature set corresponding to the acoustic sequence, and combine the vocal tract feature set and the spectral feature set to analyze the recording content information and individual voiceprint identifier.
[0007] Optionally, the step of extracting the voiceprint acoustic feature vector of the recorded content information using a pre-trained deep learning model includes: The acoustic standard in the deep learning model is used to uniformly process the recorded content information to obtain standardized acoustic data; The feature condenser in the deep learning model is used to extract essential information from the normalized acoustic data to obtain the core acoustic elements; The discrimination weights of each component in the core acoustic elements are calculated using the discrimination evaluator in the deep learning model. Based on the identification weights, the corresponding voiceprint acoustic feature vectors are generated from the core acoustic elements using the result synthesizer in the deep learning model.
[0008] Optionally, calculating the discrimination weights of each component in the core acoustic element using the discrimination evaluator in the deep learning model includes: The attribute parser in the discrimination evaluator is used to identify the phoneme markers in the core acoustic elements, and the frequency of occurrence of each phoneme marker is counted. Based on the frequency of occurrence, the distribution breadth and specificity measure of each acoustic component are calculated using the distribution metric function in the discrimination evaluator. Based on the combined distribution breadth and specificity measure, the discrimination capability value of each acoustic component is calculated; Based on the discrimination capability value, the discrimination weight corresponding to each acoustic component is determined.
[0009] Optionally, the step of extracting the voiceprint acoustic feature vector of the recorded content information using a pre-trained deep learning model to construct the dialect feature vector corresponding to the individual in the target region includes: The acoustic feature vector of the voiceprint is subjected to distribution equalization processing to obtain balanced acoustic features; Calculate the core feature value corresponding to each feature in the equal acoustic features; Based on the core feature value, the voiceprint acoustic feature vector is smoothed to obtain smoothed acoustic features; The smooth acoustic features are encoded to obtain dialect feature vectors.
[0010] Optionally, calculating the core feature value corresponding to each feature in the equalization acoustic features includes: Frequency band energy analysis is performed on the aforementioned balanced acoustic characteristics to obtain the sub-band energy distribution. Calculate the standard deviation of the stationarity corresponding to the energy distribution of the sub-band; Combining the stationarity standard deviation and the subband distribution energy, the core feature value corresponding to each feature in the equal acoustic features is calculated using the following formula: ; Where C represents the core feature value corresponding to each feature in the equal acoustic features. This represents the energy distribution of the sub-band corresponding to the k-th sub-band. This represents the standard deviation of stationarity corresponding to the k-th subband, where k represents the subband sequence number and K represents the number of subbands.
[0011] Optionally, the analysis of the voiceprint baseline features corresponding to the target dialect data includes: The target dialect data is subjected to data filtering processing to obtain standard dialect data; Extract the spectral characteristics and prosodic information from the standard dialect data; Obtain the formant parameters and fundamental frequency profile corresponding to the standard dialect data; By combining the spectral characteristics and the prosodic information, the acoustic identification features of the target dialect data are analyzed; Based on the formant parameters and the fundamental frequency profile, the pronunciation pattern features of the target dialect data are determined; By combining the acoustic identification features and the pronunciation pattern features, the voiceprint baseline features corresponding to the standard dialect data are determined.
[0012] Optionally, the step of analyzing the dialect type corresponding to the target dialect data using the dialect classifier based on the voiceprint baseline features includes: The voiceprint reference features are normalized to obtain standard voiceprint features; Extract the distinguishing feature vector from the standard voiceprint features; The distinguishing feature vector is input into the dialect classifier, and the similarity calculation unit in the dialect classifier is used to calculate the similarity measure between the distinguishing feature vector and the dialect type feature template; Based on the similarity metric, the classification decision unit in the dialect classifier is used to determine the category membership degree of the target dialect data; Based on the category membership degree, the dialect type corresponding to the target dialect data is determined.
[0013] Optionally, the step of combining the individual voiceprint identifier with the dialect type to perform voiceprint feature fusion processing on the target dialect data to obtain the adapted initial speech includes: The individual's voiceprint identifier is deconstructed to obtain the individual's physiological characteristics; The dialect types are analyzed using rules to obtain dialect phonological rules; By combining the physiological characteristics and the phonological rules, the target dialect data is reshaped to obtain an adapted initial speech.
[0014] To address the aforementioned problems, the present invention also provides a dialect generation system based on voiceprint features, the system comprising: The dialect recording and extraction module is used to collect individual dialect recordings in a target area, and extract individual voiceprint identifiers and recording content information of the target area based on the individual dialect recordings. The voiceprint feature extraction module is used to extract the voiceprint acoustic feature vector of the recorded content information using a pre-trained deep learning model, so as to construct the dialect feature vector corresponding to the individual in the target area. The dialect classification module is used to train a dialect classifier for the target region using the dialect feature vector, receive target dialect data to be processed, analyze the voiceprint baseline features corresponding to the target dialect data, and analyze the dialect type corresponding to the target dialect data based on the voiceprint baseline features using the dialect classifier. The voiceprint fusion module is used to combine the individual voiceprint identifier with the dialect type to perform voiceprint feature fusion processing on the target dialect data to obtain an adapted initial speech and analyze the regional prosodic features of the target region. The dialect generation module is used to perform prosodic optimization processing on the adapted initial speech based on the regional prosodic features, so as to generate the dialect generation result corresponding to the target region.
[0015] Compared to the problems described in the background art, this invention, by processing the individual dialect recordings, can separate speaker features and dialect content from the original speech, providing a foundation for subsequent dialect analysis and speaker recognition, and effectively improving the accuracy of dialect resource utilization. This invention, by extracting dialect acoustic features from the recording content information, can transform the original speech content into a structured feature representation, providing a unified and computable data foundation for subsequent dialect analysis, comparison, and recognition, effectively improving the efficiency of dialect processing. Furthermore, by using the dialect feature vectors to train a dialect classifier, this invention can establish a correspondence between dialect features and dialect types, improving the accuracy of dialect recognition. And by analyzing the voiceprint baseline features corresponding to the target dialect data, it can capture the target... The unique phonetic characteristics of dialects provide core feature support for subsequent dialect type determination. Furthermore, this invention combines the individual voiceprint identifier with the dialect type to perform voiceprint feature fusion processing on the target dialect data. This organically combines the speaker's individual characteristics with the dialect category characteristics, generating synthesized speech that conforms to the specific dialect's phonetic features. By analyzing the regional prosodic features of the target region, common patterns in intonation, rhythm, etc., of the dialect in that region can be obtained, providing an important basis for dialect speech synthesis and recognition. Finally, this invention optimizes the prosodic features of the initial matching speech, ensuring that the generated dialect result matches the phonetic habits of the target region, eliminating deviations between individual pronunciation and regional common prosodic features in the initial matching speech, thereby improving the accuracy of dialect generation. Therefore, the dialect generation method and system based on voiceprint features provided by this invention can improve the accuracy of dialect generation. Attached Figure Description
[0016] Figure 1 This is a flowchart illustrating a dialect generation method based on voiceprint features according to an embodiment of the present invention. Figure 2 This is a schematic diagram of the prosody optimization process in a dialect generation method based on voiceprint features provided in an embodiment of the present invention; Figure 3 This is a schematic diagram of a module for implementing a dialect generation system based on voiceprint features, provided as an embodiment of the present invention.
[0017] The objectives, features, and advantages of this invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0018] It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.
[0019] This application provides a dialect generation method based on voiceprint features. The executing entity of this dialect generation method includes, but is not limited to, at least one of the following electronic devices that can be configured to execute the method provided in this application: a server, a terminal, etc. In other words, the dialect generation method based on voiceprint features can be executed by software or hardware installed on a terminal device or a server device. The server includes, but is not limited to, a single server, a server cluster, a cloud server, or a cloud server cluster.
[0020] Reference Figure 1 The diagram shown is a flowchart illustrating a dialect generation method based on voiceprint features according to an embodiment of the present invention. In this embodiment, the dialect generation method based on voiceprint features includes: S1. Collect individual dialect recordings in the target area, and extract individual voiceprint identifiers and recording content information of the target area based on the individual dialect recordings.
[0021] This invention processes the individual dialect recordings to separate speaker features and dialect content from the original speech, providing a foundation for subsequent dialect analysis and speaker recognition, and effectively improving the accuracy of dialect resource utilization.
[0022] The individual dialect recording refers to the original recording data collected in the target area that contains the dialect speech of a specific speaker, which has not yet undergone systematic processing and feature extraction; the individual voiceprint identifier refers to the set of parameters extracted from the dialect recording to identify the speaker's voice characteristics, reflecting their timbre, pitch, pronunciation habits and other individual information; the recording content information refers to the dialect content carried by the actual speech identified from the dialect recording, including vocabulary, sentences, speech segments and other language information.
[0023] As an embodiment of the present invention, the step of extracting individual voiceprint identifiers and recording content information of the target area based on the individual dialect recording includes: The individual dialect recordings are analyzed to obtain the original audio stream and vocal tract feature set; The original audio stream is preprocessed to obtain a clean audio stream; Based on preset dialect feature requirements, the pure audio stream is subjected to acoustic unit segmentation to obtain an acoustic sequence; Extract the spectral feature set corresponding to the acoustic sequence, and combine the vocal tract feature set and the spectral feature set to analyze the recording content information and individual voiceprint identifier.
[0024] The original audio stream and the vocal tract feature set are the original signal and parsed vocal tract parameters of an individual dialect recording, respectively; the clean audio stream is the audio signal after noise reduction, removal of silence segments, and standardization; the dialect feature requirements are dialect phonological standards used to guide the segmentation of acoustic units, and the acoustic sequence is the result of segmenting the clean audio stream based on these standards: for example, it is required to include "preserving tone contours" and "distinguishing stop consonant codas," etc. After segmentation, the clean audio stream yields an acoustic sequence with phonological features; the spectral feature set is a set of frequency domain parameters extracted from the acoustic sequence, such as Mel-frequency cepstral coefficients (MFCC) and linear predictive coding coefficients (LPC).
[0025] Furthermore, audio signal analysis techniques can be used to analyze the individual dialect recordings to obtain the original audio stream and vocal tract feature set. For example, Librosa or Kaldi toolkits can be used to extract audio waveforms and acoustic parameters such as fundamental frequency and formants. Signal processing algorithms can be used to preprocess the original audio stream to obtain a clean audio stream, such as spectral subtraction for noise reduction, endpoint detection to remove silence segments, and amplitude normalization. Based on dialect feature requirements, the clean audio stream can be segmented into acoustic units using speech analysis methods to obtain an acoustic sequence. For example, Dynamic Time Warping (DTW) or Hidden Markov Models (HMMs) can be used to segment the audio into phonemes, syllables, or vowel units. Spectral analysis algorithms can be used to extract the spectral feature set corresponding to the acoustic sequence, such as calculating MFCC features or LPC coefficients. Combining the vocal tract feature set and the spectral feature set, voiceprint modeling algorithms (such as Gaussian mixture models or deep learning models) can be used to analyze the individual voiceprint identifier. At the same time, dialect recognition techniques (such as feature fusion-based classification models) can be used to parse the recording content information.
[0026] S2. Use a pre-trained deep learning model to extract the voiceprint acoustic feature vector of the recorded content information to construct the dialect feature vector corresponding to the individual in the target area.
[0027] This invention extracts dialect acoustic features from the recorded content information, transforming the original speech content into a structured feature representation. This provides a unified and computable data foundation for subsequent dialect analysis, comparison, and recognition, effectively improving the efficiency of dialect processing. The deep learning model, trained on specific data, is an artificial intelligence model capable of extracting feature vectors representing the acoustic attributes of voiceprints from the recorded content information. These voiceprint acoustic feature vectors are numerical arrays extracted from the recorded content information by the pre-trained deep learning model, used to quantitatively characterize the acoustic properties of sound at the voiceprint level (e.g., timbre, tone, rhythm).
[0028] As an embodiment of the present invention, the step of extracting the voiceprint acoustic feature vector of the recorded content information using a pre-trained deep learning model includes: The acoustic standard in the deep learning model is used to uniformly process the recorded content information to obtain standardized acoustic data; The feature condenser in the deep learning model is used to extract essential information from the normalized acoustic data to obtain the core acoustic elements; The discrimination weights of each component in the core acoustic elements are calculated using the discrimination evaluator in the deep learning model. Based on the identification weights, the corresponding voiceprint acoustic feature vectors are generated from the core acoustic elements using the result synthesizer in the deep learning model.
[0029] The acoustic standard is a component that converts the original speech content into a unified format required by the deep learning model, and consists of amplitude normalization and spectral normalization functions; the feature condenser is a component that compresses and purifies the normalized acoustic data, and is constructed using information entropy feature selection methods, such as using Mel frequency cepstral coefficients for compression representation; the discrimination evaluator is a computational unit used to measure the importance of each acoustic component in voiceprint identification, such as by maximizing inter-class distance; the result synthesizer consists of feature weighting and aggregation functions, such as linear combination based on weights.
[0030] Furthermore, as an optional embodiment of the present invention, the step of calculating the discrimination weights of each component in the core acoustic element using the discrimination evaluator in the deep learning model includes: The attribute parser in the discrimination evaluator is used to identify the phoneme markers in the core acoustic elements, and the frequency of occurrence of each phoneme marker is counted. Based on the frequency of occurrence, the distribution breadth and specificity measure of each acoustic component are calculated using the distribution metric function in the discrimination evaluator. Based on the combined distribution breadth and specificity measure, the discrimination capability value of each acoustic component is calculated; Based on the discrimination capability value, the discrimination weight corresponding to each acoustic component is determined.
[0031] The phoneme markers are categorized according to the phonetic system; the distribution breadth represents the stability of the acoustic component in different speakers or contexts; the specificity measure represents the component's ability to identify a specific voiceprint; and the discrimination value is a quantitative indicator used to ultimately measure the degree of component differentiation.
[0032] Optionally, the attribute parser is implemented by a speech processing tool library, such as LibROSA; the frequency of occurrence of the phoneme marker can be obtained by statistically analyzing the proportion of the number of times a specific phoneme appears in a speech segment to the total number of phonemes; the distribution metric function consists of a standard deviation function and an intra-class and inter-class distance ratio calculation module; the calculation process of the discrimination capability value of the acoustic component is as follows: the distribution breadth is added to the specificity measure and normalized to obtain the discrimination capability value of each component between 0 and 1, and the higher the value, the more important it is in voiceprint identification.
[0033] This invention utilizes a pre-trained deep learning model to extract the voiceprint acoustic feature vector of the recorded content information, thereby constructing the dialect feature vector corresponding to the individual in the target region. This can transform complex continuous speech signals into a compact and structured numerical representation, preserving the most distinctive dialect pronunciation patterns, and providing a basis for dialect comparison and analysis between different individuals or regions. The dialect feature vector is a numerical representation of the local dialect pronunciation characteristics in the recorded content information, covering multi-level acoustic information such as segments, prosody, and phonological structure.
[0034] As an embodiment of the present invention, the step of extracting the voiceprint acoustic feature vector of the recorded content information using a pre-trained deep learning model to construct the dialect feature vector corresponding to the individual in the target region includes: The acoustic feature vector of the voiceprint is subjected to distribution equalization processing to obtain balanced acoustic features; Calculate the core feature value corresponding to each feature in the equal acoustic features; Based on the core feature value, the voiceprint acoustic feature vector is smoothed to obtain smoothed acoustic features; The smooth acoustic features are encoded to obtain dialect feature vectors.
[0035] The balanced acoustic features are a set of features with more balanced differences between features by adjusting the numerical distribution range of each dimension of the voiceprint acoustic feature vector, thus avoiding the masking of other dialect-related features by a feature with an excessively large value in one dimension. The core feature value is a representative value that quantifies the key attributes (such as spectral peak value and pitch variation amplitude) of a single feature in the balanced acoustic features, and is used to reflect the typicality of the feature in dialect expression. The smoothed acoustic features are the result of processing that weakens the random fluctuations in features (such as feature jumps caused by instantaneous noise in the recording) so that the feature changes are more in line with the dialect pronunciation rules.
[0036] Furthermore, the distribution equalization of the voiceprint acoustic feature vector can be achieved through feature scaling: max-min normalization is used to map the feature values of each dimension to the [0,1] interval, eliminating unit differences in different acoustic indicators (such as frequency and duration) to obtain balanced acoustic features. The calculation of the core feature value in the balanced acoustic features can be achieved through statistical analysis: the mean and standard deviation of the numerical sequence of a single feature are calculated, and the median value within the range of mean ± standard deviation is taken as the core feature value (e.g., if the numerical sequence of a certain spectral feature is [2.1, 2.3, 2.2, 5.8, 2.4], after removing the outlier 5.8, the median 2.2 is taken as the core feature value). The core feature value); the feature smoothing of the voiceprint acoustic feature vector can be achieved through a moving average algorithm: using 3-5 consecutive feature values as a window, the average value of the core feature value within the window is calculated, and the feature value at the center of the window is replaced to weaken the influence of instantaneous fluctuations and obtain smooth acoustic features; the encoding of the smooth acoustic features can be achieved through an embedded layer network: the smooth acoustic features are input into a preset encoding network (such as a combination of fully connected layers and dimensionality reduction layers), and a numerical vector with a fixed dimension (such as 128 dimensions) is output, wherein the dialect-specific acoustic features (such as the tone features of the target region dialect) are enhanced through network weights to form the final dialect feature vector.
[0037] Furthermore, as another optional embodiment of the present invention, the calculation of the core feature value corresponding to each feature in the equalized acoustic features includes: Frequency band energy analysis is performed on the aforementioned balanced acoustic characteristics to obtain the sub-band energy distribution. Calculate the standard deviation of the stationarity corresponding to the energy distribution of the sub-band; By combining the standard deviation of the stability and the energy of the sub-band distribution, the core feature value corresponding to each feature in the equal acoustic features is calculated.
[0038] The sub-band distributed energy is the quantitative distribution result of acoustic energy in each sub-band after dividing the equal acoustic features according to frequency ranges, reflecting the energy proportion and distribution characteristics of different frequency components; the standard deviation of stability is a statistical measure describing the fluctuation amplitude of sub-band distributed energy in the time dimension. The smaller the value, the more stable the energy distribution, and vice versa.
[0039] Furthermore, the frequency band energy analysis of the equal acoustic characteristics can be achieved through Mel filter banks or wavelet packet decomposition, dividing the continuous spectrum into several overlapping or non-overlapping sub-bands, and calculating the total energy or energy density in each sub-band; the method for calculating the stationarity standard deviation is as follows: within a preset time window, the standard deviation of the energy sequence of each sub-band is calculated to obtain the stationarity standard deviation.
[0040] Furthermore, as another embodiment of the present invention, the core feature value corresponding to each feature in the equal acoustic features is calculated using the following formula, combining the standard deviation of the stability and the sub-band distribution energy: ; Where C represents the core feature value corresponding to each feature in the equal acoustic features. This represents the energy distribution of the sub-band corresponding to the k-th sub-band. This represents the standard deviation of stationarity corresponding to the k-th subband, where k represents the subband sequence number and K represents the number of subbands.
[0041] S3. Using the dialect feature vector, train a dialect classifier for the target region, receive the target dialect data to be processed, analyze the voiceprint baseline features corresponding to the target dialect data, and based on the voiceprint baseline features, use the dialect classifier to analyze the dialect type corresponding to the target dialect data.
[0042] This invention utilizes dialect feature vectors to train a dialect classifier, establishing a correspondence between dialect features and dialect types, thus improving the accuracy of dialect recognition. Furthermore, by analyzing the voiceprint baseline features corresponding to the target dialect data, it can capture the unique speech characteristics of the target dialect, providing core feature support for subsequent dialect type determination. The dialect classifier is a classification tool trained based on dialect feature vectors, used to distinguish different dialect types. The target dialect data is the original dialect speech data of the dialect type to be identified. The voiceprint baseline features are a set of acoustic features extracted from the target dialect data, reflecting its essential speech attributes, and covering key indicators such as fundamental frequency variation range, vowel formant position, and syllable duration interval.
[0043] As an embodiment of the present invention, the analysis of the voiceprint reference features corresponding to the target dialect data includes: The target dialect data is subjected to data filtering processing to obtain standard dialect data; Extract the spectral characteristics and prosodic information from the standard dialect data; Obtain the formant parameters and fundamental frequency profile corresponding to the standard dialect data; By combining the spectral characteristics and the prosodic information, the acoustic identification features of the target dialect data are analyzed; Based on the formant parameters and the fundamental frequency profile, the pronunciation pattern features of the target dialect data are determined; By combining the acoustic identification features and the pronunciation pattern features, the voiceprint baseline features corresponding to the standard dialect data are determined.
[0044] The standard dialect data refers to a high-quality speech data set obtained by noise removal and invalid segment extraction from the target dialect data, thus eliminating the influence of background interference and non-speech information. The spectral characteristics are the energy distribution features of the standard dialect data in different frequency ranges, including quantitative indicators such as spectral peak positions and high-frequency energy proportions. The prosodic information is a set of features reflecting the rhythm and tone changes of dialect speech, covering syllable duration, tone fluctuation amplitude, and pause intervals. The formant parameters are acoustic indicators characterizing vowel pronunciation features, including formant frequency, bandwidth, and transition characteristics. The fundamental frequency profile is the trajectory of the fundamental frequency in dialect speech over time, reflecting the tone value pattern and change rules. The acoustic identifier features are distinctive features that can distinguish different dialect systems after integrating spectral and prosodic information. The pronunciation pattern features are a set of features reflecting the unique pronunciation habits of a dialect, obtained based on formant and fundamental frequency analysis. The voiceprint baseline features are a set of core features that can represent the essential attributes of the target dialect after integrating acoustic identifiers and pronunciation pattern features.
[0045] Furthermore, the data filtering and processing of the target dialect data can be achieved through the Voice Activity Detection (VAD) algorithm, retaining valid speech segments and removing silent and noisy segments; the extraction of spectral characteristics and prosodic information can be completed using short-time Fourier transform and prosodic feature extraction tools to ensure the temporal resolution and accuracy of the features; the acquisition of formant parameters and fundamental frequency contours can be achieved through Praat speech analysis software, which supports high-precision acoustic parameter calculation; the analysis of acoustic identifier features can be achieved using principal component analysis to extract key features with significant distinguishability for different dialects; the determination of pronunciation pattern features can be combined with dialect phonetics rules to establish a mapping relationship between specific pronunciation patterns and dialect types; the determination of voiceprint baseline features can be achieved through feature splicing and normalization processing to form a feature vector with unified dimensions that can be directly used for classification.
[0046] This invention, by analyzing the dialect type corresponding to the target dialect data using the dialect classifier based on the aforementioned voiceprint reference features, can effectively identify the specific dialect type to which the target dialect data belongs, providing technical support for dialect research, protection, and application.
[0047] The dialect type refers to the dialect category to which the target dialect data belongs, such as Minnan dialect, Yue dialect, Wu dialect, etc.
[0048] As an embodiment of the present invention, the step of analyzing the dialect type corresponding to the target dialect data using the dialect classifier based on the voiceprint reference features includes: The voiceprint reference features are normalized to obtain standard voiceprint features; Extract the distinguishing feature vector from the standard voiceprint features; The distinguishing feature vector is input into the dialect classifier, and the similarity calculation unit in the dialect classifier is used to calculate the similarity measure between the distinguishing feature vector and the dialect type feature template; Based on the similarity metric, the classification decision unit in the dialect classifier is used to determine the category membership degree of the target dialect data; Based on the category membership degree, the dialect type corresponding to the target dialect data is determined.
[0049] The standard voiceprint features are a standardized feature set obtained by filtering noise and unifying the scale of the voiceprint baseline features, which can eliminate interference caused by differences in speech rate and volume among different speakers; the distinguishing feature vector is a combination of core features selected from the standard voiceprint features that can reflect the essential differences between different dialect types, such as the frequency variation range of specific tones and the unique pronunciation duration of vowels; the similarity metric is a numerical indicator that quantifies the degree of matching between the distinguishing feature vector and the dialect type feature template, with a value range of [0,1], and the larger the value, the higher the similarity; the category membership degree is the probability value of the target dialect data belonging to a certain specific dialect type, reflecting the credibility of the classification result; the dialect type is the specific branch to which the target dialect belongs, as determined by the dialect classifier based on the category membership degree, such as sub-types such as urban accents and rural accents within a region.
[0050] Furthermore, the normalization processing of the voiceprint baseline features can be achieved by normalizing the mean and variance of the Mel frequency cepstral coefficients (MFCC), mapping the feature values to a uniform distribution range; the extraction of the distinguishing feature vector can be achieved using a feature selection algorithm based on mutual information, retaining the feature components most strongly correlated with the dialect type; the calculation of the similarity metric can be achieved using the Dynamic Time Warping (DTW) algorithm, achieving accurate matching by aligning feature sequences of different lengths; the determination of the category membership degree can be combined with the probability distribution output by the classifier, using the dialect type corresponding to the highest probability value as the candidate result; the final determination of the dialect type can be achieved by setting a membership degree threshold (e.g., 0.7), directly outputting the result when the highest membership degree exceeds the threshold, otherwise triggering a secondary feature extraction and matching process.
[0051] S4. Combining the individual voiceprint identifier with the dialect type, perform voiceprint feature fusion processing on the target dialect data to obtain the adapted initial speech, and analyze the regional prosodic features of the target region.
[0052] This invention combines the individual voiceprint identifier with the dialect type to perform voiceprint feature fusion processing on the target dialect data. This organically combines the speaker's individual characteristics with the characteristics of the dialect category to generate synthesized speech that conforms to the specific dialect's speech features. By analyzing the regional prosodic features of the target area, common rules in intonation, rhythm, and other aspects of the dialect in that area can be obtained, providing an important basis for dialect speech synthesis and recognition. Furthermore, the regional prosodic features of the target area can be analyzed by comparing multiple samples of prosodic elements. For example, by comparing the slope of the tone curve and the average syllable interval of dialect samples from different towns in the target area, the unified prosodic variation rules within the area can be extracted.
[0053] The voiceprint feature fusion process is a process of integrating individual pronunciation characteristics with dialect speech features; the adapted initial speech is speech data generated after feature fusion that combines individual voiceprint characteristics and dialect category features; the regional prosodic features are a set of features that reflect the common patterns of dialects in a specific region in terms of pitch variation, syllable duration, stress distribution, etc.
[0054] As an embodiment of the present invention, the step of combining the individual voiceprint identifier and the dialect type to perform voiceprint feature fusion processing on the target dialect data to obtain an adapted initial speech includes: The individual's voiceprint identifier is deconstructed to obtain the individual's physiological characteristics; The dialect types are analyzed using rules to obtain dialect phonological rules; By combining the physiological characteristics and the phonological rules, the target dialect data is reshaped to obtain an adapted initial speech.
[0055] The individual physiological characteristics are a set of acoustic features related to the vocal organs obtained by deconstructing the individual voiceprint identifier, including the fundamental frequency range of vocal cord vibration, the frequency of vocal tract formants, and the threshold of vocal intensity. The dialect phonological rules are the speech system specifications extracted after analyzing the dialect type, covering the rules of phonological combination, tone value patterns, and syllable stress distribution criteria. The feature reshaping is a process of adapting and integrating individual physiological characteristics with dialect phonological rules, adjusting individual pronunciation parameters to conform to dialect norms while retaining unique speech characteristics. The adapted initial speech is standardized speech data generated after feature reshaping that conforms to the common features of the target dialect type while retaining the individual pronunciation distinctiveness.
[0056] Furthermore, the individual voiceprint identifier can be deconstructed using acoustic feature decomposition algorithms to extract stable acoustic parameters determined by physiological structures such as the vocal cords and vocal tract; dialect phonetics analysis can be used to parse the dialect type according to rules, and phonological rules can be summarized by combining regional phonetic records and corpora; speech feature mapping technology can be used to adaptively adjust parameters such as tone curves and syllable durations of the target dialect data by combining individual physiological characteristics with dialect phonological rules; feature reshaping can be achieved through parameter optimization algorithms to correct speech deviations that do not conform to dialect rules while maintaining the individual's pronunciation style; finally, an adapted initial speech can be generated through a speech synthesis engine to ensure that it has both dialect standardization and individual uniqueness.
[0057] S5. Based on the regional prosodic features, perform prosodic optimization processing on the adapted initial speech to generate the dialect generation result corresponding to the target region.
[0058] This invention optimizes the initial speech based on regional prosodic features, enabling the generated dialect to conform to the speech habits of the target region. This eliminates the deviation between individual pronunciations and the common regional prosodic features in the initial speech, thereby improving the accuracy of dialect generation. At the same time, by combining the common rules of regional prosodic features, it ensures that the generated dialect speech conforms to the expression habits of the target region, providing standardized regional dialect speech materials for scenarios such as dialect inheritance and voice interaction.
[0059] Optionally, the prosodic optimization of the initial speech based on regional prosodic features can be achieved by comparing the prosodic differences between the initial speech and the regional prosodic features: first, extract the actual prosodic elements such as the tone curve, syllable interval, and stress intensity of the initial speech; then, compare these elements with the corresponding common indicators in the regional prosodic features (such as the average tone fluctuation range and the proportion of common syllable intervals in the region) to identify deviations; subsequently, adjust the prosodic parameters of the initial speech according to the degree of deviation—for example, correct the tone fluctuation range to the common range of the region, adjust the rhythm of the sentence according to the proportion of syllable intervals in the region, and strengthen the pronunciation intensity of the corresponding syllables according to the regional stress distribution habits; finally, through multiple listening tests, correct the minor deviations in the optimization process to ensure that the generated dialect result retains the individual pronunciation recognition of the initial speech and fully conforms to the prosodic rules of the target region, thus obtaining the dialect generation result corresponding to the target region. Specifically, for a more intuitive understanding of the prosodic optimization process of a dialect generation method based on voiceprint features in this application, please refer to [reference needed]. Figure 2 The diagram shown is a schematic of the prosody optimization process in a dialect generation method based on voiceprint features provided by this invention. It should be noted that in this invention... Figure 2The flowchart presented is only used for the lighting process of a dialect generation method based on voiceprint features, and is not limited to the prosody optimization process of a dialect generation method based on voiceprint features in different actual application scenarios.
[0060] Compared to the problems described in the background art, this invention, by processing the individual dialect recordings, can separate speaker features and dialect content from the original speech, providing a foundation for subsequent dialect analysis and speaker recognition, and effectively improving the accuracy of dialect resource utilization. This invention, by extracting dialect acoustic features from the recording content information, can transform the original speech content into a structured feature representation, providing a unified and computable data foundation for subsequent dialect analysis, comparison, and recognition, effectively improving the efficiency of dialect processing. Furthermore, by using the dialect feature vectors to train a dialect classifier, this invention can establish a correspondence between dialect features and dialect types, improving the accuracy of dialect recognition. And by analyzing the voiceprint baseline features corresponding to the target dialect data, it can capture the target... The unique phonetic characteristics of dialects provide core feature support for subsequent dialect type determination. Furthermore, this invention combines the individual voiceprint identifier with the dialect type to perform voiceprint feature fusion processing on the target dialect data. This organically combines the speaker's individual characteristics with the dialect category characteristics, generating synthesized speech that conforms to the specific dialect's phonetic features. By analyzing the regional prosodic features of the target region, common patterns in intonation, rhythm, etc., of the dialect in that region can be obtained, providing an important basis for dialect speech synthesis and recognition. Finally, this invention optimizes the prosodic features of the initial matching speech, ensuring that the generated dialect result matches the phonetic habits of the target region, eliminating deviations between individual pronunciation and regional common prosodic features in the initial matching speech, thereby improving the accuracy of dialect generation. Therefore, the dialect generation method and system based on voiceprint features provided by this invention can improve the accuracy of dialect generation.
[0061] like Figure 3 The diagram shown is a functional block diagram of a dialect generation system based on voiceprint features according to the present invention.
[0062] The dialect generation system 300 based on voiceprint features described in this invention can be installed in an electronic device. Depending on the functions implemented, the dialect generation system may include a dialect recording and extraction module 301, a voiceprint feature extraction module 302, a dialect classification module 303, a voiceprint fusion module 304, and a dialect generation module 305. The module described in this invention can also be referred to as a unit, which refers to a series of computer program segments that can be executed by the processor of an electronic device and can perform a fixed function, and are stored in the memory of the electronic device.
[0063] In this embodiment of the invention, the functions of each module / unit are as follows: The dialect recording and extraction module 301 is used to collect individual dialect recordings in the target area and extract individual voiceprint identifiers and recording content information of the target area based on the individual dialect recordings. The voiceprint feature extraction module 302 is used to extract the voiceprint acoustic feature vector of the recording content information using a pre-trained deep learning model, so as to construct the dialect feature vector corresponding to the individual in the target area. The dialect classification module 303 is used to train a dialect classifier for the target region using the dialect feature vector, receive target dialect data to be processed, analyze the voiceprint baseline features corresponding to the target dialect data, and analyze the dialect type corresponding to the target dialect data based on the voiceprint baseline features using the dialect classifier. The voiceprint fusion module 304 is used to combine the individual voiceprint identifier with the dialect type to perform voiceprint feature fusion processing on the target dialect data to obtain an adapted initial speech, and to analyze the regional prosodic features of the target region. The dialect generation module 305 is used to perform prosodic optimization processing on the adapted initial speech based on the regional prosodic features, so as to generate the dialect generation result corresponding to the target region.
[0064] In detail, the modules in the dialect generation system 300 based on voiceprint features described in this embodiment of the invention employ the same methods as described above. Figure 1 The technique is the same as that described in the dialect generation method based on voiceprint features, and can produce the same technical effect, so it will not be elaborated here.
[0065] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention.
[0066] Finally, it should be noted that in the above embodiments, each embodiment can be combined with each other or independent. Deleting any one of them will not affect the technical implementation of other embodiments. The above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention.
Claims
1. A dialect generation method based on voiceprint features, characterized in that, The method includes: Collect individual dialect recordings in the target area, and extract individual voiceprint identifiers and recording content information of the target area based on the individual dialect recordings; The voiceprint acoustic feature vector of the recorded content information is extracted using a pre-trained deep learning model to construct the dialect feature vector corresponding to the individual in the target area; Using the dialect feature vector, a dialect classifier corresponding to the target region is trained. The target dialect data to be processed is received, and the voiceprint baseline features corresponding to the target dialect data are analyzed. Based on the voiceprint baseline features, the dialect classifier is used to analyze the dialect type corresponding to the target dialect data. By combining the individual voiceprint identifier with the dialect type, the target dialect data is subjected to voiceprint feature fusion processing to obtain the adapted initial speech, and the regional prosodic features of the target region are analyzed. Based on the regional prosodic features, the initial adapted speech is subjected to prosodic optimization processing to generate the dialect generation result corresponding to the target region.
2. The dialect generation method based on voiceprint features as described in claim 1, characterized in that, The step of extracting individual voiceprint identifiers and recording content information for the target region based on the individual dialect recording includes: The individual dialect recordings are analyzed to obtain the original audio stream and vocal tract feature set; The original audio stream is preprocessed to obtain a clean audio stream; Based on preset dialect feature requirements, the pure audio stream is subjected to acoustic unit segmentation to obtain an acoustic sequence; Extract the spectral feature set corresponding to the acoustic sequence, and combine the vocal tract feature set and the spectral feature set to analyze the recording content information and individual voiceprint identifier.
3. The dialect generation method based on voiceprint features as described in claim 1, characterized in that, The step of extracting the voiceprint acoustic feature vector of the recorded content information using a pre-trained deep learning model includes: The acoustic standard in the deep learning model is used to uniformly process the recorded content information to obtain standardized acoustic data; The feature condenser in the deep learning model is used to extract essential information from the normalized acoustic data to obtain the core acoustic elements; The discrimination weights of each component in the core acoustic elements are calculated using the discrimination evaluator in the deep learning model. Based on the identification weights, the corresponding voiceprint acoustic feature vectors are generated from the core acoustic elements using the result synthesizer in the deep learning model.
4. The dialect generation method based on voiceprint features as described in claim 3, characterized in that, The calculation of the discrimination weights of each component in the core acoustic elements using the discrimination evaluator in the deep learning model includes: The attribute parser in the discrimination evaluator is used to identify the phoneme markers in the core acoustic elements, and the frequency of occurrence of each phoneme marker is counted. Based on the frequency of occurrence, the distribution breadth and specificity measure of each acoustic component are calculated using the distribution metric function in the discrimination evaluator. Based on the combined distribution breadth and specificity measure, the discrimination capability value of each acoustic component is calculated; Based on the discrimination capability value, the discrimination weight corresponding to each acoustic component is determined.
5. The dialect generation method based on voiceprint features as described in claim 1, characterized in that, The step of extracting the voiceprint acoustic feature vector of the recorded content information using a pre-trained deep learning model to construct the dialect feature vector corresponding to the individual in the target area includes: The acoustic feature vector of the voiceprint is subjected to distribution equalization processing to obtain balanced acoustic features; Calculate the core feature value corresponding to each feature in the equal acoustic features; Based on the core feature value, the voiceprint acoustic feature vector is smoothed to obtain smoothed acoustic features; The smooth acoustic features are encoded to obtain dialect feature vectors.
6. The dialect generation method based on voiceprint features as described in claim 5, characterized in that, The calculation of the core feature value corresponding to each feature in the equalized acoustic features includes: Frequency band energy analysis is performed on the aforementioned balanced acoustic characteristics to obtain the sub-band energy distribution. Calculate the standard deviation of the stationarity corresponding to the energy distribution of the sub-band; Combining the stationarity standard deviation and the subband distribution energy, the core feature value corresponding to each feature in the equal acoustic features is calculated using the following formula: ; Where C represents the core feature value corresponding to each feature in the equal acoustic features. This represents the energy distribution of the sub-band corresponding to the k-th sub-band. This represents the standard deviation of stationarity corresponding to the k-th subband, where k represents the subband sequence number and K represents the number of subbands.
7. The dialect generation method based on voiceprint features as described in claim 1, characterized in that, The analysis of the voiceprint baseline features corresponding to the target dialect data includes: The target dialect data is subjected to data filtering processing to obtain standard dialect data; Extract the spectral characteristics and prosodic information from the standard dialect data; Obtain the formant parameters and fundamental frequency profile corresponding to the standard dialect data; By combining the spectral characteristics and the prosodic information, the acoustic identification features of the target dialect data are analyzed; Based on the formant parameters and the fundamental frequency profile, the pronunciation pattern features of the target dialect data are determined; By combining the acoustic identification features and the pronunciation pattern features, the voiceprint baseline features corresponding to the standard dialect data are determined.
8. The dialect generation method based on voiceprint features as described in claim 1, characterized in that, The step of analyzing the dialect type corresponding to the target dialect data using the dialect classifier based on the voiceprint baseline features includes: The voiceprint reference features are normalized to obtain standard voiceprint features; Extract the distinguishing feature vector from the standard voiceprint features; The distinguishing feature vector is input into the dialect classifier, and the similarity calculation unit in the dialect classifier is used to calculate the similarity measure between the distinguishing feature vector and the dialect type feature template; Based on the similarity metric, the classification decision unit in the dialect classifier is used to determine the category membership degree of the target dialect data; Based on the category membership degree, the dialect type corresponding to the target dialect data is determined.
9. The dialect generation method based on voiceprint features as described in claim 1, characterized in that, The step of combining the individual voiceprint identifier with the dialect type to perform voiceprint feature fusion processing on the target dialect data to obtain the adapted initial speech includes: The individual's voiceprint identifier is deconstructed to obtain the individual's physiological characteristics; The dialect types are analyzed using rules to obtain dialect phonological rules; By combining the physiological characteristics and the phonological rules, the target dialect data is reshaped to obtain an adapted initial speech.
10. A dialect generation system based on voiceprint features, characterized in that, The system includes: The dialect recording and extraction module is used to collect individual dialect recordings in a target area, and extract individual voiceprint identifiers and recording content information of the target area based on the individual dialect recordings. The voiceprint feature extraction module is used to extract the voiceprint acoustic feature vector of the recorded content information using a pre-trained deep learning model, so as to construct the dialect feature vector corresponding to the individual in the target area. The dialect classification module is used to train a dialect classifier for the target region using the dialect feature vector, receive target dialect data to be processed, analyze the voiceprint baseline features corresponding to the target dialect data, and analyze the dialect type corresponding to the target dialect data based on the voiceprint baseline features using the dialect classifier. The voiceprint fusion module is used to combine the individual voiceprint identifier with the dialect type to perform voiceprint feature fusion processing on the target dialect data to obtain an adapted initial speech and analyze the regional prosodic features of the target region. The dialect generation module is used to perform prosodic optimization processing on the adapted initial speech based on the regional prosodic features, so as to generate the dialect generation result corresponding to the target region.