Personalized voice content generation method

By extracting semantic segmentation information of speech text, generating tone distribution maps and speech rhythm models, combining multi-parameter balance algorithms and dynamic adjustment algorithms, the shortcomings of personalized and fine-grained customization in traditional speech synthesis technology are solved, and precise control of multiple micro parameters and stability of speech quality are achieved.

CN120126447AInactive Publication Date: 2025-06-10广州鑫研科技有限公司
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510278111.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-10
Publication Date
2025-06-10
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Traditional speech synthesis technology has shortcomings in personalization and fine-grained customization, and it is difficult to achieve precise control of multiple micro parameters while ensuring natural and smooth speech. Dynamic adjustment of parameters may introduce instability, resulting in fluctuations in speech quality.

Method used

By obtaining the speech text to be generated, semantic segmentation information is extracted, the tone change range and stress position is determined, and a preliminary tone distribution map is generated; combining pause length parameters, the pause duration is calculated, and a phonological rhythm model is generated; integrating the tone distribution map and phonological rhythm model is used to judge and adjust the conflict area between tone change and pause length; dynamically adjust the frequency of usage of tone words, and a dynamically adjusted speech parameter model is generated using a multi-parameter balance algorithm, and preliminary speech waveform data is generated through the speech synthesis engine to optimize the stability of speech output.

Benefits of technology

It realizes precise control of multiple micro parameters while ensuring natural and smooth speech, solves the problem of mutual constraints between tone and pause, optimizes the coherence and rhythm of speech, ensures the stability of speech quality, and meets users' needs for personalized voice content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120126447A_ABST
    Figure CN120126447A_ABST
Patent Text Reader

Abstract

The invention discloses a personalized voice content generation method, which comprises the following steps: acquiring a voice text to be generated, and generating a preliminary tone distribution diagram; according to a preset pause length parameter, generating a voice rhythm model including pause duration; integrating the tone distribution diagram with the voice rhythm model, and judging whether an area where tone change conflicts with pause length exists or not; extracting mood word distribution information in the text, and dynamically adjusting the use frequency of mood words by combining accent position parameters and analyzing preferences and use scenes of the user; comprehensively considering the mutual influence of tone change, pause length, accent position and mood frequency, and generating a dynamically adjusted voice parameter model; combining the dynamically adjusted voice parameter model with a text, and extracting tone, pause and accent features in a waveform; judging a natural degree index in the voice waveform data, and optimizing the stability of voice output; noise and unstable fluctuation are eliminated, and final personalized voice output is generated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of intelligent voice technology, and in particular to a method for generating personalized voice content. Background Art

[0002] With the rapid development of artificial intelligence technology, the generation of personalized voice content has gradually become a research hotspot in the field of speech processing. Although traditional speech synthesis technology can generate natural and fluent speech, it has obvious deficiencies in terms of personalization and fine-grained customization. Users' demand for personalized voice content is increasing day by day. For example, in the fields of intelligent customer service, voice assistants, audiobooks, etc., users hope that the voice can reflect a unique personal style, including microscopic parameters such as the range of pitch variation, the length of pauses, the position of stress, and the frequency of using filler words.

[0003] However, the fine-grained customization technology faces a core technical contradiction in the implementation process: how to achieve precise control of multiple microscopic parameters while ensuring the natural fluency of the voice. These parameters do not exist independently but interact and restrict each other. For example, an increase in the range of pitch variation may lead to a decrease in speech coherence, an increase in the length of pauses may affect the rhythm of the speech, and the adjustment of the stress position needs to be coordinated with the frequency of using filler words, otherwise it will appear rigid or unnatural.

[0004] In addition, when dynamically adjusting these parameters, the existing technology may introduce new instabilities, resulting in fluctuations in speech quality. For example, although some speech synthesis methods can improve the naturalness of speech through data augmentation or model optimization, they still have limitations in personalized parameter control. Other methods attempt to achieve personalization through parameter adjustment, but it is difficult to balance the overall quality of the speech while dynamically balancing the relationship between various parameters. Summary of the Invention

[0005] In order to solve the above-mentioned existing technical problems, the present invention provides a method for generating personalized voice content.

[0006] The technical solution of the present invention is realized as follows:

[0007] A method for generating personalized voice content, comprising the following steps:

[0008] Obtain the speech text to be generated, extract the semantic segmentation information in the text, determine the range of pitch variation and the position of stress for each segment, and generate a preliminary pitch distribution map;

[0009] According to the preset pause length parameter, combined with the semantic segmentation information, calculate the pause duration between each segment, and generate a speech rhythm model including the pause duration;

[0010] Integrate the pitch distribution map with the speech rhythm model to determine whether there are regions where pitch changes conflict with pause lengths;

[0011] Extract the distribution information of filler words in the text, combine it with the stress position parameters, and dynamically adjust the usage frequency of filler words by analyzing the user's preferences and usage scenarios to optimize the fluency of speech expression;

[0012] Adopt a multi-parameter balance algorithm, comprehensively consider the mutual influence of pitch changes, pause lengths, stress positions, and filler frequencies, and generate a dynamically adjusted speech parameter model;

[0013] Through the speech synthesis engine, combine the dynamically adjusted speech parameter model with the text to generate preliminary speech waveform data, and extract the pitch, pause, and stress characteristics in the waveform;

[0014] Judge the naturalness index in the speech waveform data to optimize the stability of speech output;

[0015] Post-process the optimized speech waveform data to eliminate the noise and unstable fluctuations introduced by parameter adjustment, and generate the final personalized speech output.

[0016] Furthermore, the process of generating the preliminary pitch distribution map includes:

[0017] Obtain the speech text data to be generated, and use natural language processing technology to perform semantic segmentation on the text to obtain the segmented text information;

[0018] For each semantic segment, extract the pitch change characteristics through a speech feature analysis algorithm to determine the pitch change range of each segment;

[0019] Combine the speech semantic analysis results and the speech feature extraction data, and use a stress localization algorithm to determine the stress positions in each segment;

[0020] Based on the segmented pitch change range and stress position data, use a pitch modeling algorithm to generate a preliminary pitch distribution map. If there are abnormal segments in the pitch distribution map, re-perform semantic segmentation and pitch analysis processing;

[0021] Optimize the data in the distribution map according to the integrity and accuracy requirements of the pitch distribution map.

[0022] Furthermore, the process of generating the speech rhythm model including pause durations includes:

[0023] According to the preset pause length parameters, combine with the semantic segmentation information to calculate the pause duration between each segment;

[0024] Generate a speech rhythm model including pause durations for subsequent speech synthesis processing.

[0025] Further, the process of determining whether there is a region where pitch change conflicts with pause length includes:

[0026] Obtain the data of the pitch distribution map and the speech rhythm model, and use a conflict detection algorithm to determine whether there is a conflict region between pitch change and pause length;

[0027] If there is a conflict region, extract the pitch change range data of the conflict region, and use a dynamic adjustment algorithm to correct the pitch change range.

[0028] Further, after the step of determining whether there is a region where pitch change conflicts with pause length, it also includes:

[0029] According to the corrected pitch change range data, recalculate the speech rhythm model of the conflict region to ensure that the pause length matches the pitch change;

[0030] Through a speech coherence evaluation algorithm, determine whether the adjusted speech rhythm model meets the coherence requirements. If the coherence fails to meet the standard, extract the pitch distribution map data of the non-compliant region, and use a secondary adjustment algorithm to further optimize the pitch change range;

[0031] According to the optimized pitch change range data, update the speech rhythm model to generate the final speech rhythm distribution map;

[0032] Adopt a speech synthesis algorithm to combine the final speech rhythm distribution map with the pitch distribution map to generate complete speech data;

[0033] Obtain the pitch value in the speech wave, determine whether the pitch value exceeds a preset threshold. If it exceeds, use a dynamic adjustment algorithm to adjust the pitch value;

[0034] According to the adjusted pitch value, extract the pause value, determine the correlation between the pause value and the pitch value. If there is a conflict, adjust the pause value.

[0035] Further, the process of dynamically adjusting the usage frequency of filler words includes:

[0036] Obtain text data, extract filler word distribution information, and generate a filler word distribution map;

[0037] Extract the stress position parameters from the speech data to generate a stress parameter set;

[0038] Combine user preference data and usage scenario information to construct an analysis model;

[0039] According to the analysis model, determine whether there is a conflict between the filler word distribution and the stress position. If there is a conflict, use a dynamic adjustment algorithm to adjust the filler word frequency value to generate adjusted frequency value data;

[0040] Recalculate the voice table parameters according to the adjusted frequency value data to generate optimized voice table data;

[0041] Generate the final voice expression fluency distribution map through the optimized voice table data;

[0042] Obtain the naturalness index of the voice waveform, determine whether the naturalness is lower than the preset threshold. If the naturalness is lower than the preset threshold, use the dynamic adjustment algorithm to reallocate the parameter weights;

[0043] Optimize the stability index of the voice waveform according to the reallocated weights.

[0044] Further, the process of generating the dynamically adjusted voice parameter model includes:

[0045] Obtain voice data, extract pitch change features, and generate a pitch change data set;

[0046] For the pitch change data set, combine the pause length parameter, analyze the correlation between the pause length and the pitch change. If there is a conflict between the pause length and the pitch change, use the dynamic adjustment algorithm to adjust the pause length parameter;

[0047] According to the adjusted pause length parameter, extract the stress position features and generate a stress position data set;

[0048] For the stress position data set, combine the tone frequency parameter, judge the correlation between the stress position and the tone frequency. If there is a conflict between the stress position and the tone frequency, use the balance algorithm to adjust the tone frequency parameter;

[0049] Generate a dynamically adjusted voice parameter model according to the adjusted tone frequency parameter;

[0050] Through the voice synthesis engine, combine the dynamically adjusted voice parameter model with the text to generate preliminary voice waveform data;

[0051] Extract the pitch, pause, and stress features in the waveform to obtain the final voice parameter model.

[0052] Further, the process of extracting the pitch, pause, and stress features in the waveform includes:

[0053] Load the dynamically adjusted voice parameter model through the voice synthesis engine, combine with the text data to be processed, and generate preliminary voice waveform data;

[0054] Adopt the voice feature analysis algorithm to extract the pitch features from the voice waveform data to obtain the pitch change range;

[0055] Extract pause features based on the speech waveform data to determine the pause duration between each segment;

[0056] Extract stress features from the speech waveform data through a stress localization algorithm to determine the stress positions in each segment.

[0057] Further, the process of optimizing the stability of the speech output includes:

[0058] Obtain the speech waveform data, extract the naturalness index, and obtain the naturalness value;

[0059] Judge the difference between the naturalness value and the preset threshold to determine whether it is lower than the preset threshold. If the naturalness is lower than the preset threshold, use a dynamic adjustment algorithm to reallocate the weight distribution in the parameter balancing algorithm;

[0060] Optimize the stability index of the speech waveform data according to the reallocated weights to obtain the optimized speech waveform;

[0061] Extract the optimized speech waveform and judge whether its naturalness value meets the preset threshold. If the naturalness is still lower than the preset threshold, use a weighted average algorithm to further adjust the parameter weight distribution;

[0062] Generate new speech waveform data according to the finally adjusted weights and extract its naturalness value;

[0063] Judge the difference between the naturalness of the new speech waveform and the preset threshold to determine the stability level of the speech output.

[0064] Further, the process of generating the final personalized speech output includes:

[0065] Obtain the optimized speech waveform data and extract the pitch, pause, and stress features in the waveform;

[0066] According to the extracted pitch, pause, and stress features, use a filter to eliminate high-frequency noise;

[0067] Reduce the unstable fluctuations in the waveform through a smoothing algorithm to obtain the processed speech waveform;

[0068] Extract the processed speech waveform and judge whether its naturalness meets the preset range. If the naturalness meets the requirements, generate the final personalized speech output; if the naturalness does not meet the requirements, use a dynamic adjustment algorithm to reallocate the weight distribution;

[0069] Optimize the stability index of the speech waveform according to the reallocated weights;

[0070] Obtain the optimized speech waveform and use a filter to eliminate high-frequency noise;

[0071] Reduce the unstable fluctuations in the waveform through a smoothing algorithm to generate the final personalized voice output.

[0072] Compared with the prior art, the present invention has the following beneficial effects:

[0073] 1. By extracting the semantic segmentation information of the text and determining key parameters such as the pitch change range and stress position of each segment, the present invention generates a preliminary pitch distribution map. This process utilizes natural language processing technology and speech feature analysis algorithms to ensure that the basic framework of speech generation can accurately reflect the semantic structure and speech feature requirements of the text, providing a solid foundation for subsequent personalized adjustments.

[0074] 2. By introducing a speech rhythm model, calculating the pause duration between each segment according to the preset pause length parameter and semantic segmentation information, and integrating the pitch distribution map with the speech rhythm model, the solution can determine whether there is a conflict between pitch changes and pause lengths, and correct the conflict area through a dynamic adjustment algorithm, effectively solving the problem of mutual restriction between pitch and pause, ensuring the coherence and rhythm of speech, and providing flexibility for personalized speech generation.

[0075] 3. By extracting the distribution information of modal particles in the text and dynamically adjusting the usage frequency of modal particles in combination with user preferences and usage scenarios, this personalized adjustment mechanism based on user needs enables the speech to better reflect the user's personal style and emotional expression, optimizing the fluency of speech expression. At the same time, the solution adopts a multi-parameter balance algorithm, comprehensively considering the mutual influence of pitch changes, pause lengths, stress positions, and modal frequencies, to generate a dynamically adjusted speech parameter model, ensuring the coordination between parameters and avoiding the impact of single-parameter adjustment on the overall speech quality.

[0076] 4. Aiming at the instability problem that may be introduced by dynamically adjusting parameters in the prior art, after generating the preliminary speech waveform data through a speech synthesis engine, the present invention extracts pitch, pause, and stress features, and judges the naturalness index in the speech waveform data. If the naturalness is lower than the preset threshold, the solution adopts a dynamic adjustment algorithm to reallocate the parameter weights to optimize the stability of the speech output. In addition, the speech waveform is post-processed through a filter and a smoothing algorithm to further eliminate the noise and unstable fluctuations introduced by parameter adjustment, ensuring the stability of the speech quality.

[0077] 5. Through a series of optimization measures, the present invention realizes the precise control of multiple microscopic parameters while ensuring natural and fluent speech, not only meeting the user's demand for personalized speech content, but also solving the contradiction in the prior art that it is difficult to balance the overall speech quality and personalized customization, providing a high-quality personalized speech solution for fields such as intelligent customer service, voice assistants, and audiobooks. Brief Description of the Drawings

[0078] Figure 1 This is a flowchart of the steps of a personalized voice content generation method of the present invention. Detailed Embodiments

[0079] To make the objectives, features, and advantages of the present invention more obvious and understandable, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the embodiments described below are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without making creative efforts shall fall within the protection scope of the present invention.

[0080] As Figure 1 shown, this embodiment provides a personalized voice content generation method, including the following steps:

[0081] Step 1: Obtain the voice text to be generated, extract the semantic segmentation information in the text, determine the pitch change range and stress position of each segment, and generate a preliminary pitch distribution map;

[0082] Obtain the voice text data to be generated, and perform semantic segmentation processing on the text using natural language processing technology to obtain the segmented text information;

[0083] For each semantic segment, extract the pitch change features through a voice feature analysis algorithm to determine the pitch change range of each segment;

[0084] Combine the text semantic analysis results and the voice feature extraction data, and use a stress positioning algorithm to determine the stress position in each segment;

[0085] Based on the segmented pitch change range and stress position data, use a pitch modeling algorithm to generate a preliminary pitch distribution map. If there are abnormal segments in the pitch distribution map, re-perform semantic segmentation and pitch analysis processing;

[0086] Optimize the data in the distribution map according to the integrity and accuracy requirements of the pitch distribution map.

[0087] Exemplarily, obtain the voice text data to be generated, and perform semantic segmentation processing on the text using natural language processing technology based on BERT to obtain the segmented text information. For example, a text with a length of 500 words is segmented into 10 semantic paragraphs;

[0088] For each semantic segment, extract the pitch change features through a voice feature analysis algorithm based on MFCC to determine the pitch change range of each segment. For example, the pitch range of a certain segment is between 85 Hz and 200 Hz;

[0089] Combined with the text semantic analysis results and speech feature extraction data, a rule-based stress location algorithm is used to determine the stress positions in each segment. For example, in a certain segment, 3 stress positions are identified, which are located at the 5th, 12th, and 18th syllables respectively;

[0090] Based on the segmented pitch change range and stress position data, a pitch modeling algorithm based on Gaussian mixture model is used to generate a preliminary pitch distribution map. For example, a pitch distribution curve containing 10 segments is generated;

[0091] If there are abnormal segments in the pitch distribution map, for example, the pitch range of a certain segment exceeds the normal range of 50Hz to 250Hz, then re-perform semantic segmentation and pitch analysis processing;

[0092] According to the integrity and accuracy requirements of the pitch distribution map, optimize the data in the distribution map. For example, correct the pitch range of the abnormal segment to 120Hz to 210Hz.

[0093] Step 2: According to the preset pause length parameter, combined with the semantic segmentation information, calculate the pause duration between each segment, and generate a speech rhythm model containing the pause duration;

[0094] The process of generating the speech rhythm model containing the pause duration includes:

[0095] According to the preset pause length parameter, combined with the semantic segmentation information, calculate the pause duration between each segment;

[0096] Generate a speech rhythm model containing the pause duration for subsequent speech synthesis processing.

[0097] Exemplarily, according to the preset pause length parameter (for example, set the pause length to 300ms to 500ms), combined with the semantic segmentation information, calculate the pause duration between each segment. For example, for the pause between two semantic paragraphs, calculate the pause to be 350ms according to the context semantics and preset parameters;

[0098] Generate a speech rhythm model containing the pause duration, combine the pause information with the pitch distribution map to form a complete speech rhythm curve for subsequent speech synthesis processing.

[0099] Step 3: Integrate the pitch distribution map and the speech rhythm model, and judge whether there are regions where pitch changes conflict with the pause length;

[0100] The process of judging whether there are regions where pitch changes conflict with the pause length includes:

[0101] Obtain data of the pitch distribution map and the speech rhythm model, and use a conflict detection algorithm to determine whether there is a conflict area between pitch changes and pause lengths;

[0102] If there is a conflict area, extract the pitch change range data of the conflict area, and use a dynamic adjustment algorithm to correct the pitch change range.

[0103] Exemplarily, obtain data of the pitch distribution map and the speech rhythm model, and use a conflict detection algorithm based on the dynamic time warping algorithm to determine whether there is a conflict area between pitch changes and pause lengths. If a region where the pitch change amplitude exceeds ±20 Hz and the pause length is less than 200 ms is detected, it is determined as a conflict area;

[0104] Extract the pitch change range data of the conflict area, and use a dynamic adjustment algorithm based on gradient descent to limit the pitch change amplitude within ±15 Hz and correct the pitch change range

[0105] Specifically, using a dynamic adjustment algorithm to correct the pitch change range can be described as follows:

[0106] First, identify the area where there is a conflict between pitch changes and pause lengths through a conflict detection algorithm (such as the dynamic time warping algorithm). For example, if the pitch change amplitude exceeds ±20 Hz and the pause length is less than 200 ms, it is considered that there is a conflict in this area;

[0107] Extract the pitch change range data, extract the pitch change range data from the conflict area, including the starting frequency, ending frequency, and change amplitude of the pitch, etc.;

[0108] Apply the dynamic adjustment algorithm, use the dynamic adjustment algorithm (such as the gradient descent method) to correct the pitch change range. The goal of this algorithm is to limit the pitch change amplitude within ±15 Hz to reduce the conflict with the pause length.

[0109] Calculate the corrected pitch change range. According to the dynamic adjustment algorithm, calculate the corrected pitch change range. For example, if the original pitch change range is from 85 Hz to 200 Hz, the corrected range may be from 90 Hz to 195 Hz to ensure that the change amplitude does not exceed ±15 Hz;

[0110] Suppose there is a pitch distribution map, and the pitch change range of a certain segment is from 85 Hz to 200 Hz, and the pause length is 150 ms. Since the pitch change amplitude is 115 Hz (200 Hz - 85 Hz), which exceeds the threshold of ±20 Hz, and the pause length is less than 200 ms, there is a conflict;

[0111] Use the gradient descent method to correct the pitch change range. First, calculate the midpoint frequency of the pitch change:

[0112]

[0113] Then, limit the pitch variation range within ±15 Hz, i.e.:

[0114] The corrected pitch range = 142.5 Hz ± 15 Hz = [127.5 Hz, 157.5 Hz]

[0115] The corrected pitch variation range is from 127.5 Hz to 157.5 Hz, and the variation amplitude is 30 Hz, which meets the limit of ±15 Hz, reduces the conflict with the pause length, and updates the corrected pitch variation range to the pitch distribution map for subsequent speech synthesis use.

[0116] Update the pitch distribution map, and update the corrected pitch variation range data to the pitch distribution map for subsequent speech synthesis use.

[0117] After the step of determining whether there is a conflict area between pitch variation and pause length, it further includes:

[0118] According to the corrected pitch variation range data, recalculate the speech rhythm model of the conflict area to ensure that the pause length matches the pitch variation;

[0119] Through the speech coherence evaluation algorithm, judge whether the adjusted speech rhythm model meets the coherence requirement. If the coherence does not meet the standard, extract the pitch distribution map data of the non-compliant area, and use the secondary adjustment algorithm to further optimize the pitch variation range;

[0120] According to the optimized pitch variation range data, update the speech rhythm model to generate the final speech rhythm distribution map;

[0121] Adopt the speech synthesis algorithm to combine the final speech rhythm distribution map with the pitch distribution map to generate complete speech data;

[0122] Obtain the pitch value in the speech wave, judge whether the pitch value exceeds the preset threshold. If it exceeds, use the dynamic adjustment algorithm to adjust the pitch value;

[0123] According to the adjusted pitch value, extract the pause value, judge the correlation between the pause value and the pitch value. If there is a conflict, adjust the pause value.

[0124] Exemplarily, according to the corrected pitch variation range data, recalculate the conflict area by using the speech rhythm model based on the hidden Markov model to ensure that the pause length matches the pitch variation;

[0125] Through the speech coherence evaluation algorithm based on dynamic time warping, calculate the coherence score of the adjusted speech rhythm model. If the score is lower than 0.8, it is determined that the coherence does not meet the standard;

[0126] Extract the pitch distribution map data of the non-compliant area, adopt a quadratic adjustment algorithm based on the genetic algorithm to further optimize the pitch change range, and adjust the pitch change amplitude to ±10 Hz. Update the speech rhythm model according to the optimized pitch change range data to generate the final speech rhythm distribution map;

[0127] Adopt a speech synthesis algorithm based on a deep neural network, combine the final speech rhythm distribution map with the pitch distribution map to generate complete speech data;

[0128] Obtain the pitch value in the speech wave, judge whether the pitch value exceeds the preset threshold of 250 Hz, and if it exceeds, adjust the pitch value to 240 Hz using a dynamic adjustment algorithm based on linear interpolation;

[0129] According to the adjusted pitch value, extract the pause value, judge the correlation between the pause value and the pitch value. If the pause value is lower than 150 ms and the pitch value is higher than 230 Hz, adjust the pause value to 180 ms using an algorithm based on dynamic programming.

[0130] Step 4: Extract the distribution information of filler words in the text, combine the stress position parameters, and dynamically adjust the usage frequency of filler words by analyzing the user's preferences and usage scenarios to optimize the fluency of speech expression;

[0131] The process of dynamically adjusting the usage frequency of filler words includes:

[0132] Obtain text data, extract the distribution information of filler words, and generate a filler word distribution map;

[0133] Extract the stress position parameters from the speech data to generate a stress parameter set;

[0134] Combine the user preference data and usage scenario information to construct an analysis model;

[0135] According to the analysis model, judge whether there is a conflict between the filler word distribution and the stress position. If there is a conflict, use a dynamic adjustment algorithm to adjust the filler word frequency value to generate adjusted frequency value data;

[0136] According to the adjusted frequency value data, recalculate the speech table parameters to generate optimized speech table data;

[0137] Generate the final speech expression fluency distribution map through the optimized speech table data;

[0138] Obtain the naturalness index of the speech waveform, judge whether the naturalness is lower than the preset threshold. If the naturalness is lower than the preset threshold, use a dynamic adjustment algorithm to reallocate the parameter weights;

[0139] Optimize the stability index of the speech waveform according to the reallocated weights.

[0140] Exemplarily, obtain text data, extract the distribution information of modal particles, calculate the occurrence frequency of modal particles in the text by using the frequency statistics method, and generate a modal particle distribution map;

[0141] Extract the stress position parameters from the speech data, identify the stress points in the speech waveform through the sound intensity and pitch analysis algorithms, and generate a parameter set containing stress position information;

[0142] Combine the user preference data and usage scenario information, and use a machine learning algorithm to construct an analysis model. The model inputs include the modal particle distribution, stress position, and user preference data, and the output is a relationship matrix between modal particles and stress positions;

[0143] Specifically, the process of constructing the analysis model can be described as follows:

[0144] Collect user preference data, where user preference data includes the user's preferences for aspects such as the pitch, speaking speed, and usage frequency of modal particles of the speech; for example, the user may prefer a relatively moderate usage frequency of modal particles, or prefer a more formal speech style in certain scenarios;

[0145] Collect usage scenario information, where usage scenario information includes the specific scenarios of speech generation, such as intelligent customer service, voice assistants, audiobooks, etc. Different scenarios may have different requirements for the use of modal particles. For example, intelligent customer service may require more formal modal particles, while voice assistants may require more natural and friendly modal particles;

[0146] Construct an analysis model, combine the user preference data and usage scenario information to construct an analysis model that can dynamically adjust the usage frequency of modal particles according to the input text and scenario information;

[0147] The model can be implemented using machine learning algorithms (such as linear regression, decision trees, neural networks, etc.).

[0148] Model training, use a labeled data set (including user preferences and scenario information) to train the model;

[0149] The data set can include the annotations of the usage frequency of modal particles by users in different scenarios, as well as the corresponding speech samples;

[0150] Model evaluation and optimization, evaluate the performance of the model through methods such as cross-validation to ensure that the model can accurately predict user preferences and scenario requirements;

[0151] Assume that a linear regression model is used to construct the analysis model, and the linear regression model can be expressed as:

[0152] y=β0 +β 1 x 1 +β 2 x 2 +...+β n x n

[0153] where y is the target variable, such as the frequency of using filler words, and x 1 , x 2 ..., x n are input features, such as user preference data, usage scenario information, etc., and β 0 , β 1 ..., β n are model parameters;

[0154] Suppose there are the following input features:

[0155] x 1 : User preference for filler words (a numerical value between 0 and 1, where 0 means not liking to use filler words and 1 means liking to use filler words frequently);

[0156] x 2 : Usage scenario (0 means intelligent customer service, 1 means voice assistant, 2 means audiobook);

[0157] x 3 : Text length (expressed in number of words or sentences);

[0158] The target variable y is the frequency of using filler words (expressed as the number of filler words used per minute);

[0159] The linear regression model can be expressed as:

[0160] y = β 0 +β 1 x 1 +β 2 x 2 +β 3 x 3

[0161] Using the training dataset, we can use the least squares method to estimate the model parameters, β 0 , β 1 , β 2 , β 3 ;

[0162] The goal of the least squares method is to minimize the sum of squared errors:

[0163]

[0164] where m is the number of samples in the dataset;

[0165] Suppose the following model parameters are obtained through training:

[0166] β 0 = 0.5, β 1 = 0.3, β 2 = 0.2, β 3 = 0.1;

[0167] Now, there is a new input sample:

[0168] The user's preference for filler words x 1 = 0.7 (the user likes to use filler words frequently);

[0169] Usage scenario x 2 = 1 (voice assistant);

[0170] Text length x 3 = 500 (500-word text);

[0171] According to the model, we can calculate the usage frequency y of filler words:

[0172] y = 0.5 + 0.3×0.7 + 0.2×1 + 0.1×500

[0173] y = 0.5 + 0.21 + 0.2 + 50

[0174] y = 50.91;

[0175] Therefore, the model predicts that the usage frequency of filler words in this scenario is 50.91 times per minute.

[0176] Optimize the model according to the evaluation results, and adjust the model parameters to improve the prediction accuracy.

[0177] According to the analysis model, judge whether there is a conflict between the distribution of filler words and the stress position. If the conflict detection result is true, use the dynamic adjustment algorithm to adjust the frequency value of filler words. The conflict detection threshold is set to 0.8, and generate the adjusted frequency value data;

[0178] According to the adjusted frequency value data, recalculate the voice table parameters, optimize the voice table using the weighted average algorithm, and generate the optimized voice table data;

[0179] Through the optimized voice table data, generate the final voice expression fluency distribution map. A fluency score higher than 0.9 is considered qualified;

[0180] Obtain the naturalness index of the voice waveform, calculate the naturalness using the Mel Frequency Cepstral Coefficient analysis algorithm, and judge whether the naturalness is lower than the preset threshold of 0.85. If it is lower than the threshold, use the dynamic adjustment algorithm to reallocate the parameter weights;

[0181] Optimize the stability index of the speech waveform according to the reallocated weights. A stability score higher than 0.88 is considered qualified.

[0182] Step 5: Adopt a multi-parameter balancing algorithm to comprehensively consider the mutual influence of pitch variation, pause length, stress position, and tone frequency, and generate a dynamically adjusted speech parameter model;

[0183] The process of generating the dynamically adjusted speech parameter model includes:

[0184] Obtain speech data, extract pitch variation features, and generate a pitch variation data set;

[0185] For the pitch variation data set, combine the pause length parameter, analyze the correlation between the pause length and pitch variation. If there is a conflict between the pause length and pitch variation, use a dynamic adjustment algorithm to adjust the pause length parameter;

[0186] According to the adjusted pause length parameter, extract stress position features and generate a stress position data set;

[0187] For the stress position data set, combine the tone frequency parameter, judge the correlation between the stress position and tone frequency. If there is a conflict between the stress position and tone frequency, use a balancing algorithm to adjust the tone frequency parameter;

[0188] According to the adjusted tone frequency parameter, generate a dynamically adjusted speech parameter model;

[0189] Exemplarily, obtain speech data, extract pitch variation features using Fourier transform, and generate a pitch variation data set with a frequency range of 85 Hz to 255 Hz;

[0190] For the pitch variation data set, combine the pause length parameter, use the Pearson correlation coefficient to analyze the correlation between the pause length and pitch variation. If the correlation coefficient is less than 0.5, it is determined that there is a conflict;

[0191] Use a dynamic adjustment algorithm to adjust the pause length parameter, shortening the pause length from 0.5 seconds to 0.3 seconds;

[0192] According to the adjusted pause length parameter, extract stress position features and generate a stress position data set including stress position timestamps;

[0193] For the stress position data set, combine the tone frequency parameter, use a linear regression model to judge the correlation between the stress position and tone frequency. If the regression coefficient is less than 0.6, it is determined that there is a conflict. Use a balancing algorithm to adjust the tone frequency parameter, increasing the tone frequency from 3 times per second to 4 times per second;

[0194] Generate a dynamically adjusted speech parameter model according to the adjusted tone frequency parameter. The weights of pitch variation, pause length, stress position, and tone frequency in the model are 0.3, 0.2, 0.3, and 0.2 respectively.

[0195] Step 6: Combine the dynamically adjusted speech parameter model with the text through a speech synthesis engine to generate preliminary speech waveform data, and extract pitch, pause, and stress features from the waveform.

[0196] The process of extracting pitch, pause, and stress features from the waveform includes:

[0197] Load the dynamically adjusted speech parameter model through a speech synthesis engine, combine it with the text data to be processed, and generate preliminary speech waveform data;

[0198] Adopt a speech feature analysis algorithm to extract pitch features from the speech waveform data to obtain the pitch variation range;

[0199] Extract pause features according to the speech waveform data to determine the pause duration between each segment;

[0200] Extract stress features from the speech waveform data through a stress localization algorithm to determine the stress position in each segment.

[0201] Exemplarily, load the dynamically adjusted speech parameter model through a speech synthesis engine, combine it with the text data to be processed, and generate preliminary speech waveform data;

[0202] Adopt a speech feature analysis algorithm to extract pitch features from the speech waveform data, for example, use the short-time Fourier transform to calculate the pitch variation within the frequency range of 85Hz to 255Hz;

[0203] Extract pause features according to the speech waveform data, and determine the pause duration between each segment through an energy threshold detection method, for example, mark it as a pause when the silent segment exceeds 200ms. Extract stress features from the speech waveform data through a stress localization algorithm, for example, use a peak detection method based on Mel Frequency Cepstral Coefficients (MFCC) to determine the stress position in each segment.

[0204] Step 7: Judge the naturalness index in the speech waveform data and optimize the stability of the speech output;

[0205] The process of optimizing the stability of the speech output includes:

[0206] Obtain the speech waveform data, extract the naturalness index, and obtain the naturalness value;

[0207] Judge the difference between the naturalness value and the preset threshold, and determine whether it is lower than the preset threshold. If the naturalness is lower than the preset threshold, use a dynamic adjustment algorithm to re - distribute the weight distribution in the parameter balance algorithm;

[0208] Optimize the stability index of the speech waveform data according to the re - distributed weights to obtain the optimized speech waveform;

[0209] Extract the optimized speech waveform, and judge whether its naturalness value meets the preset threshold. If the naturalness is still lower than the preset threshold, use a weighted average algorithm to further adjust the parameter weight distribution;

[0210] Generate new speech waveform data according to the finally adjusted weights, and extract its naturalness value;

[0211] Judge the difference between the naturalness of the new speech waveform and the preset threshold to determine the stability level of the speech output.

[0212] Exemplarily, obtain speech waveform data, extract the naturalness index, and use an analysis method based on Mel - Frequency Cepstral Coefficients (MFCC) to calculate the naturalness value. For example, the naturalness score is 0.85 by calculating waveform smoothness and spectral continuity;

[0213] Judge the difference between the naturalness value and the preset threshold of 0.90, and determine whether it is lower than the preset threshold. If the naturalness is lower than the preset threshold, use a dynamic adjustment algorithm based on gradient descent to re - distribute the weight distribution in the parameter balance algorithm. For example, adjust the spectral weight from 0.6 to 0.7, and adjust the waveform smoothness weight from 0.4 to 0.3;

[0214] Optimize the stability index of the speech waveform data according to the re - distributed weights, and use a stability evaluation function to calculate the optimized waveform stability score. For example, it is improved from 0.78 to 0.82;

[0215] Extract the optimized speech waveform, and judge whether its naturalness value meets the preset threshold. For example, calculate the naturalness score again as 0.88. If the naturalness is still lower than the preset threshold, use a weighted average algorithm to further adjust the parameter weight distribution. For example, adjust the spectral weight to 0.75 and the smoothness weight to 0.25;

[0216] Generate new speech waveform data according to the finally adjusted weights, and extract its naturalness value. For example, calculate the naturalness score as 0.91;

[0217] Judge the difference between the naturalness of the new speech waveform and the preset threshold to determine the stability level of the speech output. For example, the stability score is 0.85, which meets the preset range of 0.8 to 0.9.

[0218] Step 8: Post-process the optimized voice waveform data to eliminate the noise and unstable fluctuations introduced by parameter adjustment, and generate the final personalized voice output.

[0219] The process of generating the final personalized voice output includes:

[0220] Obtain the optimized voice waveform data, and extract the pitch, pause, and stress features in the waveform;

[0221] According to the extracted pitch, pause, and stress features, use a filter to eliminate high-frequency noise;

[0222] Reduce the unstable fluctuations in the waveform through a smoothing algorithm to obtain the processed voice waveform;

[0223] Extract the processed voice waveform, and judge whether its naturalness meets the preset range. If the naturalness meets the requirements, generate the final personalized voice output; if the naturalness does not meet the requirements, use a dynamic adjustment algorithm to reallocate the weight distribution;

[0224] Optimize the stability index of the voice waveform according to the reallocated weights;

[0225] Obtain the optimized voice waveform, and use a filter to eliminate high-frequency noise;

[0226] Reduce the unstable fluctuations in the waveform through a smoothing algorithm to generate the final personalized voice output.

[0227] Exemplarily, obtain the optimized voice waveform data, and extract the pitch, pause, and stress features in the waveform. For example, extract the pitch frequency range from 85 Hz to 255 Hz through Fourier transform, determine the pause position using short-time energy analysis, and identify the stress features using the fundamental frequency extraction algorithm;

[0228] According to the extracted pitch, pause, and stress features, use a filter to eliminate high-frequency noise. For example, use a Butterworth low-pass filter with a cut-off frequency of 4000 Hz and an attenuation slope set to 24 dB / octave;

[0229] Reduce the unstable fluctuations in the waveform through a smoothing algorithm. For example, use a Savitzky-Golay filter with a window size of 21 and a polynomial order of 3 to obtain the processed voice waveform;

[0230] Extract the processed speech waveform and determine whether its naturalness meets the preset range. For example, calculate the cosine similarity between the Mel cepstral coefficients and the target naturalness model, and set the threshold to 0.85. If the naturalness meets the requirements, generate the final personalized speech output. For example, encode the waveform data into 16-bit PCM format with a sampling rate of 44.1 kHz. If the naturalness does not meet the requirements, use a dynamic adjustment algorithm to reallocate the weight distribution. For example, optimize the weight coefficients of pitch, pause, and stress based on the gradient descent method, and set the learning rate to 0.01;

[0231] According to the reallocated weights, optimize the stability index of the speech waveform. For example, calculate the variance of the waveform energy distribution and ensure that it is less than the preset threshold of 0.05;

[0232] Obtain the optimized speech waveform and use a filter to eliminate high-frequency noise. For example, use a Chebyshev type II filter with a cut-off frequency of 3500 Hz and a stopband attenuation of 40 dB;

[0233] Reduce the unstable fluctuations in the waveform through a smoothing algorithm. For example, use a median filter with a window size of 15 to generate the final personalized speech output.

Claims

1. A method for generating personalized voice content, characterized in that: The following steps are involved: Obtain the speech text to be generated, extract the semantic segmentation information in the text, determine the pitch change range and stress position of each segment, and generate a preliminary pitch distribution map; According to the preset pause length parameter and in combination with the semantic segmentation information, the pause duration between each segment is calculated to generate a speech rhythm model including the pause duration; Integrate the pitch distribution map with the speech rhythm model to determine whether there are areas where pitch changes conflict with pause lengths; Extract the distribution information of modal particles in the text, combine it with the stress position parameters, analyze the user's preferences and usage scenarios, dynamically adjust the usage frequency of modal particles, and optimize the fluency of speech expression; A multi-parameter balancing algorithm is used to comprehensively consider the mutual influence of pitch change, pause length, stress position and tone frequency to generate a dynamically adjusted speech parameter model; The dynamically adjusted speech parameter model is combined with the text through the speech synthesis engine to generate preliminary speech waveform data and extract the pitch, pause and stress features in the waveform; Determine the naturalness index in the speech waveform data and optimize the stability of speech output; The optimized speech waveform data is post-processed to eliminate the noise and unstable fluctuations introduced by parameter adjustment and generate the final personalized speech output.

2. The method for generating personalized voice content according to claim 1, characterized in that: The process of generating a preliminary pitch distribution map includes: Acquire the speech text data to be generated, use natural language processing technology to perform semantic segmentation on the text, and obtain segmented text information; For each semantic segment, extract the pitch change features and determine the pitch change range of each segment; Combining the text semantic analysis results with the speech feature extraction data, determine the stress position in each segment; Based on the segmented pitch change range and stress position data, a pitch modeling algorithm is used to generate a preliminary pitch distribution map. If there are abnormal segments in the pitch distribution map, semantic segmentation and pitch analysis are performed again. According to the requirements of completeness and accuracy of the tone distribution map, the data in the distribution map is optimized.

3. The method for generating personalized voice content according to claim 1, characterized in that: The process of generating a speech rhythm model including pause duration comprises: According to the preset pause length parameters and combined with the semantic segmentation information, the pause duration between each segment is calculated; Generate a speech rhythm model including pause duration for subsequent speech synthesis processing.

4. The method for generating personalized voice content according to claim 1, characterized in that: The process of determining whether there is a region where the tone change conflicts with the pause length includes: Obtain data of the pitch distribution map and speech rhythm model to determine whether there is a conflicting area between pitch change and pause length; If there is a conflicting area, the pitch change range data of the conflicting area is extracted, and the pitch change range is corrected using a dynamic adjustment algorithm.

5. The method for generating personalized voice content according to claim 1, characterized in that: After the step of determining whether there is a region where the tone change conflicts with the pause length, the method further includes: According to the corrected pitch change range data, the speech rhythm model of the conflicting area is recalculated to ensure that the pause length matches the pitch change; Determine whether the adjusted speech rhythm model meets the coherence requirement. If the coherence does not meet the requirement, extract the pitch distribution map data of the area that does not meet the requirement, and use a secondary adjustment algorithm to further optimize the pitch change range. According to the optimized pitch change range data, the speech rhythm model is updated to generate a final speech rhythm distribution map; The final speech rhythm distribution map is combined with the pitch distribution map to generate complete speech data; the pitch value in the speech wave is obtained to determine whether the pitch value exceeds a preset threshold, and if so, the pitch value is adjusted using a dynamic adjustment algorithm; According to the adjusted pitch value, the pause value is extracted, and the correlation between the pause value and the pitch value is determined. If there is a conflict, the pause value is adjusted.

6. The method for generating personalized voice content according to claim 1, characterized in that: The process of dynamically adjusting the usage frequency of modal particles includes: Acquire text data, extract modal particle distribution information, and generate a modal particle distribution map; Extracting stress position parameters from speech data to generate stress parameter sets; Combine user preference data and usage scenario information to build an analysis model; According to the analysis model, determine whether there is a conflict between the distribution of modal particles and the stress position. If there is a conflict, use a dynamic adjustment algorithm to adjust the frequency value of the modal particles and generate adjusted frequency value data; Recalculate the voice table parameters according to the adjusted frequency value data to generate optimized voice table data; Generate the final speech expression fluency distribution diagram through the optimized speech table data; Obtain the naturalness index of the speech waveform and determine whether the naturalness is lower than a preset threshold. If the naturalness is lower than the preset threshold, a dynamic adjustment algorithm is used to reallocate parameter weights. According to the reallocated weights, the stability index of the speech waveform is optimized.

7. The method for generating personalized voice content according to claim 1, characterized in that: The process of generating the dynamically adjusted speech parameter model includes: Acquire speech data, extract pitch change features, and generate a pitch change data set; For the pitch change dataset, combined with the pause length parameter, the correlation between the pause length and the pitch change is analyzed. If there is a conflict between the pause length and the pitch change, a dynamic adjustment algorithm is used to adjust the pause length parameter. According to the adjusted pause length parameter, the stress position feature is extracted to generate a stress position data set; For the stress position data set, combined with the tone frequency parameter, determine the correlation between the stress position and the tone frequency. If there is a conflict between the stress position and the tone frequency, adjust the tone frequency parameter. generating a dynamically adjusted speech parameter model according to the adjusted tone frequency parameters; The dynamically adjusted speech parameter model is combined with the text through the speech synthesis engine to generate preliminary speech waveform data; The pitch, pause and stress features in the waveform are extracted to obtain the final speech parameter model.

8. The method for generating personalized voice content according to claim 1, characterized in that: The process of extracting the pitch, pause and stress features in the waveform includes: The dynamically adjusted speech parameter model is loaded through the speech synthesis engine and combined with the text data to be processed to generate preliminary speech waveform data; Extract pitch features from speech waveform data to obtain a pitch variation range; Extract pause features based on speech waveform data and determine the pause duration between each segment; The stress features are extracted from the speech waveform data and the stress position in each segment is determined.

9. The method for generating personalized voice content according to claim 1, characterized in that: The process of optimizing the stability of speech output includes: Acquire speech waveform data, extract naturalness index, and obtain naturalness value; Determine the difference between the naturalness value and a preset threshold value to determine whether it is lower than the preset threshold value. If the naturalness is lower than the preset threshold value, use a dynamic adjustment algorithm to reallocate the weight distribution in the parameter balancing algorithm. According to the reallocated weights, the stability index of the speech waveform data is optimized to obtain an optimized speech waveform; Extract the optimized speech waveform and determine whether its naturalness value meets the preset threshold. If the naturalness is still lower than the preset threshold, further adjust the parameter weight distribution. According to the final adjusted weight, new speech waveform data is generated and its naturalness value is extracted; The difference between the naturalness of the new speech waveform and the preset threshold is judged to determine the stability level of the speech output.

10. The method for generating personalized voice content according to claim 1, characterized in that: The process of generating the final personalized voice output includes: Obtain optimized speech waveform data and extract pitch, pause and stress features in the waveform; Based on the extracted pitch, pause and stress features, filters are used to remove high-frequency noise; The unstable fluctuations in the waveform are reduced by a smoothing algorithm to obtain a processed speech waveform; Extract the processed speech waveform and determine whether its naturalness meets the preset range. If the naturalness meets the requirements, generate the final personalized speech output; if the naturalness does not meet the requirements, use a dynamic adjustment algorithm to redistribute the weight distribution; According to the reallocated weights, the stability index of the speech waveform is optimized; Obtain the optimized speech waveform and use a filter to eliminate high-frequency noise; The smoothing algorithm is used to reduce unstable fluctuations in the waveform and generate the final personalized speech output.

Citation Information

Cited By

  • Intelligent chart dynamic generation method based on voice recognition and multi-modal interaction

    CN120670583A

  • Deep learning-based teaching voice naturalness optimization method and system

    CN121583236A