Voice generation method and device, equipment and medium

Through the joint encoding processing of characters and pinyin and the context-driven modeling of semantics, the problem of insufficient accuracy and personalized expression of the existing TTS system in polyphonic and ambiguity processing is solved, and speech cloning with higher fidelity and strong expression capabilities is achieved.

CN120279883APending Publication Date: 2025-07-08PING AN TECH (SHENZHEN) CO LTD

Patent Information

Application Number
CN202510643898.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-19
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

When dealing with ambiguity of financial-specific terms or institutional names, existing TTS systems have mispronunciation or semantic deviations in polyphonic pronunciation, which cannot effectively distinguish the polyphonic phenomenon in medical terms, and lack the complementary advantages of characters and pinyin information, resulting in inaccurate speech synthesis results and insufficient personalized expression.

Method used

By receiving the prompt speech, converting it into initial text and performing joint encoding processing of characters and pinyin, character pinyin fusion features are generated, and features are extracted using text encoder and speaker encoder, and target speech is generated by combining the generation adversarial decoder to generate target speech, so as to realize joint modeling of characters and pinyin and context semantic driving.

Benefits of technology

It improves the accuracy and personalized expression ability of pronunciation of polyphonic characters, reduces the fidelity and expression of speech cloning, and adapts to complex speech input scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279883A_ABST
    Figure CN120279883A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of speech synthesis, can be applied to business scenes such as financial science and technology, medical health and the like, and discloses a speech generation method, device, equipment and medium. And inputting a text speech language model to generate an intermediate code in combination with a speech code extracted based on a codebook generation mode, extracting a speaker feature vector in the prompt speech, decoding the intermediate code and the speaker feature vector by a generative adversarial decoder, and outputting a target speech. According to the method, fine-grained text representation is established by fusing character semantics and pinyin pronunciation information, and voice cloning is completed through unified voice language modeling and an adversarial generation mechanism in combination with codebook-driven acoustic coding and speaker personality characteristics, so that the naturalness, similarity and pronunciation accuracy of generated voices are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of speech synthesis, and particularly to a speech generation method, apparatus, device, and storage medium. Background Art

[0002] In current text-to-speech (TTS) systems, there are still multiple key technical aspects that cannot meet the actual requirements of complex business speech generation scenarios, especially showing obvious deficiencies in applications that require personalized expression and accurate semantic presentation.

[0003] In the field of fintech business, application scenarios such as intelligent customer service, voice risk control, and wealth management voice assistants are highly sensitive to the naturalness, pronunciation accuracy, and user personality expression of speech synthesis. However, when existing TTS systems process ambiguous financial terms or institutional names, they often make mispronunciations of polyphonic characters or semantic deviations due to the lack of context judgment ability. In addition, the insufficient voice similarity in voice cloning makes it difficult to unify the expression style of financial service voices in different role scenarios, affecting user trust and business consistency.

[0004] In the field of healthcare business, applications such as electronic medical record announcements, doctor-assisted voice interactions, and rehabilitation voice prompts have high requirements for the semantic accuracy and personality restoration ability of speech synthesis results. Existing TTS systems often cannot effectively distinguish the polyphonic phenomena in medical terms. For example, the pronunciations of words such as "xing" (walking / behavior) or "fa" (inflammation / occurrence) in different contexts may mislead patients in pathological announcements or voice interactions. At the same time, personalized voice cloning technology also faces the problem of rough speaker feature extraction in scenarios such as doctor-patient communication simulation and elderly care voice, resulting in the generated voice lacking warmth and realism.

[0005] From the perspective of the overall technical system, the existing speech synthesis system still has a relatively weak joint modeling mechanism between characters and pinyin, and cannot effectively utilize the complementary advantages of information between the character layer and the pinyin layer, resulting in insufficient expression ability when dealing with semantic ambiguity or long-tail words. At the same time, the quantization strategy used in the acoustic coding process is usually a fixed design, lacking the ability to dynamically select the compression method based on the speech content, resulting in a double bottleneck of compression loss and insufficient restoration in speech coding. In addition, the combination method of text semantics and speech coding generally uses simple splicing, lacking structured fusion modeling, and unable to fully capture the multi-modal collaborative features required for speech synthesis, affecting the naturalness and emotional expression of the final speech. Summary of the Invention

[0006] The main objective of the present invention is to provide a voice generation method, apparatus, device, and storage medium, aiming to solve the technical problems in the prior art that there is a lack of a context semantic-driven fusion mechanism in character and pinyin information modeling, resulting in incorrect pronunciation of polyphonic characters and inaccurate mapping of text semantics to voice expressions.

[0007] To achieve the above objective, the present invention provides a voice generation method, including:

[0008] Receiving a prompt voice containing target voice features;

[0009] Converting the prompt voice into an initial text through a voice recognition module;

[0010] Performing joint encoding processing of characters and pinyin on the initial text to generate a character-pinyin fusion feature;

[0011] Processing the character-pinyin fusion feature through a text encoder to generate a text feature;

[0012] Determining a codebook generation method and extracting a voice code in the prompt voice based on the codebook generation method;

[0013] Inputting the text feature and the voice code into a text-to-speech language model to generate an intermediate code;

[0014] Extracting a speaker feature vector in the prompt voice through a speaker encoder;

[0015] Performing decoding processing on the intermediate code and the speaker feature vector through a generative adversarial decoder to generate a target voice.

[0016] Furthermore, to achieve the above objective, the present invention provides a voice generation apparatus, including:

[0017] A voice input module, configured to receive a prompt voice containing target voice features;

[0018] A voice recognition module, configured to convert the prompt voice into an initial text through the voice recognition module;

[0019] A pinyin encoding module, configured to perform joint encoding processing of characters and pinyin on the initial text to generate a character-pinyin fusion feature;

[0020] A text encoding module, configured to process the character-pinyin fusion feature through a text encoder to generate a text feature;

[0021] A voice quantization module, configured to determine a codebook generation method and extract a voice code in the prompt voice based on the codebook generation method;

[0022] A multi-modal fusion module for inputting the text features and the voice encoding into a text-to-speech language model to generate an intermediate encoding;

[0023] A speaker analysis module for extracting a speaker feature vector from the prompt voice through a speaker encoder;

[0024] A speech synthesis module for decoding the intermediate encoding and the speaker feature vector through a generative adversarial decoder to generate a target voice.

[0025] Furthermore, to achieve the above object, the present invention also provides a computer device, which includes a memory, a processor, and a voice generation program stored on the memory and executable on the processor. When the voice generation program is executed by the processor, the steps of the voice generation method as described above are implemented.

[0026] Furthermore, to achieve the above object, the present invention also provides a computer-readable storage medium, on which a voice generation program is stored. When the voice generation program is executed by a processor, the steps of the voice generation method as described above are implemented.

[0027] Beneficial effects: The present invention relates to the technical field of speech synthesis and can be applied to business scenarios such as fintech and healthcare. A voice generation method is disclosed, including: receiving a prompt voice containing target voice features, converting the prompt voice into an initial text, performing joint encoding processing on the initial text for characters and pinyin to generate a character-pinyin fusion feature, processing the character-pinyin fusion feature with a text encoder to generate text features, determining a codebook generation method and extracting the voice encoding in the prompt voice, inputting the text features and the voice encoding into a text-to-speech language model to generate an intermediate encoding, extracting a speaker feature vector from the prompt voice through a speaker encoder, and decoding the intermediate encoding and the speaker feature vector through a generative adversarial decoder to generate a target voice. By constructing a joint encoding mechanism for characters and pinyin, modeling in combination with context semantics, and jointly inputting text features, voice encoding, and speaker features into a multi-modal language model and a generative decoder for processing, the present invention realizes voice cloning output with higher fidelity, stronger expressive ability, and lower data dependence. Description of the Drawings

[0028] The present invention will be further described below in conjunction with the drawings and embodiments. In the drawings:

[0029] Figure 1 is a schematic diagram of an application environment of the voice generation method in an embodiment of the present invention;

[0030] Figure 2 is a schematic flowchart of an embodiment of the voice generation method of the present invention;

[0031] Figure 3 Schematic diagram of functional modules of a preferred embodiment of the voice generation device of the present invention;

[0032] Figure 4 Schematic diagram of a structure of a computer device in an embodiment of the present invention;

[0033] Figure 5 Another schematic diagram of a structure of a computer device in an embodiment of the present invention. Detailed implementation manners

[0034] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0035] The voice generation method provided by the embodiment of the present invention can be applied to an application environment such as Figure 1 where the client communicates with the server through a network. The server can receive a prompt voice containing target voice features through the client, convert the prompt voice into an initial text, perform joint encoding processing of characters and pinyin on the initial text to generate a character-pinyin fusion feature, use a text encoder to process the character-pinyin fusion feature to generate a text feature, determine a codebook generation method and extract the voice encoding in the prompt voice, input the text feature and the voice encoding into a text-voice language model to generate an intermediate encoding, extract a speaker feature vector in the prompt voice through a speaker encoder, and decode the intermediate encoding and the speaker feature vector through a generative adversarial decoder to generate a target voice. The present invention realizes a voice cloning output with higher fidelity, stronger expressive ability and lower data dependence by constructing a joint encoding mechanism of characters and pinyin, modeling in combination with context semantics, and jointly inputting the text feature, the voice encoding and the speaker feature into a multi-modal language model and a generative decoder for processing. Among them, the client can be but is not limited to various personal computers, laptop computers, smart phones, tablet computers and portable wearable devices. The server can be implemented by an independent server or a server cluster composed of multiple servers. The present invention will be described in detail below through specific embodiments.

[0036] Please refer to Figure 2 , Figure 2 which is a schematic flowchart of an embodiment of the voice generation method provided by the present invention. It should be noted that although the logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than here.

[0037] As Figure 2 shown, the voice generation method proposed by the present invention includes the following steps:

[0038] S10, receiving a prompt voice containing target voice features;

[0039] In this embodiment, the operation of receiving the prompt voice is the starting point of the voice cloning process. Its purpose is to obtain the data entry for subsequent modeling under the condition of no text or known text but without associated target speaker features. The prompt voice is an audio data containing complete voice features, which must cover information such as the speaker's intonation, speech rate, voiceprint, pronunciation habits, etc. This audio can be a natural speech recording, a reading material, or other representative voice segments. The receiving behavior is not limited to the acquisition by local input devices such as audio sensors like microphones, and can also obtain online audio streams or stored voice files through remote data transfer protocols such as RTSP, MQTT, and HTTP API. To ensure that the collected voice data has sufficient trainability and structural feature integrity, a silent segment detection module is usually combined in the receiving stage to automatically remove non-speech segments, or an audio energy threshold judgment mechanism is used to retain only valid speech. The reception of voice signals is generally encapsulated in a standard format, and the current mainstream is the WAV, FLAC, or PCM encoding form. The sampling rate is recommended to be controlled within the range of 16 kHz to 48 kHz, and the quantization bit depth is not less than 16 bits to retain the spectral structure details. In addition, the received prompt voice also needs to be accompanied by meta-information such as timestamps, acquisition source identifiers, device models, and user IDs. These information can be used for conditional encoding, identity matching, or multi-speaker sample management during the training process. To improve the processing efficiency, the receiving process can also be combined with the pre-feature extraction module of the edge device. For example, a lightweight convolutional neural network is deployed to extract Mel spectrogram or short-time Fourier transform features, and feature pre-screening is completed in the receiving stage and the transmission band is compressed

[0040] The original voice signal of the user can be collected through the microphone module deployed on the edge terminal and encoded into a PCM format data stream in real time, and uploaded to the model server through the local bus or asynchronous task for subsequent processing. It is also possible to transmit the recorded prompt voice file to the backend based on the HTTP POST protocol. This file can be sourced from an existing voice library or temporary recordings. When deployed in a low-resource scenario, a low-power microphone array can be used, and the audio reception process can be dynamically started and stopped through the voice activity detection module. In an environment where multiple users access simultaneously, each prompt voice needs to be accompanied by a unique user identifier and acquisition time information, and a mapping relationship with the speaker features is established through the identity identification mechanism for use in the subsequent speaker modeling module. In addition, in scenarios where noise resistance needs to be improved, a noise reduction pre-processing module can be added before reception, such as a voice enhancement network based on spectral masking or spatial filtering technology, so as to improve the quality of the voice signal and reduce subsequent recognition and cloning errors.

[0041] Example illustration: In the field of medical and health services, medical staff can input their own voices as prompt voices through wearable terminals, and the system automatically captures their voice features for subsequent synthesis of personalized voice prompt content, such as customized rehabilitation plan announcements, medication reminders, or nursing guidance announcements, to improve patient acceptance and interaction naturalness.

[0042] In the field of fintech services, the prompt voices of bank customer managers or insurance advisors are input through the CRM system. After the system extracts their voice features, they can be used to generate automated voice responses or customer follow-up voices with a consistent style, reducing the cost of manual recording and maintaining the voice interaction experience familiar to customers. In the field of manufacturing voice interaction devices, manufacturers can achieve multi-role switching of voice assistants in terminal devices by batch importing prompt voices, thus meeting the diverse needs of different people for voice individuality.

[0043] By receiving the prompt voice containing the complete target voice features, the subsequent text recognition, feature fusion, and voice generation modules are provided with the basic corpus for constructing the expression pattern of the target speaker, thus establishing a controllable pronunciation synthesis channel for arbitrary voice content. It not only provides a personalized voice sample entry for the voice cloning system, but also standardizes the audio format, sampling parameters, and voice validity through the receiving module, improving the overall system's modeling accuracy and cross-scenario deployment ability. It effectively solves the bottleneck in the existing system that high-fidelity voice cloning cannot be achieved in the absence of high-quality pronunciation samples.

[0044] S20, convert the prompt voice into initial text through a voice recognition module;

[0045] In this embodiment, the process of converting the prompt speech into the initial text belongs to the key link of mapping from acoustic signals to discrete language symbols. Its core goal is to restore the word information carried in the prompt speech as the language basis for subsequent text feature modeling and speech synthesis. After receiving and preprocessing, the prompt speech must first be performed through an end-to-end speech recognition module to perform acoustic modeling and decoding operations. This module generally adopts a structure based on a deep neural network, including an input layer, an acoustic feature extraction layer, an encoder layer, an attention alignment module, and an output decoder. The input is a feature representation of the time domain waveform or frequency domain transformation of the prompt speech, such as an acoustic feature matrix such as Mel spectrum, MFCC or FBANK. The encoder can use structures such as bidirectional LSTM, Conformer or Transformer to model the time correlation and context dependency of speech. The attention mechanism or CTC connection timing prediction structure is used to map speech segments to text segments, thereby identifying text information without relying on precise alignment. The output layer uses a language modeling module combined with a decoding strategy (such as beam search) to generate an initial text result, which maintains the pronunciation order in the original speech and retains the language expression structure of polyphones, long-tail words, and sound change phenomena. To improve recognition accuracy, pinyin modeling, external language model-assisted correction and contextual vocabulary constraint strategies can be integrated to ensure that the speech recognition results can restore the speaker's intentions to the greatest extent possible.

[0046] In one specific implementation, the speech recognition module adopts an end-to-end deep model structure, and the acoustic model based on Conformer is responsible for extracting time and spectrum features, while combining with an externally trained language model to assist in predicting word sequences. The input prompt voice is sent to the Conformer encoder after MFCC processing, and the character-level output sequence is generated after alignment by the attention module, and restored to Chinese text through dictionary mapping. In another implementation, the speech recognition module can adopt the CTC decoding method to avoid explicit modeling alignment path, so that the recognition system has stronger robustness when facing unknown speech speed or sentence punctuation habits. For scenarios containing professional terms or user-customized corpora, a word graph search mechanism can be introduced in the decoding stage to allow multiple recognition paths to be inferred concurrently, and the optimal recognition result can be selected through confidence scoring. It is also possible to execute a word order correction strategy based on pinyin reversal after the recognition result is generated. For example, in the scenario where the incorrect output of "Bank" is "Line Line", the correct text is restored through pinyin alignment and language model scoring. The multi-module serial structure of the recognition process supports deployment on the server side for large-scale reasoning, and is also adapted to lightweight terminal models for local fast recognition.

[0047] Example illustration: In the field of healthcare business, doctors often describe the diagnosis process or postoperative suggestions through voice input. These voice messages have extremely strong individual characteristics and a high density of professional terms. To achieve precise cloning of the doctor's voice, this system first collects their prompt voice and converts it into an initial text with complete semantics and alignable voice features through a voice recognition module. This initial text not only serves as the language index for subsequent text-to-speech mapping but also retains the specificity of medical professional terms and language structures, enabling the synthesized voice to accurately reproduce the speaker's language style and professional expressions, thus serving the personalized voice synthesis needs in doctor-patient interaction and remote consultation.

[0048] In the field of fintech business, when financial advisors or account managers explain products to clients, their voices often contain information expressions closely related to interest rates, terms, risk preferences, etc. After receiving their prompt voice, this system first accurately identifies it as the initial text to retain the semantic features in the language that reflect professionalism and trustworthiness and lay a precise language foundation for subsequent text feature modeling and voice synthesis. This conversion step not only makes the voice content more structured but also ensures that the original speaker's expression logic and intonation style are maintained in the synthesized voice, achieving a real and controllable voice cloning output and enhancing voice individuality and trust in scenarios such as financial remote services and voice risk control broadcasts.

[0049] By converting the prompt voice into the initial text, not only is a language foundation established for the joint encoding processing of characters and pinyin, but also the problems of information redundancy and alignment difficulties faced when directly modeling text semantics from acoustic features are avoided. This method converts the original continuous signal into discrete language units, thus significantly improving the accuracy of subsequent semantic modeling, pinyin ambiguity resolution, and text feature generation. Especially when facing prompt voices with polyphonic characters, long-tail words, or large intonation variations, the semantic integrity and controllability of text restoration can be improved through context modeling and language correction mechanisms, enhancing the system's adaptability to complex voice inputs.

[0050] S30, perform joint encoding processing on the initial text to generate character-pinyin fusion features;

[0051] In the present embodiment, the initial text is a language content in text form output by the speech recognition module, which contains a continuous sequence of characters, but does not carry clear speech pronunciation information itself. In order to build a speech cloning system with contextual pronunciation control capabilities, it is necessary to extract a composite expression that has both text semantics and pronunciation attributes from the initial text. The "joint coding process of characters and pinyin" in this step is aimed at this. Characters are the basic units that constitute the language content, and pinyin is the phoneme-level pronunciation code that characterizes Chinese speech. The essence of the joint processing of the two is to fuse semantics with phonological information, thereby enhancing the model's modeling capabilities for polyphones, long-tail words, weak syllables, and other phenomena.

[0052] The joint encoding process usually starts with the establishment of a character-pinyin mapping relationship library, which comes from a many-to-one mapping resource based on the Chinese Pinyin standard, which can be built through open source Pinyin annotation tools such as pypinyin. For each character, the system extracts possible pronunciations based on the context, and then uses the language model to determine the most reasonable pronunciation combination in the context. For example, the word "行" is pronounced "háng" in "银行" and "xíng" in "走". The system needs to perform dynamic pronunciation selection within the sentence-level context.

[0053] Next, each Chinese character is converted into a dense character encoding vector through the character embedding layer, usually with a dimension ranging from 128 to 512. Simultaneously, the pinyin embedding layer converts the pinyin corresponding to each character into a pronunciation vector representation, which may contain sub-components such as initials, finals, and tones, or it can be directly encoded as an overall pinyin representation. In order to fuse these two embedded representations, the system usually designs a weighted fusion module to dynamically assign weight coefficients of characters and pinyin in the joint representation based on contextual semantic information. This weight can be either linearly predicted by semantic features or automatically learned by the attention mechanism.

[0054] Finally, the character code and the pinyin code perform a linear weighted operation according to the fusion weight, output a single fusion vector, and input it into the subsequent text encoder. This fusion feature not only has the ability to express the semantic dimension of the text, but also explicitly models the pronunciation attributes, providing control variables for subsequent pronunciation decisions and speech synthesis. The technical key in this process lies in the context-aware fusion weight calculation method and the collaborative modeling ability of semantic-phonological information.

[0055] In one embodiment, the joint encoding of characters and pinyin uses a pre-trained Chinese text embedding model as the source of character encoding, combined with a custom-built pinyin embedding table, and the pinyin embedding can use a one-hot vector or a low-dimensional dense representation. The contextual semantics is extracted by a lightweight Transformer model, and the contextual attention weight corresponding to each character is output as a fusion weight reference.

[0056] The contribution ratio of character encoding to pinyin encoding can also be dynamically adjusted by introducing a trainable gating mechanism. For example, a sigmoid function is set as the gating function, which is input by the context vector and outputs a scalar in the range of [0,1] as the retention factor of pinyin information, and then the factor is weightedly fused with the character encoding and pinyin encoding.

[0057] In another implementation, the joint coding structure can also support the pinyin decomposition mechanism, that is, splitting each pinyin into initial consonants, finals and tones and encoding them separately, and then combining them through convolution or attention mechanisms to enhance the system's ability to parse complex pronunciations, which is especially suitable for modeling problems with blurred syllable boundaries in dialect scenarios.

[0058] Example description: In medical and health scenarios, doctors often use terms such as "blood pressure", "heart rate" and professional abbreviations such as "MRI", "CT", etc., and their pronunciation may change due to changes in context. Through the joint encoding of characters and pinyin, not only can "blood pressure" be accurately marked as "xuèyā" instead of the common mispronunciation "xuěyā", but the inter-word rhythm and speech flow characteristics can also be retained in combination with the pinyin sequence, making the generated speech synthesis results more professional and credible in medical expression.

[0059] In the field of financial technology business, account managers use terms such as "interest rate", "annualization", and "bond type" when explaining product solutions. The word "rate" has two pronunciations: "lǜ" and "shuài". The joint encoding mechanism can accurately determine the use of "lǜ" based on the complete context, avoiding misreading of professional terms as pronunciations that do not match the semantics, while maintaining the consistency of intonation in the entire speech, thereby ensuring the quality of audio broadcasting and improving customer trust and product understanding.

[0060] Through the fusion of characters and pinyin, not only the semantic integrity of the original text is retained, but also Chinese pronunciation information is explicitly introduced, thereby enhancing the model's ability to control pronunciation. Especially in the Chinese language environment where polyphones are frequent and pinyin ambiguities are more, this processing method can effectively reduce pronunciation errors, significantly improving the naturalness and accuracy of voice cloning. It also provides downstream text encoders with richer information dimensions, which helps to improve the emotional expressiveness and consistency of the overall synthesized speech.

[0061] S40, processing the character and pinyin fusion features through a text encoder to generate text features;

[0062] In this embodiment, the character-pinyin fusion feature is a multi-modal semantic pronunciation expression formed by combining the semantic information of characters and the pronunciation information of pinyin in the initial text. This feature does not yet have a unified semantic encoding form for the speech generation model. Therefore, it is necessary to map this fusion feature into a fixed-dimensional text feature that can express sentence semantics, language structure, and context dependence through a text encoder to support subsequent speech encoding matching and speech generation tasks.

[0063] The text encoder is usually constructed based on the Transformer architecture in this processing flow and has strong context modeling capabilities. This module first receives the character-pinyin fusion feature as input, which has integrated the semantic expression of the text and the pronunciation characteristics of pinyin through joint embedding. It is usually a two-dimensional matrix, with rows representing the character sequence positions and columns representing the dimensions of each fusion vector.

[0064] In the input stage, the text encoder introduces positional encoding information through the embedding layer, enabling the model to perceive the relative order between characters and solving the problem of lack of perception of order information in sequence modeling. Subsequently, the correlation between each character site and other characters globally is processed through multiple layers of multi-head self-attention mechanisms, and the system can learn complex language features such as word sense disambiguation, context mapping, and emotion modification. The multi-head attention mechanism allows the model to extract association relationships in parallel from multiple subspaces, enhancing the modeling ability for the implicit hierarchical structure in the language modality.

[0065] After the attention processing, the text encoder also includes a feed-forward neural network layer for non-linear transformation and dimension reconstruction of the encoding at each position. This layer usually adopts a two-layer structure. The first layer increases the dimension for feature expansion, and the second layer reduces the dimension to return to the target feature dimension, and combines activation functions such as GELU to achieve complex feature transformation.

[0066] During the entire encoding process, each layer is connected to a residual connection and a layer normalization mechanism to keep the network gradient stable and improve the training efficiency. The final output is a set of text feature vectors with a unified dimension, which not only retains the sound-meaning combination of the character-pinyin fusion feature but also incorporates the context structure expression of the language, providing a stable input for the subsequent multi-modal modeling module.

[0067] In one implementation, the text encoder is fine-tuned based on the BERT structure. After inputting the character-pinyin fusion feature, it extracts layer-by-layer representations for each character position through the multi-head attention and feed-forward network layers in the Transformer module, and outputs a text feature vector with a dimension of 768. This method can utilize the weight transfer ability of existing large-scale corpus training to improve the generalization performance in few-shot tasks.

[0068] It is also possible to construct a lightweight Transformer structure. For example, a small model architecture with 4 layers of encoders, 4 heads of attention, and a hidden dimension of 256 can be adopted to adapt to the deployment requirements of embedded speech synthesis in resource-constrained scenarios. At the same time, a residual suppression module based on a gating mechanism can be introduced, and a gating unit is connected after the attention layer to control the retention degree of different types of semantic paths, which is suitable for the expression of low-frequency words, long words, or rare pinyin combinations in professional fields.

[0069] In another implementation, the text encoder can adopt a non-Transformer structure, such as a bidirectional LSTM combined with a convolutional enhancement module, to perform temporal modeling on the character-pinyin fusion features. This method is relatively weaker in modeling long-range dependencies in the sequence than the Transformer, but has a speed advantage in extremely low-latency inference systems.

[0070] Example illustration: In the scenario of medical and health data broadcasting, the text may contain sensitive or key information such as "abnormal body temperature" and "blood sugar fluctuations". Since the word "abnormal" may express different tones such as warning, suggestion, or statement in different sentence patterns, the text encoder can capture potential variables such as the strength of the tone and semantic logic through comprehensive modeling of the context semantics, thereby providing emotional expression clues for the speech synthesis module and improving the accuracy and affinity of the speech broadcast.

[0071] In the field of fintech business, professional terms such as "annualized rate of return" and "risk exposure" may appear in the text of customer investment instructions. If simple character input is used, it often causes pronunciation ambiguity of the word "rate" or unbalanced sentence rhythm. By modeling the full-sentence logical structure and professional semantics on the basis of the character-pinyin fusion features, the text encoder ensures the correct expression of professional terms and coherent speech output, thus significantly improving the trust and compliance in applications such as customer compliance notification and risk warning voice push.

[0072] By processing the character-pinyin fusion features through the text encoder, not only are the input features converted into a representation vector of a unified length and dimension, but also the language context structure, syntactic dependency information, and pronunciation difference characteristics between words are retained. This conversion effectively solves the problem that the character-pinyin fusion features do not have controllable modeling ability in the original state, enabling the subsequent speech generation system to still maintain clear pronunciation and consistent semantics when facing scenarios of polyphonic words, long-tailed words, and context ambiguity. This processing method also improves the generalization ability of the model on new words and combined words, enhancing the naturalness and expressiveness of the overall speech output.

[0073] S50, determine the codebook generation method, and extract the speech encoding in the prompt speech based on the codebook generation method;

[0074] In this embodiment, the codebook generation method is used to specify the method for compressing and encoding the acoustic speech markers in the prompt speech. The core lies in providing a stable and low-redundancy speech coding expression for subsequent generation tasks. A codebook is a collection of discrete speech units, representing several cluster centers in space. The input acoustic features are quantized into the closest codebook vectors to represent the audio content.

[0075] In this process, it is first necessary to evaluate the optional quantization methods to determine the optimal codebook generation method. The two commonly used methods are vector quantization and finite scalar quantization. Vector quantization is based on clustering strategies (such as k-means or Product Quantization), encoding with multiple dimensions as a whole unit, and has the advantage of maintaining global similarity; finite scalar quantization quantizes each dimension of the feature separately, with a more concise structure and faster encoding speed.

[0076] To evaluate the effects of these methods, it is a necessary step to construct a test data set containing acoustic speech marker samples. This data set can extract features from a large amount of speech data by calling a pre-trained acoustic front-end network, such as Mel spectrogram, cepstral parameters, or continuous features output by Conformer.

[0077] In a specific implementation, the test data generates two codebooks through vector quantization and finite scalar quantization methods respectively, and calculates the compression rate and reconstruction error of each codebook for the test data. The compression rate refers to the ratio of the length of the encoded data to the length of the original features. The reconstruction error is commonly expressed by the mean square error (MSE) or a perceptual distance metric (such as Mel-spectrogram distance).

[0078] The system obtains the performance score of each quantization method based on the weighted combination of these two metrics, so as to select the best codebook generation method. After the selected method, the same quantization and encoding process is performed on the acoustic speech markers extracted from the prompt speech, and finally the speech coding is obtained as the audio representation input of the subsequent generation model.

[0079] In one implementation, the selection of the codebook generation method is achieved through the VQ-VAE structure. The acoustic features in the training set are input into the encoder module, and k-means is used for clustering to obtain a codebook of fixed length. During training, the quantization boundaries are optimized based on the minimum reconstruction error criterion. This method is suitable for generation tasks that require high fidelity for audio content, such as medical information broadcasting.

[0080] The FSQ (Finite Scalar Quantization) method can also be used. Discrete value sets are set for each dimension of the acoustic features respectively, and dimensional discretization is performed one by one according to the minimum distance principle. This method has more advantages in low-resource deployment scenarios, such as a local voice cloning system running on a mobile terminal.

[0081] A dynamic selection strategy can also be introduced. By recording the performance of the two quantization methods for different categories of audio in real time during the training phase, a more matching codebook generation method can be selected according to the timbre characteristics of the input prompt speech during inference, so as to reduce the model complexity while ensuring the quality.

[0082] Example: In the medical and health business scenario, when the system needs to push personalized voice reports to patients, such as "The blood pressure value has fluctuated greatly recently. Please pay attention to rest", the tone, rhythm and tension in the original speech need to be retained as much as possible. By encoding the speech through the vector quantization method, more context structure features can be retained, making the generated speech have a higher emotional expression degree and enhancing the perceived value of the voice broadcast for patients.

[0083] In the field of fintech, such as automatically generating annual income reports or risk warning broadcasts by voice, the system needs to quickly process multiple audio files and perform voice cloning. At this time, by compressing the speech encoding through the finite scalar quantization method, the encoding efficiency can be greatly improved, reducing the server load while ensuring clarity, which helps the concurrent execution of large-scale voice personalization services.

[0084] By selecting the codebook generation method and encoding the prompt speech based on the optimal method, the compactness and fidelity of the speech encoding are significantly improved. This strategy not only compresses the data volume of the speech representation, but also reduces the computational burden of the downstream decoder, while keeping the timbre clear and details complete in speech synthesis. In addition, when facing complex audio content or speaker changes, this method has stronger adaptability, improving the voice cloning stability of the system in multi-speaker and multi-situation scenarios.

[0085] S60, input the text feature and the speech encoding into a text-to-speech language model to generate an intermediate encoding;

[0086] In this embodiment, the text-to-speech language model is a multi-modal deep network structure that integrates text semantic information and speech feature information. Its main task is to jointly model the input text feature and speech encoding to generate an intermediate encoding that expresses semantic content and carries pronunciation instructions. This intermediate encoding not only carries the information of character semantics and pronunciation order, but also provides a clear and structured speech generation guiding signal for speech synthesis in the decoding stage.

[0087] In actual processing, the text feature is the result of modeling the character-pinyin fusion feature by a text encoder, usually a sequence of context-related vectors with a fixed dimension, while the speech encoding is a discrete speech representation compressed and extracted from the prompt speech, reflecting audio information such as pronunciation style, rhythm and speech rate. The purpose of jointly inputting these two types of features into the language model is to establish a mapping relationship between pronunciation control and semantic accuracy.

[0088] The combined input method can be accomplished through mechanisms such as simple vector concatenation, weighted fusion, and attention alignment. In the embedding stage, linear transformations can be performed on the text features and speech encodings respectively to match their dimensions and then perform joint processing. Subsequently, a multi-layer Transformer structure or a lightweight multi-modal attention module is used to model the interaction between the two, and the positional encoding mechanism is used to maintain the traceability of the text sequence structure.

[0089] The generated intermediate encoding is a sequence of multi-dimensional vectors, where each time step corresponds to a speech frame or sub-word level speech unit, and elements such as semantics, tone, speaking style, and language rhythm are fused in the encoding. This encoding will be used as the input for the subsequent speech generation decoder, directly determining the controllability and personalized effect of the generated speech.

[0090] In one implementation, the text-to-speech language model adopts a Transformer architecture, where the text feature input is mapped to the same vector dimension as the speech encoding through a linear layer, and the speech encoding is promoted to a learnable continuous representation through an embedding layer. After the two are fused, they are fed into a multi-layer self-attention module to perform cross-modal information alignment.

[0091] A two-stream structure can also be adopted, where the text features and speech encodings are respectively input into two branches, and cross-attention mechanisms are used for joint modeling, and finally an intermediate encoding sequence is generated in the fusion layer. This structure is suitable for scenarios where the speech encoding dimension is low but contains tone changes, such as generating emotional response speech.

[0092] In a deployment-constrained environment, a lightweight attention gating structure can also be used to perform the intervention of the speech encoding on the text only at key positions, so as to reduce the number of model parameters while maintaining the individual expression ability of the tone color.

[0093] By inputting text features and speech encodings into a language model to generate intermediate encodings, a mapping channel between semantics and speech is constructed, greatly improving the comprehensive control ability of semantic accuracy and tone color restoration in the speech synthesis stage. It enables the system to output intermediate encodings with clear structure, natural rhythm, and consistent speaking style on the premise of only providing prompt speech and text, laying a foundation for high-quality voice cloning, especially maintaining pronunciation coherence and speaker feature consistency even under zero-shot conditions.

[0094] Example illustration: In the field of medical and health, when a doctor inputs the text "Your average blood sugar level this month is slightly higher than the normal value", the system needs to use a synthesized voice with a gentle tone and friendly intonation for broadcasting. The intermediate encoding plays a bridging role of "understanding semantic content and generating appropriate speech expressions" in this process, ensuring that the generated voice not only retains the doctor's pronunciation style but also does not lose semantic clarity.

[0095] In the field of fintech business, when the system automatically generates voice prompts such as "You are approaching the credit card repayment date. Please handle it in time", through an intermediate coding control mechanism, the emphasized words (such as "approaching", "please handle it in time") can be slightly emphasized in intonation, making the reminder have an appropriate sense of urgency. This tone regulation ability of speech synthesis stems from the effective integration of text semantics and speech coding, making the speech output not only accurate but also expressive.

[0096] S70, extract the speaker feature vector in the prompt voice through a speaker encoder;

[0097] In this embodiment, the extraction process of the speaker feature vector is centered around the prompt voice. The goal is to accurately model the speaker's personality characteristics contained in the prompt voice, so that the subsequent speech generation can present parameter characteristics such as timbre, pronunciation method, and language rhythm that are highly consistent with the speaker. This process relies on the speaker encoder to complete. The encoder needs to have strong expressive power and temporal modeling ability in its structural design, and be able to capture long-term pronunciation characteristics rather than short-term speech content from the original or preprocessed speech.

[0098] The prompt voice usually needs to perform standard preprocessing operations before being input into the encoder, including framing, windowing, spectral analysis, or filter bank transformation (such as Mel spectrum extraction), to generate a structured sequence of speech frames. These frame sequences are input into the first-stage module of the speaker encoder, such as a convolutional neural network layer, to extract local acoustic features from the time-frequency dimension, including indicators such as formant positions, spectral envelopes, and short-term energies, forming a low-order perceptual basis of the speech.

[0099] The encoder further models the long-term dependencies between speech frames through a multi-head self-attention mechanism, especially forming an encoded representation of the speaker's stability characteristics in different dimensions such as pronunciation rhythm, intensity distribution, and accent changes. The parallel path structure of the attention module helps to capture the pronunciation commonalities across time frames, thus suppressing the interference caused by short-term variations in speech content.

[0100] The local acoustic features and global dependency features are fused through a residual connection module, retaining local details while enhancing global consistency. Finally, the fused features are compressed into a fixed-dimensional speaker feature vector through a fully connected network. This feature representation has reconstructability, supporting the subsequent decoder to generate the target voice with a timbre style based on it.

[0101] In one implementation, the prompt speech is processed into a sequence of Mel-spectrogram feature frames and then input into the speaker encoder, which uses a hybrid modeling method based on the Conformer structure. The convolution module is responsible for local feature extraction, the self-attention module handles long-distance dependencies, and each layer output is passed upward through the residual path. The final output sequence features are compressed into a vector through time average pooling, enter the multi-layer perceptron to complete dimensionality reduction and output a 128-dimensional speaker feature vector.

[0102] In scenarios where speech data changes dramatically or the audio noise background is complex, a multi-scale convolution module can be used to replace the traditional convolution path, so that the encoder can extract redundant features in different time windows to improve robustness; or a dual-branch structure can be introduced, one for pronunciation rhythm and the other for intonation, which are finally spliced ​​and merged into a speaker feature vector.

[0103] A self-supervised training framework can also be introduced to use speaker similarity discrimination loss to train the encoder, thereby improving its ability to distinguish speaker features and supporting continuous optimization on unlabeled speech data.

[0104] Example description: In the field of medical health, when developing a speech reconstruction system for aphasia patients, you can collect the patient's past recordings as prompt voices, extract the speaker's feature vectors, and use them to clone their voices for use in human-computer interaction systems. This allows the patient to use "his own voice" through the system even if he cannot speak in person, thereby maintaining the continuity of his identity expression.

[0105] In the financial technology business field, when the system needs to send voice information such as "account change notification" and "credit reminder" in the user's accustomed tone or the customer service representative's fixed voice style, the speaker encoder can be used to extract the prompt voice features of the specified employee or user, so that each voice output maintains sound consistency, thereby enhancing the user's trust and intimacy with the service.

[0106] By performing speaker feature extraction on the prompt voice, the personal features such as timbre, pronunciation habits, and intonation changes implicit in the original audio are efficiently compressed into a structured vector representation, thereby providing strong personalized control capabilities for the subsequent voice cloning process. The existence of this vector enables the system to replicate the expression style of a specific speaker even without multiple rounds of samples, improve the individual consistency and naturalness of the synthesized voice, and meet the high requirements for timbre restoration in actual scenarios.

[0107] S80, decoding the intermediate code and the speaker feature vector by generating an adversarial decoder to generate a target speech.

[0108] In this embodiment, during the process of generating the target speech, the generative adversarial decoder undertakes the key mapping task from the representation layer to the perceptual layer. The input of this stage consists of two core vectors, namely the intermediate encoding and the speaker feature vector. The intermediate encoding mainly carries the semantic and rhythmic information of the synthetic sentence, reflecting "what to say"; while the speaker feature vector defines "who is speaking" and "how to speak", reflecting individual timbre, pronunciation style and intonation characteristics.

[0109] To achieve information fusion, first, these two vectors are concatenated to construct a unified decoder input feature. The concatenation operation not only realizes information juxtaposition in the vector dimension, but also ensures that semantic information and personalized features are synchronously encoded in the time dimension, thus providing sufficient conditions for the generation stage. The generator part of the generative adversarial decoder receives this concatenated vector and performs a forward calculation process to map the high-dimensional semantic representation into a playable time-domain speech signal.

[0110] The architecture of the generator generally adopts a deep convolutional residual network and combines an upsampling mechanism to gradually reconstruct the low-resolution speech encoding into a complete waveform. Technologies such as noise-aware gating and frequency-domain feature fusion channels may be introduced inside to ensure the balance of speech in terms of naturalness, clarity and expressiveness. At the same time, to ensure the perceptual quality of the output speech, the training of the generator needs to refer to discriminative feedback at multiple scales, so that the distribution of the generated speech gradually approaches the statistical characteristics of real speech samples.

[0111] The output of the generative adversarial decoder is a continuous time-domain speech waveform, which has complete audible speech content and personalized pronunciation style. This process no longer depends on the phase reconstruction algorithm or the post-processing module, but directly outputs a terminal audio stream that can be transmitted and played.

[0112] In one implementation, the generator network is designed based on the BigVGAN structure, and its input is a 256-dimensional vector sequence formed by concatenating the intermediate encoding and the speaker feature vector. This sequence is decoded layer by layer through multiple one-dimensional convolutions, residual connections and LeakyReLU activation functions, and finally restored to an audio waveform with variable length through the upsampling module and the reconstruction module at the tail of the decoder. In the training stage, a multi-scale discriminator is used to provide feedback from dimensions such as the time domain, frequency domain and perceptual spectrum, so that the synthesized speech not only sounds close to real speech, but also has a highly restorative ability in terms of spectral envelope and transition details.

[0113] In another implementation, the decoder can be set as a dual-channel path. One path controls the backbone structure based on semantic rhythm encoding, and the other path controls the fine-tuning path such as formant center of gravity and F0 fundamental frequency by the speaker vector. The dual-channel outputs are fused at the end to ensure the independent regulation ability of content and timbre, and improve the cross-speaker generalization effect.

[0114] The parameter configuration of the decoder can also be adjusted according to different deployment platforms (such as mobile devices or edge servers). For example, reducing the depth of the convolutional layer and replacing the upsampling mechanism with linear interpolation can reduce the computational load while maintaining the generation quality.

[0115] Example illustration: In modern voice cloning technology, accurately restoring the voice characteristics of the target speaker and being able to accurately imitate their intonation, emotional expression, and personalized pronunciation features is a challenging task.

[0116] In the financial field, voice synthesis technology has been widely applied in scenarios such as intelligent customer service, voice banking, and customer service. In traditional voice synthesis systems, the generated voice often lacks natural intonation and emotional expression, resulting in a poor customer experience, especially in scenarios that require personalized interaction such as voice customer service. By accurately extracting the timbre and voice characteristics of the target speaker and through high-precision voice synthesis, it is possible to provide more efficient and personalized voice services for financial institutions. In the specific implementation process, after the financial institution receives the customer's input through the voice recognition module and converts it into the initial text, the technical solution generates the character-pinyin fusion features to accurately simulate the customer's accent, intonation, speech rate, and other voice characteristics. During this process, the pre-trained language model embedding layer converts this text information into high-dimensional semantic vectors, which are further deeply processed by the text encoder to ensure that the customer's tone and emotion can be accurately conveyed during the voice synthesis process. Through the combination of the generative adversarial decoder and the speaker feature vector, not only the cloning of the customer's voice is achieved, but also a more natural and fluent voice interaction can be provided.

[0117] In the field of healthcare, the application of intelligent speech systems is developing rapidly. Especially in scenarios such as telemedicine, intelligent health consultations, and patient education, personalized speech synthesis technology can greatly enhance the patient experience and the accuracy of information transmission. However, traditional speech synthesis technology still has significant limitations when dealing with complex medical terms, professional knowledge, and emotional expressions. Through precise voice cloning technology, more expressive and natural speech synthesis services can be provided in the healthcare field, especially with significant advantages for patient voice interaction and the interpretation of medical data. In specific implementations, conversations between patients and doctors, the interpretation of medical reports, etc. are all completed through high-quality speech generation technology. First, the patient's questions or medical record entries are converted into initial text by a speech recognition module. The system processes them through a combined encoding of characters and pinyin to ensure accurate handling of medical terms and polyphonic characters. The pre-trained language model embedding layer generates high-dimensional semantic vectors to ensure the capture of semantic information and context dependencies in the professional field. Then, through the multi-head self-attention mechanism of the text encoder, the system can accurately generate speech features containing unique information in the medical field. Finally, through the combination of a generative adversarial decoder and the speaker feature vector, the generated speech can not only accurately express professional knowledge but also simulate the speech features of doctors or patients, making the speech more realistic and personalized.

[0118] By inputting the intermediate encoding and the speaker feature vector into the generative adversarial decoder, a personalized, semantically complete high-fidelity speech waveform is directly output, achieving the goal of end-to-end voice cloning. This processing method eliminates the redundant vocoder module and post-processing operations in traditional TTS, shortens the inference link, and improves the speech generation speed and overall expressiveness. At the same time, since the speaker feature vector is introduced as a control condition, the system can also maintain a stable tone output under zero-shot conditions, achieving a faithful reproduction of voice personality.

[0119] The present invention relates to the field of speech synthesis technology, which can be applied to business scenarios such as financial technology and medical health, and discloses a speech generation method, including: receiving a prompt speech containing target speech features, converting the prompt speech into an initial text, performing a joint encoding process of characters and pinyin on the initial text, generating character pinyin fusion features, using a text encoder to process the character pinyin fusion features to generate text features, determining a codebook generation method and extracting speech encoding in the prompt speech, inputting text features and speech encoding into a text speech language model to generate an intermediate encoding, extracting a speaker feature vector in the prompt speech through a speaker encoder, and decoding the intermediate encoding and speaker feature vector through a generative adversarial decoder to generate a target speech. The present invention constructs a joint encoding mechanism of characters and pinyin, combines context semantics for modeling, and inputs text features, speech encoding, and speaker features into a multimodal language model and a generative decoder for processing, thereby achieving speech cloning output with higher fidelity, stronger expressiveness, and lower data dependence.

[0120] In one embodiment, the above step S30 includes:

[0121] S301, obtaining at least one pinyin corresponding to each character in the initial text from a preset character-pinyin mapping relationship library;

[0122] S302, converting characters in the initial text into character codes through a character embedding layer;

[0123] S303, converting the pinyin corresponding to the characters in the initial text into pinyin codes through a pinyin embedding layer;

[0124] S304, extracting contextual semantic features of the character encoding and the phonetic encoding through a pre-trained language model, and determining a fusion weight of the character encoding and the phonetic encoding;

[0125] S305, performing linear weighted processing on the character code and the pinyin code of each character according to the fusion weight to generate a character-pinyin fusion feature.

[0126] In this embodiment, the voice restoration capability of text input is enhanced, especially in a language environment where polyphones and semantic ambiguity are frequent (such as Chinese), to construct a more controllable and accurate text representation. The first operation is to obtain the pinyin of each character in the initial text from the character-pinyin mapping relationship library. The relationship library pre-sorts the one-to-many or many-to-one mapping rules between characters and pinyin. It can adopt the pinyin rule system based on the National Language Commission standards, or it can be adaptively adjusted according to the context. For example, the pinyin of "行" in different semantic structures should be distinguished as "xíng" or "háng". This mapping behavior can be dynamically inferred based on part-of-speech tagging or dependency syntax analysis models.

[0127] Next, each character is sent to the character embedding layer to generate character encodings. The character embedding layer converts the discrete character inputs into real-valued vectors of a fixed dimension through a look-up table mechanism or a neural network mapping mechanism. The vectors can represent the statistical co-occurrence attributes and semantic field positions of the characters. The embedding space can be constructed based on the Token Embedding method in Skip-gram or Transformer.

[0128] In parallel, the pinyin embedding layer also converts the obtained pinyin items into vector form. As prior phonetic information, pinyin often contains components such as initials, finals, and tones. This layer can perform a decomposed encoding of the pinyin structure (e.g., three-way split embedding of initials-finals-tones), and then map it to a unified embedding space through a multi-layer perceptron to ensure the same dimension as the character encoding.

[0129] Subsequently, a pre-trained language model is used to extract the context semantic features of the above encoding pairs, and based on this, the fusion weights between the character encoding and the pinyin encoding at each character position are determined. This model can be based on the Transformer architecture and adopt a multi-head attention mechanism to capture the semantic context at each character position in the sentence, generating the fusion decision weights at the corresponding positions. The weight values can dynamically adjust the contribution ratio between the character and the pinyin. Especially for polyphonic characters, technical terms, or written language constructs, the context information of the language model can significantly improve the accuracy of pinyin selection.

[0130] Finally, according to the aforementioned fusion weights, a linear weighted operation is performed on the character encoding and the pinyin encoding at each character position to form the character-pinyin fusion feature. The fusion method can adopt dot product of the weight vector followed by normalization, or enhance the context fluidity through a residual structure. This fusion feature has the ability to carry both semantic and pronunciation information simultaneously and is the input basis for the subsequent text encoder and speech generation network.

[0131] In one implementation, the character embedding layer and the pinyin embedding layer respectively construct two embedding matrices. The dimension of the character embedding matrix is [V_char, d], and the dimension of the pinyin embedding matrix is [V_pinyin, d], where V is the vocabulary size and d is the embedding dimension. The input initial text is mapped to a sequence of character indices, and each index takes the corresponding vector in the embedding matrix. After obtaining the mapping relationship, the pinyin part is also converted into a sequence of pinyin indices and the pinyin encoding vectors are obtained by looking up the table.

[0132] The determination of the fusion weights can adopt a small attention network. After concatenating the character encoding and the pinyin encoding as input, a two-layer perceptron outputs the weight α ∈ [0, 1], representing the proportion of the pinyin. The final fusion feature can be obtained by the formula: F = α * P + (1 - α) * C, where C is the character encoding, P is the pinyin encoding, and F is the fusion feature.

[0133] In another implementation, shared positional encoding and a unified encoder model can be used for character and pinyin embedding vectors to naturally fuse context features. This structure can effectively reduce the parameter scale and improve the accuracy of character-pinyin fusion contrast, and is suitable for scenarios with lightweight deployment requirements.

[0134] In this embodiment, by introducing pinyin embedding on the basis of character encoding and performing dynamic fusion under the guidance of context semantics, not only the pronunciation controllability of polyphonic characters and long-tail words is enhanced, but also the model has better language restoration ability under zero-shot or few-shot speaker conditions. This method eliminates the pronunciation uncertainty purely based on glyphs in traditional character-level modeling, and at the same time fuses context awareness ability to improve the naturalness and semantic accuracy of the generated speech.

[0135] In one embodiment, the above step S50 includes:

[0136] S501, constructing a test data set containing acoustic speech marker samples;

[0137] S502, encoding the acoustic speech marker samples in the test data set by vector quantization to generate a first codebook, and analyzing the compression rate and reconstruction error of the first codebook;

[0138] S503, encoding the acoustic speech marker samples in the test data set by finite scalar quantization to generate a second codebook, and analyzing the compression rate and reconstruction error of the second codebook;

[0139] S504, comparing the compression rates and reconstruction errors of the first codebook and the second codebook, and selecting the optimal quantization method as the codebook generation method based on the comparison results;

[0140] S505, based on the codebook generation method, generating a speech encoding according to the acoustic speech markers of the prompt speech.

[0141] In this embodiment, in order to implement a more efficient and high-fidelity speech encoding strategy, it is necessary to generate optional codebooks based on multiple quantization methods before applying the model, and select the optimal codebook generation method through comparative analysis. In this process, first, a test data set containing acoustic speech marker samples needs to be constructed. This test data set should cover a wide range of speaker timbre characteristics, pronunciation methods, speech rates, and acoustic variations in different contexts to ensure the universality of the quantization evaluation. The speech marker samples can be obtained by extracting acoustic features from large-scale speech data. The acoustic features generally include Mel spectrogram, linear prediction cepstral coefficients (LPCC), fundamental frequency contour, energy envelope, etc. The extraction method can use methods such as sliding window framing, Fourier transform, or convolutional network for processing.

[0142] Next, the acoustic speech labels of each speech frame in the test dataset are encoded using vector quantization. Vector quantization maps high-dimensional continuous vectors to the nearest vectors in a predefined finite vector set to achieve compression. This set is the first codebook, and the generation method can be based on k-means clustering, perceptual hashing, aggregated self-attention structure, or Product Quantization, etc. Analyzing the compression rate of the first codebook requires calculating the ratio between the original vector size and the quantization index size. The reconstruction error can be measured by the mean squared error (MSE) or the signal-to-noise ratio (SNR) to measure the gap between the restored samples and the original speech labels.

[0143] At the same time, the same dataset is encoded using finite scalar quantization to generate a second codebook. Scalar quantization independently quantizes each dimension of the acoustic label, resulting in a lower-complexity but possibly less accurate compression result. The construction of this codebook is usually based on non-uniform interval partitioning, histogram equalization distribution, or entropy coding mechanism, which is suitable for deployment requirements sensitive to the model inference time.

[0144] Compare the compression rates and reconstruction errors of the first codebook and the second codebook. A weighted scoring function can be used to sum the compression rate and the reconstruction error multiplied by their respective weights to obtain an overall evaluation metric. According to the scoring results, select the quantization method with the best compression efficiency and reconstruction quality as the final codebook generation method. This method can also consider the model adaptation scenario, for example, giving priority to the compression rate for edge devices and giving priority to the restoration quality for cloud models.

[0145] After the selection is completed, encode the acoustic speech labels of the prompt speech according to this codebook generation method. The prompt speech needs to pass through a feature extraction module first to obtain an acoustic speech label representation with the same dimension as the test dataset, and then perform an encoding mapping operation through the selected codebook to obtain the final speech encoding result. This speech encoding will be used as the conditional input of the speech generation model, undertaking the function of compressing the speech content information and maintaining the ability to represent speech details.

[0146] In one implementation, the speech labels of the test dataset are extracted through 64-dimensional Mel spectrogram coefficients. The k-means algorithm is used to cluster all frame vectors, and the number of cluster centers is set to 512 as the first codebook, encoding each frame as an index; during reconstruction, the index is mapped back to the corresponding center vector. The compression rate is the ratio of the original 64-dimensional floating-point vector to the 9-bit discrete index, and the error is the MSE.

[0147] In another implementation, the finite scalar quantization method discretizes each Mel-spectrum dimension by dividing it into 8 intervals to form a second codebook. The analysis results show that the first codebook has a lower reconstruction error, while the second codebook has a higher compression rate. If it is for the deployment requirements of edge terminals, the second codebook can be selected; if it is used for high-quality synthesis in the training phase, the first codebook is selected.

[0148] When generating voice encoding, each frame of the prompt voice is mapped to the closest vector or quantization value in the selected codebook, and the compression information is retained in the form of a codebook index or quantization code string, which can be used as a conditional input in the subsequent voice synthesis model.

[0149] In this embodiment, by constructing multiple codebooks and performing index comparison and analysis in the preprocessing stage, the most suitable quantization method can be effectively selected according to the task objective, so that the voice encoding not only has a high compression rate but also has a low reconstruction error, improving the overall quality of the voice synthesis system. The voice encoding generated by this method not only has a controllable compression effect but also has good voice restoration ability, providing stable and high-quality input conditions for the subsequent multi-modal generation model.

[0150] In one embodiment, the above step S70 includes:

[0151] S701, performing frame windowing processing on the prompt voice through the speaker encoder to generate a sequence of voice frames;

[0152] S702, extracting local acoustic features of the sequence of voice frames through the convolutional neural network layer of the speaker encoder;

[0153] S703, processing the sequence of voice frames through the multi-head self-attention mechanism of the speaker encoder to capture the global dependencies of the sequence of voice frames;

[0154] S704, performing residual connection processing on the local acoustic features and the global dependencies through the residual connection module of the speaker encoder to generate a fused feature vector;

[0155] S705, performing dimensionality reduction processing on the fused feature vector through the fully connected layer of the speaker encoder to generate the speaker feature vector.

[0156] In this embodiment, in order to accurately extract the personality feature information related to the speaker identity in the prompt voice, this module designs a speaker encoding process consisting of temporal modeling, local structure extraction, and fusion optimization. First, the prompt voice needs to be framed and windowed, dividing the continuous voice signal into frame segments of the same length with short-term stationarity. Each frame can be realized through a window with a length of 20 - 25 milliseconds, and the sliding step size is set to control the overlap degree between frames, such as 10 milliseconds, to retain a sufficiently fine-grained temporal structure. The windowing method can use a Hamming window or a Gaussian window to reduce the spectral leakage phenomenon, thereby constructing a balanced sequence of speech frames.

[0157] This sequence of speech frames is then input into the convolutional neural network layer in the speaker encoder to extract the local acoustic features at the frame level. The convolutional operation slides multiple kernels in the time dimension to extract short-term spectral patterns, including speaker-related signals such as fundamental frequency variations and formant distributions. Usually, multiple layers of one-dimensional convolution, activation functions, and batch normalization operations are configured to gradually enhance the acoustic representation ability and achieve spatial compression and robust modeling.

[0158] In order to capture the long-term dependencies between frames in the global context, the sequence of speech frames is simultaneously input into the multi-head self-attention mechanism module of the speaker encoder. This mechanism constructs multiple attention heads and calculates the correlation weights of each frame vector with other frames, thereby explicitly modeling the dynamic connection of the speech on the overall time axis. The Query-Key-Value calculation structure in the attention mechanism can incorporate long-range speech trends into the encoding modeling based on position independence, which helps to capture long-term characteristics such as speech rate, intonation, and pauses between sentences.

[0159] The residual connection module in the speaker encoder is responsible for fusing the local acoustic features obtained from the convolutional network and the global dependencies generated by the self-attention mechanism. The fusion method retains the independent expression capabilities of both in terms of spatial distribution and global relationship through the residual addition operation, and uses normalization and non-linear activation layers to enhance the learnability of the fused features, preventing feature degradation or dimensional mismatch problems. The output after fusion is the intermediate speech representation that aggregates multi-scale features.

[0160] Finally, the fused feature vector is dimensionally reduced through the fully connected layer of the speaker encoder. This processing compresses the high-dimensional speech frame features into a speaker feature vector of a fixed length, representing the individual speech style, vocal tract physiological characteristics, and timbre representation information contained in the current speech. The dimensional reduction structure can include linear transformation, normalization, and compression layers, and the dimension of the output speaker feature vector can be controlled at a standard vector length such as 256 or 512, which is convenient for subsequent speech generation modules to use.

[0161] In one implementation, the window length of the speech frame sequence is set to 25 milliseconds, the frame shift is 10 milliseconds, and the Hamming window function is used for windowing to ensure the balance between time-domain stationarity and frequency-domain smoothness. The convolutional neural network part is set with three one-dimensional convolutions, and the number of channels is 128, 256, and 512 respectively. The size of each convolutional kernel is 5, and the stride is 1. Combined with the ReLU activation function and LayerNorm, local feature extraction is realized.

[0162] The multi-head self-attention module is configured with 4 attention heads, and the dimension of each head is 64. Based on the Transformer encoding structure, a Query-Key-Value matrix is constructed, and attention weighting is performed on the frame dimension to output the global time series relationship vector. After the residual connection module adds the output vectors of the convolutional network and the attention mechanism element-wise, it is followed by BatchNorm and the ReLU activation function. The fused vector is mapped to a 256-dimensional speaker feature vector through a fully connected network with Dropout and linear projection.

[0163] In another implementation, the multi-head attention and convolutional operations adopt a parallel rather than a serial structure. After independently extracting features respectively, they are uniformly fused at the residual connection to improve the performance of the model in diverse speaker contexts. In addition, the vector after residual fusion can adopt the attention-weighted fusion method, and the local and global feature weights can be adaptively adjusted through the learned fusion gating mechanism to achieve stronger personalized vector compression.

[0164] In this embodiment, by using the combination method of local structure modeling and global context relationship, the speaker feature information can be stably extracted under different speech rates, accents, and pitch ranges. The residual fusion strategy strengthens the effective integration of time-series multi-level features and improves the feature representation ability; through the final fully connected dimensionality reduction, the generated speaker feature vector has high speech identity recognition ability and good transferability, providing stable support for tone preservation and speaker consistency in the subsequent generation stage.

[0165] In one embodiment, the above step S80 includes:

[0166] S801, performing feature splicing on the intermediate encoding and the speaker feature vector to generate decoder input features;

[0167] S802, performing forward calculation on the decoder input features through the generator network of the generative adversarial decoder to generate a time-domain speech waveform;

[0168] S803, outputting the time-domain speech waveform as the target speech.

[0169] In this embodiment, to generate the target speech, it is necessary to effectively fuse the intermediate encoding representing semantic content with the speaker feature vector representing the speaker's personality characteristics to form a complete generation input. The intermediate encoding is usually generated by a text-to-speech language model and carries content information such as syllable-level semantic expressions and prosodic intentions, while the speaker feature vector contains static speech style attributes such as pronunciation style, vocal tract structure, and timbre. Performing feature concatenation on the two aims to construct a unified vector sequence so that the subsequent generator network can capture both the pronunciation content and timbre characteristics simultaneously.

[0170] The way of feature concatenation is generally dimension concatenation at the vector level, keeping the time dimension consistent and expanding in the feature dimension direction. For example, a middle encoding sequence of length T and a vector sequence of dimension d1, after being broadcast and concatenated with a speaker vector of dimension d2, form a two-dimensional tensor of T×(d1 + d2), and this tensor is used as the input to the decoder.

[0171] The generator network is usually implemented based on the generative adversarial network (GAN) architecture. Its structure can include modules such as stacked convolutional layers, residual blocks, and skip connections, which are responsible for converting the concatenated input features into a continuous time-domain speech waveform. The generator network performs a forward calculation process, that is, layer-by-layer mapping, activation, and convolution operations on the input feature sequence, and synthesizes the complete speech through a multi-scale decoding path. This process is a one-time forward propagation calculation and does not require backpropagation training. It can be used for real-time inference after pre-training.

[0172] Finally, the time-domain speech waveform output by the generator is the target speech, which has a natural intonation, coherent rhythm, and the timbre style of the target speaker. The output format is usually single-channel floating-point waveform data at 16kHz or 24kHz, which can be directly used for speech playback or subsequent processing systems.

[0173] In one implementation, the intermediate encoding generated by the text-to-speech language model is set as a tensor of length T and dimension 512, and the dimension of the speaker feature vector is set as 256. By broadcasting and replicating the speaker vector in the time axis direction to match the time length of T, and then concatenating along the feature dimension direction, a decoder input feature of T×768 dimensions is formed.

[0174] The generator network structure can adopt a decoder framework designed based on BigVGAN2, including four layers of upsampling convolutional modules, and each layer is configured with a transposed convolution, a convolutional residual block, and a noise injection module. Each residual block contains a convolutional layer, normalization, and a LeakyReLU activation function. Through scale widening and time dimension recovery, the 768-dimensional input feature is finally mapped into a speech waveform with the same duration as the original speech. The generated speech waveform is represented in 32-bit floating-point format and the sampling rate is 22050Hz.

[0175] In another implementation, the UNet structure can be used to replace BigVGAN2. The skip connection method is adopted to directly transfer the low-level semantic information encoded in the middle to the high-level output, enhancing the local stability of the speech. Meanwhile, to improve the robustness of the model, a noise vector perturbation can be added to the input features of the decoder to enhance the generation stability under variable speech rates or repeated semantic inputs.

[0176] In this embodiment, through the fusion modeling of the middle encoding and the speaker features, this process can restore the voice characteristics of the target speaker while maintaining clear semantic expression, realizing personalized and highly natural speech generation. The forward generator structure has a strong time-domain restoration ability, can generate smooth, clear, and non-distorted speech waveforms, and does not rely on post-processing steps. The entire process design is adapted to the requirements of real-time synthesis and is applicable to low-latency interactive speech scenarios.

[0177] In one embodiment, before the above step S80, the following steps are further included:

[0178] S8001, constructing a training data set containing the real speech waveforms of the target speaker;

[0179] S8002, generating a random noise vector, and concatenating the random noise vector with the speaker feature vector of the target speaker to generate generator input features;

[0180] S8003, through the generator network of the generative adversarial decoder, performing a non-linear transformation on the generator input features to generate a synthetic speech waveform;

[0181] S8004, through the multi-scale discriminator network of the generative adversarial decoder, respectively performing multi-resolution analysis on the synthetic speech waveform and the real speech waveforms in the training data set to determine the value of the adversarial loss function;

[0182] S8005, based on the value of the adversarial loss function, alternately updating the parameters of the generator network and the multi-scale discriminator network until the Mel cepstral coefficient distance between the synthetic speech waveform generated by the generator network and the real speech waveform is lower than a preset threshold, obtaining a pre-trained generative adversarial decoder.

[0183] In this embodiment, in order to use the generative adversarial decoder for high-quality speech generation in the inference stage, a generator network with speech style learning ability needs to be constructed through the pre-training process. This process first needs to construct a training data set covering the speech style characteristics of the target speaker. Each piece of data in the training data set is the real speech waveform of the target speaker, and the sampling rate, bit width, and speech duration need to be processed according to unified standards to ensure the consistency of feature expression, constituting the real sample set in the training process.

[0184] The training input is generated by concatenating a random noise vector and the speaker feature vector of the target speaker. The concatenation process is carried out by connecting along the feature dimension to form the generator input feature with a fixed structure. The noise vector, as a latent variable, carries diverse information in the latent speech space, ensuring that the generator output has rich speech variation capabilities; the speaker feature vector guides the generator to learn speaker characteristics such as the target timbre and intonation. The generator input feature formed by the combination of the two is used to drive the speech synthesis process.

[0185] The generator network performs a non-linear transformation on the input feature, and the output is an analog speech waveform. Its structure usually adopts a multi-layer transposed convolution network to gradually restore the time dimension and amplitude features. The generator can include cross-layer connections, attention modules, spectral transformation modules, etc. to enhance the modeling ability.

[0186] The discriminator network discriminates and compares the synthesized speech waveform and the real speech waveform. The discriminator network is a multi-scale structure, which evaluates the distribution differences of the two types of samples at the original waveform scale, downsampled scale, and frame-level frequency domain scale respectively, so as to improve the discrimination ability. The discriminator can include multiple parallel paths, and each path uses different sampling rates, convolution kernel sizes, and stride designs to jointly model the structural differences between real and fake samples from the global and local dimensions.

[0187] During the training process, the adversarial loss function is used as the optimization objective. The generator aims to minimize the difference between the synthesized samples and the real samples in the discriminator output, while the discriminator optimizes its discrimination accuracy. The optimization process adopts an alternating update strategy, that is, fixing the generator parameters to optimize the discriminator network to improve its true / false recognition ability; then fixing the discriminator parameters to optimize the generator network to minimize the possibility that the synthesized samples are recognized as fake.

[0188] To quantify the convergence performance of the generator, the Mel Cepstral Distortion (MCD) is introduced as a quality evaluation index. This index measures the spectral fidelity of the speech by calculating the frame-level differences between the generated speech and the real speech on the Mel spectrum. When the average MCD value during the training process is lower than the preset threshold, it indicates that the generator has been able to perceptually approximate the real speech, and the training can be regarded as converged. The current generator network parameters are saved as the pre-trained generator for use in the inference stage.

[0189] In one implementation, the sampling rate of the training dataset is set to 22050Hz, covering target speaker samples in multiple emotional states and pronunciation scenarios. The dimension of the random noise vector is set to 128, and the dimension of the speaker feature vector is set to 256. After concatenation, a 384-dimensional input is formed. The generator uses a fully convolutional network with a residual structure to reconstruct a waveform segment of length T through four layers of transposed convolution and ReLU activation units. The multi-scale discriminator contains three parallel networks that separately process the original waveform, 2x downsampled, and mel spectrogram image inputs, and are jointly optimized through cross-entropy loss and adversarial loss.

[0190] In another approach, a two-stage training process can be adopted. In the first stage, the discriminator structure is fixed and only the perceptual loss of the generator is optimized. In the second stage, the adversarial loss is enabled to participate in the optimization, and the discriminator network structure parameters are dynamically adjusted. After the training cycle ends, if the MCD is lower than 4.5dB for five consecutive rounds, the training is aborted and the generator parameters are frozen.

[0191] Example Illustration: In the field of healthcare, in response to the need for intelligent assisted voice announcements of doctor's opinions, the entire training process of the generative adversarial decoder can be applied to simulate the language style of a specific doctor and achieve personalized voice output. First, the hospital information system will collect the voice records of the target doctor in common diagnosis and treatment scenarios, such as oral condition notifications, examination advice explanations, postoperative education, etc., with the patient's authorization, and construct these data into a training dataset containing real voice waveforms and diagnosis and treatment context labels. Then, based on these real voice waveforms, the system extracts the acoustic features of each voice (such as Mel spectrogram, formant trajectory, fundamental frequency contour) through short-time Fourier transform, and extracts the speaker feature vector in combination with the identity information of the doctor. During the training process, the system periodically generates synthetic voice samples consistent with the target doctor's style, generates an initial noise tensor through random sampling, and concatenates it with the speaker feature vector of the doctor to form the input feature of the generator. The generator adopts a fully convolutional architecture based on the stacking of convolutional residual modules, and gradually restores the concatenated features to the time-domain voice waveform through multiple layers of non-linear mapping. The generated voice waveform will be fed into the multi-scale discriminator module, which uses a shallow convolutional network to evaluate the consistency between the generated voice and the real voice in terms of time structure, frequency texture, and energy distribution at three scales: the original sampling rate, half the downsampling rate, and one-fourth the downsampling rate. The system calculates the adversarial loss value between the generated result of each round and the corresponding real voice at the three scales, and uses the mean square error to measure the difference in Mel cepstral distance (MCD) between the two. In the discriminator training step, the system fixes the parameters of the generator and only updates the convolutional kernel weights and normalization biases of the discriminator to make it more accurate in distinguishing real and generated voices; while in the generator training step, the system fixes the parameters of the discriminator and optimizes the parameters of the generator through backpropagation, so that the generated result gradually "deceives" the discriminator and reduces the MCD distance from the real voice. When the average MCD of multiple validation batches continues to be lower than the threshold of 3.8 dB (set based on the critical value distinguishable by the human ear), it is considered that the generator has reached the deployable quality level, and the system saves the model parameters of the generator as the pre-trained generator. This pre-trained generator is encapsulated into the generative adversarial decoder for voice cloning processing of structured diagnosis and treatment suggestions in subsequent actual announcement tasks.

[0192] In the field of fintech, this training mechanism is also applicable to building a customer service system with a unique voice style. Financial enterprises can collect the voice data of high-performing customer service representatives within the enterprise in typical scenarios, such as bill explanations, product descriptions, risk reminders, etc., and label scene tags for the voice segments in combination with the business context. Through the above training process, the system uses these voice samples to generate speaker feature vectors, and splices them with a randomly initialized noise tensor as the input to drive the generator to simulate the intonation, speaking rhythm, and semantic stress of the representative. The multi-scale discriminator scores the output waveform separately from low frequency (stability) to high frequency (naturalness), making the finally generated financial voice more realistic and authoritative in expression. In actual deployment, the system can be embedded in an intelligent customer service platform. When a user asks a question through text or structured input, such as "What fees will be incurred for overdue loans?", the platform first generates a corresponding intermediate encoding through the natural language processing module, calls the pre-trained generator to fuse with the speaker features of the customer service representative, and synthesizes a voice content with a gentle voice, clear rhythm, and kind tone in real time to broadcast the response answer to the user. This method greatly improves the professionalism of the voice output and the user acceptance, and solves the problems of the rigid synthetic voice and large perceived distance of traditional customer service.

[0193] In this embodiment, the generator network is enabled to master the ability to map from a random variable and a speaker feature to a high-quality voice waveform, and the naturalness and timbre consistency of the output result are improved through the feedback of the discriminator. The multi-scale discriminator structure effectively alleviates the problem of the discriminator overfitting to a specific pattern and improves the modeling ability for complex timbre changes. Using the Mel cepstral coefficients as the evaluation index can ensure that the generated result approximates the real voice both in the spectral domain and in the perceptual quality, thus providing a stable and high-fidelity decoder model for subsequent inference.

[0194] In one embodiment, the above step S40 includes:

[0195] S401, mapping the character pinyin fusion feature to a high-dimensional semantic vector through a pre-trained language model embedding layer;

[0196] S402, performing context-dependent modeling on the high-dimensional semantic vector through the multi-head self-attention mechanism layer of the text encoder to generate context-related sequence features;

[0197] S403, performing a non-linear transformation on the context-related sequence features through the feed-forward neural network layer of the text encoder to generate a text feature vector;

[0198] S404, performing layer normalization on the text feature vector to generate the text feature.

[0199] In this embodiment, the pre-trained language model embedding layer is an important component in the field of natural language processing (NLP). It is located at the input end of the language model and is mainly used to map discrete symbols (such as characters, words, or pinyin) into a continuous, high-dimensional vector space. These vectors can capture semantic information, enabling the model to understand and process the vocabulary and its context in language.

[0200] A pre-trained language model is a deep learning model trained on a large amount of text data with the aim of learning the representation of language (also called "language model"). These models learn the relationships between words, grammatical structures, and semantic information from a large-scale corpus through unsupervised learning. These models are usually based on the Transformer architecture.

[0201] The embedding layer is a component of a neural network used to convert discrete words (or characters, pinyin, etc.) into a fixed-size vector. The output of the embedding layer is a high-dimensional vector that carries the semantic information of each input symbol. In traditional NLP, words are usually mapped to a dense vector space through word embeddings (such as Word2Vec or GloVe), but in modern pre-trained language models, the embedding layer is the first part of the entire model, responsible for mapping the input symbols into a high-dimensional space.

[0202] The embedding layer of a pre-trained language model not only converts the input symbols into vectors but also uses pre-trained parameters to enhance the semantic expression ability of these vectors. Specifically, the pre-trained embedding layer will use the language knowledge learned from the large-scale corpus during training to generate vector representations that can carry information in multiple aspects such as grammar and semantics. For example, in the BERT model, the input to the embedding layer includes not only the embeddings of words (such as the vectors obtained by looking up the word IDs in the vocabulary), but also the position information of the word (the position of the word in the sentence), and the relationship between sentences (if it is a multi-sentence input). After being processed by the model parameters obtained through pre-training, these information can generate high-dimensional vectors with rich semantics.

[0203] High-dimensional here refers to mapping the fused features of input characters and pinyin through the embedding layer to generate a numerically represented form with higher dimensions. Specifically, the high-dimensional vector representation can contain hundreds to thousands of dimensions, and can accurately describe the semantic relationships and context information of the fused features of characters and pinyin in a multi-dimensional space. For example, in the BERT model in natural language processing, each input word or character is converted into a 768-dimensional high-dimensional vector through the embedding layer. This vector can carry rich semantic information, including word meaning, grammatical structure, context, etc. The characteristic of high-dimensional is that it can combine local features (such as the phonetic features of pinyin) and global features (such as the influence of other characters in the sentence) at the same time, improving the expression ability of the model.

[0204] In the multi-head self-attention mechanism layer, high-dimensional semantic vectors are further processed to capture the context dependencies between different positions. The multi-head mechanism means that through multiple parallel attention heads, different subspace information is respectively focused on, so as to capture finer-grained semantic relationships. This mechanism can provide the model with an understanding of different long-range dependencies between characters. The high-dimensional semantic vectors are used as inputs during this process, and the relationships between positions are dynamically adjusted through attention weights. The output sequence features contain rich context information, enhancing the ability to handle polyphonic characters and ambiguous words.

[0205] Through the feed-forward neural network layer, the context-related features originally generated by the multi-head self-attention mechanism will undergo non-linear transformation to further learn complex feature relationships, and finally compress them into text feature vectors of a fixed dimension. This dimension compression step ensures that the features can be effectively used in subsequent tasks and prevents computational efficiency problems caused by excessive feature dimensions. In this layer, the high-dimensional vectors are mapped to a vector that better meets the requirements of downstream tasks. The compression process helps to reduce redundant information while ensuring that the text features can adapt to the specific task requirements.

[0206] Layer normalization is a regularization method used to ensure the stability of the output of each layer in the neural network during training. This is very important for solving the problems of vanishing gradients or exploding gradients, and at the same time it can also accelerate the training process of the model. After being processed by the feed-forward network, the high-dimensional vectors need to go through layer normalization to normalize their eigenvalue, ensuring that the model can be better optimized. This step helps to stabilize the update of model parameters and improve the convergence speed during training.

[0207] Through the above steps in this embodiment, the character pinyin fusion features can be effectively converted into standardized text features that can be used for downstream speech generation tasks. The introduction of high-dimensional semantic vectors enhances the representation ability of the model, enabling the model to better understand and process complex language phenomena such as polyphonic characters and ambiguous words. The text encoder effectively captures the context relationships between characters through the multi-head self-attention mechanism, thereby improving the naturalness and expressiveness of voice cloning. The use of layer normalization further stabilizes the training process, accelerates the convergence of the model, and provides good performance guarantees for actual deployment.

[0208] In one embodiment, a voice generation device is provided, and this voice generation device corresponds one-to-one with the voice generation method in the above embodiment. Refer to Figure 3 , Figure 3Schematic diagram of functional modules of a preferred embodiment of the voice generation device of the present invention. Voice input module 10, voice recognition module 20, pinyin encoding module 30, text encoding module 40, voice quantization module 50, multimodal fusion module 60, speaker analysis module 70, and voice synthesis module 80. The detailed description of each functional module is as follows:

[0209] The voice input module 10 is used to receive a prompt voice containing target voice features;

[0210] The voice recognition module 20 is used to convert the prompt voice into an initial text through the voice recognition module;

[0211] The pinyin encoding module 30 is used to perform joint encoding processing of characters and pinyin on the initial text to generate character-pinyin fusion features;

[0212] The text encoding module 40 is used to process the character-pinyin fusion features through a text encoder to generate text features;

[0213] The voice quantization module 50 is used to determine a codebook generation method and extract a voice code in the prompt voice based on the codebook generation method;

[0214] The multimodal fusion module 60 is used to input the text features and voice codes into a text-voice language model to generate an intermediate code;

[0215] The speaker analysis module 70 is used to extract a speaker feature vector in the prompt voice through a speaker encoder;

[0216] The voice synthesis module 80 is used to perform decoding processing on the intermediate code and the speaker feature vector through a generative adversarial decoder to generate a target voice.

[0217] In one embodiment, the pinyin encoding module 30 is specifically used for:

[0218] Obtain at least one pinyin corresponding to each character in the initial text from a preset character-pinyin mapping relationship library;

[0219] Convert the characters in the initial text into character codes through a character embedding layer;

[0220] Convert the pinyin corresponding to the characters in the initial text into pinyin codes through a pinyin embedding layer;

[0221] Extract the context semantic features of the character codes and pinyin codes through a pre-trained language model, and determine the fusion weight of the character codes and pinyin codes;

[0222] Perform linear weighted processing on the character encoding and pinyin encoding of each character according to the fusion weight to generate a character-pinyin fusion feature.

[0223] In one embodiment, the voice quantization module 50 is specifically configured to:

[0224] Construct a test data set containing acoustic speech marker samples;

[0225] Encode the acoustic speech marker samples in the test data set by vector quantization to generate a first codebook, and analyze the compression rate and reconstruction error of the first codebook;

[0226] Encode the acoustic speech marker samples in the test data set by finite scalar quantization to generate a second codebook, and analyze the compression rate and reconstruction error of the second codebook;

[0227] Compare the compression rates and reconstruction errors of the first codebook and the second codebook, and select the optimal quantization method as the codebook generation method based on the comparison results;

[0228] Based on the codebook generation method, generate a voice encoding according to the acoustic speech marker of the prompt voice.

[0229] In one embodiment, the speaker analysis module 70 is specifically configured to:

[0230] Perform frame windowing processing on the prompt voice through the speaker encoder to generate a sequence of voice frames;

[0231] Extract local acoustic features of the voice frame sequence through the convolutional neural network layer of the speaker encoder;

[0232] Process the voice frame sequence through the multi-head self-attention mechanism of the speaker encoder to capture the global dependency relationship of the voice frame sequence;

[0233] Perform residual connection processing on the local acoustic features and the global dependency relationship through the residual connection module of the speaker encoder to generate a fused feature vector;

[0234] Perform dimensionality reduction processing on the fused feature vector through the fully connected layer of the speaker encoder to generate the speaker feature vector.

[0235] In one embodiment, the voice synthesis module 80 is specifically configured to:

[0236] Perform feature splicing on the intermediate encoding and the speaker feature vector to generate a decoder input feature;

[0237] Through the generator network of the generative adversarial decoder, perform forward calculation on the decoder input features to generate a time-domain speech waveform;

[0238] Output the time-domain speech waveform as the target speech.

[0239] In one embodiment, the speech synthesis module 80 is specifically configured to:

[0240] Construct a training data set including the real speech waveforms of the target speaker;

[0241] Generate a random noise vector, and splice the random noise vector with the speaker feature vector of the target speaker to generate generator input features;

[0242] Through the generator network of the generative adversarial decoder, perform a non-linear transformation on the generator input features to generate a synthetic speech waveform;

[0243] Through the multi-scale discriminator network of the generative adversarial decoder, perform multi-resolution analysis on the synthetic speech waveform and the real speech waveforms in the training data set respectively to determine the adversarial loss function value;

[0244] Based on the adversarial loss function value, alternately update the parameters of the generator network and the multi-scale discriminator network until the Mel cepstral coefficient distance between the synthetic speech waveform generated by the generator network and the real speech waveform is lower than a preset threshold, and obtain a pre-trained generative adversarial decoder.

[0245] In one embodiment, the text encoding module 40 is specifically configured to:

[0246] Map the character-pinyin fusion features to high-dimensional semantic vectors through a pre-trained language model embedding layer;

[0247] Perform context-dependent modeling on the high-dimensional semantic vectors through the multi-head self-attention mechanism layer of the text encoder to generate context-related sequence features;

[0248] Perform a non-linear transformation on the context-related sequence features through the feed-forward neural network layer of the text encoder to generate text feature vectors;

[0249] Perform layer normalization processing on the text feature vectors to generate the text features.

[0250] In one embodiment, a computer device is provided. The computer device can be a server, and its internal structure diagram can be as Figure 4As shown in the figure. The computer device includes a processor, a memory, a network interface, and a database connected via a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile and / or volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external client via a network connection. When the computer program is executed by the processor, it realizes the functions or steps on the server side of a voice generation method.

[0251] In one embodiment, a computer device is provided. The computer device can be a client, and its internal structure diagram can be as Figure 5 shown in the figure. The computer device includes a processor, a memory, a network interface, a display screen, and an input device connected via a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external server via a network connection. When the computer program is executed by the processor, it realizes the functions or steps on the client side of a voice generation method

[0252] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the following steps are realized:

[0253] Receive a prompt voice containing target voice features;

[0254] Convert the prompt voice into an initial text through a voice recognition module;

[0255] Perform joint encoding processing of characters and pinyin on the initial text to generate a character-pinyin fusion feature;

[0256] Process the character-pinyin fusion feature through a text encoder to generate a text feature;

[0257] Determine the codebook generation method, and extract the voice encoding in the prompt voice based on the codebook generation method;

[0258] Input the text feature and the voice encoding into a text-voice language model to generate an intermediate encoding;

[0259] Extract the speaker feature vector in the prompt voice through a speaker encoder;

[0260] The intermediate encoding and the speaker feature vector are decoded by a generative adversarial decoder to generate a target speech.

[0261] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:

[0262] Receiving a prompt speech containing target speech features;

[0263] Converting the prompt speech into an initial text by a speech recognition module;

[0264] Performing a joint encoding process of characters and pinyin on the initial text to generate a character-pinyin fusion feature;

[0265] Processing the character-pinyin fusion feature by a text encoder to generate a text feature;

[0266] Determining a codebook generation method, and extracting a speech code in the prompt speech based on the codebook generation method;

[0267] Inputting the text feature and the speech code into a text-to-speech language model to generate an intermediate encoding;

[0268] Extracting a speaker feature vector in the prompt speech by a speaker encoder;

[0269] The intermediate encoding and the speaker feature vector are decoded by a generative adversarial decoder to generate a target speech.

[0270] It should be noted that for the functions or steps that can be implemented by the above computer-readable storage medium or computer device, reference can be made to the relevant descriptions on the server side and the user side in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.

[0271] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.

[0272] Those skilled in the art can clearly understand that for the convenience and brevity of description, only the above-mentioned division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules as needed, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above.

[0273] It should be noted that if there are software tools or components of other companies in the embodiments of the present application, they are only used for illustrative introduction and do not represent actual use. The above-mentioned embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included in the protection scope of the present invention.

Claims

1. A voice generation method, characterized in that, Including the following steps: Receiving a prompt voice containing target voice features; Converting the prompt voice into an initial text through a voice recognition module; Performing joint encoding processing of characters and pinyin on the initial text to generate character-pinyin fusion features; Processing the character-pinyin fusion features through a text encoder to generate text features; Determining a codebook generation method and extracting a voice encoding in the prompt voice based on the codebook generation method; Inputting the text features and the voice encoding into a text-speech language model to generate an intermediate encoding; Extracting a speaker feature vector in the prompt voice through a speaker encoder; Performing decoding processing on the intermediate encoding and the speaker feature vector through a generative adversarial decoder to generate a target voice.

2. The voice generation method according to claim 1, wherein Performing joint encoding processing of characters and pinyin on the initial text to generate character-pinyin fusion features, including: Obtaining at least one pinyin corresponding to each character in the initial text from a preset character-pinyin mapping relationship library; Converting the characters in the initial text into character encodings through a character embedding layer; Converting the pinyin corresponding to the characters in the initial text into pinyin encodings through a pinyin embedding layer; Extracting context semantic features of the character encoding and the pinyin encoding through a pre-trained language model and determining a fusion weight of the character encoding and the pinyin encoding; Performing linear weighted processing on the character encoding and the pinyin encoding of each character according to the fusion weight to generate character-pinyin fusion features.

3. The voice generation method according to claim 1, wherein, Determining a codebook generation method and extracting a voice encoding in the prompt voice based on the codebook generation method, including: Constructing a test data set containing acoustic voice token samples; Encoding the acoustic voice token samples in the test data set through a vector quantization method to generate a first codebook and analyzing the compression rate and reconstruction error of the first codebook; Encoding the acoustic voice token samples in the test data set through a finite scalar quantization method to generate a second codebook and analyzing the compression rate and reconstruction error of the second codebook; Comparing the compression rates and reconstruction errors of the first codebook and the second codebook, and selecting an optimal quantization method as the codebook generation method based on the comparison result; Based on the codebook generation method, generating a voice encoding according to the acoustic voice tokens of the prompt voice.

4. The voice generation method according to claim 1, wherein Extracting a speaker feature vector in the prompt voice through a speaker encoder, including: Performing frame windowing processing on the prompt voice through the speaker encoder to generate a voice frame sequence; Extracting local acoustic features of the voice frame sequence through a convolutional neural network layer of the speaker encoder; Processing the voice frame sequence through a multi-head self-attention mechanism of the speaker encoder to capture global dependencies of the voice frame sequence; Performing residual connection processing on the local acoustic features and the global dependencies through a residual connection module of the speaker encoder to generate a fusion feature vector; Performing dimensionality reduction processing on the fusion feature vector through a fully connected layer of the speaker encoder to generate the speaker feature vector.

5. The voice generation method according to claim 1, wherein Performing decoding processing on the intermediate encoding and the speaker feature vector through a generative adversarial decoder to generate a target voice, including: Perform feature concatenation on the intermediate encoding and the speaker feature vector to generate decoder input features; Perform forward calculation on the decoder input features through the generator network of the generative adversarial decoder to generate a time-domain speech waveform; Output the time-domain speech waveform as the target speech.

6. The voice generation method according to claim 1, characterized in that, Before generating the target speech by decoding the intermediate encoding and the speaker feature vector through the generative adversarial decoder, it further includes: Construct a training data set containing the real speech waveforms of the target speaker; Generate a random noise vector, and concatenate the random noise vector with the speaker feature vector of the target speaker to generate generator input features; Perform non-linear transformation on the generator input features through the generator network of the generative adversarial decoder to generate a synthetic speech waveform; Perform multi-resolution analysis on the synthetic speech waveform and the real speech waveforms in the training data set respectively through the multi-scale discriminator network of the generative adversarial decoder to determine the adversarial loss function value; Alternately update the parameters of the generator network and the multi-scale discriminator network based on the adversarial loss function value until the Mel cepstral coefficient distance between the synthetic speech waveform generated by the generator network and the real speech waveform is lower than a preset threshold, and obtain a pre-trained generative adversarial decoder.

7. The voice generation method according to claim 1, wherein Processing the character-pinyin fusion feature through a text encoder to generate text features, including: Mapping the character-pinyin fusion feature into a high-dimensional semantic vector through a pre-trained language model embedding layer; Performing context-dependent modeling on the high-dimensional semantic vector through the multi-head self-attention mechanism layer of the text encoder to generate context-related sequence features; Performing non-linear transformation on the context-related sequence features through the feed-forward neural network layer of the text encoder to generate a text feature vector; Performing layer normalization processing on the text feature vector to generate the text features.

8. A voice generation device, characterized in that, The speech generation device includes: A speech input module for receiving a prompt speech containing target speech features; A speech recognition module for converting the prompt speech into an initial text through the speech recognition module; A pinyin encoding module for performing joint encoding processing of characters and pinyin on the initial text to generate a character-pinyin fusion feature; A text encoding module for processing the character-pinyin fusion feature through a text encoder to generate text features; A speech quantization module for determining a codebook generation method and extracting speech codes in the prompt speech based on the codebook generation method; A multi-modal fusion module for inputting the text features and speech codes into a text-speech language model to generate an intermediate encoding; A speaker analysis module for extracting the speaker feature vector in the prompt speech through a speaker encoder; A speech synthesis module for decoding the intermediate encoding and the speaker feature vector through a generative adversarial decoder to generate a target speech.

9. A computer device, characterized in that, The computer device includes a memory, a processor, and a voice generation program stored in the memory and executable on the processor. When the voice generation program is executed by the processor, it implements the steps of the voice generation method according to any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, A voice generation program is stored on the storage medium. When the voice generation program is executed by the processor, it implements the steps of the voice generation method according to any one of claims 1-7.

Citation Information

Patent Citations

  • Speech synthesis method, device and equipment and storage medium

    CN112735373A

  • Speech synthesis front-end processing method and device, equipment and storage medium

    CN118800212A

  • Emotional voice conversion method and device based on rhythm prediction, equipment and medium

    CN119207371A

  • Voice synthesis method and device based on gated attention mechanism, equipment and medium

    CN119314463A

Cited By

  • Personalized customized AI voiceprint cloning method and system and storage medium thereof

    CN120783725A

  • Personalized customized ai voiceprint cloning method, system and storage medium thereof

    CN120783725B

  • Speech synthesis method and system for human-computer interaction

    CN120877706A

  • Aviation speech generation method and system based on generative adversarial network

    CN120913539A

  • Text-to-speech synthesis method and device, equipment and medium

    CN121011170A