Controllable zero-sample speech conversion method, device, equipment and medium
By employing self-supervised speech learning and stream matching techniques, the problem of insufficient style controllability in zero-sample speech conversion is solved, achieving high-quality personalized speech synthesis applicable to fields such as medical rehabilitation and fintech.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- PING AN TECH (SHENZHEN) CO LTD
- Filing Date
- 2025-08-28
- Publication Date
- 2026-08-04
AI Technical Summary
Existing zero-sample speech conversion technology struggles to achieve personalized, high-fidelity, and consistent speech conversion without labeled data, particularly in terms of speech rate, rhythm, and emotion control, which limits its application in fields such as medical rehabilitation and fintech.
Self-supervised speech representation of unlabeled data is obtained through self-supervised speech learning. Content feature vectors and prosodic style vectors are extracted, converted into discrete tokens and masked. Stream matching is performed in conjunction with user style embedding, and target Mel spectrograms are generated and speech waveforms are reconstructed to achieve high-quality personalized speech synthesis.
Without requiring additional labeled data, it achieves natural and realistic control over content, prosody, and timbre, improving the naturalness and fidelity of the generated speech and enhancing the model's generalization ability under different speech inputs and styles.
Smart Images

Figure CN121034280B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of speech semantics technology, and in particular to a controllable zero-sample speech conversion method, apparatus, device, and medium. Background Technology
[0002] While current zero-shot voice conversion (ZS-VC) techniques can synthesize speech from speakers who haven't been seen, they still have many shortcomings. Existing methods largely focus on timbre conversion, lacking control over fine-grained style features such as speech rate, prosody, stress, and emotion, resulting in poor performance in terms of naturalness and consistency with the target style. Furthermore, speech content and prosodic features are often intertwined, making independent control difficult and limiting flexible manipulation of attributes like speech rate and pitch. Although some methods attempt to introduce discrete representations or contextual learning, they still fall short in fine-grained prosodic control, capturing the unique style of the target speaker, and generalization capabilities. Especially in non-autoregressive models, prosodic modeling and alignment remain challenging, and high-fidelity systems often rely on complex multi-stage training, increasing the difficulty of development and deployment.
[0003] In the healthcare field, these methods can be used for rehabilitation training or voice assistance, but existing methods generally suffer from insufficient style controllability. For example, for stroke or aphasic patients, the system struggles to simultaneously ensure speech clarity and personalized style reproduction, resulting in deviations in the naturalness and comprehensibility of the generated speech. Furthermore, the difficulty in completely decoupling content and prosodic features means that while enhancing the intelligibility of the patient's pronunciation, the original speech rate or emotional characteristics are often sacrificed, limiting their practical value in clinical rehabilitation and telemedicine.
[0004] In the fintech sector, voice authentication can be used for intelligent customer service and identity verification, but existing systems still have significant shortcomings in addressing fraud and impersonation risks. Due to limited ability to capture and generalize the target speaker's style, the system may generate speech that "sounds similar but is not entirely consistent," making it difficult to provide a realistic and natural customer service interaction experience, and potentially allowing attackers to bypass authentication using speech conversion techniques. Furthermore, the inadequacy of non-autoregressive models in prosodic control results in a lack of stability and consistency in the generated speech across interactive scenarios, increasing risks in financial security applications.
[0005] Therefore, in the current technology, the existing ZS-VC technology still has significant shortcomings in terms of style controllability, content and prosody decoupling, and realistic imitation of the target style. Under the condition of unlabeled speech data, it is difficult to achieve personalized, high-fidelity, and style-consistent zero-sample speech conversion. Summary of the Invention
[0006] This invention provides a controllable zero-sample speech conversion method, apparatus, device, and medium. Its main purpose is to solve the problem of difficulty in achieving personalized, high-fidelity, and style-consistent zero-sample speech conversion under conditions of unlabeled speech data.
[0007] In a first aspect, to achieve the above objective, the present invention provides a controllable zero-sample speech conversion method, comprising: Acquire a number of unlabeled speech data, perform self-supervised speech learning on the unlabeled speech data, and obtain self-supervised speech representation; Extract the content feature vector and prosodic style vector of the self-supervised speech representation, and convert the content feature vector and the prosodic style vector into discrete content tokens and discrete prosodic tokens, respectively. The discrete prosodic token is masked to obtain the target prosodic token; Obtain the reference speech of the target user and extract the user style embedding from the reference speech; Stream matching is performed on the discrete content token, the target prosodic token, and the user style embedding to generate a target Mel spectrogram; The target Mel spectrogram is reconstructed and optimized to obtain zero-sample speech conversion results.
[0008] Secondly, the present invention also provides a controllable zero-sample speech conversion device, comprising: The self-supervised speech learning module is used to acquire several unlabeled speech data, perform self-supervised speech learning on the unlabeled speech data, and obtain a self-supervised speech representation. The content and prosodic token generation module is used to extract the content feature vector and prosodic style vector of the self-supervised speech representation, and convert the content feature vector and the prosodic style vector into discrete content tokens and discrete prosodic tokens, respectively. The prosody token masking module is used to generate a mask for the discrete prosody tokens to obtain the target prosody tokens; The user style embedding extraction module is used to obtain the reference speech of the target user and extract the user style embedding from the reference speech. The target Mel spectrogram generation module is used to perform stream matching on the discrete content token, the target prosodic token, and the user style embedding to generate a target Mel spectrogram. The speech waveform reconstruction module is used to reconstruct and optimize the speech waveform of the target Mel spectrogram to obtain zero-sample speech conversion results.
[0009] Thirdly, the present invention also provides an electronic device, the electronic device comprising: At least one processor; and, A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, which enables the at least one processor to perform the controllable zero-sample speech conversion method described above.
[0010] Fourthly, the present invention also provides a computer-readable storage medium storing at least one computer program, which is executed by a processor in an electronic device to implement the controllable zero-sample speech conversion method described above.
[0011] This invention acquires several unlabeled speech data sets and performs self-supervised speech learning on them. It utilizes random masking and the prediction mechanism of the self-supervised model to improve the model's robustness and contextual modeling ability, thereby obtaining a more generalizable and semantically expressive self-supervised speech representation. The invention extracts the content feature vector and prosodic style vector from the self-supervised speech representation and converts them into discrete content tokens and discrete prosodic tokens, respectively. This preserves the semantic information of the speech while accurately expressing speaking style and prosodic features, improving the model's generalization ability while reducing its dependence on large-scale labeled data. The discrete prosodic tokens are then masked to generate target prosodic tokens. This not only enables automatic completion when prosodic information is missing or incomplete but also improves generation efficiency and parallelism with the support of non-autoregressive generation. To ensure the harmony between prosody and semantics, the system acquires reference speech from the target user and extracts user style embeddings from the reference speech. Pre-emphasis enhances high-frequency information, making subsequent feature extraction more sensitive to the user's unique timbre and speaking style. Multi-layer convolution captures deep features of speech in the time-frequency domain, combined with average pooling to obtain robust initial style embeddings. Stream matching is performed on the discrete content token, the target prosody token, and the user style embeddings to generate a target Mel spectrogram. The generation path is guided by a composite condition of semantics, prosody, and user style, achieving highly consistent control over content, prosody, and personalized style. Speech waveform reconstruction and optimization are performed on the target Mel spectrogram to obtain zero-sample speech conversion results that are natural and realistic in content, prosody, and timbre, achieving high-quality personalized speech synthesis without additional labeled data. Attached Figure Description
[0012] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments of the present invention will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0013] Figure 1 This is a schematic diagram of an application environment for a controllable zero-sample speech conversion method according to an embodiment of the present invention; Figure 2 This is a flowchart illustrating a controllable zero-sample speech conversion method according to an embodiment of the present invention. Figure 3 This is a flowchart illustrating the self-supervised speech learning process in a controllable zero-shot speech conversion method according to an embodiment of the present invention. Figure 4 This is a schematic diagram of a controllable zero-sample speech conversion device according to an embodiment of the present invention; Figure 5 A schematic diagram of the structure of an electronic device for implementing a controllable zero-sample speech conversion method according to an embodiment of the present invention; Figure 6 This is another structural schematic diagram of an electronic device that implements a controllable zero-sample speech conversion method according to an embodiment of the present invention.
[0014] The objectives, features, and advantages of this invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0015] To enable those skilled in the art to better understand the technical solutions of this disclosure, and to fully understand and implement the process of how this disclosure applies technical means to solve technical problems and achieve corresponding technical effects, the technical solutions in the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this disclosure, not all embodiments. The embodiments of this disclosure and the various features within them can be combined with each other without conflict, and the resulting technical solutions are all within the protection scope of this disclosure. All other embodiments obtained by those skilled in the art based on the embodiments of this disclosure without creative effort should fall within the protection scope of this disclosure.
[0016] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this disclosure are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this disclosure described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, apparatus, product, or device that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or devices.
[0017] This application provides a controllable zero-shot speech conversion method. The execution subject of this method includes, but is not limited to, at least one electronic device that can be configured to execute the device provided in this application, such as a server or a terminal. In other words, the controllable zero-shot speech conversion method can be executed by software or hardware installed on a terminal device or a server device. The server includes, but is not limited to, a single server, a server cluster, a cloud server, or a cloud server cluster. The server can be an independent server or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms.
[0018] This invention provides a controllable zero-sample speech conversion method, which can be applied to applications such as... Figure 1In this application environment, the client communicates with the server via a network. The server can obtain several unlabeled speech data from the client, perform self-supervised speech learning on the unlabeled speech data, and improve the robustness and context modeling ability of the model by using random masking and the prediction mechanism of the self-supervised model, thereby obtaining a more generalizable and semantically expressive self-supervised speech representation. The server extracts the content feature vector and prosodic style vector of the self-supervised speech representation, and converts the content feature vector and the prosodic style vector into discrete content tokens and discrete prosodic tokens, respectively. This preserves the semantic information of the speech while accurately expressing the speaking style and prosodic features, improving the model's generalization ability while reducing its dependence on large-scale labeled data. The server then performs masking on the discrete prosodic tokens to obtain the target prosodic token. This not only enables automatic completion when prosodic information is missing or incomplete, but also improves generation efficiency and parallelism with the support of non-autoregressive generation, ensuring the consistency of prosody and semantics. The system coordinates and obtains the target user's reference speech, extracts the user style embedding from the reference speech, and pre-emphasizes and enhances high-frequency information, making subsequent feature extraction more sensitive to the user's unique timbre and speaking style. Multi-layer convolution captures deep features of speech in the time-frequency domain, and average pooling is used to obtain a robust initial style embedding. Stream matching is performed on the discrete content token, the target prosodic token, and the user style embedding to generate a target Mel spectrogram. The generation path is guided by a composite condition of semantics, prosody, and user style, achieving highly consistent control over content, prosody, and personalized style. Speech waveform reconstruction and optimization are performed on the target Mel spectrogram to obtain a zero-shot speech conversion result that is natural and realistic in content, prosody, and timbre, achieving high-quality personalized speech synthesis without additional labeled data. Finally, the zero-shot speech conversion result is output and fed back to the user client. The client can be, but is not limited to, various personal computers, laptops, smartphones, tablets, and portable wearable devices. The server can be implemented using a standalone server or a server cluster consisting of multiple servers. The following detailed description of specific embodiments further illustrates the invention.
[0019] The following is an explanation of this invention. This invention transforms noise vectors into target Mel spectrograms through stream matching, and guides the generation path using a combination of semantic, prosodic, and user-style conditions, achieving highly consistent control over content, prosodic, and personalized style. This not only improves the naturalness and fidelity of the generated speech, but also supports multi-dimensional conditional constraints, making speech synthesis more accurate and flexible, and enhancing the model's generalization ability under different speech inputs and styles. Finally, the target Mel spectrogram is reconstructed and optimized to obtain zero-sample speech conversion results that are natural and realistic in content, prosodic, and timbre, achieving high-quality personalized speech synthesis without the need for additional labeled data.
[0020] Reference Figure 2 The diagram shown is a flowchart illustrating a controllable zero-shot speech conversion method according to an embodiment of the present invention. In this embodiment, the controllable zero-shot speech conversion method includes: S1. Obtain several unlabeled speech data, and perform self-supervised speech learning on the unlabeled speech data to obtain a self-supervised speech representation.
[0021] In this embodiment of the invention, the speech data is denoised and normalized, acoustic waveform features are extracted, some features are randomly masked, and the masked features are input into a preset self-supervised model for mask prediction and analysis. The prediction results are then mapped to a self-supervised speech representation.
[0022] In specific healthcare scenarios, it can be used for voice recording analysis during patient follow-up. For example, it can be used to denoise, extract features, and perform self-supervised modeling on unlabeled voice data from doctor's consultations or patient self-reports, thereby automatically learning health status features in the voice and assisting in early disease screening, rehabilitation monitoring, and clinical decision support. In specific fintech scenarios, it can be applied to customer service or risk control scenarios. For example, it can perform self-supervised representation learning on voice data during customer inquiries, voice customer service interactions, or remote identity verification. It can automatically capture speaking habits, emotional characteristics, and potential risk signals in voice without extensive annotation, thereby improving the responsiveness of intelligent customer service and the accuracy of financial risk identification.
[0023] Figure 3 This is a flowchart illustrating the self-supervised speech learning process in a controllable zero-sample speech conversion method according to an embodiment of the present invention.
[0024] In this embodiment of the invention, the step of performing self-supervised speech learning on the unlabeled speech data to obtain a self-supervised speech representation includes: The unlabeled speech data is denoised to obtain denoised speech data; The denoised speech data is normalized to obtain normalized speech data; Extract the acoustic waveform features of the normalized speech data; Randomly mask a portion of the acoustic waveform features to obtain masked waveform features; The mask waveform features are analyzed using a pre-defined self-supervised model, and the mask analysis results are mapped to a self-supervised speech representation.
[0025] In detail, the original unlabeled speech data is preprocessed using methods such as filtering, spectral subtraction, or deep learning denoising models to remove background noise and interference components while retaining the effective information of the main speech content, thus obtaining clearer and purer denoised speech data. Audio sample values are scaled to a fixed range or the overall energy distribution is equalized to ensure consistency in volume and dynamic range across different speech data, resulting in normalized speech data.
[0026] By using methods such as short-time Fourier transform (STFT), Mel frequency cepstral coefficients (MFCC), Mel spectrum, or filter bank energy, time-domain and frequency-domain features are extracted from speech waveforms to generate acoustic waveform features that can reflect speech energy distribution, frequency structure, and temporal changes.
[0027] By randomly selecting a portion of time segments or frequency ranges from acoustic waveform features and replacing the feature values with mask symbols or gaps, an incomplete feature representation is formed. This allows the model to infer the masked content based on contextual information during training, thus obtaining masked waveform features. This provides a learning task for prediction and reconstruction for self-supervised learning.
[0028] The generated mask waveform features are input into a pre-defined self-supervised model. The model's contextual modeling capabilities are used to predict and reconstruct the masked feature regions. The prediction results are aligned or mapped with the original feature space, and a high-level vector representation that can characterize the overall content and potential structure of the speech is extracted, thus obtaining a self-supervised speech representation.
[0029] Without the need for manual annotation, effective speech representations are automatically learned from a large amount of raw speech. The quality of input data is ensured through noise reduction and normalization. The robustness and context modeling ability of the model are improved by using random masking and the prediction mechanism of self-supervised model, thereby obtaining self-supervised speech representations with greater generalization and semantic expressive power. This provides high-quality feature support for downstream tasks such as speech recognition, speaker verification, and sentiment analysis, and significantly reduces the cost of manual annotation.
[0030] S2. Extract the content feature vector and prosodic style vector of the self-supervised speech representation, and convert the content feature vector and the prosodic style vector into discrete content tokens and discrete prosodic tokens, respectively.
[0031] In this embodiment of the invention, a content encoder is used to extract the content features of speech, and Gaussian noise is added to enhance robustness. Discrete content tokens are obtained through vector quantization. A prosodic encoder is used to extract the prosodic style features of speech, and a style-supervised update loss function is combined to constrain the prosodic quantizer, so as to better preserve the prosodic features during quantization and obtain discrete prosodic tokens.
[0032] In specific healthcare scenarios, it can be used for intelligent processing of voice medical records and rehabilitation follow-ups. For example, it can decompose a patient's voice expression into discrete content tokens and discrete prosodic tokens, thereby capturing disease-related semantic information and prosodic features such as tone and rhythm, helping doctors to understand the patient's condition more accurately and assisting in the assessment of emotional state and the tracking of rehabilitation progress.
[0033] In specific fintech scenarios, it can be applied to voice customer service and risk monitoring. For example, customer interaction voice can be mapped into discrete content tokens and prosody tokens, which can extract business semantics for automatic understanding and business processing, and identify potential emotional fluctuations or abnormal voice patterns through prosodic style for customer satisfaction analysis and fraud risk warning.
[0034] In this embodiment of the invention, the step of extracting the content feature vector and prosodic style vector of the self-supervised speech representation, and converting the content feature vector and the prosodic style vector into discrete content tokens and discrete prosodic tokens, respectively, includes: The self-supervised speech representation content feature vector is extracted using a preset content encoder; Gaussian noise is added to the content feature vector to obtain the content constraint vector; The content constraint vector is vector quantized to obtain discrete content tokens; The prosodic style vector of the self-supervised speech representation is extracted using a preset prosodic encoder; Obtain the style analysis loss function, add style supervision to the style analysis loss function, and obtain the update loss function; The preset prosody vector quantizer is constrained using the update loss function to obtain a constrained prosody vector quantizer. The prosodic style vector is vector-quantized using the constrained prosodic vector quantizer to obtain a discrete prosodic token.
[0035] In detail, the self-supervised speech representation is input into a pre-defined content encoder. The encoder models and abstracts the temporal and frequency features of the speech, and extracts a high-dimensional vector representation that reflects the semantic information of the speech, namely the content feature vector. This vector is used to capture the language content in the speech while ignoring the differences in speaking style or prosody as much as possible, thus providing a semantic basis for downstream speech understanding and generation tasks.
[0036] Random Gaussian noise is introduced into the extracted content feature vector to slightly perturb the vector, thereby generating a content constraint vector. This enhances the robustness of the features, prevents the model from overfitting to the training data, and in the subsequent vector quantization process, enables the model to learn a smoother and more generalizable discrete content representation.
[0037] The content constraint vector, after being infused with Gaussian noise, is input into the vector quantization module. By matching it with a preset discrete codebook, the continuous high-dimensional vector is mapped to its corresponding discrete representation, thereby generating a discrete content token. This quantization can transform the continuous content features of speech into a processable symbolic sequence, providing a discretization basis for downstream speech synthesis, speech conversion, or semantic analysis.
[0038] The self-supervised speech representation is input into a pre-defined prosodic encoder. The encoder analyzes the prosodic features of the speech, such as rhythm, pitch, stress, and pauses, and extracts a high-dimensional vector that can represent speaking style and emotional expression, namely the prosodic style vector, which provides basic features for subsequent prosodic quantification and style modeling.
[0039] By introducing style supervision signals, such as labeled speaking styles or emotion categories, into the original style analysis loss function, and calculating the difference between the predicted prosodic features and the supervision target, an updated loss function is obtained. This can constrain learning when training the prosodic quantizer or related models, so that the extracted prosodic vectors can more accurately reflect the style features of speech.
[0040] The updated style analysis loss function is used as a constraint objective and applied to the training of the pre-defined prosodic vector quantizer. By minimizing the error between the predicted prosodic features and style supervision, the quantizer can more accurately preserve the rhythm, pitch and emotion features of speech when mapping the prosodic style vector into a discrete representation, thus obtaining the constrained prosodic vector quantizer.
[0041] The extracted prosodic style vector is input into a prosodic vector quantizer trained with style constraints. By matching continuous prosodic features with discrete codebooks in the quantizer, the vector is discretized, thereby generating discrete prosodic tokens. This symbolically represents the prosodic and style information of speech, making it easier for downstream speech conversion tasks to use.
[0042] By modeling through both content and prosody, self-supervised speech representation is efficiently transformed into discrete tokens: content features are obtained as robust discrete content tokens through Gaussian perturbation and quantization, while prosodic features are quantized into discrete prosodic tokens under style supervision constraints. This approach preserves the semantic information of speech while accurately expressing speaking style and prosodic features, improving the model's generalization ability while reducing its dependence on large-scale labeled data.
[0043] S3. Perform mask generation on the discrete prosody token to obtain the target prosody token.
[0044] In this embodiment of the invention, a portion of the prosodic token is randomly masked, and the speech content features generated by the discrete content token are fused with the original prosodic information to form a mask cue. The cue information is used to predict the distribution of candidate tokens at the mask positions under a non-autoregressive generation framework, and the prediction results are filled back into the mask positions, thereby completing the reconstruction of the prosodic features, obtaining the target prosodic token, and realizing the completion and optimization of speech prosodic information.
[0045] In specific healthcare scenarios, it can be applied to speech training in rehabilitation therapy. For example, it can mask and reconstruct the prosodic tokens of a patient's speech to identify and complete the prosodic loss or abnormality caused by the disease, assisting doctors in analyzing the recovery of the patient's speech rhythm and intonation, and improving the objectivity and accuracy of rehabilitation assessment.
[0046] In specific fintech scenarios, it can be used for intelligent customer service voice generation. By predicting and reconstructing the mask of prosodic tokens, the synthesized voice can have a natural and fluent tone and style while ensuring the accuracy of the content, thereby improving the customer interaction experience. It can also detect tone differences in risk scenarios to help identify potential anomalies or fraudulent behaviors.
[0047] In this embodiment of the invention, the step of masking the discrete prosodic token to obtain the target prosodic token includes: Masking a random portion of the discrete prosody token yields a masked prosody token; Generate speech content features based on the discrete content tokens; Content fusion is performed on the discrete prosody token and the speech content features to obtain mask position prompt information; The candidate token distribution is obtained by using the mask position hint information to generate the mask prosodic tokens non-autoregressively. The candidate token distribution is filled into the mask position of the mask prosody token to reconstruct the target prosody token.
[0048] In detail, several positions are selected from the complete prosodic token sequence according to a certain proportion or probability. The real prosodic tokens at these positions are removed and replaced with mask symbols, thus forming a masked prosodic token sequence. By searching the codebook vector or through embedding layer mapping, the symbolized discrete content tokens are restored to continuous representations, thereby obtaining speech content features that can characterize speech semantics and articulation structure.
[0049] By fusing two types of information—discrete prosodic tokens and speech content features generated from content tokens—through feature alignment, attention mechanisms, or feature concatenation, semantic content and prosodic style are expressed complementaryly in the same representation space. The fused result can provide contextual clues for mask location, indicating the semantic environment and style trend of missing prosody, thus providing accurate reference for subsequent mask prediction.
[0050] The fused mask position hints are input into a non-autoregressive generation model, which simultaneously predicts at all mask positions and directly outputs the candidate token distribution for each mask position. Unlike the stepwise generation autoregressive approach, non-autoregressive generation can fill multiple missing positions in parallel, significantly improving generation efficiency and providing diverse candidate prosodic tokens for each mask position while maintaining contextual consistency.
[0051] The candidate token distribution output by the non-autoregressive model is used as the completion result and sequentially filled into the masked prosodic token sequence at the masked positions. The specific filling value is determined by selecting the token with the highest probability or by using a sampling strategy, thereby restoring the complete prosodic token sequence. Through this reconstruction process, the originally missing prosodic information is filled in, and the target prosodic token is finally obtained.
[0052] By masking discrete prosodic tokens and combining them with content features for prediction and reconstruction, this approach not only enables automatic completion when prosodic information is missing or incomplete, but also improves generation efficiency and parallelism through non-autoregressive generation, ensuring consistency between prosody and semantics. The resulting target prosodic tokens retain the original prosodic style while possessing greater integrity and naturalness, thereby enhancing performance in tasks such as speech synthesis, conversion, and style transfer.
[0053] S4. Obtain the reference speech of the target user and extract the user style embedding from the reference speech.
[0054] In this embodiment of the invention, silence is removed and the sampling rate is unified. High-frequency components are pre-emphasized to enhance speech features. Mel spectrograms are extracted and timbre and speaking style features are obtained through multi-layer convolution. These features are aggregated using average pooling to form an initial style embedding. After embedding normalization, a user style embedding that can stably represent the timbre and speaking style of the target user is obtained.
[0055] In specific healthcare scenarios, it can be applied to remote rehabilitation or psychological intervention. For example, by extracting user style embeddings from patients' voices, it can capture timbre and speaking style features to identify voice abnormalities caused by disease or emotional changes, thereby assisting doctors in tracking disease progression, assessing emotions, and monitoring rehabilitation effects.
[0056] In specific fintech scenarios, it can be applied to voice identity authentication and risk control. For example, it can extract the style embedding of a user's voice to generate unique voiceprint features, which can be used for identity recognition in voice interaction, remote account opening or transaction verification. At the same time, it can be combined with speaking style features to monitor abnormal behavior and improve the security and reliability of the financial system.
[0057] In this embodiment of the invention, extracting the user style embedding from the reference speech includes: The silent segments are removed from the reference speech to obtain the effective speech segments; The effective speech segments are sampled in a normalized manner to obtain normalized speech segments; The high-frequency features of the normalized speech segment are pre-emphasized to obtain an enhanced speech segment; Extract the reference Mel spectrogram of the enhanced speech segment; Multi-layer convolution is performed on the reference Mel spectrogram to obtain the timbre and speaking style features of the target user; The timbre features and the speech style features are averaged and pooled to obtain the initial style embedding; The initial style embedding is normalized to obtain the user style embedding.
[0058] In detail, silence detection is performed on the reference speech signal. Methods such as energy thresholding, endpoint detection, or speech activity detection (VAD) are used to identify silent or invalid parts in the speech and remove them, retaining only the valid segments containing speech information, thereby reducing redundant data interference.
[0059] Normalized speech segments are obtained by unifying valid speech segments collected from different sources or devices to a preset sampling rate and resampling the audio signal through upsampling or downsampling to ensure that the speech data maintains consistency in temporal resolution and frequency distribution.
[0060] Applying pre-emphasis filtering to a normalized speech segment typically involves using a first-order high-pass filter to enhance the high-frequency components of the speech signal, thereby compensating for high-frequency attenuation caused by vocal tract characteristics during recording. After pre-emphasis, the high-frequency details of the speech are enhanced, improving speech clarity and feature discrimination, resulting in an enhanced speech segment more suitable for subsequent spectrum extraction and modeling.
[0061] The enhanced speech segment is transformed into a spectral representation through Short Time Fourier Transform (STFT). Then, a Mel filter bank is used to weight and map the spectral energy, resulting in a reference Mel spectrogram that conforms to human auditory perception characteristics, effectively representing the energy distribution and frequency structure of speech. Through layer-by-layer convolution, nonlinear activation, and feature extraction operations, the local and global patterns of speech in the time-frequency domain of the reference Mel spectrogram are gradually captured. High-dimensional vector representations characterizing the target user's timbre and speaking style are then separated and extracted, laying the foundation for constructing a unique user-specific speech style embedding.
[0062] By weighting and averaging the extracted timbre and speech style features along the time or frequency dimension, redundant information is compressed while global statistical features are preserved, thereby generating an initial style embedding representation that comprehensively reflects the overall timbre and speech style of the target user. The mean and variance of the initial style embedding representation are calculated and adjusted to a uniform numerical distribution range to eliminate scale differences between different speakers or speech segments, thus obtaining a more stable and comparable user style embedding representation.
[0063] By removing silence and normalizing the sampling rate, the validity and consistency of the input speech segments can be guaranteed. Pre-emphasis enhances high-frequency information, making subsequent feature extraction more sensitive to the user's unique timbre and speaking style. Multi-layer convolution captures deep features of speech in the time-frequency domain, and average pooling is used to obtain a robust initial style embedding. The final embedding normalization can eliminate the differences in scale and amplitude between different speech segments, making the obtained user style embedding more stable, robust, and transferable, which is beneficial to subsequent speech conversion tasks.
[0064] S5. Perform stream matching on the discrete content token, the target prosody token, and the user style embedding to generate a target Mel spectrogram.
[0065] In this embodiment of the invention, noise vectors are sampled from a standard Gaussian distribution as the starting point for stream matching. Discrete content tokens, target prosody tokens, and user style embeddings are mapped to semantic embedding vectors, rhythm and prosody control vectors, and global condition vectors, respectively, and then fused into composite conditions. Ordinary differential equations are used to learn the path of the noise vectors. Under the guidance of composite conditions, the distribution is gradually transformed, and finally, a target Mel spectrogram that meets the requirements of content, prosody, and style is generated.
[0066] In specific healthcare scenarios, it can be used for personalized speech rehabilitation or psychological intervention. By embedding the patient's speech content, prosodic features, and personal style into the stream matching to generate a Mel spectrogram, it can achieve natural and fluent speech reconstruction that preserves the patient's personal speech characteristics, assisting doctors in evaluating the effectiveness of rehabilitation training or monitoring emotional states.
[0067] In specific fintech scenarios, it can be applied to intelligent voice customer service and voice identity verification. By conditionally generating user voice content, rhythm and style features, it can achieve high-fidelity, natural voice synthesis or style-preserving voice reproduction, while also assisting in the identification of abnormal voice patterns, thereby improving customer experience and the security of financial transactions.
[0068] In this embodiment of the invention, the step of performing stream matching on the discrete content token, the target prosodic token, and the user style embedding to generate a target Mel spectrogram includes: Obtain a standard Gaussian distribution, sample a noise vector from the standard Gaussian distribution, and use the noise vector as the starting point for flow matching; Conditional embedding is performed on the discrete content token, the target prosody token, and the user style embedding to obtain a semantic embedding vector, a rhythm and prosody control vector, and a global conditional vector. The semantic embedding vector, the rhythm and prosody control vector, and the global condition vector are fused to obtain a composite condition. Obtain the ordinary differential equation, perform path learning on the flow matching starting point based on the ordinary differential equation, and guide the path learning process using the composite condition to finally obtain the target Mel spectrum map.
[0069] In detail, a standard Gaussian distribution with zero mean and one variance is defined. A high-dimensional noise vector is randomly sampled from the standard Gaussian distribution and used as the starting point for the flow matching process. The high-dimensional noise vector provides initial randomness and diversity, providing the basic input for subsequent path learning to generate the target Mel spectrum map through flow matching under conditional guidance.
[0070] Discrete content tokens, target prosody tokens, and user style embeddings are input into a pre-defined embedding network or encoder. The symbolic content tokens are mapped to semantic embedding vectors, the prosody tokens are mapped to rhythm and prosody control vectors, and the style embeddings are mapped to global conditional vectors, thereby transforming discrete and continuous features into high-dimensional vector representations that can be used for flow matching condition guidance.
[0071] Semantic embedding vectors, rhythm and prosody control vectors, and global condition vectors are fused through feature concatenation, weighted summation, or attention mechanisms to form a unified composite condition vector. The composite condition vector can simultaneously contain speech content, prosody control, and user style information, providing comprehensive condition guidance for the stream matching generation process.
[0072] A flow matching dynamic is established based on a pre-defined ordinary differential equation (ODE). Starting with a noise vector sampled from a Gaussian distribution, a continuous path from noise to the target distribution is learned by solving the ODE. During path learning, a composite conditional vector is introduced to guide the generation process, simultaneously constraining the final generated sequence by speech content, prosodic control, and user style features. After iterative transformation, the target Mel spectrogram that meets the conditions is obtained. The calculation formula is shown below:
[0073]
[0074] in, Represents a semantic embedding vector. Represents the rhythm and prosody control vector. Represents the global condition vector. Indicates a compound condition. The path learning time is indicated as Time-Mel spectrum state, =0 Indicates the starting point of the stream matching. Indicates the path learning time. Represents a parameterized vector field. Indicates a parameter.
[0075] By progressively transforming the noise vector into the target Mel spectrogram through stream matching, and guiding the generation path with a combination of semantic, prosodic, and user-style conditions, a high degree of consistent control over content, prosodic, and personalized style can be achieved. This not only improves the naturalness and fidelity of the generated speech, but also supports multi-dimensional conditional constraints, making speech synthesis more accurate and flexible, and enhancing the model's generalization ability under different speech inputs and styles.
[0076] S6. Reconstruct and optimize the speech waveform of the target Mel spectrogram to obtain the zero-sample speech conversion result.
[0077] In this embodiment of the invention, the target Mel spectrogram is denormalized and numerically cropped to obtain a safe and stable spectral representation. The cropped Mel spectrogram is converted into a time-domain waveform by a neural vocoder, and the waveform is subjected to noise reduction to improve sound quality. The waveform is then fine-tuned for prosodic consistency and enhanced for clarity, so that the speech retains its original content while maintaining a natural style and harmonious rhythm, ultimately generating a high-fidelity, zero-sample speech conversion result.
[0078] In specific healthcare scenarios, it can be used for remote speech rehabilitation and assisted diagnosis. For example, the patient's speech input can be processed through Mel spectrum reconstruction, noise reduction, and prosody fine-tuning to generate natural and fluent speech samples, which can be used to analyze the progress of speech recovery, assess pronunciation ability, or psychological state. At the same time, it does not require a large amount of labeled data, thus enabling personalized rehabilitation training.
[0079] In specific fintech scenarios, it can be applied to intelligent customer service and voice synthesis risk control. For example, customer voice can be converted into high-fidelity, consistent voice output through zero-sample conversion for multi-channel voice interaction or simulated voice verification. At the same time, the naturalness of the voice and the reliability of recognition can be improved through prosody and style fine-tuning, which helps to improve customer experience and transaction security.
[0080] In this embodiment of the invention, the step of reconstructing and optimizing the speech waveform of the target Mel spectrogram to obtain a zero-sample speech conversion result includes: The target Mel spectrum is denormalized to obtain the restored Mel spectrum; Numerical safe cropping and dynamic range limiting are applied to the restored Mel spectrum to obtain the cropped Mel spectrum; The cropped Mel spectrum is reconstructed using a preset neural vocoder to obtain a time-domain waveform; The time-domain waveform is denoised to obtain a denoised waveform; The denoised waveform is then fine-tuned for rhythm consistency to obtain the fine-tuned waveform; The fine-tuned waveform is then enhanced for clarity to obtain zero-sample speech conversion results.
[0081] In detail, the target Mel spectrogram is restored from the normalized numerical range to the physical magnitude of the original spectrum. The amplitude and energy distribution are recovered through inverse normalization, so that the Mel spectrogram can reflect the true frequency characteristics of the actual speech signal, providing an accurate spectral basis for subsequent waveform reconstruction and optimization.
[0082] The amplitude of the restored Mel spectrogram exceeding a preset threshold is limited to a safe range, and the dynamic range is normalized to prevent extreme values or noise from interfering with the waveform reconstruction of the subsequent neural vocoder, thereby obtaining a stable and safe-to-process cropped Mel spectrogram.
[0083] By using a neural network model trained in a neural vocoder, the spectral features of the cropped Mel spectrogram are mapped back to the time domain signal. By learning the nonlinear conversion relationship between frequency and waveform, the reconstruction from the Mel spectrum to a continuous audio waveform is achieved, thus obtaining a playable time-domain speech waveform.
[0084] Background noise and interference signals in the reconstructed time-domain waveform are removed by spectral subtraction, neural network denoising models, or filters, resulting in a clear, noise-suppressed denoised waveform. This provides a cleaner speech foundation for subsequent prosodic fine-tuning and speech enhancement. By analyzing and optimizing prosodic features such as rhythm, pitch, stress, and pauses, the waveform maintains consistency with the target prosodic and speaking style while preserving the original content, thus obtaining a prosodic, natural, and fluent fine-tuned waveform.
[0085] By enhancing the speech details and pronunciation clarity of the waveform after prosody fine-tuning through high-frequency enhancement, amplitude equalization, or deep neural network enhancement models, high-fidelity, natural and fluent zero-sample speech conversion results are generated, so that the final speech achieves optimized effects in terms of content, prosody and sound quality.
[0086] By ensuring the stability and security of the Mel spectrum through inverse normalization and numerical pruning, the neural vocoder achieves high-fidelity spectrum-to-waveform conversion, noise reduction improves speech clarity, prosody fine-tuning maintains consistency in speech rhythm and style, and clarity enhancement further optimizes sound quality, so that the final zero-sample speech conversion result is natural and realistic in content, prosody and timbre, achieving high-quality personalized speech synthesis without additional labeled data.
[0087] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0088] like Figure 4 The diagram shown is a functional block diagram of a controllable zero-sample speech conversion device provided in an embodiment of the present invention.
[0089] In this embodiment of the disclosure, a controllable zero-sample speech conversion device is provided, which corresponds one-to-one with the controllable zero-sample speech conversion method described in the above embodiments. For example... Figure 4 As shown, this controllable zero-shot speech conversion device 100 can be installed in an electronic device. According to its functions, the controllable zero-shot speech conversion device 100 includes a self-supervised speech learning module 101, a content and prosody token generation module 102, a prosody token masking module 103, a user style embedding extraction module 104, a target Mel spectrogram generation module 105, and a speech waveform reconstruction module 106. Detailed descriptions of each functional module are as follows: The self-supervised speech learning module 101 is used to acquire several unlabeled speech data, perform self-supervised speech learning on the unlabeled speech data, and obtain a self-supervised speech representation. The content and prosodic token generation module 102 is used to extract the content feature vector and prosodic style vector of the self-supervised speech representation, and convert the content feature vector and the prosodic style vector into discrete content tokens and discrete prosodic tokens, respectively. The prosody token masking module 103 is used to generate a mask for the discrete prosody token to obtain the target prosody token. User style embedding extraction module 104 is used to obtain the reference speech of the target user and extract the user style embedding in the reference speech. The target Mel spectrogram generation module 105 is used to perform stream matching on the discrete content token, the target prosody token and the user style embedding to generate a target Mel spectrogram. The speech waveform reconstruction module 106 is used to reconstruct and optimize the speech waveform of the target Mel spectrogram to obtain zero-sample speech conversion results.
[0090] In one embodiment, the self-supervised speech learning module 101 performs self-supervised speech learning on the unlabeled speech data to obtain a self-supervised speech representation, including: The unlabeled speech data is denoised to obtain denoised speech data; The denoised speech data is normalized to obtain normalized speech data; Extract the acoustic waveform features of the normalized speech data; Randomly mask a portion of the acoustic waveform features to obtain masked waveform features; The mask waveform features are analyzed using a pre-defined self-supervised model, and the mask analysis results are mapped to a self-supervised speech representation.
[0091] In one embodiment, the content and prosodic token generation module 102 extracts the content feature vector and prosodic style vector of the self-supervised speech representation, and converts the content feature vector and the prosodic style vector into discrete content tokens and discrete prosodic tokens, respectively, including: The self-supervised speech representation content feature vector is extracted using a preset content encoder; Gaussian noise is added to the content feature vector to obtain the content constraint vector; The content constraint vector is vector quantized to obtain discrete content tokens; The prosodic style vector of the self-supervised speech representation is extracted using a preset prosodic encoder; Obtain the style analysis loss function, add style supervision to the style analysis loss function, and obtain the update loss function; The preset prosody vector quantizer is constrained using the update loss function to obtain a constrained prosody vector quantizer. The prosodic style vector is vector-quantized using the constrained prosodic vector quantizer to obtain a discrete prosodic token.
[0092] In one embodiment, the prosody token masking module 103 performs mask generation on the discrete prosody tokens to obtain a target prosody token, including: Masking a random portion of the discrete prosody token yields a masked prosody token; Generate speech content features based on the discrete content tokens; Content fusion is performed on the discrete prosody token and the speech content features to obtain mask position prompt information; The candidate token distribution is obtained by using the mask position hint information to generate the mask prosodic tokens non-autoregressively. The candidate token distribution is filled into the mask position of the mask prosody token to reconstruct the target prosody token.
[0093] In one embodiment, the user style embedding extraction module 104, when performing the extraction of user style embedding from the reference speech, includes: The silent segments are removed from the reference speech to obtain the effective speech segments; The effective speech segments are sampled in a normalized manner to obtain normalized speech segments; The high-frequency features of the normalized speech segment are pre-emphasized to obtain an enhanced speech segment; Extract the reference Mel spectrogram of the enhanced speech segment; Multi-layer convolution is performed on the reference Mel spectrogram to obtain the timbre and speaking style features of the target user; The timbre features and the speech style features are averaged and pooled to obtain the initial style embedding; The initial style embedding is normalized to obtain the user style embedding.
[0094] In one embodiment, the target Mel spectrogram generation module 105 generates a target Mel spectrogram by performing stream matching on the discrete content token, the target prosodic token, and the user style embedding, including: Obtain a standard Gaussian distribution, sample a noise vector from the standard Gaussian distribution, and use the noise vector as the starting point for flow matching; Conditional embedding is performed on the discrete content token, the target prosody token, and the user style embedding to obtain a semantic embedding vector, a rhythm and prosody control vector, and a global conditional vector. The semantic embedding vector, the rhythm and prosody control vector, and the global condition vector are fused to obtain a composite condition. Obtain the ordinary differential equation, perform path learning on the flow matching starting point based on the ordinary differential equation, and guide the path learning process using the composite condition to finally obtain the target Mel spectrum map.
[0095] In one embodiment, the speech waveform reconstruction module 106 performs speech waveform reconstruction and optimization on the target Mel spectrogram to obtain a zero-sample speech conversion result, including: The target Mel spectrum is denormalized to obtain the restored Mel spectrum; Numerical safe cropping and dynamic range limiting are applied to the restored Mel spectrum to obtain the cropped Mel spectrum; The cropped Mel spectrum is reconstructed using a preset neural vocoder to obtain a time-domain waveform; The time-domain waveform is denoised to obtain a denoised waveform; The denoised waveform is then fine-tuned for rhythm consistency to obtain the fine-tuned waveform; The fine-tuned waveform is then enhanced for clarity to obtain zero-sample speech conversion results.
[0096] In this invention, a controllable zero-sample speech conversion device is first developed by acquiring several unlabeled speech data sets and performing self-supervised speech learning on these data sets. The robustness and contextual modeling capabilities of the model are improved using random masking and the prediction mechanism of the self-supervised model, thereby obtaining a more generalizable and semantically expressive self-supervised speech representation. The content feature vector and prosodic style vector of the self-supervised speech representation are extracted and converted into discrete content tokens and discrete prosodic tokens, respectively. This process preserves the semantic information of the speech while accurately expressing speaking style and prosodic features, improving the model's generalization ability while reducing its dependence on large-scale labeled data. The discrete prosodic tokens are then masked to generate target prosodic tokens, which not only automatically complete prosodic information when it is missing or incomplete, but also benefit from non-autoregressive generation. To improve generation efficiency and parallelism, ensuring consistency between prosody and semantics, the system acquires reference speech from the target user and extracts user style embeddings from it. High-frequency information is pre-emphasized to enhance subsequent feature extraction, making it more sensitive to the user's unique timbre and speaking style. Multi-layer convolution captures deep features of speech in the time-frequency domain, combined with average pooling to obtain robust initial style embeddings. Stream matching is performed on the discrete content token, the target prosody token, and the user style embeddings to generate a target Mel spectrogram. The generation path is guided by a composite condition of semantics, prosody, and user style, achieving highly consistent control over content, prosody, and personalized style. Finally, the target Mel spectrogram is reconstructed and optimized to obtain zero-shot speech conversion results. These results are natural and realistic in content, prosody, and timbre, achieving high-quality personalized speech synthesis without additional labeled data. Specific limitations of a controllable zero-shot speech conversion device can be found in the limitations of a controllable zero-shot speech conversion method described above, and will not be repeated here. Each module in the aforementioned controllable zero-shot speech conversion device can be implemented entirely or partially through software, hardware, or a combination thereof. The above modules can be embedded in the processor of a computer device in hardware form or independent of it, or they can be stored in the memory of a computer device in software form, so that the processor can call and execute the operations corresponding to each of the above modules.
[0097] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 5As shown, the computer device includes a processor, memory, network interface, and database connected via a system bus. The processor provides computational and control capabilities. The memory includes non-volatile and / or volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The network interface is used to communicate with external clients via a network connection. When the computer program is executed by the processor, it implements the functions or steps of a controllable zero-sample speech conversion method on the server side.
[0098] In one embodiment, a computer device is provided, which may be a client, and its internal structure diagram may be as follows: Figure 6 As shown, the computer device includes a processor, memory, network interface, display screen, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The network interface is used to communicate with an external server via a network connection. When the computer program is executed by the processor, it implements the client-side functions or steps of a controllable zero-sample speech conversion method.
[0099] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to perform the following steps: Acquire a number of unlabeled speech data, perform self-supervised speech learning on the unlabeled speech data, and obtain self-supervised speech representation; Extract the content feature vector and prosodic style vector of the self-supervised speech representation, and convert the content feature vector and the prosodic style vector into discrete content tokens and discrete prosodic tokens, respectively. The discrete prosodic token is masked to obtain the target prosodic token; Obtain the reference speech of the target user and extract the user style embedding from the reference speech; Stream matching is performed on the discrete content token, the target prosodic token, and the user style embedding to generate a target Mel spectrogram; The target Mel spectrogram is reconstructed and optimized to obtain zero-sample speech conversion results.
[0100] In the several embodiments provided by this invention, it should be understood that the disclosed devices and apparatuses can be implemented in other ways. For example, the system embodiments described above are merely illustrative; for instance, the division of modules is only a logical functional division, and other division methods may be used in actual implementation.
[0101] Furthermore, the functional modules in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or in the form of hardware plus software functional modules.
[0102] Therefore, the embodiments should be considered exemplary and non-limiting in all respects, and the scope of the invention is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of equivalents of the claims are intended to be embraced within the invention. No appended diagram markings in the claims should be construed as limiting the scope of the claims.
[0103] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention.
[0104] In some embodiments of this example, a computer-readable storage medium is provided, on which a computer program is stored, characterized in that the computer program, when executed by a processor, implements the steps of the method described in the above embodiments.
[0105] The readable storage medium of the present invention stores a computer program, which, when executed by a processor of an electronic device, can perform the following: Acquire a number of unlabeled speech data, perform self-supervised speech learning on the unlabeled speech data, and obtain self-supervised speech representation; Extract the content feature vector and prosodic style vector of the self-supervised speech representation, and convert the content feature vector and the prosodic style vector into discrete content tokens and discrete prosodic tokens, respectively. The discrete prosodic token is masked to obtain the target prosodic token; Obtain the reference speech of the target user and extract the user style embedding from the reference speech; Stream matching is performed on the discrete content token, the target prosodic token, and the user style embedding to generate a target Mel spectrogram; The target Mel spectrogram is reconstructed and optimized to obtain zero-sample speech conversion results.
[0106] It should be noted that the functions or steps that can be implemented by the computer-readable storage medium or computer device described above can be referred to the relevant descriptions on the server side and client side in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.
[0107] Computer-readable storage media may also store at least one computer-executable program / instruction, such as computer-readable instructions. Computer-readable storage media include, but are not limited to, volatile memory and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) and / or cache memory. Computer-readable storage media may include, for example, read-only memory (ROM), hard disk, flash memory, etc. For example, a non-transitory computer-readable storage medium may be connected to a computing device such as a computer, and then, when the computing device executes the computer-readable instructions stored on the computer-readable storage medium, the various methods described above can be performed.
[0108] In addition, the computer device may include (but is not limited to) a data bus, an input / output (I / O) bus, a display, and input / output devices (e.g., keyboard, mouse, speakers, etc.).
[0109] The processor can communicate with external devices via the I / O bus through wired or wireless networks.
[0110] In one embodiment, the at least one computer-executable instruction may also be compiled into or comprise a software product / computer program product, wherein one or more computer-executable instructions are executed by a processor to perform the steps of the various functions and / or methods in the embodiments described herein.
[0111] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0112] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0113] In the embodiments provided in this disclosure, it should be understood that the disclosed apparatus and methods can also be implemented in other ways. The apparatus embodiments described above are merely illustrative; for example, the flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of apparatus, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than those marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram and / or flowchart, and combinations of blocks in block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.
[0114] It should be noted that, in this disclosure, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element limited by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
[0115] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.
[0116] It should be noted that if any software tools or components not belonging to our company appear in the embodiments of this application, they are merely for illustrative purposes and do not represent actual use.
Claims
1. A controllable zero-sample speech conversion method, characterized in that, The method includes: Acquire a number of unlabeled speech data, perform self-supervised speech learning on the unlabeled speech data, and obtain self-supervised speech representation; Extracting the content feature vector and prosodic style vector of the self-supervised speech representation, and converting the content feature vector and the prosodic style vector into discrete content tokens and discrete prosodic tokens respectively, includes: extracting the content feature vector of the self-supervised speech representation using a preset content encoder; adding Gaussian noise to the content feature vector to obtain a content constraint vector; performing vector quantization on the content constraint vector to obtain a discrete content token; extracting the prosodic style vector of the self-supervised speech representation using a preset prosodic encoder; obtaining a style analysis loss function, adding style supervision to the style analysis loss function to obtain an update loss function; constraining a preset prosodic vector quantizer using the update loss function to obtain a constrained prosodic vector quantizer; and performing vector quantization on the prosodic style vector using the constrained prosodic vector quantizer to obtain a discrete prosodic token. The discrete prosodic token is masked to obtain the target prosodic token; Obtain the reference speech of the target user and extract the user style embedding from the reference speech; The process of performing flow matching on the discrete content token, the target prosodic token, and the user style embedding to generate a target Mel spectrogram includes: obtaining a standard Gaussian distribution; sampling a noise vector from the standard Gaussian distribution and using the noise vector as the starting point for flow matching; performing conditional embedding on the discrete content token, the target prosodic token, and the user style embedding to obtain a semantic embedding vector, a rhythm and prosodic control vector, and a global condition vector; performing conditional fusion on the semantic embedding vector, the rhythm and prosodic control vector, and the global condition vector to obtain a composite condition; obtaining an ordinary differential equation; performing path learning on the starting point for flow matching based on the ordinary differential equation; and using the composite condition to guide the path learning process to finally obtain the target Mel spectrogram. The target Mel spectrogram is reconstructed and optimized to obtain zero-sample speech conversion results.
2. The controllable zero-sample speech conversion method as described in claim 1, characterized in that, The step of performing self-supervised speech learning on the unlabeled speech data to obtain a self-supervised speech representation includes: The unlabeled speech data is denoised to obtain denoised speech data; The denoised speech data is normalized to obtain normalized speech data; Extract the acoustic waveform features of the normalized speech data; Randomly mask a portion of the acoustic waveform features to obtain masked waveform features; The mask waveform features are analyzed using a pre-defined self-supervised model, and the mask analysis results are mapped to a self-supervised speech representation.
3. The controllable zero-sample speech conversion method as described in claim 1, characterized in that, The process of masking the discrete prosodic tokens to obtain the target prosodic token includes: Masking a random portion of the discrete prosody token yields a masked prosody token; Generate speech content features based on the discrete content tokens; Content fusion is performed on the discrete prosody token and the speech content features to obtain mask position prompt information; The candidate token distribution is obtained by using the mask position hint information to generate the mask prosodic tokens non-autoregressively. The candidate token distribution is filled into the mask position of the mask prosody token to reconstruct the target prosody token.
4. The controllable zero-sample speech conversion method as described in claim 1, characterized in that, The step of extracting the user style embedding from the reference speech includes: The silent segments are removed from the reference speech to obtain the effective speech segments; The effective speech segments are sampled in a normalized manner to obtain normalized speech segments; The high-frequency features of the normalized speech segment are pre-emphasized to obtain an enhanced speech segment; Extract the reference Mel spectrogram of the enhanced speech segment; Multi-layer convolution is performed on the reference Mel spectrogram to obtain the timbre and speaking style features of the target user; The timbre features and the speech style features are averaged and pooled to obtain the initial style embedding; The initial style embedding is normalized to obtain the user style embedding.
5. The controllable zero-sample speech conversion method as described in claim 1, characterized in that, The step of reconstructing and optimizing the speech waveform of the target Mel spectrogram to obtain zero-sample speech conversion results includes: The target Mel spectrum is denormalized to obtain the restored Mel spectrum; Numerical safe cropping and dynamic range limiting are applied to the restored Mel spectrum to obtain the cropped Mel spectrum; The cropped Mel spectrum is reconstructed using a preset neural vocoder to obtain a time-domain waveform; The time-domain waveform is denoised to obtain a denoised waveform; The denoised waveform is then fine-tuned for rhythm consistency to obtain the fine-tuned waveform; The fine-tuned waveform is then enhanced for clarity to obtain zero-sample speech conversion results.
6. A controllable zero-sample speech conversion device, used to implement the controllable zero-sample speech conversion method as described in any one of claims 1 to 5, characterized in that, The device includes: The self-supervised speech learning module is used to acquire several unlabeled speech data, perform self-supervised speech learning on the unlabeled speech data, and obtain a self-supervised speech representation. The content and prosodic token generation module is used to extract the content feature vector and prosodic style vector of the self-supervised speech representation, and convert the content feature vector and the prosodic style vector into discrete content tokens and discrete prosodic tokens, respectively. The prosody token masking module is used to generate a mask for the discrete prosody tokens to obtain the target prosody tokens; The user style embedding extraction module is used to obtain the reference speech of the target user and extract the user style embedding from the reference speech. The target Mel spectrogram generation module is used to perform stream matching on the discrete content token, the target prosody token, and the user style embedding to generate a target Mel spectrogram. The speech waveform reconstruction module is used to reconstruct and optimize the speech waveform of the target Mel spectrogram to obtain zero-sample speech conversion results.
7. An electronic device, characterized in that, The electronic device includes: At least one processor; and, A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the controllable zero-sample speech conversion method as described in any one of claims 1 to 5.
8. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the controllable zero-sample speech conversion method as described in any one of claims 1 to 5.