Voice style migration method and device, equipment and medium

By extracting and separating speech features, using time parameters and flow matching models for speech style transfer, the problem of lack of cyclic consistency constraints in the existing technology is solved, and the stability and naturalness of speech style transfer are improved.

CN120220653APending Publication Date: 2025-06-27PING AN TECH (SHENZHEN) CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510417726.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-03
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The prior art lacks cyclic consistency constraints in the process of speech style transfer, which makes it difficult for the speech after style conversion to maintain semantic and style consistency, and when transferring unseen speaker styles, the generation quality is degraded and the stability is insufficient.

Method used

By extracting the characteristics of source and reference speech, separating the content and style characteristics, and using time parameters to perform linear interpolation processing, intermediate characteristics are generated. Then, the intermediate features are input into the stream matching model, the intermediate reference features are generated and the source features are reconstructed, the loop consistency loss is calculated, and the parameters of the stream matching model are optimized to improve the stability of style migration.

Benefits of technology

Ensure that the speech after style transfer can restore the original speech in the reverse process, improve the stability of semantics and style, improve the naturalness and quality of generated speech, reduce dependence on large-scale annotation data, and improve the generalization ability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120220653A_ABST
    Figure CN120220653A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of voice processing, can be applied to business scenes such as financial science and technology and medical health, and discloses a voice style migration method which comprises the following steps: extracting features of a source voice and a reference voice, and separating content features and style features; performing linear interpolation on the initial source features based on the time parameters to generate intermediate features; inputting the intermediate features into a stream matching model to generate reference features and reconstruction features; calculating circulation consistency loss, and optimizing flow matching model parameters based on the loss; and applying the optimized model to style migration to generate a migration voice waveform. According to the method, the semantic and style consistency of the voice is ensured through the loop consistency loss constraint style migration process, the conversion smoothness is improved in combination with the time interpolation processing, the cross-speaker style migration is realized by using the stream matching model, the style adaptability of the unseen speaker is improved, the model optimization reduces the dependence on the annotation data, and the user experience is improved. And the stability and naturalness of the generated voice are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of speech processing technology, and in particular, to a speech style transfer method, device, equipment and storage medium. Background Art

[0002] In the field of speech synthesis, research on speaker style adaptation mainly focuses on style transfer based on multi-speaker models, style adaptation based on self-supervised learning, and style transfer based on generative models. However, these existing technologies still have significant deficiencies in terms of style adaptation ability, speech consistency guarantee, and computational resource consumption, which limit their feasibility in actual application scenarios.

[0003] In the field of healthcare business, speech technology can be used in applications such as health consultation, telemedicine, and rehabilitation assistance. Existing style adaptation methods based on self-supervised learning, such as HuBERT, WavLM, etc., usually learn general speech representations through pre-trained models and then adjust the speech style. However, such models are usually large in scale and slow in inference speed, making it difficult to meet the requirements of real-time speech interaction. In addition, the current technology is difficult to ensure the consistency of semantics and style after speech style conversion, resulting in a decrease in speech naturalness in scenarios such as patient consultation and health broadcast, which may affect the transmission and understanding of information.

[0004] In the field of fintech business, intelligent speech synthesis is widely used in scenarios such as intelligent customer service, financial assistants, and risk warnings. However, existing style transfer methods often rely on a large amount of labeled data to extract style embeddings or conditional encodings from multi-speaker corpora for style feature transfer, such as methods like GlobalStyleToken (GST) and ProsodyTransfer. But when these methods transfer unseen speaker styles, the generation quality and style consistency significantly decrease, resulting in the synthesized speech being difficult to maintain a stable tone and clear expression in the financial customer service scenario. In addition, generative models (such as variational autoencoders or normalizing flows) can provide a certain degree of style control in the field of financial speech reporting, but their instability may affect the reliability of the speech and thus affect the user experience.

[0005] In the core technology field of speech synthesis, existing methods generally lack cyclic consistency constraints, making it difficult for the speech after style transfer to be restored to the original speech through the inverse process. This defect may cause the generated speech to lose some key information, affecting semantic accuracy. In addition, current style transfer methods based on generative models require high computational resources and are limited by computing power when deployed on the cloud or embedded devices, making it difficult to meet the real-time application requirements of high concurrency and multiple terminals. Summary of the Invention

[0006] The main objective of the present invention is to provide a voice style transfer method, device, equipment, and storage medium, aiming to solve the technical problems in the prior art that during the voice style transfer process, there is a lack of cyclic consistency constraints, resulting in the difficulty of maintaining semantic and style consistency in the voice after style conversion, and when transferring the styles of unseen speakers, the generation quality deteriorates and the stability is insufficient.

[0007] To achieve the above objective, the present invention provides a voice style transfer method, including:

[0008] Extract the initial source features of the source voice, and separate the content features and the first style features from the initial source features;

[0009] Extract the initial reference features of the reference voice, and separate the second style features from the initial reference features;

[0010] Perform linear interpolation processing on the initial source features according to the time parameter to generate intermediate source features;

[0011] Input the intermediate source features, the second style features, and the content features into the flow matching model to generate intermediate reference features;

[0012] Input the intermediate reference features, the first style features, and the content features into the flow matching model to generate reconstructed source features;

[0013] Determine the first cyclic consistency loss between the reconstructed source features and the initial source features;

[0014] Perform linear interpolation processing on the initial reference features according to the time parameter to generate verification intermediate features;

[0015] Input the verification intermediate features, the second style features, and the content features into the flow matching model to generate verification reference features;

[0016] Determine the second cyclic consistency loss between the verification reference features and the intermediate reference features;

[0017] Jointly optimize the parameters of the flow matching model with the first cyclic consistency loss and the second cyclic consistency loss;

[0018] Input the initial source features, the content features, and the second style features into the optimized flow matching model to generate a transferred voice waveform.

[0019] Furthermore, to achieve the above objective, the present invention provides a voice style transfer device, including:

[0020] A source speech feature extraction module, configured to extract initial source features of a source speech, and separate content features and first style features from the initial source features;

[0021] A reference speech style extraction module, configured to extract initial reference features of a reference speech, and separate second style features from the initial reference features;

[0022] A time interpolation processing module, configured to perform linear interpolation processing on the initial source features according to time parameters to generate intermediate source features;

[0023] A flow matching mapping module, configured to input the intermediate source features, the second style features, and the content features into a flow matching model to generate intermediate reference features;

[0024] A reconstructed feature generation module, configured to input the intermediate reference features, the first style features, and the content features into the flow matching model to generate reconstructed source features;

[0025] A consistency loss analysis module, configured to determine a first cycle consistency loss between the reconstructed source features and the initial source features;

[0026] A reference feature interpolation module, configured to perform linear interpolation processing on the initial reference features according to the time parameters to generate verification intermediate features;

[0027] A verification feature mapping module, configured to input the verification intermediate features, the second style features, and the content features into the flow matching model to generate verification reference features;

[0028] A verification loss analysis module, configured to determine a second cycle consistency loss between the verification reference features and the intermediate reference features;

[0029] A model optimization module, configured to jointly optimize parameters of the flow matching model according to the first cycle consistency loss and the second cycle consistency loss;

[0030] A speech generation module, configured to input the initial source features, the content features, and the second style features into the optimized flow matching model to generate a migrated speech waveform.

[0031] Further, to achieve the above object, the present invention further provides a computer device, where the computer device includes a memory, a processor, and a speech style migration program stored on the memory and executable on the processor. When the speech style migration program is executed by the processor, the steps of the speech style migration method as described above are implemented.

[0032] Further, to achieve the above object, the present invention also provides a computer-readable storage medium, on which a voice style transfer program is stored. When the voice style transfer program is executed by a processor, the steps of the voice style transfer method as described above are implemented.

[0033] Beneficial effects: The present invention relates to the technical field of speech processing and can be applied to business scenarios such as fintech and healthcare. A voice style transfer method is disclosed, including: extracting the initial source features of the source speech and separating the content features and the first style features; extracting the initial reference features of the reference speech and separating the second style features; performing linear interpolation processing on the initial source features based on time parameters to generate intermediate source features; inputting the intermediate source features, the second style features, and the content features into a flow matching model to generate intermediate reference features; inputting the intermediate reference features, the first style features, and the content features into the flow matching model to generate reconstructed source features; calculating the first cycle consistency loss between the reconstructed source features and the initial source features; performing linear interpolation processing on the initial reference features based on time parameters to generate verification intermediate features; inputting the verification intermediate features, the second style features, and the content features into the flow matching model to generate verification reference features; calculating the second cycle consistency loss between the verification reference features and the intermediate reference features; jointly optimizing the parameters of the flow matching model with the first cycle consistency loss and the second cycle consistency loss; inputting the initial source features, the content features, and the second style features into the optimized flow matching model to generate a transferred speech waveform. By introducing the cycle consistency loss, the present invention ensures that the speech after style transfer can recover the original speech in the reverse process, improves the stability of semantics and style, uses time parameters for linear interpolation processing, enables the model to have a smooth transition ability during the style conversion process, enhances the naturalness of the generated speech, realizes the adaptive transfer of cross-speaker styles through the flow matching model, enables the styles of unseen speakers to still maintain high-quality conversion effects, combines the first cycle consistency loss and the second cycle consistency loss to optimize the model parameters, improves the stability of style conversion, reduces the dependence on large-scale labeled data, and enhances the generalization ability of the model. Description of the Drawings

[0034] The following will further illustrate the present invention in conjunction with the drawings. In the drawings:

[0035] Figure 1 is a schematic diagram of an application environment of the voice style transfer method in an embodiment of the present invention;

[0036] Figure 2 is a schematic flowchart of an embodiment of the voice style transfer method of the present invention;

[0037] Figure 3 is a schematic diagram of the functional modules of a preferred embodiment of the voice style transfer device of the present invention;

[0038] Figure 4 A schematic structural diagram of a computer device in an embodiment of the present invention;

[0039] Figure 5 Another schematic structural diagram of a computer device in an embodiment of the present invention. Detailed implementation manners

[0040] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0041] The voice style transfer method provided by the embodiment of the present invention can be applied in an application environment such as Figure 1 . In this environment, the client communicates with the server through a network. The server can extract the initial source features of the source voice from the client, separate the content features and the first style features; extract the initial reference features of the reference voice, separate the second style features; perform linear interpolation processing on the initial source features based on time parameters to generate intermediate source features; input the intermediate source features, the second style features and the content features into the flow matching model to generate intermediate reference features; input the intermediate reference features, the first style features and the content features into the flow matching model to generate reconstructed source features; calculate the first cycle consistency loss between the reconstructed source features and the initial source features; perform linear interpolation processing on the initial reference features based on time parameters to generate verification intermediate features; input the verification intermediate features, the second style features and the content features into the flow matching model to generate verification reference features; calculate the second cycle consistency loss between the verification reference features and the intermediate reference features; jointly optimize the parameters of the flow matching model with the first cycle consistency loss and the second cycle consistency loss; input the initial source features, the content features and the second style features into the optimized flow matching model to generate a transferred voice waveform. By introducing the cycle consistency loss, the present invention ensures that the voice after style transfer can be restored to the original voice in the reverse process, improving the stability of semantics and style. Using time parameters for linear interpolation processing enables the model to have a smooth transition ability during the style conversion process, enhancing the naturalness of the generated voice. The cross-speaker style adaptive transfer is achieved through the flow matching model, enabling high-quality conversion effects for unseen speaker styles. Combining the first cycle consistency loss and the second cycle consistency loss to optimize the model parameters improves the stability of style conversion, reduces the dependence on large-scale labeled data, and enhances the generalization ability of the model. Among them, the client can be, but is not limited to, various personal computers, laptop computers, smartphones, tablet computers, and portable wearable devices. The server can be implemented by an independent server or a server cluster composed of multiple servers. The present invention will be described in detail below through specific embodiments.

[0042] Please refer to Figure 2 , Figure 2Schematic flowchart of an embodiment of the voice style transfer method provided by the present invention. It should be noted that although the logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than here.

[0043] As Figure 2 shown, the voice style transfer method proposed by the present invention includes the following steps:

[0044] S10, extract the initial source features of the source voice, and separate the content features and the first style features from the initial source features;

[0045] In this embodiment, in order to accurately extract the features of the source voice during the voice style transfer process, it is necessary to preprocess the source voice signal and further separate its core components, including content features and style features. The initial features of the source voice contain multi-level information of the voice signal, including frequency components, temporal structure, and the personalized style of the speaker. In order to ensure the accuracy of the voice content during the style transfer process and at the same time ensure the independence of the style features, it is necessary to adopt a high-precision feature extraction method to extract features in different dimensions from the original voice signal.

[0046] After the source voice signal is sampled and digitized, it is first divided into a series of short-time frames, and each frame represents the voice waveform data within a small time range. This process usually uses the short-time Fourier transform (STFT) for frequency-domain analysis to extract the time-frequency information of the voice and at the same time remove redundant background noise and unnecessary components. The short-time Fourier transform can decompose the voice signal into different frequency components, so that different voice features can be extracted separately, improving the accuracy of subsequent feature analysis.

[0047] In order to ensure that voice features at different time scales can be effectively extracted, a multi-scale feature extraction strategy can be adopted. By performing multi-scale convolution operations with convolutional kernels of different resolutions, information in the low frequency, high frequency, and intermediate frequency can be extracted respectively. The low-frequency part mainly contains the overall timbre of the voice, the high-frequency part mainly reflects the clarity and transient details of the voice, and the intermediate-frequency part combines the characteristics of both. The low-frequency components are usually related to the individual's vocal tract characteristics, while the high-frequency part determines the voice clarity and the emotional expression in the voice. Therefore, during the multi-scale feature extraction process, filters of multiple scales can be used to decompose the frequency bands, and different pooling methods can be adopted to perform temporal pooling processing on the low-frequency features to capture long-term phoneme information, and at the same time use the channel attention mechanism to weight the high-frequency features to enhance the transient detail performance of the voice.

[0048] After the initial source features are extracted, it is necessary to further separate the content features and style features. Content features refer to the core information in speech, such as phoneme sequences and the basic semantic structure of speech, while style features refer to the personalized features of speech, including the speaker's timbre, prosody, intonation, energy distribution, etc. The extraction of content features usually uses a pre-trained speech content encoder, which is trained using a large-scale corpus and can accurately extract phoneme sequences from the initial source features. This process maps the speech signal to the phoneme space to obtain a speech representation that is not affected by the speaker's style.

[0049] The extraction of style features requires extracting components from the initial source features that can represent the personalized information of speech. Usually, the fundamental frequency trajectory and energy envelope can be used as representatives of style features. The fundamental frequency trajectory reflects the pitch changes in speech and mainly affects the ups and downs of intonation, such as the rising tone in interrogative sentences and the falling tone in declarative sentences. The energy envelope reflects the energy distribution of the speech signal over time and determines the stress pattern, intensity change, and emotional expression of speech. The Mel filter bank energy analysis method can be used to calculate the energy envelope and use it as the first style feature. In addition, to ensure the consistency of content features and style features in the time dimension, a multi-head attention mechanism can be used to perform time alignment on the phoneme sequence and the fundamental frequency trajectory, so as to ensure that there are no mismatches in speech rhythm or intonation in the finally generated style-transferred speech.

[0050] Example illustration: In the field of medical and health, intelligent speech systems can be used for medical record entry, remote consultation, and the generation of voice health reports. Existing speech synthesis systems often have difficulty accurately imitating the speech styles of different doctors, resulting in understanding barriers for patients during voice communication. By separating content features and style features, it is possible to maintain the clarity of medical terms while generating personalized speech according to the speech style of doctors, improving the user experience of telemedicine. In addition, in a health broadcast system, by adjusting the style features, the speech synthesis system can be adapted to the auditory preferences of users of different ages. For example, a slower speaking speed and a clearer speech style can be adopted for the elderly to improve understandability.

[0051] In the financial field, intelligent voice customer service needs to handle various financial business consultations simultaneously, and the professionalism and consistency of the speech style are crucial for the user experience. By separating content features and style features, it is possible to maintain the standard expression of financial terms while adjusting the style features to adapt to different application scenarios. For example, a more gentle and persuasive speech style can be adopted when introducing financial products, and a more serious speech style can be used in risk control reminders and emergency notification scenarios, thereby enhancing user trust. In addition, in a multilingual financial service scenario, through style transfer technology, the speech content in different languages can be converted into a version with a consistent style in the target language to provide a more natural cross-language voice interaction experience.

[0052] By extracting the initial features of the source speech and using methods such as multi-scale feature extraction, fundamental frequency trajectory analysis, and energy envelope calculation to separate the content features and style features, the integrity of the content information can be ensured during the style transfer process, while maintaining the personalized features of the speech. By using the multi-head attention mechanism for time alignment, the matching degree of different features in the time dimension is guaranteed, improving the naturalness and stability of the style-transferred speech. It can reduce the dependence on manual annotation, improve the adaptability of the model to unseen speakers, and at the same time reduce the risk of semantic loss during the style transfer process.

[0053] S20, extract the initial reference features of the reference speech and separate the second style features from the initial reference features;

[0054] In this embodiment, in order to ensure that the target speech can accurately retain the style features of the reference speech during the style transfer process, it is necessary to extract the features of the reference speech and further separate the personalized style information irrelevant to the speech content. The initial features of the reference speech contain multi-dimensional information, including frequency structure, temporal pattern, and individual speaking habits. The extraction and separation of these information are crucial for achieving high-quality style transfer.

[0055] The processing flow of the reference speech is similar to the feature extraction of the source speech. First, the input speech data is normalized to eliminate the influence of sampling noise and different recording environments on the signal. Usually, the signal of the reference speech will undergo a short-time Fourier transform (STFT) to extract the frequency-domain features, so that different speech components can be effectively represented in the frequency space.

[0056] After the frequency-domain data of the reference speech is extracted, it is necessary to further use the multi-scale feature extraction method to capture the style information in different frequency ranges. The multi-scale feature extraction can use multiple convolution kernels of different sizes to perform feature mapping on different time windows, so that the short-time speech features (such as instantaneous energy changes) and long-time speech features (such as overall intonation trends) can be extracted. Through multi-scale convolution processing, low-frequency, medium-frequency, and high-frequency information can be obtained respectively. The low-frequency part mainly reflects the basic timbre features of the speaker, the medium-frequency part contains prosody information, and the high-frequency part reflects the clarity and personalized expression of the speech.

[0057] To ensure that style features are independent of speech content, a dedicated style encoder is usually employed to model the initial features of the reference speech. The goal of the style encoder is to remove the content information from the input speech signal and retain only the part related to the speaker's personalized expression. This encoder can adopt an adversarial training method to force the separation of content features and style features, thus ensuring that the extracted style information does not contain specific speech content. During this process, normalization transformations such as Adaptive Instance Normalization (AdaIN) can be used to enable the conversion of speech features of different styles in the same feature space and achieve style invariance.

[0058] During the specific process of splitting style features, the fundamental frequency trajectory, energy envelope, and prosodic information are usually extracted as key style elements. The fundamental frequency trajectory mainly determines the pitch change pattern of speech. There are obvious differences in the intonation patterns of different speakers. For example, some speakers tend to use a higher fundamental frequency for expression, while others have a lower intonation. In addition, the energy envelope reflects the distribution of the energy of the speech signal over time and directly affects the stress pattern and prosody features of speech. For example, some speakers are accustomed to using stronger stress to emphasize key information, while others have a more gentle pronunciation.

[0059] Based on these style feature extraction methods, statistical modeling or neural network methods can be further used for style characterization. For example, variational autoencoders (VAEs) or normalizing flow models can be used to model style features in the form of probability distributions and generate continuous style representation vectors. In this way, style features can be transferred between different speakers while maintaining the naturalness and coherence of the original speech.

[0060] In practical applications, to ensure the stability of style features, time alignment processing is usually required for the extracted style information. Since the duration of the reference speech may not match that of the source speech, directly using the style features of the original reference speech may lead to unstable styles or failed migrations. Dynamic Time Warping (DTW) or alignment methods based on the attention mechanism can be adopted to align the extracted style features with the time axis of the target speech, thereby improving the stability of style transfer.

[0061] Example description: In the field of healthcare, the voice style of doctors plays an important role in voice health broadcasts, remote diagnosis, and medical inquiry scenarios. For example, when patients are undergoing remote consultations, they are usually more inclined to accept the doctor's calm and clear voice style, while in emotional comfort or rehabilitation advice, they may prefer to hear a friendly voice style. By extracting the style features of the reference voice, the doctor's speaking style can be transferred to the speech synthesis system, making the automatically generated medical voice more realistic, enhancing the patient's sense of trust, and improving the experience of remote medical interaction. In addition, in medical training scenarios, style transfer technology can be used to adapt the explanation styles of different experts, so that medical courses can present different styles of explanations while keeping the content unchanged to meet the needs of different learners.

[0062] In the financial field, the style of customer service voice directly affects the user experience. In scenarios such as bank customer service, insurance services, and securities investment advisors, the voice style needs to be adjusted according to business needs. For example, in scenarios such as risk warning and credit card collection, the voice needs to adopt a more serious and professional style, while when recommending smart financial advisors or value-added services, a more friendly and trustworthy expression is required. By extracting the style features of the reference voice and keeping the content consistent during the style migration process, the customer service system can switch different voice styles in different business scenarios, thereby improving user acceptance and satisfaction. In addition, in the intelligent voice assistant application of financial technology, through style feature modeling, the voice styles of different brands or companies can be integrated into the customer service system, making the brand image more consistent and improving users' recognition of the service.

[0063] By extracting the initial features of the reference speech and separating the second style features, the style characteristics of the target speaker can be retained without affecting the speech content information, so that the speech after style transfer can accurately reproduce the rhythm, timbre and expression of the reference speech. The use of multi-scale feature extraction and style encoder for modeling helps to separate content features and style features, avoiding the problem of content confusion during style transfer.

[0064] S30, performing linear interpolation processing on the initial source feature according to the time parameter to generate an intermediate source feature;

[0065] In this embodiment, in order to achieve a smooth transition of speech features and reduce the problem of unnatural speech caused by sudden style changes, it is necessary to perform linear interpolation on the initial source features to generate intermediate source features that can maintain continuity during the style conversion process. Linear interpolation is a method that uses the linear relationship between known data points to predict unknown data points and is widely used in data smoothing and time series modeling. In the speech style transfer task, linear interpolation not only helps to maintain the gradual change of features during the style conversion process but also improves the naturalness of the generated speech by smoothing the feature changes.

[0066] The initial source features are the basic features extracted from the source speech and contain key contents such as phoneme information, fundamental frequency trajectory, and energy envelope. Before performing linear interpolation, it is necessary to determine the time parameter, which is used to control the smoothness of the interpolation and the weight distribution of the features during the transition process. The time parameter can be dynamically adjusted based on the intensity of the target style transfer to ensure that there is no over-smoothing or style loss during the transition process.

[0067] The basic process of linear interpolation processing includes the following key steps: First, divide the time axis of the initial source features into multiple interpolation intervals, and each interpolation interval is defined as the transition area between two adjacent time points. Within each interpolation interval, a random interpolation ratio is selected according to the time parameter, and this ratio is used to control the offset degree of the current time point between the initial state and the target state. Then, a reference feature vector is defined, which is usually a zero vector or calculated from the mean of the standardized speech dataset and is used to construct the intermediate state during the interpolation process.

[0068] During the interpolation calculation process, the initial source features are weighted according to the weight of the time parameter. Among them, the weight of the initial source features is called the time parameter weight, and the weight of the reference feature vector is called the complementary weight, and the sum of the two is 1. Through this weight distribution mechanism, it is ensured that the features after linear interpolation can smoothly transition between the two states and gradually tend to the target style. Subsequently, weighted calculations are performed on the initial source features and the reference feature respectively to generate the weighted initial source features and the weighted reference feature. Finally, the two are linearly superimposed to obtain the final intermediate source features.

[0069] To improve the stability of linear interpolation, an adaptive interpolation strategy can be adopted, that is, adjust the time parameter weight based on the dynamic range of the initial source features, so that the interpolation process is smoother in the segments with faster phoneme conversion, while the interpolation change is smaller in the segments with stable prosody. In addition, non-linear interpolation methods such as exponential interpolation or weighted moving average can be introduced to enhance the continuity of style changes during the interpolation process, thus avoiding sudden changes.

[0070] By performing linear interpolation on the initial source features during the style transfer process, the abrupt changes during the style conversion can be effectively reduced, enhancing the fluency and naturalness of the generated speech. The interpolation ratio is controlled by a time parameter, making the smoothness of the features adjustable during the transfer process to meet the requirements of different style transfer intensities. By adopting a weighted weight allocation method, the feature interpolation process becomes more stable, avoiding unnatural breakpoints or style mismatches during speech generation.

[0071] S40, input the intermediate source feature, the second style feature, and the content feature into the flow matching model to generate an intermediate reference feature;

[0072] In this embodiment, to achieve high-quality speech style transfer, it is necessary to use a flow matching model to model the intermediate features of the source speech and the target style features, and generate an intermediate reference feature that conforms to the target style. The flow matching model is a deep learning method based on probability distribution transformation, which can map the input speech features to the target distribution, so that the generated features can not only retain the content information of the speech but also match the target style features, thus ensuring the stability and consistency of the style conversion.

[0073] The intermediate source feature is a feature obtained by linear interpolation from the initial source feature, which has achieved a smooth transition in the time dimension, thus reducing the abrupt change phenomenon during the style transfer process. The second style feature comes from the reference speech and contains information such as the personalized prosody, fundamental frequency distribution, and energy envelope of the target speaker, while the content feature is the core speech content information extracted from the source speech to ensure that the semantic information does not change during the style transfer process.

[0074] Before the features are input into the flow matching model, feature fusion is required so that the flow matching model can process multiple input features simultaneously. The method of feature fusion can adopt the channel splicing method, splicing the intermediate source feature, the second style feature, and the content feature into a unified input vector according to the channel dimension, thus maintaining the independence of each feature. In addition, a weighted fusion method can also be adopted, assigning different weights to each feature to meet different requirements for style fidelity and semantic stability in different scenarios.

[0075] The flow matching model usually consists of multiple invertible neural network layers (INN). The core idea of this model is to use a series of invertible transformations to map the input data to a probability distribution, making the features of the source speech gradually approach the feature distribution of the target style. After the features are fused, they will first be input into a fully connected layer to improve the expression ability of the features and map them to a high-dimensional space. In the high-dimensional space, the features are mapped through a series of invertible transformations, so that the intermediate source feature gradually approaches the second style feature while maintaining the semantic information.

[0076] During the calculation process of the reversible neural network layer, the output of each layer can be restored to the input through reversible transformation, ensuring that key information is not lost during the style transformation process. In this way, even if the speech after style conversion needs to be reversed back to the original style, the original features can be restored through reversible transformation, thereby improving the stability of style conversion.

[0077] To further enhance the stability of the generated intermediate reference features, a regularization constraint (such as Kullback-Leibler (KL) divergence constraint) can be added during the calculation process of the flow matching model. Adding the KL divergence constraint during the calculation process of the flow matching model is mainly to make the style distribution of the generated intermediate reference features as close as possible to the distribution of the second style features, thereby improving the stability of style conversion. This process needs to be carried out during the model training stage, and the specific implementation method is as follows:

[0078] First, it is necessary to model the distribution of the second style features. To characterize the style features in the reference speech, an additional style encoder can be used to convert the second style features into a standardized distribution. The neural network will calculate the parameters describing this distribution, such as the mean and the range of variation, based on the input second style features. These parameters can be used to measure the variation trend of the target style in different speech segments, making subsequent style matching more accurate.

[0079] Then, a similar modeling is performed on the intermediate reference features generated by the flow matching model. After being trained, the flow matching model will transform the input intermediate source features, content features, and second style features into a new feature representation, which needs to be as close as possible to the distribution of the second style features. To ensure the style consistency between the two, it is necessary to calculate the difference between their distributions.

[0080] During the training process, calculate the deviation between the distribution of the generated intermediate reference features and the distribution of the second style features, and add this deviation as an additional loss term to the model optimization objective. To reduce this deviation, a sampling-based method can be used to calculate the average deviation through multiple speech samples, thereby improving the stability of the model. During the calculation process, if it is found that the generated feature distribution is significantly different from the distribution of the second style features, the model will adjust the network parameters through backpropagation, making the generated style features gradually approach the second style features.

[0081] In terms of the optimization strategy, variational inference can be used for training, enabling the flow matching model to better adapt to the changes of different second style features. In addition, at the initial stage of training, the style conversion intensity of the model can be appropriately constrained to avoid excessive style deviation in the early stage, which may affect the naturalness of the finally generated speech. As the training progresses, gradually relax the constraints so that the model can freely learn the ability to adapt to different styles.

[0082] In this way, KL divergence constraint is added to the flow matching model, making the generated intermediate reference features closer to the second style features in style, improving the stability of style transfer, and avoiding the occurrence of style distortion or mutation. Finally, this constraint mechanism can ensure that the speech after style conversion not only conforms to the target style, but also maintains the clarity and consistency of semantics.

[0083] By inputting the intermediate source features, the second style features and the content features into the flow matching model, it can effectively ensure that the content and style of the speech are consistent during the style conversion process. The features are fused by means of channel splicing to ensure that the model can consider the roles of multiple features simultaneously and enhance the stability of style transfer. The flow matching model uses invertible neural network layers for distribution mapping, enabling the converted features to retain the semantic information of the speech while matching the target style and improving the naturalness of the converted speech.

[0084] S50, input the intermediate reference features, the first style features and the content features into the flow matching model to generate reconstructed source features;

[0085] In this embodiment, in order to ensure that the converted speech can not only maintain the content information of the original speech, but also retain the style features of the input speech, it is necessary to input the intermediate reference features, the first style features and the content features into the flow matching model to generate reconstructed source features. The core objective of this process is to enable the generated reconstructed source features to restore the style of the source speech and ensure that the speech content remains unchanged through the probability transformation ability of the flow matching model, thereby constructing a cyclic consistency training mechanism and improving the stability of style transfer.

[0086] The core role of the Flow Matching Model in the style transfer task is to make the input features approach the target style features through probability density transformation while maintaining the content consistency. This model uses an invertible neural network (INN, Invertible Neural Network) for distribution mapping and combines normalizing flow for probability density learning to achieve high-quality style conversion.

[0087] The flow matching model can include the following core modules. Each module undertakes specific conversion or optimization tasks and jointly completes the probability transformation process of style transfer:

[0088] Feature input and fusion module: This module is responsible for receiving the intermediate source features, the second style features and the content features, and performing feature preprocessing on the input. The output of this module is the fused feature vector, which serves as the input to the flow matching model.

[0089] Invertible Neural Network Layer (INN): The invertible neural network (INN) is the core component of the flow matching model. Its main function is to perform an invertible transformation without losing information, enabling the features to be transformed between different styles while still being able to recover the original features. The invertible transformation allows the input features to be restored to the original features after being transformed by the model (improving consistency). It enables the features of different styles to learn the corresponding transformation relationships through mapping, ensuring that the features after style transfer can match the target distribution. It adopts a dual-path architecture (such as the NICE, RealNVP, or Glow structure), and by decomposing the feature flow, it ensures that the model does not lose information during style transformation. During the transformation process, each layer of the network remains invertible, enabling the output features to be accurately reconstructed.

[0090] Residual Transformation Module: Based on the invertible neural network, residual learning is added to enhance the ability to represent details of the transformed speech features. It avoids losing small but crucial style information during the style transfer process, such as personalized features in timbre, emotional expressions, etc. It enables the generated features to not only conform to the target style distribution but also retain the personalized information of the source speech to a certain extent, improving the stability of the transformation.

[0091] Regularization Constraint Module: The main function of this module is to ensure that the generated style features do not deviate from the style distribution of the reference speech, improving the consistency of the transformation effect. The main constraint methods can include: KL divergence regularization, which calculates the difference between the probability distributions of the intermediate reference features and the second-style features and optimizes the model to make their distributions as close as possible; cycle consistency loss, which calculates the error between the reconstructed source features and the initial source features to ensure that the flow matching model does not lose the original information of the speech; style adversarial training, which uses adversarial learning to make the generated features more style-matched while keeping the speech content stable.

[0092] Fully Connected Layer (Feature Enhancement Module): Since there may be a problem of dimension mismatch after feature fusion, a fully connected layer (FC Layer) is needed to adjust the input features to meet the input requirements of the flow matching model. It enables the input features to be mapped in the feature space of different scales, enhancing the separability. It enhances the feature expression ability, enabling the model to capture complex style information changes.

[0093] Normalizing Flow: The main function of this layer is to map features to a standardized distribution, enabling the model to perform style transfer more easily during the transformation process. By adjusting the probability density of the input features through normalizing flow, the style distribution of the source speech can be made to approach the target style distribution. This makes the generated features more stable and less prone to style mutations or distortions. Standard normalizing flow techniques such as RealNVP, Glow, or NICE structures are used to perform manifold transformations on the input features according to the set probability density, thus achieving smooth style transfer.

[0094] Output Generation Module: This module is responsible for outputting the finally transformed features for subsequent speech synthesis tasks. It generates and reconstructs the source features to ensure that the features after style transfer can still match the original speech content. This enables the subsequent speech synthesis process to utilize these transformed features to generate speech signals that conform to the target style.

[0095] The input to the style matching model consists of three parts. The intermediate reference feature is the feature generated in the previous stage, which has been adjusted in the direction of the target style and reflects the style distribution of the target speaker. The first style feature is the style feature extracted from the source speech, representing the personalized expression of the source speaker, including information such as fundamental frequency trajectory and energy envelope. The content feature is the core speech information separated from the source speech to ensure that the converted speech does not change the original semantic structure.

[0096] Before these features are input into the style matching model, they first need to be processed by feature fusion. The feature fusion method can use channel concatenation to connect the three features along the channel dimension to ensure that the style matching model can utilize all input features simultaneously. Additionally, weighted fusion can also be used, where different weights are assigned according to the importance of different features. For example, during style conversion, higher weights may need to be assigned to the content features to ensure the consistency of the speech content, while reasonably constraining the influence of the style features to avoid speech distortion caused by excessive style conversion.

[0097] After the input of the fused features enters the style matching model, it will go through a series of invertible transformation layers for probability density mapping. The role of the invertible transformation layer is to learn the conversion relationship between different style features, enabling the style information of the original speech to be restored through the transformation process without affecting the speech content. During the invertible transformation, the output of each layer can be restored to the input through reverse calculation, ensuring that the model does not lose key information. This mechanism allows the model to maintain high-quality reconstruction capabilities in style transfer tasks. Even after multiple conversions between different styles, it can still generate stable speech features.

[0098] In order to further improve the accuracy of the reconstructed source features, a cycle consistency loss can be added during the training process of the stream matching model to minimize the difference between the generated reconstructed source features and the original source speech features. In addition, adversarial training methods can be used to enable the stream matching model to better learn the distribution of style conversion, so that the converted speech can retain the rhythm and rhythm features of the source speech while having the naturalness of the target style. Finally, the features transformed by the stream matching model are the reconstructed source features, which can not only maintain the content information of the source speech, but also restore the style features of the source speaker as much as possible, providing a basis for the subsequent calculation of cycle consistency loss.

[0099] Example description: In the field of medical health, doctors may need to adjust their voice styles in different scenarios. For example, when explaining the condition, a clear and formal expression is required, while in the process of psychological comfort, a gentler and more friendly voice style is required. In order to achieve this voice style conversion, the original voice of the doctor can be used as the source voice, from which content features (medical information expression) and first style features (doctor's intonation, rhythm and speaking habits) are extracted. Then, by using the target doctor or emotional voice as the reference voice, the second style features (a gentler or more professional speaking style) are extracted. The features of the source voice are transformed through the stream matching model to generate intermediate reference features that match the target style, and finally synthesize personalized doctor voices that meet the needs of specific medical scenarios. For example, in a remote consultation system, patients can choose the doctor's voice style so that the system automatically adjusts the style of voice broadcasting to improve the patient's trust and comfort. In addition, in the field of medical education, the voices of different medical experts can be used as reference voices to adjust the expression style of the teaching content to be more stable, delicate or vivid to meet the needs of different learners.

[0100] In the financial field, an intelligent customer service system needs to switch different voice styles according to business scenarios. For example, in financial consulting, an amiable and relaxed voice style is required, while in risk warnings or collection notices, a more formal and rigorous voice style is needed. To achieve this voice style adaptation, the default voice of the customer service system can be used as the source voice, from which content features (the core semantics of business answers) and first style features (the basic voice style of the customer service) are extracted. Then, a recording with the target voice style is selected as the reference voice, and second style features (e.g., a more formal and faster speaking style) are extracted. Through a flow matching model, the features of the source voice are transformed, enabling the customer service system to generate personalized customer service voices that meet the needs of users according to business scenarios. For example, when a customer consults about financial products, the customer service voice can present a gentle and professional style, while in the scenarios of overdue reminders and risk warnings, it can automatically adjust to a more serious tone to improve the efficiency of business communication and the customer experience. In addition, in multilingual financial services, the flow matching model can be used to transfer styles between customer service voices in different languages, ensuring that the voice expressions in cross-language services are natural and fluent, and improving user satisfaction and brand consistency.

[0101] By inputting the intermediate reference features, the first style features, and the content features into the flow matching model and generating reconstructed source features, it can effectively ensure that the content information and style features of the voice remain consistent during the style transfer process. The feature fusion method is adopted to enable the model to consider the influence of multiple features simultaneously and enhance the stability of style conversion. The flow matching model uses reversible transformation for distribution mapping, enabling the transformed features to retain the semantic information of the voice while restoring the style of the source speaker and improving the authenticity of the converted voice.

[0102] S60, determining a first cycle consistency loss between the reconstructed source features and the initial source features;

[0103] In this embodiment, it is crucial to ensure that the converted voice can not only maintain the target style but also be consistent with the original voice in content. For this purpose, it is necessary to calculate the first cycle consistency loss between the reconstructed source features and the initial source features to measure whether the converted voice can accurately restore to the original voice features. The role of this loss term is to constrain the flow matching model, enabling the transformed features to retain the key information of the source voice and preventing content loss or style drift during the voice conversion process.

[0104] The main process of calculating the first cycle consistency loss includes feature normalization, time-domain and frequency-domain error calculation, weight assignment, time alignment processing, and comprehensive loss calculation, which are specifically as follows:

[0105] First, standardize the reconstructed source features and the initial source features to eliminate the numerical bias between different batches of speech data. Standardization usually adopts zero-mean unit-variance normalization, that is, zeroize the mean of each feature channel and scale it to unit variance, so that the numerical distribution of all features is more balanced, preventing the model training from being unstable due to some feature values being too large or too small.

[0106] After feature standardization, calculate the error between the reconstructed source features and the initial source features. Error calculation mainly includes time-domain error and frequency-domain error to ensure that both the content and spectral information of the speech can be effectively preserved:

[0107] The time-domain error measures the difference between the two in the time dimension. Usually, absolute difference calculation is adopted, that is, compare the amplitude values of the two at each time point to obtain the overall time-domain change.

[0108] The frequency-domain error measures the difference between the two in the spectral space. Usually, logarithmic spectral distance is used, which can more accurately reflect the change of the speech frequency structure and avoid the error compensation problem caused by relying only on the time-domain error.

[0109] Since the time-domain error and the frequency-domain error have different impacts on the speech quality, appropriate weights need to be assigned to these two errors. During the style transfer process, the frequency structure of the speech is usually more critical than the short-time waveform. Therefore, usually a larger weight is assigned to the frequency-domain error, while the weight of the time-domain error is relatively small. This can ensure that the overall timbre and intonation of the converted speech do not change significantly, and the overall quality will not be affected by the slight differences at individual time points.

[0110] Next, to ensure the consistency of the time dynamic characteristics, it is also necessary to perform time alignment processing on the initial source features and the reconstructed source features. The purpose of time alignment is to correct the speech rhythm deviation caused by style transfer, so that the reconstructed speech features can be aligned with the original speech in the time dimension, thereby improving the fluency of the converted speech. The time alignment method can adopt dynamic time warping (DTW) or attention alignment mechanism to automatically adjust the time difference between features, making the finally calculated loss more stable.

[0111] Finally, calculate the final first-cycle consistency loss. The calculation method is to add the time-domain error (multiplied by the first weight), the frequency-domain error (multiplied by the second weight), and the spectral convergence loss after time alignment to form a comprehensive loss value. The smaller this loss value is, the closer the reconstructed source features are to the features of the original source speech, and the less content is lost during the style conversion process, thereby improving the naturalness and accuracy of the speech.

[0112] By calculating the first cycle consistency loss, it can be ensured that the speech after style transfer is consistent with the source speech in terms of content, preventing the loss or error of speech content. Through normalization processing, the numerical deviation between different speech samples can be eliminated, improving the stability of loss calculation. By combining the time-domain error and frequency-domain error calculations, the changes in speech features can be comprehensively measured, avoiding the problem of error compensation caused by a single feature. Through time alignment processing, the speech rhythm deviation caused by style conversion can be effectively corrected, ensuring that the converted speech sounds smoother and more natural.

[0113] S70, perform linear interpolation processing on the initial reference feature according to the time parameter to generate a verification intermediate feature;

[0114] In this embodiment, in order to ensure the stability and reversibility of style conversion, cycle consistency verification is required. The core of this verification lies in performing linear interpolation processing on the initial reference feature to generate a verification intermediate feature, so as to ensure that during different style conversion processes, the content features and style features of the speech can maintain a reasonable corresponding relationship.

[0115] Linear interpolation is a method that uses the linear relationship between known data points to infer unknown data points, and is widely used in time series data smoothing and feature transition. Linear interpolation is used to adjust the initial reference feature so that it can gradually approach the target style during the style conversion process, avoiding sudden changes during the style transfer process.

[0116] Specifically, the initial reference feature is the key feature extracted from the reference speech, which contains style information such as the timbre, prosody, fundamental frequency trajectory, and energy envelope of the target speaker. In order to generate a verification intermediate feature, interpolation calculation needs to be performed on the initial reference feature based on the time parameter, so that the feature changes during the verification process can conform to the natural style transfer process.

[0117] First, it is necessary to divide the time axis of the initial reference feature, and divide the entire speech feature sequence into multiple interpolation intervals. The number of interpolation intervals can be dynamically determined according to the duration of the speech segment, ensuring that the interpolation process can cover the entire speech feature.

[0118] Then, within each interpolation interval, select the interpolation ratio according to the time parameter. The time parameter determines the smoothness and change rate of the interpolation. Generally, a uniform sampling method can be used to make the interpolation ratio evenly distributed between 0 and 1, or according to the style conversion requirements, dynamically adjust the interpolation ratio to make the style transition more natural.

[0119] Next, define a reference feature vector, which is usually a zero-valued vector or a mean vector calculated based on a standard speech dataset. The role of the reference feature vector is to provide the starting state for interpolation, enabling the interpolation calculation to smoothly adjust between the initial reference feature and the reference state.

[0120] When calculating the interpolation, assign interpolation weights to the initial reference feature and the reference feature, where: the weight of the initial reference feature is called the time parameter weight, representing the transition ratio from the reference feature to the target feature; the weight of the reference feature is called the complementary weight, and the sum of the complementary weight and the time parameter weight is equal to 1.

[0121] Then, multiply the initial reference feature by the time parameter weight to obtain the weighted initial reference feature, and multiply the reference feature by the complementary weight to obtain the weighted reference feature. Finally, linearly superimpose these two parts to obtain the verified intermediate feature.

[0122] To further optimize the interpolation process, an adaptive interpolation strategy can be adopted, adjusting the time parameter according to the dynamic range of the initial reference feature, so that the interpolation is smoother in segments with faster rhythm changes, and the interpolation changes less in segments with stable rhythm. In addition, non-linear interpolation methods (such as exponential interpolation or Gaussian weighted interpolation) can be introduced to enhance the naturalness of feature transition and avoid overly abrupt feature changes caused by linear interpolation.

[0123] By linearly interpolating the initial reference feature, the sudden changes during the style conversion process can be effectively reduced, making the feature transition after style transfer smoother. Using the time parameter to control the interpolation ratio enables the degree of change of the feature during the style transfer process to be dynamically adjusted according to specific requirements, suitable for scenarios with different style transfer intensities. Adopting the interpolation weight assignment mechanism ensures the stability of the interpolation calculation and avoids unnatural breakpoints or sudden changes during the style transfer process.

[0124] S80, input the verified intermediate feature, the second style feature, and the content feature into the flow matching model to generate a verified reference feature;

[0125] In this embodiment, to verify the stability and consistency of the conversion process, it is necessary to input the verified intermediate feature, the second style feature, and the content feature into the flow matching model to generate a verified reference feature. The core goal of this process is to ensure that the speech features after style conversion still conform to the target style and maintain the corresponding relationship with the original speech in terms of content, thereby completing the cycle consistency verification and improving the reliability of style transfer.

[0126] The verified intermediate feature is a feature obtained by linearly interpolating the initial reference feature, which is smoothed on the time axis to make the style conversion process more natural and reduce sudden changes. The second style feature is the style information extracted from the target reference speech, representing the personalized features such as the timbre, prosody, and fundamental frequency of the target speaker. The content feature is the core speech information extracted from the original speech to ensure that the semantic content of the speech does not change during the conversion process.

[0127] Before the input stream is matched with the model, these three features need to be feature fused to ensure that the model can comprehensively consider all input information. The ways of feature fusion include:

[0128] Channel concatenation: The verified intermediate feature, the second style feature, and the content feature are concatenated along the channel dimension into a multi-dimensional feature vector, enabling the model to process multiple feature dimensions simultaneously.

[0129] Attention-weighted fusion: Use the attention mechanism to calculate the importance weights of different features and dynamically adjust their contributions to the generation of the final feature.

[0130] The fused feature vector is input into the stream matching model. The core calculation process of the model includes the following steps:

[0131] Fully connected layer mapping: First, the feature is mapped through the fully connected layer to enhance the feature expression ability and ensure that the input meets the calculation requirements of the stream matching model.

[0132] Invertible neural network transformation (INN): Through a series of invertible transformations, the feature can be mapped between different styles and ensure that the transformed feature can still be restored to the original feature. The main role of the invertible transformation is to learn the distribution mapping of style changes, enabling the verified reference feature to match the target style feature while still maintaining content consistency.

[0133] Residual connection and skip connection: During the invertible transformation process, the residual connection or skip connection method is adopted to enable the feature to maintain a smooth change between different levels, preventing information loss or excessive transformation.

[0134] Flow transformation (Normalizing Flow): Use flow transformation technology to adjust the probability density of the feature, making the generated verified reference feature closer to the style distribution of the target speaker and ensuring that the transformed feature is more natural in speech style.

[0135] After the above steps, the stream matching model finally outputs the verified reference feature, which is close to the target speaker in style and still conforms to the semantic information of the original speech in content. This feature will be used as an important input for the subsequent calculation of the cycle consistency loss to measure the reversibility and stability of the style transfer process.

[0136] By inputting the verification intermediate feature, the second style feature, and the content feature into the flow matching model and generating the verification reference feature, it can be ensured that the speech feature after style transfer can match the target speaker in style and remain consistent in semantic content. By adopting the method of channel splicing and attention-weighted fusion, the model can learn the relationship between style and content features more precisely, improving the stability of style conversion. The flow matching model utilizes reversible neural networks and flow transformation techniques to ensure the reversibility of the style conversion process, enabling the generated verification reference feature to be used for subsequent calculation of the cyclic consistency loss and optimizing the training quality of the style transfer model.

[0137] S90, determining a second cyclic consistency loss between the verification reference feature and the intermediate reference feature;

[0138] In this embodiment, the second cyclic consistency loss is used to measure the similarity between the verification reference feature and the intermediate reference feature, thereby ensuring the reversibility and stability of the style transfer process. The core objective of this loss term is to constrain the flow matching model so that the features after style conversion can still return to a reasonable target style distribution, ensuring that the speech after style transfer conforms to the target style without affecting the semantic integrity of the speech.

[0139] The main process of calculating the second cyclic consistency loss includes feature standardization, time-domain and frequency-domain error calculation, weight assignment, time alignment processing, and final loss calculation, as follows:

[0140] First, perform zero-mean unit-variance standardization on the verification reference feature and the intermediate reference feature to eliminate the numerical deviation between different speech samples, enabling the influence of amplitude changes to be ignored when calculating the loss and only focusing on the similarity of speech style features. This step ensures the stability of the calculation, preventing excessive gradient changes or unstable convergence caused by different ranges of feature values.

[0141] Then, calculate the time-domain error and the frequency-domain error respectively to comprehensively measure the changes in speech features:

[0142] The time-domain error reflects the differences in features in the time dimension and is usually calculated by frame-by-frame absolute difference, that is, comparing the feature values of the two at each time point to obtain the overall time-domain error trend. This error can measure the similarity of the rhythm and energy envelope of the converted speech.

[0143] The frequency-domain error is used to measure the difference in the spectral structure of speech and is usually calculated using the logarithmic spectral distance. Since the timbre and prosody of speech are more stable in the frequency domain, the frequency-domain error can more accurately measure the consistency of speech style features, ensuring that the timbre does not shift significantly after style transfer.

[0144] Due to the different degrees of influence of time-domain error and frequency-domain error, appropriate weights need to be assigned to control their contributions to the overall loss. Generally speaking, to ensure style consistency, a relatively large weight is usually assigned to the frequency-domain error, while the weight of the time-domain error is relatively small, so as to ensure that the converted speech still conforms to the timbre characteristics of the target speaker and does not cause a decrease in speech naturalness due to excessive constraints on time changes.

[0145] Next, time alignment processing needs to be performed on the verification reference feature and the intermediate reference feature to ensure that there is no rhythm misalignment during the style conversion process. The method of time alignment can adopt dynamic time warping (DTW), which makes the time relationship after speech conversion consistent by dynamically adjusting the matching method of the time axis. In addition, an attention alignment mechanism can also be adopted to automatically adjust the time offset between features by using a deep learning model, making the converted features more stable.

[0146] Finally, calculate the second cycle consistency loss, which is obtained by adding three parts: the time-domain error (multiplied by the first weight), the frequency-domain error (multiplied by the second weight), and the spectral convergence loss after time alignment. The smaller this loss value is, the smaller the difference between the verification reference feature and the intermediate reference feature, the stronger the reversibility of the style transfer process, and thus the stability of the model and the naturalness of the generated speech are improved.

[0147] By calculating the second cycle consistency loss, it can be ensured that the speech after style transfer can maintain stability in the target style and will not lose its reversibility due to excessive style changes. Using zero-mean unit-variance normalization can eliminate the numerical deviation between different samples and improve the stability of loss calculation. Combining the calculation of time-domain error and frequency-domain error can comprehensively measure the change of style features, ensuring that the converted speech not only conforms to the target style but also maintains the consistency of timbre. Adopting time alignment processing can effectively correct the rhythm offset caused by style conversion during the style transfer process and improve the fluency of speech. Finally, by optimizing this loss term, the robustness of the flow matching model can be improved, so that the generated speech can still maintain high-quality style consistency and semantic integrity after style transfer.

[0148] S100, optimize the parameters of the flow matching model by combining the first cycle consistency loss and the second cycle consistency loss;

[0149] In this embodiment, in order to ensure that the converted speech can not only match the target style but also maintain consistency with the original speech, it is necessary to optimize the flow matching model. The core goal of optimization is to jointly constrain the model using the first cycle consistency loss and the second cycle consistency loss, so that the speech after style conversion can maintain stability in content and style features, avoiding information loss or style drift.

[0150] The first cycle consistency loss is used to measure the difference between the reconstructed source features and the initial source features to ensure that the content features are not lost during the style transfer process. The second cycle consistency loss, on the other hand, is used to measure the similarity between the verification reference features and the intermediate reference features to ensure the reversibility of the style transfer. The joint optimization of these two can guarantee that the model can not only retain semantic integrity but also ensure the stability of style conversion during the style transfer process.

[0151] The process of optimizing the flow matching model involves loss calculation, gradient update, parameter optimization, and convergence judgment. The specific implementation is as follows:

[0152] First, the first cycle consistency loss and the second cycle consistency loss are weighted and summed according to a preset weight ratio to generate the total cycle consistency loss. Under different task scenarios, the weights of these two losses can be dynamically adjusted. For example, in scenarios where content preservation is prioritized (such as medical explanations or legal document readings), the weight of the first cycle consistency loss should be higher, while in scenarios where stricter style changes are required (such as movie dubbing or personalized speech synthesis), the weight of the second cycle consistency loss should be higher.

[0153] Then, based on the total cycle consistency loss, the gradients in the flow matching model are calculated. Gradient calculation involves multiple core modules:

[0154] Invertible Neural Network Layer (INN): Calculate gradients to ensure that the flow matching model does not change the semantic information of the original speech while maintaining style features.

[0155] Fully Connected Layer: Calculate gradients to optimize the parameters of feature transformation and improve the model's adaptability to different speech styles.

[0156] Attention Module: Calculate gradients to ensure the matching degree of style features and improve the stability of the model during the style transfer process.

[0157] To improve the stability of training, the calculated gradients are clipped layer by layer by threshold, that is, if the gradient change of a certain layer is too large and exceeds the preset threshold range, it will be clipped to prevent the model training from being unstable due to gradient explosion. In addition, during the training process, a momentum optimization strategy can be introduced to make the gradient update smoother and improve the convergence speed of the model.

[0158] Next, an adaptive optimizer (such as Adam, RMSprop, or LAMB) is used to optimize the clipped gradients and update the weight parameters of the flow matching model. The adaptive optimizer can dynamically adjust the learning rate, making the training process more stable and avoiding local optimal solutions.

[0159] During the optimization process, to improve the generalization ability of the model, a learning rate decay mechanism can be adopted, that is, gradually reduce the learning rate as the number of training iterations increases. The specific methods include:

[0160] Exponential decay: The learning rate is decreased exponentially after each iteration, enabling rapid convergence in the initial stage of training and fine-tuning the parameters in the later stage.

[0161] Piecewise decay: After reaching the preset training stage, manually adjust the learning rate to adapt to the training requirements of different stages.

[0162] Finally, when the following two conditions are met, the model determines convergence and terminates the optimization: The fluctuation of the average total cycle consistency loss for multiple consecutive training batches is less than the preset threshold, indicating that the loss has stabilized and no longer significantly decreases; the current learning rate has reached the set minimum learning rate threshold to ensure that the optimizer does not cause unstable convergence due to too high a learning rate.

[0163] Through the above optimization process, the flow matching model can achieve high-quality style transfer while ensuring content consistency, making the finally generated speech conform to the target style and maintain a fluent and natural speech expression.

[0164] By jointly optimizing the parameters of the flow matching model using the first cycle consistency loss and the second cycle consistency loss, it can ensure the stability of the content information and style features of the speech during the style transfer process. Using gradient clipping can avoid the problem of gradient explosion and improve the training stability. Using an adaptive optimizer can dynamically adjust the learning rate, improve the model convergence speed, and avoid local optima. Through learning rate decay, it can ensure more stable fine-tuning in the later stage of training and improve the generalization ability of the model.

[0165] S110, input the initial source feature, the content feature, and the second style feature into the optimized flow matching model to generate a transferred speech waveform.

[0166] In this embodiment, the ultimate goal is to generate a speech waveform that conforms to the target style, while ensuring that the content information of the speech remains unchanged and the transferred speech has high naturalness and stability. To achieve this goal, it is necessary to input the initial source feature, the content feature, and the second style feature into the optimized flow matching model, and generate a transferred speech waveform through a series of transformations.

[0167] The initial source feature is the low-level feature extracted from the source speech, containing key information such as timbre and prosody. The content feature is the semantic-related information separated from the source speech to ensure that the speech content remains consistent after style transfer. The second style feature is the style attribute extracted from the target style speech, containing information such as the pitch, rhythm, and energy envelope of the target speaker.

[0168] The optimized flow matching model already has strong style transfer ability through pre-training and the constraint of cyclic consistency loss. The main function of the model is to transform the input features so that they conform to the feature distribution of the target style while keeping the content unchanged. The specific processing flow includes the following core steps:

[0169] First, fuse the input features to ensure that the model can consider both the content and style information of the speech at the same time. The way of feature fusion can adopt channel concatenation, that is, concatenate the initial source feature, content feature and second style feature along the channel dimension to generate a fused feature vector. It can also adopt attention-weighted fusion, that is, dynamically adjust their contribution ratios in the final migrated speech according to the importance of the content feature and style feature to ensure the stability of style transfer.

[0170] Then, input the fused features into the flow matching model, and go through the following transformation process:

[0171] Invertible Neural Network (INN) transformation: Under the action of the invertible neural network, perform an invertible mapping on the input features, so that the speech can still maintain invertibility after style transfer, ensuring a smooth change in the speech style. The role of the invertible transformation is to ensure that the input features do not lose key information when migrating to the target style, and at the same time ensure that they can be restored to the original space, improving the stability of the model.

[0172] Normalizing Flow: The normalizing flow technique is used to adjust the probability density distribution of the features to align it with the distribution of the target style speech. This step ensures that the transformed speech features can match the speech features of the target speaker in style, while avoiding style drift or information loss.

[0173] Residual connection and skip connection: Inside the flow matching model, introduce a residual connection so that the input features still retain some original information during the transformation process, improving the stability of style transfer and reducing the degradation of speech quality that may be caused by excessive transformation.

[0174] After completing the style transformation, the generated speech features are still in the form of time-frequency domain features and need to be further converted into speech waveforms. This process includes:

[0175] Inverse Short-Time Fourier Transform (ISTFT): Perform an inverse transformation on the time-frequency domain features to restore the time-domain signal.

[0176] Phase optimization: Use a phase recovery algorithm (such as the Griffin-Lim iterative algorithm) or a deep learning-based phase prediction model to improve the naturalness and clarity of the speech and reduce speech distortion.

[0177] Range compression: The generated speech signal is amplitude-normalized to ensure that the output waveform conforms to the volume and dynamic range of the target speech, so that the speech can maintain a stable sound quality on different playback devices or in different environments.

[0178] Finally, after the above processing, the flow matching model generates a migrated speech waveform that matches the characteristics of the target speaker in style while maintaining the same content information as the original speech, ensuring the naturalness, clarity, and stability of the generated speech.

[0179] By inputting the initial source features, content features, and second style features into the optimized flow matching model, high-quality speech style transfer can be achieved, enabling the converted speech to have the target style characteristics while maintaining the original content information. Feature fusion is adopted to ensure that the model can comprehensively consider speech content and style attributes, improving the stability of the conversion. Using reversible neural networks and flow transformations can make the style transfer process reversible, prevent information loss, and ensure the stability of the speech under the target style distribution. Combining phase optimization and range compression improves the clarity and naturalness of the generated speech, making it maintain a consistent listening experience on different devices and in different scenarios.

[0180] The present invention relates to the technical field of speech processing and can be applied to business scenarios such as fintech and healthcare. It discloses a speech style transfer method, including: extracting the features of the source speech and the reference speech, separating the content features and the style features; performing linear interpolation on the initial source features based on time parameters to generate intermediate features; inputting the intermediate features into a flow matching model to generate reference features and reconstructed features; calculating the cyclic consistency loss and optimizing the parameters of the flow matching model based on this loss; using the optimized model for style transfer to generate a migrated speech waveform. The present invention constrains the style transfer process through the cyclic consistency loss to ensure the semantic and style consistency of the speech, combines time interpolation processing to improve the conversion smoothness, and uses a flow matching model to achieve cross-speaker style transfer, improving the style adaptation ability to unseen speakers. Model optimization reduces the dependence on labeled data and improves the stability and naturalness of the generated speech.

[0181] In one embodiment, the above step S10 includes:

[0182] S101, performing frame windowing processing on the source speech and obtaining time-frequency domain data through short-time Fourier transform;

[0183] S102, inputting the time-frequency domain data into a multi-scale feature extraction module to generate initial source features;

[0184] S103, extracting the phoneme sequence from the initial source features through a pre-trained speech content encoder;

[0185] S104, extract the fundamental frequency trajectory from the initial source features through the fundamental frequency extraction module;

[0186] S105, extract the energy envelope from the initial source features through the Mel filter bank energy analysis module, and use the energy envelope as the first style feature;

[0187] S106, perform time alignment processing on the phoneme sequence and the fundamental frequency trajectory through the multi-head attention mechanism;

[0188] S107, perform channel fusion on the time-aligned phoneme sequence and fundamental frequency trajectory to generate the content features.

[0189] In this embodiment, it is necessary to extract feature data from the source speech that can characterize the content information and style features to ensure that both the semantic consistency of the speech and the target style transfer can be maintained during the style transfer process. Specifically, it is necessary to extract the initial source features from the source speech and further decompose them into content features and the first style features.

[0190] First, perform frame windowing processing on the source speech to facilitate the temporal analysis of the audio data. Since the speech signal is a continuously changing time series, direct global analysis may lose local features. Therefore, it is necessary to use the sliding window method to divide the speech data into short-time frames and superimpose a window function (such as Hamming window or Hanning window) on each frame to reduce the inter-frame spectral leakage problem.

[0191] After frame windowing, perform short-time Fourier transform (STFT) on the windowed speech signal to convert the time-domain signal to the time-frequency domain. STFT provides the time evolution information of the speech signal at different frequency components, enabling subsequent feature extraction modules to perform in-depth analysis in both the time dimension and the frequency dimension.

[0192] Then, input the time-frequency domain data into the multi-scale feature extraction module to extract the initial source features. The multi-scale feature extraction module uses convolutional filters of different scales to capture the low-frequency energy distribution (such as speech formants) and high-frequency transient features (such as plosives, fricatives) respectively. In addition, the low-frequency features are pooled over time to obtain long-term speech information, while the high-frequency features are enhanced by channel attention weighting to enhance transient details, so that the generated initial source features can comprehensively characterize the time-frequency structure of the speech.

[0193] Extract the content features and the first style features from the initial source features, specifically including:

[0194] Extract the phoneme sequence through the pre-trained speech content encoder:

[0195] The speech content encoder uses a pre-trained end-to-end speech recognition model, such as DeepSpeech, Conformer, or HuBERT, to extract phoneme-level content features from the initial source features. The model utilizes convolutional networks and self-attention mechanisms to learn stable speech semantic representations without relying on explicit text annotations, thereby ensuring the semantic consistency of the speech after style transfer.

[0196] The fundamental frequency (F0) of speech is a key feature affecting pitch and intonation, mainly determined by vocal cord vibration. The fundamental frequency extraction module is usually based on the autocorrelation function (ACF) or neural network prediction, and is used to extract the fundamental frequency trajectory from the initial source features and characterize the pitch change pattern of the speaker.

[0197] The energy envelope of the speech signal describes the time-varying speech energy distribution, reflecting the speaker's volume control, stress distribution, and emotional intensity. By calculating the energy distribution in different frequency bands using a Mel filter bank, speech energy information consistent with human ear perception can be obtained and used to characterize the speaking style. Therefore, the energy envelope is defined as the first style feature for controlling the loudness and expression mode of the speech.

[0198] Since the time scales of the phoneme sequence and the fundamental frequency trajectory are different, alignment processing is required. The multi-head attention mechanism (Multi-HeadAttention) is used to calculate the correlation between the phoneme sequence and the fundamental frequency trajectory, and automatically adjust the time alignment method to make the content information and the fundamental frequency information have a good match in the time dimension.

[0199] After completing the time alignment, the phoneme sequence and the fundamental frequency trajectory are fused at the channel level to construct the final content features. This fusion process can adopt the method of feature concatenation or weighted summation to ensure that the fused features can retain both semantic information and combine fundamental frequency information to achieve a complete speech representation.

[0200] Finally, the extracted content features are mainly used to ensure that the semantic information of the speech remains unchanged, while the first style feature (energy envelope) is mainly used to control the style features of the speech, enabling the subsequent style transfer process to adjust the speech style while ensuring content stability.

[0201] In this embodiment, by extracting content features and first style features from the initial source features, high-precision decomposition of the source speech can be achieved, ensuring semantic consistency of the speech after style transfer, while allowing independent adjustment of the style features. By using the short-time Fourier transform, the time-frequency information of the speech can be efficiently obtained, and combined with multi-scale feature extraction, the robustness of feature extraction can be ensured. By pre-training a speech content encoder to extract phoneme sequences, style information can be effectively removed, improving the stability of the content features. Combining the fundamental frequency trajectory can accurately characterize the pitch changes of the speech, while the energy envelope, as the style feature, can control the emotional expression of the speech.

[0202] In one embodiment, step S30 above includes:

[0203] S301, divide the time axis of the initial source features into multiple interpolation intervals according to the duration of the source speech;

[0204] S302, randomly select a time parameter value within each interpolation interval, where the time parameter value represents the transition ratio from the initial state to the target state;

[0205] S303, use a preset reference feature vector as the initial zero-value feature;

[0206] S304, assign a time parameter weight to the initial source features according to the time parameter value, and assign a complementary weight to the initial zero-value feature;

[0207] S305, multiply the initial source features by the time parameter weight to generate weighted initial source features;

[0208] S306, multiply the initial zero-value feature by the complementary weight to generate weighted initial zero-value features;

[0209] S307, superimpose the weighted initial source features and the weighted initial zero-value features to generate the intermediate source features.

[0210] In this embodiment, the initial source features need to undergo interpolation calculations to generate intermediate source features for style conversion, so as to ensure that the speech features change smoothly during the style transfer process and avoid sound quality loss caused by mutations. To achieve this goal, linear interpolation processing based on time parameters is adopted. By controlling the interpolation method of the features, the converted features conform to the dynamic change trend of natural speech on the time axis and ensure the stability of the time-frequency domain information.

[0211] First, according to the duration of the source speech, the time axis of the initial source features is divided into multiple interpolation intervals. The number and length of the interpolation intervals can be dynamically adjusted according to the duration of the speech to ensure that the interpolation calculation can cover the entire speech features. In long-duration speech, a fixed interval division can be adopted to make the length of each interval equal; in short-duration speech or rapidly changing speech segments, an adaptive division strategy can be adopted to improve the accuracy of the interpolation calculation.

[0212] Within each interpolation interval, a time parameter value is randomly selected to control the transition ratio of the interpolation calculation. The time parameter value usually varies between 0 and 1, indicating the degree of change from the initial state to the target state. Adopting a random sampling method helps to enhance the generalization ability of the data, so that the features after style transfer do not overly rely on a certain fixed time ratio and improves the robustness of the model.

[0213] To ensure the rationality of the interpolation calculation, a preset reference feature vector is defined as the initial zero-value feature. The reference feature vector can be a zero vector (i.e., a vector with all dimensions being zero) or a mean vector calculated from a standard speech dataset. Using the reference feature vector as the initial state enables the starting point of the interpolation calculation to align with the feature distribution of the standard speech data and improves the stability of the interpolation calculation.

[0214] When performing the interpolation calculation, interpolation weights need to be assigned to the initial source features and the initial zero-value features. The time parameter weight is determined by the time parameter value and represents the contribution ratio of the initial source features at the current moment to the intermediate source features. The complementary weight is the supplementary part of the time parameter weight, that is, the sum of the two is always equal to 1 to ensure the balance of the interpolation calculation.

[0215] Then, the initial source features are multiplied by the time parameter weight to obtain the weighted initial source features; at the same time, the initial zero-value features are multiplied by the complementary weight to obtain the weighted initial zero-value features. These two weighted features are linearly superimposed to finally obtain the intermediate source features, which can smoothly transition on the time axis and thus provide a stable feature input for subsequent style conversion.

[0216] After generating the intermediate source features, it is necessary to further check whether their spectral distribution satisfies the time-frequency continuity constraint. Time-frequency continuity is an important characteristic of speech signals, that is, in the frequency domain, the features of adjacent time frames should not change abruptly but should maintain smooth changes to conform to the dynamic characteristics of natural speech. Therefore, it is necessary to analyze the spectral distribution of the intermediate source features. If it is found that the spectral changes are too drastic or discontinuous, it means that there is a problem with the interpolation calculation, which may lead to unnatural breaks or mutations in the speech after style transfer.

[0217] If the spectral distribution of the intermediate source feature does not satisfy the time-frequency continuity constraint, it is necessary to increase the division density of the interpolation interval and resample the time parameter values. Increasing the density of the interpolation interval means dividing the time axis more finely, making the interpolation calculation smoother and reducing the mutation phenomenon caused by large-span interpolation. In addition, resampling the time parameter values can make the interpolation process more in line with the dynamic characteristics of the target style and improve the smoothness and naturalness of feature conversion.

[0218] In this embodiment, through the linear interpolation processing based on time parameters, the change of the initial source feature can be effectively smoothed, ensuring the stability of the speech feature during the style conversion process. Using random time parameter values can enhance the generalization ability of the model, making the features after style transfer more diverse. Using the benchmark feature vector as the initial zero-value feature can make the starting point of the interpolation calculation more stable, avoiding distortion during the interpolation transition process. In addition, through the time-frequency continuity constraint, it can be ensured that the converted speech feature changes smoothly in the time-frequency domain, improving the naturalness of the speech. If the interpolation calculation does not meet the continuity requirement, the division density of the interpolation interval can be dynamically adjusted to make the interpolation calculation more refined, thereby improving the final style transfer effect.

[0219] In one embodiment, the above step S40 includes:

[0220] S401, concatenate the intermediate source feature, the second style feature, and the content feature along the channel dimension to generate a multi-modal fusion feature;

[0221] S402, perform normalization processing on the multi-modal fusion feature, and input the normalized multi-modal fusion feature into a fully connected layer to generate a high-dimensional hidden layer representation;

[0222] S403, perform bidirectional feature mapping on the high-dimensional hidden layer representation through a reversible neural network layer to generate a preliminary transfer feature;

[0223] S404, perform residual stacking on the preliminary transfer feature and the multi-modal fusion feature to generate an enhanced transfer feature;

[0224] S405, perform moving average filtering processing on the enhanced transfer feature along the time axis to generate the intermediate reference feature.

[0225] In this embodiment, in order to generate stable intermediate reference features that conform to the target style, the intermediate source features, the second style features, and the content features need to be input into the flow matching model. Through multi-layer feature processing, the features after style transfer are made to maintain consistency and high quality in terms of content, style, and speech naturalness. Among them, the key steps include feature splicing, normalization, hidden layer feature mapping, residual enhancement, and temporal filtering to ensure that the generated intermediate reference features can stably exist under the speech distribution of the target style.

[0226] First, the intermediate source features, the second style features, and the content features are spliced along the channel dimension to generate a multi-modal fusion feature. Different features play different roles in the style transfer process: the intermediate source features represent the source speech features after interpolation smoothing to ensure the smoothness of the style conversion; the second style features represent the style features of the target speaker, including information such as timbre and prosody, which are the key to style conversion; the content features are used to ensure that the converted speech still maintains the semantic information of the source speech and prevent the style transfer from affecting the intelligibility of the speech content.

[0227] After these features are spliced along the channel dimension, a multi-modal fusion feature is formed, which simultaneously contains the content information, style information of the speech, and the source speech features after interpolation processing. Since the numerical ranges and distributions of these features may vary greatly, it is necessary to normalize the multi-modal fusion feature to eliminate the numerical bias between different features and improve the computational stability of the model. Subsequently, the normalized multi-modal fusion feature is input into the fully connected layer to generate a high-dimensional hidden layer representation.

[0228] High-Dimensional Latent Representation refers to that when the flow matching model processes the multi-modal fusion feature, it is mapped to a higher-dimensional feature space through the fully connected layer to enhance the expressive ability of the feature and improve the model's learning ability of style information.

[0229] The range of the high dimension depends on the model structure and task requirements, usually between 128 dimensions and 2048 dimensions. A lower dimension (such as 128 dimensions) is suitable for smaller-scale models, while a higher dimension (such as 2048 dimensions) is suitable for large-scale speech generation tasks, such as multi-speaker style transfer tasks.

[0230] The criteria for defining the high dimension are mainly based on:

[0231] The computational ability of the model: If the computing resources are limited, a lower dimension can be adopted to reduce the computational overhead.

[0232] The complexity of the style information: If the target style involves multiple speakers or complex prosodic features, a higher-dimensional feature space is required to store this style information.

[0233] Experimental verification: Usually through model training and testing on the validation dataset, dimensions that can achieve a balance between accuracy and computational cost are selected.

[0234] After generating the high-dimensional hidden layer representation, the model performs bidirectional feature mapping on the high-dimensional hidden layer representation through an invertible neural network layer (INN), that is, transforms the input features to the target style distribution and ensures that the transformation process is invertible, so as to achieve controllability and stability of style conversion in subsequent cycle consistency calculations.

[0235] Next, the preliminary transferred features and the multi-modal fusion features are superimposed residually to generate enhanced transferred features. The role of residual superposition is to retain the original feature information while enhancing the expression ability of the target style, so that the original speech information will not be lost during the style conversion process. In addition, residual connections can also improve the stability of training and prevent information loss or overfitting.

[0236] Finally, a moving average filter is applied to the enhanced transferred features along the time axis to generate intermediate reference features. The role of the moving average filter is to eliminate possible noise or mutations during the conversion process, making the finally generated intermediate reference features smoother in the time-frequency domain, and improving the naturalness and stability of the speech. Here, the time axis refers to the time dimension of the speech features, that is, the temporal evolution process of the speech signal, usually with the frame as the basic unit.

[0237] In speech signal processing, speech data is usually represented in the form of short-time frames, that is, the speech signal is segmented into small time segments, each frame corresponding to a small segment of speech data, and then feature extraction is performed on each frame. For example, after short-time Fourier transform (STFT) or Mel spectrum processing, each frame will contain a series of spectral features reflecting the speech characteristics at that moment, and these frames arranged in time order constitute the complete speech sequence.

[0238] In the style transfer task, the enhanced transferred features generated by the model are arranged along the time axis, that is, each time step corresponds to the feature representation of a speech frame, and the features between different time steps need to maintain a certain smoothness to avoid sudden changes or instability during the style conversion process. Therefore, the moving average filter is to perform smoothing processing along the time dimension of the speech frames, making the feature changes between adjacent frames more natural and smooth, thus avoiding abrupt style changes or noise interference. The implementation method of the moving average filter is to perform weighted averaging on the features of multiple adjacent frames on the time axis to smooth the feature values between adjacent time frames. This process ensures the smooth transition of speech features on the time axis, prevents drastic changes caused by style conversion, and makes the generated speech sound more natural and stable.

[0239] Example: Suppose the duration of a piece of speech is 2 seconds, and the feature extraction is performed using a window length of 25 ms. Then, the speech signal will be segmented into 80 frames (2000 ms / 25 ms = 80 frames). After generating the transfer features, the feature values of each frame may fluctuate significantly at different time steps. By performing a moving average filter along the time axis, such as using a 5-frame moving window (2 frames before and after + the current frame), the speech style can be smoothly transitioned within a range of 125 ms (5 × 25 ms), without generating abrupt style changes.

[0240] In this embodiment, by inputting the intermediate source feature, the second style feature, and the content feature into the stream matching model, and through processing steps such as feature fusion, normalization, high-dimensional hidden layer mapping, invertible neural network mapping, residual enhancement, and time filtering, the style features of the speech can be effectively extracted and transformed, and it is ensured that the content information of the speech is not affected by the style transfer. The introduction of the high-dimensional hidden layer representation enhances the expressive ability of the features, enabling the model to learn more complex style information and improving the fineness and accuracy of the style conversion. By adopting the invertible neural network layer, the stability of the conversion process can be ensured, so that the generated intermediate reference feature can maintain natural speech characteristics under the distribution of the target style.

[0241] In one embodiment, the above step S60 includes:

[0242] S601, performing zero-mean unit-variance normalization processing on the reconstructed source feature and the initial source feature respectively to obtain the reconstructed source feature and the initial source feature after the normalization processing;

[0243] S602, determining the time-domain absolute difference and the log Mel spectral distance between the reconstructed source feature after the normalization processing and the initial source feature after the normalization processing;

[0244] S603, assigning a first weight to the time-domain absolute difference and a second weight to the log Mel spectral distance, where the second weight is greater than the first weight;

[0245] S604, performing time warping alignment processing on the Mel spectrum of the initial source feature and the Mel spectrum of the reconstructed source feature;

[0246] S605, determining the spectral convergence loss between the Mel spectrum of the initial source feature after the time warping alignment processing and the Mel spectrum of the reconstructed source feature after the time warping alignment processing;

[0247] S606, adding the product of the time-domain absolute difference and the first weight, the product of the log Mel spectral distance and the second weight, and the spectral convergence loss to generate a first cycle consistency loss.

[0248] In this embodiment, in order to ensure that the converted speech can conform to the feature distribution of the target style while maintaining semantic information, it is necessary to constrain the style conversion process of the model through the first cycle consistency loss, so that the reconstructed source features can be consistent with the initial source features. The core objective of the first cycle consistency loss is to measure whether the speech after style conversion still retains the key features of the source speech, and optimize the model parameters through this loss function to ensure that the finally generated speech is both natural and conforms to the target style.

[0249] To calculate the first cycle consistency loss, it is first necessary to perform zero-mean unit-variance normalization on the reconstructed source features and the initial source features, that is, by removing the mean and normalizing the variance, so that these two features can be compared on the same numerical scale. The purpose of this normalization process is to eliminate the difference in the numerical range of the data, ensure that the contributions of different features are not affected by the numerical size during loss calculation, and improve the stability of model training.

[0250] After the normalization process, it is necessary to calculate the time-domain absolute difference between the normalized reconstructed source features and the normalized initial source features, that is, calculate the point-by-point difference between the two features in the time dimension to measure the deviation of the speech signal on the time axis. At the same time, it is also necessary to calculate the log Mel spectrogram distance to measure the similarity between the two features in the frequency dimension. The calculation of the log Mel spectrogram distance is usually based on the log Mel filter bank, which emphasizes low-frequency information while reducing the influence of high-frequency noise, making the calculated loss more in line with the human auditory perception characteristics.

[0251] Since the importance of the time-domain absolute difference and the log Mel spectrogram distance is different in the style transfer process, different weights need to be assigned to them. Among them, the weight of the log Mel spectrogram distance should be higher than that of the time-domain absolute difference, because the Mel spectrogram structure of the speech has a greater impact on the style, while the absolute error in the time domain mainly affects the instantaneous amplitude change of the speech.

[0252] Subsequently, it is necessary to perform time warping alignment on the Mel spectrograms of the initial source features and the reconstructed source features. Since the rhythm of the speech signal may change during the conversion process, directly comparing misaligned spectrograms will lead to inaccurate loss calculation. Therefore, it is necessary to use dynamic time warping (DTW) or attention-based time alignment methods to align the two features in the time dimension to ensure the accuracy of loss calculation.

[0253] After time warping alignment, the spectral convergence loss can be calculated to measure the similarity between the Mel spectrogram of the aligned initial source features and the Mel spectrogram of the aligned reconstructed source features. The spectral convergence loss is usually calculated using the L2 norm or KL divergence to measure the distance between two spectrograms and ensure that the speech after style conversion still conforms to the structural characteristics of the source speech.

[0254] Finally, the product of the time-domain absolute difference and the first weight, the product of the logarithmic Mel spectrogram distance and the second weight, and the spectral convergence loss are added together to generate the first cycle consistency loss. This loss value is used to optimize the parameters of the flow matching model, making the converted speech features closer to the original speech and improving the quality and stability of style transfer.

[0255] To further improve the stability of the loss value, mean normalization needs to be performed on the first cycle consistency loss, that is, by calculating the average loss of all samples and normalizing the current loss value, so that the scale of the loss is not affected by the sample size, thereby improving the convergence speed and training stability of model optimization.

[0256] In this embodiment, by calculating the first cycle consistency loss, the content retention ability in the style transfer process can be effectively constrained to ensure that the speech after style conversion still conforms to the feature distribution of the original speech. Using zero-mean unit-variance normalization can eliminate the influence of the eigenvalue range and improve the stability of loss calculation. By weighted calculation of the time-domain absolute difference and the logarithmic Mel spectrogram distance, the time-domain and frequency-domain information of the speech signal can be comprehensively considered to improve the similarity of the speech after style transfer. Using time warping alignment can correct the possible rhythm changes during style conversion, making the converted speech more natural.

[0257] In one embodiment, the above step S100 includes:

[0258] S1001, adding the first cycle consistency loss and the second cycle consistency loss according to a preset weight ratio to generate a total cycle consistency loss;

[0259] S1002, determining the gradients of the reversible neural network layer, fully connected layer, and attention module in the flow matching model based on the total cycle consistency loss;

[0260] S1003, performing layer-by-layer threshold clipping on the gradients;

[0261] S1004, using an adaptive optimizer to train the flow matching model according to the clipped gradients to update the weight parameters of the flow matching model;

[0262] S1005, exponentially decay the learning rate of the adaptive optimizer according to the current training iteration number until the minimum learning rate threshold is reached;

[0263] S1006, if the fluctuation of the average total cycle consistency loss of multiple consecutive training batches is less than a preset threshold and the current learning rate has reached the minimum learning rate threshold, then determine that the model has converged and terminate the optimization, generating an optimized flow matching model.

[0264] In this embodiment, the optimization process of the flow matching model is a key link to ensure that the generated speech can not only maintain content consistency but also match the target style. To improve the training stability and optimization effect of the model, it is necessary to combine the first cycle consistency loss and the second cycle consistency loss, and through a multi-level optimization strategy, improve the performance of the flow matching model so that it can stably convert between different speech styles.

[0265] First, add the first cycle consistency loss and the second cycle consistency loss according to a preset weight ratio to generate a total cycle consistency loss. The first cycle consistency loss is used to constrain the similarity between the reconstructed source feature and the initial source feature, while the second cycle consistency loss is used to constrain the similarity between the verification reference feature and the intermediate reference feature. These two losses work together to ensure that the model can remain stable during the style conversion process between the input and the output. The weight ratio of different losses can be adjusted according to experimental results. For example, if more attention is paid to content consistency, the weight of the first cycle consistency loss is large; if more attention is paid to style fidelity, the weight of the second cycle consistency loss is large.

[0266] After calculating the total cycle consistency loss, it is necessary to determine the gradients of the reversible neural network layer, the fully connected layer, and the attention module in the flow matching model based on this loss. The calculation of gradients is usually based on the backpropagation algorithm (Backpropagation), where: the reversible neural network layer is responsible for the bidirectional mapping of style conversion, and its gradient is used to adjust the parameters of the reversible transformation; the fully connected layer is used for feature learning and fusion, and its gradient affects the overall parameter optimization of the model; the attention module determines the weight distribution of different features during the conversion process, and its gradient is used to optimize the attention mechanism of style conversion.

[0267] After calculating the gradients, it is necessary to perform layer-by-layer threshold clipping on the gradients to ensure that the gradient norm does not exceed a preset upper limit. The purpose of gradient clipping is to prevent the problem of gradient explosion. Especially in deep learning models, if the gradient value is too large, it may lead to instability in the optimization process and even cause the training to not converge. Specifically, first set an upper limit for the gradient norm. If the magnitude of the gradient calculated for a certain layer exceeds this upper limit, then normalize the gradient to shrink it to the allowed range. The normalization method is to scale the gradient according to its own magnitude so that the adjusted gradient magnitude is exactly equal to the preset maximum value. This method can ensure that the gradients will not be overly amplified during training, thereby improving the stability of training and preventing numerical overflow or abnormal oscillations in the model during the optimization process.

[0268] After gradient clipping, use an adaptive optimizer to train the flow matching model according to the clipped gradients to update the weight parameters of the model. Adaptive optimizers (such as Adam, Adagrad, or RMSProp) have the ability to dynamically adjust the learning rate and can adaptively adjust the parameter update step size based on historical gradient information to improve the convergence speed and optimization effect of training.

[0269] During the training process, the learning rate of the model cannot be fixed but should be exponentially decayed according to the current training iteration count until it reaches the minimum learning rate threshold. The role of exponential decay of the learning rate is as follows: maintain a relatively large learning rate at the beginning of training to ensure that the model can quickly converge to a better direction; reduce the learning rate in the later stage of training to reduce the amplitude of gradient updates to avoid oscillations in the model during the later stage of optimization and make the final model more stable.

[0270] Finally, during the model training process, it is necessary to monitor the fluctuations of the total cycle consistency loss. If the average total cycle consistency loss fluctuations of multiple consecutive training batches are less than the preset threshold and the current learning rate has reached the minimum learning rate threshold, then it can be determined that the model has converged, stop the training, and generate the optimized flow matching model. At this time, the model already has stable style transfer capabilities and can perform natural conversions between different speech styles.

[0271] In this embodiment, by jointly optimizing the first cycle consistency loss and the second cycle consistency loss, the style transfer ability of the flow matching model can be improved, so that the generated speech can not only maintain content consistency but also have style features that highly match the speech distribution of the target speaker. Using gradient clipping can prevent gradient explosion during training and improve the stability of training. Using an adaptive optimizer can adaptively adjust the parameter learning rate and improve the optimization efficiency. Using an exponentially decaying learning rate can achieve fast convergence in the early stage of training and stable convergence in the later stage of training, thereby improving the generation quality of the final model.

[0272] In one embodiment, the above step S110 includes:

[0273] S1101, concatenate the initial source feature, the content feature, and the second style feature along the channel dimension to generate a fused feature vector;

[0274] S1102, perform bidirectional mapping on the fused feature vector through the invertible neural network layer in the optimized flow matching model to generate a time-frequency domain transfer feature;

[0275] S1103, perform an inverse short-time Fourier transform on the time-frequency domain transfer feature to generate a preliminary transferred speech waveform;

[0276] S1104, perform phase optimization and range compression processing on the preliminary transferred speech waveform to generate a final transferred speech waveform.

[0277] In this embodiment, in order to generate a speech waveform that conforms to the target style, it is necessary to input the initial source feature, the content feature, and the second style feature into the optimized flow matching model. Through multi-layer feature fusion, mapping transformation, frequency domain processing, and waveform optimization, it is ensured that the finally generated speech not only conforms to the target style but also maintains the accuracy and naturalness of the speech content. The entire process involves key steps such as feature concatenation, bidirectional mapping, frequency domain transformation, phase optimization, and waveform compression to ensure the quality and stability of the transferred speech waveform.

[0278] First, concatenate the initial source feature, the content feature, and the second style feature along the channel dimension to generate a fused feature vector. Among them: the initial source feature provides the basic timbre information of the source speech; the content feature ensures that the semantic information of the speech is not lost, so that the generated speech is consistent with the input speech in content; the second style feature is used to control the style characteristics of the target speaker, so that the output speech conforms to the target style in terms of prosody, timbre, etc.

[0279] Since these features may have different numerical ranges and distributions, before concatenation, it is usually necessary to normalize each feature to ensure that the concatenated fused feature vector has better stability during the calculation process inside the model.

[0280] Subsequently, perform bidirectional mapping on the fused feature vector through the invertible neural network layer in the optimized flow matching model to generate a time-frequency domain transfer feature. The role of the invertible neural network layer is to ensure the reversibility of the style conversion process and ensure that the model can still retain the key features of the input speech after style transformation. Bidirectional mapping means that: in the forward calculation process, the input fused feature vector is converted into the feature distribution of the target style; in the reverse calculation process, the features after style conversion can be restored to ensure that the generated speech does not lose information or deform.

[0281] Through the mapping of the reversible neural network layer, the generated time-frequency domain transfer features not only conform to the target style but also maintain the integrity of the speech content, providing a reliable basis for subsequent waveform reconstruction.

[0282] After generating the time-frequency domain transfer features, it is necessary to perform the inverse short-time Fourier transform (ISTFT) on these features to generate the preliminary transferred speech waveform. The short-time Fourier transform (STFT) is used to convert the time-domain signal to the frequency domain, while the ISTFT has the opposite effect, that is: through ISTFT, the time-frequency domain features are restored to the time-domain signal to form a playable speech waveform; this process needs to retain the phase information of the original speech to avoid waveform distortion.

[0283] Since the phase information of the speech may be affected during the ISTFT process, it is necessary to optimize the phase of the preliminary transferred speech waveform. The purpose of phase optimization is to re-estimate the phase information of the speech through the Griffin-Lim algorithm or a neural network-based phase recovery model to reduce speech distortion and improve the sound quality; by adjusting the phase, the speech becomes more natural in terms of auditory perception, reducing unstable vibrations or sudden pitch changes.

[0284] Finally, in order to ensure that the dynamic range of the final speech conforms to the optimal range of human ear hearing, it is necessary to perform range compression processing on the preliminary transferred speech waveform to generate the final transferred speech waveform. The functions of range compression include: reducing abnormal amplitudes to avoid phenomena of overly large or small amplitudes caused by style conversion; optimizing the loudness distribution of the speech to ensure reasonable consistency in volume for speech in different styles;

[0285] Adapting to the dynamic range of the playback device so that the speech can maintain good clarity and audibility when played on different audio devices (such as mobile phones, headphones, speakers).

[0286] Through the above series of processes, the finally generated transferred speech waveform can maintain natural speech characteristics in the target style and meet the speech quality requirements in practical applications.

[0287] In this embodiment, by inputting the initial source features, content features, and second style features into the optimized flow matching model and going through processing steps such as feature fusion, reversible neural network mapping, inverse short-time Fourier transform, phase optimization, and range compression, it can be ensured that the finally generated speech not only conforms to the target style but also remains consistent with the source speech in content. The bidirectional mapping of the reversible neural network layer improves the stability of style transfer and avoids information loss during the style conversion process. The inverse short-time Fourier transform ensures the restoration of the speech signal in the time domain and improves the audibility of the speech. Phase optimization reduces the sound quality loss caused by style conversion and makes the speech clearer and more natural. Range compression ensures that the dynamic range of the speech conforms to auditory habits and improves compatibility and audibility on playback devices.

[0288] In one embodiment, a voice style transfer device is provided, which corresponds one-to-one with the voice style transfer method in the above embodiment. Referring to Figure 3 , Figure 3 FIG. is a schematic diagram of functional modules of a preferred embodiment of the voice style transfer device of the present invention. A source voice feature extraction module 10, a reference voice style extraction module 20, a time interpolation processing module 30, a flow matching mapping module 40, a reconstructed feature generation module 50, a consistency loss analysis module 60, a reference feature interpolation module 70, a verification feature mapping module 80, a verification loss analysis module 90, a model optimization module 100, and a voice generation module 110. The detailed description of each functional module is as follows:

[0289] The source voice feature extraction module 10 is configured to extract an initial source feature of the source voice and separate a content feature and a first style feature from the initial source feature;

[0290] The reference voice style extraction module 20 is configured to extract an initial reference feature of the reference voice and separate a second style feature from the initial reference feature;

[0291] The time interpolation processing module 30 is configured to perform linear interpolation processing on the initial source feature according to a time parameter to generate an intermediate source feature;

[0292] The flow matching mapping module 40 is configured to input the intermediate source feature, the second style feature, and the content feature into a flow matching model to generate an intermediate reference feature;

[0293] The reconstructed feature generation module 50 is configured to input the intermediate reference feature, the first style feature, and the content feature into the flow matching model to generate a reconstructed source feature;

[0294] The consistency loss analysis module 60 is configured to determine a first cycle consistency loss between the reconstructed source feature and the initial source feature;

[0295] The reference feature interpolation module 70 is configured to perform linear interpolation processing on the initial reference feature according to the time parameter to generate a verification intermediate feature;

[0296] The verification feature mapping module 80 is configured to input the verification intermediate feature, the second style feature, and the content feature into the flow matching model to generate a verification reference feature;

[0297] The verification loss analysis module 90 is configured to determine a second cycle consistency loss between the verification reference feature and the intermediate reference feature;

[0298] A model optimization module 100, configured to jointly optimize the parameters of the flow matching model by using the first cycle consistency loss and the second cycle consistency loss;

[0299] A voice generation module 110, configured to input the initial source feature, the content feature, and the second style feature into the optimized flow matching model to generate a migrated voice waveform.

[0300] In one embodiment, the source voice feature extraction module 10 is specifically configured to:

[0301] Perform frame adding and windowing processing on the source voice, and obtain time-frequency domain data through short-time Fourier transform;

[0302] Input the time-frequency domain data into a multi-scale feature extraction module to generate an initial source feature;

[0303] Extract a phoneme sequence from the initial source feature through a pre-trained voice content encoder;

[0304] Extract a fundamental frequency trajectory from the initial source feature through a fundamental frequency extraction module;

[0305] Extract an energy envelope from the initial source feature through a Mel filter bank energy analysis module, and use the energy envelope as a first style feature;

[0306] Perform time alignment processing on the phoneme sequence and the fundamental frequency trajectory through a multi-head attention mechanism;

[0307] Perform channel fusion on the time-aligned phoneme sequence and fundamental frequency trajectory to generate the content feature.

[0308] In one embodiment, the time interpolation processing module 30 is specifically configured to:

[0309] Divide the time axis of the initial source feature into multiple interpolation intervals according to the duration of the source voice;

[0310] Randomly select a time parameter value in each interpolation interval, where the time parameter value represents the transition ratio from the initial state to the target state;

[0311] Use a preset reference feature vector as an initial zero-value feature;

[0312] Assign a time parameter weight to the initial source feature and a complementary weight to the initial zero-value feature according to the time parameter value;

[0313] Multiply the initial source feature by the time parameter weight to generate a weighted initial source feature;

[0314] Multiply the initial zero-value feature by the complementary weight to generate a weighted initial zero-value feature;

[0315] Superimpose the weighted initial source features and the weighted initial zero-value features to generate the intermediate source features.

[0316] In one embodiment, the flow matching mapping module 40 is specifically configured to:

[0317] Concatenate the intermediate source features, the second style features, and the content features along the channel dimension to generate multimodal fusion features;

[0318] Perform normalization processing on the multimodal fusion features, and input the normalized multimodal fusion features into a fully connected layer to generate a high-dimensional hidden layer representation;

[0319] Perform bidirectional feature mapping on the high-dimensional hidden layer representation through a reversible neural network layer to generate preliminary transfer features;

[0320] Perform residual stacking on the preliminary transfer features and the multimodal fusion features to generate enhanced transfer features;

[0321] Perform moving average filtering on the enhanced transfer features along the time axis to generate the intermediate reference features.

[0322] In one embodiment, the consistency loss analysis module 60 is specifically configured to:

[0323] Perform zero-mean unit-variance standardization processing on the reconstructed source features and the initial source features respectively to obtain the standardized reconstructed source features and initial source features;

[0324] Determine the time-domain absolute difference and the log Mel spectrum distance between the standardized reconstructed source features and the standardized initial source features;

[0325] Assign a first weight to the time-domain absolute difference and a second weight to the log Mel spectrum distance, where the second weight is greater than the first weight;

[0326] Perform time warping alignment processing on the Mel spectrum of the initial source features and the Mel spectrum of the reconstructed source features;

[0327] Determine the spectral convergence loss between the Mel spectrum of the initial source features after time warping alignment processing and the Mel spectrum of the reconstructed source features after time warping alignment processing;

[0328] Add the product of the time-domain absolute difference and the first weight, the product of the log Mel spectrum distance and the second weight, and the spectral convergence loss to generate a first cycle consistency loss.

[0329] In one embodiment, the model optimization module 100 is specifically configured to:

[0330] Add the first cycle consistency loss and the second cycle consistency loss according to a preset weight ratio to generate a total cycle consistency loss;

[0331] Determine the gradients of the reversible neural network layer, the fully connected layer, and the attention module in the flow matching model based on the total cycle consistency loss;

[0332] Perform layer-by-layer threshold clipping on the gradients;

[0333] Use an adaptive optimizer to train the flow matching model according to the clipped gradients to update the weight parameters of the flow matching model;

[0334] Exponentially decay the learning rate of the adaptive optimizer according to the current training iteration number until the minimum learning rate threshold is reached;

[0335] If the fluctuations of the average total cycle consistency loss of multiple consecutive training batches are less than a preset threshold and the current learning rate has reached the minimum learning rate threshold, determine that the model has converged and terminate the optimization to generate an optimized flow matching model.

[0336] In one embodiment, the speech generation module 110 is specifically configured to:

[0337] Concatenate the initial source feature, the content feature, and the second style feature along the channel dimension to generate a fused feature vector;

[0338] Perform bidirectional mapping on the fused feature vector through the reversible neural network layer in the optimized flow matching model to generate time-frequency domain migration features;

[0339] Perform an inverse short-time Fourier transform on the time-frequency domain migration features to generate a preliminary migrated speech waveform;

[0340] Perform phase optimization and range compression processing on the preliminary migrated speech waveform to generate a final migrated speech waveform.

[0341] In one embodiment, a computer device is provided. The computer device can be a server, and its internal structure diagram can be as Figure 4As shown. The computer device includes a processor, a memory, a network interface, and a database connected via a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile and / or volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external client via a network connection. When the computer program is executed by the processor, it realizes the functions or steps on the server side of a voice style transfer method.

[0342] In one embodiment, a computer device is provided. The computer device may be a client, and its internal structure diagram may be as Figure 5 As shown. The computer device includes a processor, a memory, a network interface, a display screen, and an input device connected via a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external server via a network connection. When the computer program is executed by the processor, it realizes the functions or steps on the client side of a voice style transfer method

[0343] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the following steps are realized:

[0344] Extract the initial source features of the source voice, and separate the content features and the first style features from the initial source features;

[0345] Extract the initial reference features of the reference voice, and separate the second style features from the initial reference features;

[0346] Perform linear interpolation processing on the initial source features according to the time parameter to generate intermediate source features;

[0347] Input the intermediate source features, the second style features, and the content features into a flow matching model to generate intermediate reference features;

[0348] Input the intermediate reference features, the first style features, and the content features into the flow matching model to generate reconstructed source features;

[0349] Determine the first cycle consistency loss between the reconstructed source features and the initial source features;

[0350] Perform linear interpolation processing on the initial reference feature according to the time parameter to generate a verification intermediate feature;

[0351] Input the verification intermediate feature, the second style feature, and the content feature into the flow matching model to generate a verification reference feature;

[0352] Determine the second cycle consistency loss between the verification reference feature and the intermediate reference feature;

[0353] Jointly optimize the parameters of the flow matching model with the first cycle consistency loss and the second cycle consistency loss;

[0354] Input the initial source feature, the content feature, and the second style feature into the optimized flow matching model to generate a migrated speech waveform.

[0355] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:

[0356] Extract the initial source feature of the source speech, and separate the content feature and the first style feature from the initial source feature;

[0357] Extract the initial reference feature of the reference speech, and separate the second style feature from the initial reference feature;

[0358] Perform linear interpolation processing on the initial source feature according to the time parameter to generate an intermediate source feature;

[0359] Input the intermediate source feature, the second style feature, and the content feature into the flow matching model to generate an intermediate reference feature;

[0360] Input the intermediate reference feature, the first style feature, and the content feature into the flow matching model to generate a reconstructed source feature;

[0361] Determine the first cycle consistency loss between the reconstructed source feature and the initial source feature;

[0362] Perform linear interpolation processing on the initial reference feature according to the time parameter to generate a verification intermediate feature;

[0363] Input the verification intermediate feature, the second style feature, and the content feature into the flow matching model to generate a verification reference feature;

[0364] Determine the second cycle consistency loss between the verification reference feature and the intermediate reference feature;

[0365] Optimize the parameters of the flow matching model by combining the first cycle consistency loss and the second cycle consistency loss;

[0366] Input the initial source feature, the content feature, and the second style feature into the optimized flow matching model to generate a migrated speech waveform.

[0367] It should be noted that for the functions or steps that can be implemented by the above computer-readable storage medium or computer device, reference can be made to the relevant descriptions on the server side and the user side in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.

[0368] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above method embodiments. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.

[0369] Those skilled in the art can clearly understand that for the convenience and brevity of description, only the above division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0370] It should be noted that if there are software tools or components of other companies in the embodiments of this application, they are only used for illustrative introduction and do not represent actual use. The above-described embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included in the protection scope of the present invention.

Claims

1. A speech style transfer method, characterized in that: The following steps are involved: Extracting initial source features of the source speech, and separating content features and first style features from the initial source features; Extracting initial reference features of the reference speech, and separating second style features from the initial reference features; Performing linear interpolation processing on the initial source features according to the time parameter to generate intermediate source features; Inputting the intermediate source feature, the second style feature and the content feature into a stream matching model to generate an intermediate reference feature; Inputting the intermediate reference feature, the first style feature and the content feature into the stream matching model to generate a reconstructed source feature; determining a first cycle consistency loss between the reconstructed source features and the original source features; Performing linear interpolation processing on the initial reference feature according to the time parameter to generate a verification intermediate feature; Inputting the verification intermediate feature, the second style feature and the content feature into the stream matching model to generate a verification reference feature; determining a second cycle consistency loss between the verification reference feature and the intermediate reference feature; optimizing the parameters of the flow matching model by combining the first cycle consistency loss and the second cycle consistency loss; The initial source feature, the content feature and the second style feature are input into the optimized stream matching model to generate a migration speech waveform.

2. The method for voice style transfer according to claim 1, characterized in that: Extracting initial source features of the source speech, and separating content features and first style features from the initial source features, including: Perform frame division and windowing processing on the source speech, and obtain time-frequency domain data through short-time Fourier transform; Inputting the time-frequency domain data into a multi-scale feature extraction module to generate initial source features; Extracting a phoneme sequence from the initial source features by a pre-trained speech content encoder; Extracting a fundamental frequency trajectory from the initial source features by a fundamental frequency extraction module; extracting an energy envelope from the initial source feature through a Mel filter bank energy analysis module, and using the energy envelope as a first style feature; Performing time alignment processing on the phoneme sequence and the fundamental frequency trajectory through a multi-head attention mechanism; The phoneme sequence after time alignment processing is channel-fused with the fundamental frequency track to generate the content feature.

3. The method for voice style transfer according to claim 1, wherein: The initial source feature is subjected to linear interpolation processing according to the time parameter to generate an intermediate source feature, including: Dividing the time axis of the initial source feature into a plurality of interpolation intervals according to the duration of the source speech; Randomly selecting a time parameter value in each interpolation interval, wherein the time parameter value represents a transition ratio from an initial state to a target state; Using the preset reference feature vector as the initial zero-value feature; Assigning a time parameter weight to the initial source feature according to the time parameter value, and assigning a complementary weight to the initial zero-value feature; Multiplying the initial source feature by the time parameter weight to generate a weighted initial source feature; Multiplying the initial zero-value feature by the complementary weight to generate a weighted initial zero-value feature; The weighted initial source feature and the weighted initial zero-value feature are superimposed to generate the intermediate source feature.

4. The method for voice style transfer according to claim 1, wherein: Inputting the intermediate source feature, the second style feature and the content feature into a stream matching model to generate an intermediate reference feature includes: splicing the intermediate source feature, the second style feature and the content feature according to the channel dimension to generate a multimodal fusion feature; Normalizing the multimodal fusion features, and inputting the normalized multimodal fusion features into a fully connected layer to generate a high-dimensional hidden layer representation; Performing bidirectional feature mapping on the high-dimensional hidden layer representation through a reversible neural network layer to generate preliminary migration features; Performing residual superposition of the preliminary migration feature and the multimodal fusion feature to generate an enhanced migration feature; The enhanced migration features are subjected to sliding average filtering along the time axis to generate the intermediate reference features.

5. The method for voice style transfer according to claim 1, wherein: Determining a first cycle consistency loss between the reconstructed source feature and the original source feature includes: Performing zero mean unit variance normalization processing on the reconstructed source features and the initial source features respectively to obtain the reconstructed source features and the initial source features after the normalization processing; Determine the time domain absolute difference and the logarithmic Mel-spectral distance between the normalized reconstructed source features and the normalized original source features; Assigning a first weight to the time domain absolute difference, and assigning a second weight to the logarithmic Mel spectrum distance, wherein the second weight is greater than the first weight; Performing time warping and alignment processing on the Mel spectrum of the initial source feature and the Mel spectrum of the reconstructed source feature; Determining a spectral convergence loss between a Mel-spectrogram of an initial source feature after time warping and alignment and a Mel-spectrogram of a reconstructed source feature after time warping and alignment; The product of the time-domain absolute difference and the first weight, the product of the logarithmic Mel-spectrum distance and the second weight, and the spectral convergence loss are added to generate a first cycle consistency loss.

6. The method for voice style transfer according to claim 1, wherein: Optimizing the parameters of the flow matching model by combining the first cycle consistency loss and the second cycle consistency loss includes: Adding the first cycle consistency loss and the second cycle consistency loss according to a preset weight ratio to generate a total cycle consistency loss; Determining gradients of a reversible neural network layer, a fully connected layer, and an attention module in the stream matching model based on the total cycle consistency loss; Performing layer-by-layer threshold clipping on the gradient; Using an adaptive optimizer to train the flow matching model according to the clipped gradient to update the weight parameters of the flow matching model; Exponentially decaying the learning rate of the adaptive optimizer according to the current number of training iterations until a minimum learning rate threshold is reached; If the average total cycle consistency loss fluctuation of multiple consecutive training batches is less than a preset threshold and the current learning rate has reached the minimum learning rate threshold, the model is determined to have converged and the optimization is terminated to generate an optimized flow matching model.

7. The method for voice style transfer according to claim 1, wherein: Inputting the initial source feature, the content feature and the second style feature into the optimized stream matching model to generate a migration speech waveform, including: splicing the initial source feature, the content feature and the second style feature according to the channel dimension to generate a fused feature vector; Bidirectionally mapping the fused feature vector through the reversible neural network layer in the optimized flow matching model to generate time-frequency domain migration features; Performing short-time inverse Fourier transform on the time-frequency domain migration features to generate a preliminary migration speech waveform; The preliminary migration speech waveform is subjected to phase optimization and range compression processing to generate a final migration speech waveform.

8. A speech style transfer device, characterized in that: The speech style transfer device comprises: A source speech feature extraction module, used to extract initial source features of the source speech, and separate content features and a first style feature from the initial source features; A reference speech style extraction module, used to extract initial reference features of the reference speech and separate second style features from the initial reference features; A time interpolation processing module, used for performing linear interpolation processing on the initial source features according to time parameters to generate intermediate source features; A stream matching mapping module, configured to input the intermediate source feature, the second style feature and the content feature into a stream matching model to generate an intermediate reference feature; A reconstruction feature generation module, used for inputting the intermediate reference feature, the first style feature and the content feature into the stream matching model to generate a reconstruction source feature; a consistency loss analysis module, configured to determine a first cycle consistency loss between the reconstructed source feature and the initial source feature; A reference feature interpolation module, used to perform linear interpolation processing on the initial reference feature according to the time parameter to generate a verification intermediate feature; A verification feature mapping module, used for inputting the verification intermediate feature, the second style feature and the content feature into the stream matching model to generate a verification reference feature; a verification loss analysis module, configured to determine a second cycle consistency loss between the verification reference feature and the intermediate reference feature; A model optimization module, configured to optimize the parameters of the flow matching model by combining the first cycle consistency loss and the second cycle consistency loss; The speech generation module is used to input the initial source feature, the content feature and the second style feature into the optimized stream matching model to generate a migration speech waveform.

9. A computer device, characterized in that: The computer device includes a memory, a processor, and a speech style transfer program stored in the memory and executable on the processor. When the speech style transfer program is executed by the processor, the steps of the speech style transfer method according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium, characterized in that: The storage medium stores a speech style transfer program, and when the speech style transfer program is executed by the processor, the steps of the speech style transfer method according to any one of claims 1 to 7 are implemented.

Citation Information

Cited By

  • AIGC cross-medium-based meta-universe scene dynamic generation method

    CN121614035A