Training method of dialect voice cloning model, dialect voice cloning method and device, terminal equipment and storage medium

By introducing Lora module and low-rank matrix decomposition technology into the speech cloning model, the problem of insufficient realism and adaptability in handling dialect speech is solved, and more efficient training and more realistic dialect speech cloning effects are achieved.

CN119943027AActive Publication Date: 2025-05-06GUANGDONG KAMFU TECH CO LTD
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202411849270.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-16
Publication Date
2025-05-06
Estimated Expiration
2044-12-16

AI Technical Summary

Technical Problem

When handling local dialects, the existing speech cloning model is difficult to accurately capture and reproduce the unique characteristics of dialect pronunciation, resulting in the lack of realism in the cloned dialect pronunciation, and the model is poorly adaptable, requiring a large amount of dialect data for retraining, which is inefficient.

Method used

A training method for dialect speech cloning model is proposed, using text encoder, speech word segmenter, large language submodel, stream matching submodel, Lora module and vocoder to improve the model's ability to capture and reproduce dialect speech features through iterative training and low-rank matrix decomposition technology.

Benefits of technology

It improves the realism of dialect speech cloning and the training efficiency of the model, reduces the need for training data, enhances the adaptability of the model, and can effectively capture and reproduce dialect-specific pronunciation features.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119943027A_ABST
    Figure CN119943027A_ABST
Patent Text Reader

Abstract

The invention discloses a training method of a dialect speech cloning model, a dialect speech cloning method and device, terminal equipment and a storage medium. The training method comprises the steps of obtaining a dialect audio sample and corresponding text information; parameters of a text encoder, a voice word segmentation device, a large language sub-model and a stream matching sub-model are set as fixed parameters, and a first low-rank matrix and a second low-rank matrix of a Lora module are initialized; performing iterative training on the dialect voice clone model by taking the dialect audio sample and the text information as input and the clone voice corresponding to the text information as output until a loss function is converged; during training, under the condition that a loss function is not converged, the first low-rank matrix and the second low-rank matrix are updated, and the product of the updated first low-rank matrix and the updated second low-rank matrix serves as a weight matrix updating quantity and is added into the large language sub-model. According to the invention, the reality sense of dialect cloning and the training efficiency of the model can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of voice cloning technology, and in particular to a method for training a dialect voice cloning model, a dialect voice cloning method, a device, a terminal device and a storage medium. Background Art

[0002] In today's information society, speech synthesis and cloning technology have become an important part of the field of human-computer interaction and are widely used in many industries such as customer service, entertainment, and education. In particular, speech cloning technology can replicate the speech characteristics of a specific individual and achieve the generation of personalized speech.

[0003] However, existing voice cloning models are often optimized for common languages ​​(such as Mandarin). Although they perform well in cloning common voices such as standard Mandarin, they do not retain enough specific voice features of local dialects, such as tones and accents, resulting in a lack of authenticity in the cloned dialect voices. Especially for Cantonese dialects such as Foshan dialect, their unique speech rhythm, pronunciation and vocabulary usage make it difficult for traditional voice cloning models to accurately capture and reproduce. In addition, existing voice cloning models have poor adaptability when dealing with different dialects, and often require a large amount of dialect data for retraining, which is inefficient. Summary of the invention

[0004] The embodiments of the present invention provide a dialect speech cloning model training method, a dialect speech cloning method, an apparatus, a terminal device and a storage medium, which can improve the realism of dialect cloning and the training efficiency of the model.

[0005] An embodiment of the present invention provides a method for training a dialect speech cloning model, which is applicable to the dialect speech cloning model. The dialect speech cloning model includes: a text encoder, a speech word segmenter, a large language sub-model, a stream matching sub-model, a Lora module and a vocoder; the training method includes:

[0006] Get dialect audio samples and corresponding text information;

[0007] Set the parameters of the text encoder, speech segmenter, large language sub-model, and stream matching sub-model to corresponding fixed parameters, and initialize the first low-rank matrix and the second low-rank matrix of the Lora module;

[0008] Taking the dialect audio sample and the corresponding text information as input and the cloned speech corresponding to the text information as output, the initial dialect speech cloning model is iteratively trained until the loss function of the dialect speech cloning model converges; wherein, during training, the cloned speech output by the vocoder is compared with the dialect audio sample, and the loss function value is calculated according to the comparison result. When the loss function value has not converged, the first low-rank matrix and the second low-rank matrix of the Lora module are updated, and the product of the updated first low-rank matrix and the second low-rank matrix is ​​used as the weight matrix update amount and added to the large language sub-model.

[0009] Furthermore, the obtaining of the dialect audio sample and the corresponding text information includes:

[0010] Acquire initial dialect audio, and perform standardization processing on the initial dialect audio according to a preset sampling rate and format to obtain standardized dialect audio;

[0011] Performing noise reduction processing on the standardized dialect audio to obtain noise-reduced dialect audio;

[0012] The denoised dialect audio is cut according to a preset time length to obtain a number of dialect audio clips;

[0013] By using speaker recognition technology, voices of different speakers in the dialect audio segment are separated to obtain separated dialect audio segments, so that each separated dialect audio segment corresponds to the voice of a single speaker;

[0014] Removing the silent part from the separated dialect audio segment to obtain the dialect audio sample;

[0015] Perform text transcription on the dialect audio samples to generate corresponding text information.

[0016] Furthermore, the initial dialect audio is subjected to standardization processing according to a preset sampling rate and format to obtain a standardized dialect audio, including:

[0017] The initial dialect audio is converted into a 16000 Hz WAV format to obtain a standardized dialect audio.

[0018] Furthermore, the preset duration is 5 to 10 seconds.

[0019] Based on the above method embodiment, the present invention provides a corresponding device embodiment;

[0020] An embodiment of the present invention provides a training device for a dialect speech cloning model, comprising: a data acquisition module, a parameter setting module, and an iterative training module;

[0021] The data acquisition module is used to acquire dialect audio samples and corresponding text information;

[0022] The parameter setting module is used to set the parameters of the text encoder, the speech segmenter, the large language sub-model and the stream matching sub-model to corresponding fixed parameters, and initialize the first low-rank matrix and the second low-rank matrix of the Lora module;

[0023] The iterative training module is used to take the dialect audio sample and the corresponding text information as input, and the cloned speech corresponding to the text information as output, and iteratively train the initial dialect speech cloning model until the loss function of the dialect speech cloning model converges; wherein, during training, the cloned speech output by the vocoder is compared with the dialect audio sample, and the loss function value is calculated according to the comparison result. When the loss function value does not converge, the first low-rank matrix and the second low-rank matrix of the Lora module are updated, and the product of the updated first low-rank matrix and the second low-rank matrix is ​​used as the weight matrix update amount and added to the large language sub-model.

[0024] Furthermore, the data acquisition module includes: a standardization processing unit, a noise reduction unit, an audio cutting unit, an audio separation unit, a silence removal unit and a text transcription unit;

[0025] The standardization processing unit is used to obtain the initial dialect audio, perform standardization processing on the initial dialect audio according to a preset sampling rate and format, and obtain standardized dialect audio;

[0026] The noise reduction unit is used to perform noise reduction processing on the standardized dialect audio to obtain noise-reduced dialect audio;

[0027] The audio cutting unit is used to cut the de-noised dialect audio according to a preset time length to obtain a plurality of dialect audio segments;

[0028] The audio separation unit is used to separate the voices of different speakers in the dialect audio segment by speaker recognition technology to obtain separated dialect audio segments, so that each separated dialect audio segment corresponds to the voice of a single speaker;

[0029] The silence removal unit is used to remove the silence part in the separated dialect audio segment to obtain the dialect audio sample;

[0030] The text transcription unit is used to transcribe the dialect audio sample to generate corresponding text information.

[0031] Based on the above method embodiment, the present invention provides a corresponding terminal device embodiment;

[0032] Another embodiment of the present invention provides a terminal device, including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein when the processor executes the computer program, the method for training a dialect speech cloning model as described in any one of the embodiments is implemented.

[0033] Based on the above method embodiment, the present invention provides a storage medium embodiment;

[0034] Another embodiment of the present invention provides a storage medium, which includes a stored computer program, wherein when the computer program is running, the device where the storage medium is located is controlled to execute a method for training a dialect speech cloning model as described in any one of the above embodiments.

[0035] In addition, another embodiment of the present invention further provides a dialect voice cloning method, comprising:

[0036] Obtain the dialect speech to be cloned and the text to be synthesized;

[0037] The dialect speech to be cloned and the text to be synthesized are input into a dialect speech cloning model trained according to the training method of the dialect speech cloning model, so that the dialect speech cloning model extracts the semantic features of the text to be synthesized through a text encoder, extracts the timbre features of the dialect speech to be cloned through a speech segmenter, inputs the semantic features and timbre features into a large language sub-model, generates a speech token sequence corresponding to the text to be synthesized through the large speech sub-model, converts the token sequence into a Mel-spectrogram through a stream matching sub-model, performs sound wave synthesis through a vocoder based on the Mel-spectrogram, and generates a cloned speech with the same timbre as the dialect speech to be cloned and corresponding to the text content of the text to be synthesized.

[0038] The following beneficial effects are achieved by implementing the present invention:

[0039] The invention discloses a training method, a dialect voice cloning method, a device, a terminal device and a storage medium for a dialect voice cloning model. The training method first obtains a dialect audio sample and corresponding text information; then sets the parameters of a text encoder, a speech segmenter, a large language sub-model and a stream matching sub-model to corresponding fixed parameters, and initializes the first low-rank matrix and the second low-rank matrix of a Lora module; then, the dialect audio sample and the corresponding text information are input, and the cloned voice corresponding to the text information is output, and the initial dialect voice cloning model is iteratively trained until the loss function of the dialect voice cloning model converges; during training, the cloned voice output by the vocoder is compared with the dialect audio sample, and the loss function value is calculated according to the comparison result. When the loss function value does not converge, the first low-rank matrix and the second low-rank matrix of the Lora module are updated, and the product of the updated first low-rank matrix and the second low-rank matrix is ​​used as the weight matrix update amount and added to the large language sub-model. During the training process, the parameters of the four modules of the text encoder, the speech segmenter, the large voice sub-model and the stream matching model are kept unchanged. On this basis, we focus on training the Lora module and use the training samples to optimize its first low-rank matrix and second low-rank matrix. The product of the optimized first low-rank matrix and the second low-rank matrix is ​​used as the weight matrix update amount, and the weight matrix update amount of the model is the product of the two low-rank matrices. Since the dimensions of the two low-rank matrices are much smaller than the original weight matrix update amount, this strategy significantly reduces the number of parameters required in the training process and the demand for training data. It can maintain a high cloning quality even when data resources are limited, and improve the training efficiency of the model. In addition, this low-rank matrix decomposition method is not only simple in structure, but also enhances the adaptability of the model to specific dialect cloning, effectively captures and reproduces the unique speech features of the dialect, and improves the regional characteristics and realism of dialect speech cloning. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] Figure 1 The figure is a flow chart of a method for training a dialect speech cloning model provided by one embodiment of the present invention.

[0041] Figure 2 It is a schematic diagram of the structure of a dialect speech cloning model provided by one embodiment of the present invention.

[0042] Figure 3 It is a schematic diagram of the principle of fine-tuning the Lora module in a training method for a dialect speech cloning model provided by an embodiment of the present invention.

[0043] Figure 4 It is a schematic diagram of the structure of a training device for a dialect speech cloning model provided by one embodiment of the present invention.

[0044] Figure 5 The figure is a schematic diagram of the principle of a dialect speech cloning method provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0045] In order to make the purpose, technical solutions and advantages of this application clearer, the technical solutions in this application will be clearly and completely described below in conjunction with the drawings in the embodiments of this application. Obviously, the described embodiments are part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0046] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by technicians in the technical field to which this application belongs; the terms used herein are only for the purpose of describing specific embodiments and are not intended to limit this application; the terms "including" and "having" in the specification and claims of this application and the above-mentioned figure descriptions and any variations thereof are intended to cover non-exclusive inclusions.

[0047] In the description of the embodiments of the present application, the technical terms "first", "second", etc. are only used to distinguish different objects, and cannot be understood as indicating or implying relative importance or implicitly indicating the number, specific order or primary and secondary relationship of the indicated technical features. In the description of the embodiments of the present application, the meaning of "multiple" is more than two, unless otherwise clearly and specifically defined.

[0048] Reference to "embodiments" herein means that a particular feature, structure, or characteristic described in conjunction with the embodiments may be included in at least one embodiment of the present application. The appearance of the phrase in various locations in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment that is mutually exclusive with other embodiments. It is explicitly and implicitly understood by those skilled in the art that the embodiments described herein may be combined with other embodiments.

[0049] In the description of the embodiments of the present application, the term "and / or" is only a description of the association relationship of the associated objects, indicating that there may be three relationships. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. In addition, the character " / " in this article generally indicates that the associated objects before and after are in an "or" relationship.

[0050] In the description of the embodiments of the present application, the term "multiple" refers to more than two (including two). Similarly, "multiple groups" refers to more than two groups (including two groups), and "multiple pieces" refers to more than two pieces (including two pieces).

[0051] In the description of the embodiments of the present application, unless otherwise clearly specified and limited, technical terms such as "installed", "connected", "connected", "fixed" and the like should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection or an indirect connection through an intermediate medium, and it can be the internal connection of two elements or the interaction relationship between two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the embodiments of the present application can be understood according to the specific circumstances.

[0052] like Figure 1 and Figure 2 As shown, an embodiment of the present invention provides a training method for a dialect speech cloning model, which is applicable to the dialect speech cloning model. The dialect speech cloning model includes: a text encoder, a speech segmenter, a large language sub-model, a stream matching sub-model, a Lora module and a vocoder. The training method includes the following steps:

[0053] S1. Obtain dialect audio samples and corresponding text information.

[0054] In a preferred embodiment, the obtaining of the dialect audio sample and the corresponding text information includes:

[0055] Acquire initial dialect audio, and perform standardization processing on the initial dialect audio according to a preset sampling rate and format to obtain standardized dialect audio;

[0056] Performing noise reduction processing on the standardized dialect audio to obtain noise-reduced dialect audio;

[0057] The denoised dialect audio is cut according to a preset time length to obtain a number of dialect audio clips;

[0058] By using speaker recognition technology, voices of different speakers in the dialect audio segment are separated to obtain separated dialect audio segments, so that each separated dialect audio segment corresponds to the voice of a single speaker;

[0059] Removing the silent part from the separated dialect audio segment to obtain the dialect audio sample;

[0060] Perform text transcription on the dialect audio samples to generate corresponding text information.

[0061] In a preferred embodiment, the standardization processing of the initial dialect audio according to a preset sampling rate and format to obtain the standardized dialect audio includes: converting the initial dialect audio into a 16000 Hz WAV format to obtain the standardized dialect audio.

[0062] In a preferred embodiment, the preset duration is 5 to 10 seconds.

[0063] Instructively, taking Foshan dialect as an example, more than 500 hours of Foshan dialect audio materials can be collected from online video platforms to obtain the above-mentioned initial dialect audio.

[0064] Then, the collected Foshan dialect audio materials were uniformly converted into WAV format with a sampling rate of 16,000 Hz to obtain the above-mentioned standardized dialect audio; through standardization processing, the audio collected from different online video platforms was converted into a unified format to maintain data consistency for subsequent processing and model training.

[0065] Since the collected audio data is often accompanied by background noise, in order to avoid interference caused by these noises during the training process, we use noise reduction technology to perform noise reduction on the standardized dialect audio to obtain the above-mentioned noise-reduced dialect audio, so as to extract the pure vocal part and ensure the audio quality.

[0066] Next, in order to improve training efficiency and reduce computing resource consumption, the denoised dialect audio is cut into short audio clips of 5 to 10 seconds to obtain the above-mentioned dialect audio clips to meet the needs of model training.

[0067] For the obtained dialect audio clips, speaker recognition technology is used to separate the voices of different speakers in the dialect audio clips of multiple speakers to obtain separated dialect audio clips, ensuring that the separated dialect audio clips only contain the voice of a single speaker to improve the accuracy of timbre recognition. It is understandable that if there is only one speaker in the dialect audio clip, then separation processing is not required.

[0068] Through silence detection technology, the silent parts in the separated dialect audio clips are identified and removed to obtain the final dialect audio samples. The audio data is further optimized through silence processing and the data volume is reduced, thereby improving the efficiency and quality of training data.

[0069] Finally, accurate speech recognition is performed on each dialect audio sample, and the recognized audio content is transcribed into text information, so that the text information corresponding to each dialect audio sample can be generated.

[0070] By processing the dialect audio samples used for training through the above steps, the quality of the training samples can be greatly improved, thereby improving the accuracy and efficiency of the trained model.

[0071] S2. Set the parameters of the text encoder, speech segmenter, large language sub-model, and stream matching sub-model to corresponding fixed parameters, and initialize the first low-rank matrix and the second low-rank matrix of the Lora module.

[0072] Specifically, in the present invention, a pre-trained model without the Lora module is first constructed, and the pre-trained model includes the above-mentioned text encoder, speech segmenter, large language sub-model, stream matching sub-model and HiFi GAN vocoder;

[0073] Next, the pre-trained model is trained using audio samples of a specific dialect (such as audio samples of the Foshan dialect) and corresponding text information until the pre-trained model converges. When the pre-trained model converges, the parameters of the text encoder, the parameters of the speech segmenter, the parameters of the large language sub-model, and the parameters of the stream matching sub-model are used as the corresponding fixed parameters of the present invention.

[0074] Next, a dialect speech cloning model is constructed, which includes a Lora module, a text encoder, a speech segmenter, a large language sub-model, a stream matching sub-model and a Hifi GAN vocoder. The text encoder, the speech segmenter, the large language sub-model and the stream matching sub-model are set to the corresponding fixed parameters and remain unchanged. Then, the first low-rank matrix and the second low-rank matrix of the Lora module are initialized to obtain an initial dialect speech cloning model.

[0075] S3. Taking the dialect audio sample and the corresponding text information as input and the cloned speech corresponding to the text information as output, the initial dialect speech cloning model is iteratively trained until the loss function of the dialect speech cloning model converges; wherein, during training, the cloned speech output by the vocoder is compared with the dialect audio sample, and the loss function value is calculated according to the comparison result. When the loss function value has not converged, the first low-rank matrix and the second low-rank matrix of the Lora module are updated, and the product of the updated first low-rank matrix and the second low-rank matrix is ​​used as the weight matrix update amount and added to the large language sub-model.

[0076] After obtaining the initial dialect speech cloning model, the dialect audio sample and the corresponding text information are used as input, and the cloned speech corresponding to the text information is used as output. The initial dialect speech cloning model is iteratively trained until the loss function of the dialect speech cloning model converges, and the trained dialect speech cloning model can be obtained.

[0077] Specifically, Figure 3 As shown in the figure, in order to save computing resources to the maximum extent during the training process, we introduced two low-rank matrices A (i.e., the first low-rank matrix) and B (i.e., the first low-rank matrix) of the Lora module based on the aforementioned large speech sub-model to approximate the original weight matrix W of the large speech sub-model. Since the sizes of A and B are much smaller than W, this significantly reduces the number of parameters and computing costs while maintaining the performance of the model, thereby achieving efficient utilization of resources.

[0078] In actual application, by modifying the weight matrix in the linear layer Δ W to achieve:

[0079] (W+ΔW)x=Wx+ABx

[0080] Here, W represents the original weight matrix, and Δ W represents the weight matrix update amount. In the LoRA framework, the weight matrix update amount ΔW is decomposed into the product of two low-rank matrices A and B. Since the dimensions of matrices A and B are much smaller than ΔW, this strategy significantly reduces the number of parameters required during training.

[0081] This low-rank matrix decomposition method is not only simple in structure, but also enhances the adaptability of the model to specific new tasks or data sets while maintaining the original capabilities of the pre-trained model. This method effectively balances the generalization ability of the model with rapid adaptation to new situations, providing an effective way to efficiently fine-tune the pre-trained model.

[0082] Compared with traditional models, the speech cloning model based on the Lora framework we proposed shows significant advantages. This model significantly reduces the demand for training data, so that it can maintain a high synthesis quality when data resources are limited. At the same time, due to the reduction in model complexity, its performance requirements for running devices are more tolerant, thereby broadening the application scope of the model. In addition, the training process of this model is more efficient and the time required is greatly shortened, which not only improves R&D efficiency but also reduces the consumption of computing resources. In summary, the speech synthesis model based on Lora shows excellent performance in data dependence, device compatibility, and training efficiency.

[0083] Based on the implementation of the above training method, another embodiment of the present invention provides a training device for a dialect speech cloning model.

[0084] like Figure 4 As shown, an embodiment of the present invention provides a training device for a dialect speech cloning model, including: a data acquisition module, a parameter setting module and an iterative training module;

[0085] The data acquisition module is used to acquire dialect audio samples and corresponding text information;

[0086] The parameter setting module is used to set the parameters of the text encoder, the speech segmenter, the large language sub-model and the stream matching sub-model to corresponding fixed parameters, and initialize the first low-rank matrix and the second low-rank matrix of the Lora module;

[0087] The iterative training module is used to take the dialect audio sample and the corresponding text information as input, and the cloned speech corresponding to the text information as output, and iteratively train the initial dialect speech cloning model until the loss function of the dialect speech cloning model converges; wherein, during training, the cloned speech output by the vocoder is compared with the dialect audio sample, and the loss function value is calculated according to the comparison result. When the loss function value does not converge, the first low-rank matrix and the second low-rank matrix of the Lora module are updated, and the product of the updated first low-rank matrix and the second low-rank matrix is ​​used as the weight matrix update amount and added to the large language sub-model.

[0088] In a preferred embodiment, the data acquisition module includes: a standardization processing unit, a noise reduction unit, an audio cutting unit, an audio separation unit, a silence removal unit and a text transcription unit;

[0089] The standardization processing unit is used to obtain the initial dialect audio, perform standardization processing on the initial dialect audio according to a preset sampling rate and format, and obtain standardized dialect audio;

[0090] The noise reduction unit is used to perform noise reduction processing on the standardized dialect audio to obtain noise-reduced dialect audio;

[0091] The audio cutting unit is used to cut the de-noised dialect audio according to a preset time length to obtain a plurality of dialect audio segments;

[0092] The audio separation unit is used to separate the voices of different speakers in the dialect audio segment by speaker recognition technology to obtain separated dialect audio segments, so that each separated dialect audio segment corresponds to the voice of a single speaker;

[0093] The silence removal unit is used to remove the silence part in the separated dialect audio segment to obtain the dialect audio sample;

[0094] The text transcription unit is used to transcribe the dialect audio sample to generate corresponding text information.

[0095] It should be noted that the device embodiments described above are merely schematic, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. In addition, in the accompanying drawings of the device embodiments provided by the present invention, the connection relationship between the modules indicates that there is a communication connection between them, which may be specifically implemented as one or more communication buses or signal lines. A person of ordinary skill in the art may understand and implement it without paying any creative effort.

[0096] Those skilled in the art can clearly understand that for the sake of convenience and brevity, the specific working process of the device described above can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here.

[0097] Another preferred embodiment of the present invention provides a terminal device, including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein when the processor executes the computer program, a method for training a dialect speech cloning model as described in any one of the above embodiments is implemented.

[0098] The terminal device may be a computing device such as a desktop computer, a notebook, a PDA, a cloud server, etc. The terminal device may include, but is not limited to, a processor and a memory.

[0099] The processor may be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc. The processor is the control center of the terminal device, and uses various interfaces and lines to connect various parts of the entire terminal device.

[0100] The memory can be used to store the computer program, and the processor realizes various functions of the terminal device by running or executing the computer program stored in the memory and calling the data stored in the memory. The memory can mainly include a program storage area and a data storage area, wherein the program storage area can store an operating system, at least one application required for a function, etc.; the data storage area can store data created according to the use of the mobile phone, etc. In addition, the memory can include a high-speed random access memory, and can also include a non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a smart memory card (Smart Med i aCard, SMC), a secure digital (Secure Digital, SD) card, a flash card (F l ash Card), at least one disk storage device, a flash memory device, or other volatile solid-state storage devices.

[0101] Another preferred embodiment of the present invention provides a storage medium, which includes a stored computer program, wherein when the computer program is running, the device where the storage medium is located is controlled to execute any one of the methods for training a dialect speech cloning model described in the present invention.

[0102] The storage medium is a computer-readable storage medium, and the computer program is stored in the computer-readable storage medium. When the computer program is executed by the processor, the steps of each of the above-mentioned method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal and software distribution medium.

[0103] Based on the above embodiment of the method for training a dialect speech cloning model, another embodiment of the present invention provides a dialect speech cloning method, comprising:

[0104] Obtain the dialect speech to be cloned and the text to be synthesized;

[0105] The dialect speech to be cloned and the text to be synthesized are input into a dialect speech cloning model trained by the training method of the dialect speech cloning model according to any one of claims 1 to 4, so that the dialect speech cloning model extracts semantic features of the text to be synthesized through a text encoder, extracts timbre features of the dialect speech to be cloned through a speech segmenter, inputs the semantic features and timbre features into a large language sub-model, generates a speech token sequence corresponding to the text to be synthesized through the large speech sub-model, converts the token sequence into a Mel-spectrogram through a stream matching sub-model, performs sound wave synthesis through a vocoder based on the Mel-spectrogram, and generates a cloned speech with the same timbre as the dialect speech to be cloned and corresponding to the text content of the text to be synthesized.

[0106] Indicatively, Figure 5 As shown, the Foshan dialect speech to be cloned and the text to be synthesized are input into the above-mentioned trained dialect speech cloning model. The model fuses the processing results of the speech segmenter and the text encoder, and introduces the trained Lora parameters, which are sent to the large speech sub-model for deep processing. Subsequently, the processed data flows into the streaming matching model for matching optimization. Finally, the system outputs the Foshan dialect clone speech corresponding to the text to be synthesized, and the audio content of the output Foshan dialect clone speech is the text to be synthesized. Assuming that the text content of the text to be synthesized is "I am from Foshan", and the input Foshan dialect speech to be cloned is a Foshan dialect of any content, then the clone speech generated at this time is the speech of "I am from Foshan" formed in the form of Foshan dialect.

[0107] The above is a preferred embodiment of the present invention. It should be pointed out that a person skilled in the art can make several improvements and modifications without departing from the principle of the present invention. These improvements and modifications are also considered to be within the scope of protection of the present invention.

Claims

1. A method for training a dialect speech cloning model, characterized in that: Applicable to a dialect speech cloning model, the dialect speech cloning model comprising: a text encoder, a speech word segmenter, a large language sub-model, a stream matching sub-model, a Lora module and a vocoder; The training method comprises: Get dialect audio samples and corresponding text information; Set the parameters of the text encoder, speech segmenter, large language sub-model, and stream matching sub-model to corresponding fixed parameters, and initialize the first low-rank matrix and the second low-rank matrix of the Lora module; Taking the dialect audio sample and the corresponding text information as input and the cloned speech corresponding to the text information as output, the initial dialect speech cloning model is iteratively trained until the loss function of the dialect speech cloning model converges; wherein, during training, the cloned speech output by the vocoder is compared with the dialect audio sample, and the loss function value is calculated according to the comparison result. When the loss function value has not converged, the first low-rank matrix and the second low-rank matrix of the Lora module are updated, and the product of the updated first low-rank matrix and the second low-rank matrix is ​​used as the weight matrix update amount and added to the large language sub-model.

2. The method for training a dialect speech cloning model as claimed in claim 1, characterized in that: The obtaining of the dialect audio sample and the corresponding text information includes: Acquire initial dialect audio, and perform standardization processing on the initial dialect audio according to a preset sampling rate and format to obtain standardized dialect audio; Performing noise reduction processing on the standardized dialect audio to obtain noise-reduced dialect audio; The denoised dialect audio is cut according to a preset time length to obtain a number of dialect audio clips; By using speaker recognition technology, voices of different speakers in the dialect audio segment are separated to obtain separated dialect audio segments, so that each separated dialect audio segment corresponds to the voice of a single speaker; Removing the silent part from the separated dialect audio segment to obtain the dialect audio sample; Perform text transcription on the dialect audio samples to generate corresponding text information.

3. The method for training a dialect speech cloning model as claimed in claim 2, characterized in that: The step of performing standardization processing on the initial dialect audio according to a preset sampling rate and format to obtain the standardized dialect audio includes: The initial dialect audio is converted into a 16000 Hz WAV format to obtain a standardized dialect audio.

4. The method for training a dialect speech cloning model as claimed in claim 3, characterized in that: The preset duration is 5 to 10 seconds.

5. A training device for a dialect speech cloning model, characterized in that: Suitable for training a dialect speech cloning model, the dialect speech cloning model includes: a text encoder, a speech word segmenter, a large language sub-model, a stream matching sub-model, a Lora module and a vocoder; The training device comprises: a data acquisition module, a parameter setting module and an iterative training module; The data acquisition module is used to acquire dialect audio samples and corresponding text information; The parameter setting module is used to set the parameters of the text encoder, the speech segmenter, the large language sub-model and the stream matching sub-model to corresponding fixed parameters, and initialize the first low-rank matrix and the second low-rank matrix of the Lora module; The iterative training module is used to take the dialect audio sample and the corresponding text information as input, and the cloned speech corresponding to the text information as output, and iteratively train the initial dialect speech cloning model until the loss function of the dialect speech cloning model converges; wherein, during training, the cloned speech output by the vocoder is compared with the dialect audio sample, and the loss function value is calculated according to the comparison result. When the loss function value does not converge, the first low-rank matrix and the second low-rank matrix of the Lora module are updated, and the product of the updated first low-rank matrix and the second low-rank matrix is ​​used as the weight matrix update amount and added to the large language sub-model.

6. The training device for a dialect speech cloning model as claimed in claim 5, characterized in that: The data acquisition module includes: a standardization processing unit, a noise reduction unit, an audio cutting unit, an audio separation unit, a silence removal unit and a text transcription unit; The standardization processing unit is used to obtain the initial dialect audio, perform standardization processing on the initial dialect audio according to a preset sampling rate and format, and obtain standardized dialect audio; The noise reduction unit is used to perform noise reduction processing on the standardized dialect audio to obtain noise-reduced dialect audio; The audio cutting unit is used to cut the de-noised dialect audio according to a preset time length to obtain a plurality of dialect audio segments; The audio separation unit is used to separate the voices of different speakers in the dialect audio segment by speaker recognition technology to obtain separated dialect audio segments, so that each separated dialect audio segment corresponds to the voice of a single speaker; The silence removal unit is used to remove the silence part in the separated dialect audio segment to obtain the dialect audio sample; The text transcription unit is used to transcribe the dialect audio sample to generate corresponding text information.

7. A terminal device, characterized in that: The method comprises a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein the processor implements the method for training a dialect speech cloning model as claimed in any one of claims 1 to 4 when executing the computer program.

8. A storage medium, characterized in that: The storage medium includes a stored computer program, wherein when the computer program is executed, the device where the storage medium is located is controlled to execute the method for training a dialect speech cloning model as claimed in any one of claims 1 to 4.

9. A dialect voice cloning method, characterized in that: include: Obtain the dialect speech to be cloned and the text to be synthesized; The dialect speech to be cloned and the text to be synthesized are input into a dialect speech cloning model trained by the training method for a dialect speech cloning model according to any one of claims 1 to 4, so that the dialect speech cloning model extracts semantic features of the text to be synthesized through a text encoder, extracts timbre features of the dialect speech to be cloned through a speech segmenter, inputs the semantic features and timbre features into a large language sub-model, generates a speech token sequence corresponding to the text to be synthesized through the large speech sub-model, converts the token sequence into a Mel-spectrogram through a stream matching sub-model, performs sound wave synthesis through a vocoder based on the Mel-spectrogram, and generates a cloned speech with the same timbre as the dialect speech to be cloned and corresponding to the text content of the text to be synthesized.

Citation Information

Patent Citations

  • Voice cloning model training method, readable storage medium and voice cloning method

    CN111696521A

  • Model training method, dialect recognition method, device, server and storage medium

    CN112634867A

  • Speech clone model training method and device, speech synthesis method and device and related equipment

    CN116403558A

  • Voice cloning method and system and electronic equipment

    CN116913301A

  • Automatic voice processing method based on language large model and electronic equipment

    CN119091864A