Method and device for generating translation speech with specified tone, and storage medium
By receiving target timbre audio and using neural networks for timbre embedding and mapping adjustment, the problem of AI translation systems being unable to reproduce the speaker's timbre is solved, generating personalized target timbre translated speech and improving the user experience.
Patent Information
- Application Number
- CN202511467182.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-14
- Publication Date
- 2026-01-02
AI Technical Summary
Existing AI translation systems cannot effectively reproduce the original speaker's vocal characteristics during the speech synthesis process, resulting in overly mechanical and rigid output speech that affects user experience and the coherence of emotional communication.
By receiving the target timbre audio, using a preset timbre extraction algorithm and neural network to extract the timbre embedding vector, and combining it with a preset translation algorithm and timbre mapping adjustment processing, the target timbre translated speech is generated, preserving the original speaker's personalized voice characteristics.
The generated target voice translation closely resembles the original speaker's voice, overcoming the problems of mechanical stiffness and monotony, improving the personalization of speech translation, and enhancing the user experience.
Smart Images

Figure CN121260147A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of voice generation, and in particular to a method for generating translated voice with specified tone, a device and a storage medium. BACKGROUND
[0002] In the prior art, with the rapid development of artificial intelligence (AI) technology, AI translation systems have been widely used in the field of voice translation, enabling real-time voice communication across languages. However, the current mainstream AI translation systems, after completing voice translation, usually generate voice output in the target language using preset voice synthesis technology. This synthesized voice often uses a standardized and mechanized voice model, lacking the personalized voice characteristics and tone expressions of the original speaker, resulting in translated voice that sounds stiff and monotonous, with obvious "robot" characteristics, severely affecting the user's auditory experience and the coherence of emotional communication.
[0003] The existing AI translation system generally uses traditional voice synthesis technology based on parameter synthesis or splicing synthesis in the voice synthesis link. Although these technologies can achieve basic voice output, they cannot effectively restore the original speaker's tone and other personalized characteristics. Therefore, in view of the technical problem that the output voice of current voice translation is too mechanical and stiff to effectively restore the speaker's tone, a new technology is needed to solve the current problem. SUMMARY
[0004] The main purpose of the present application is to solve the technical problem that the output voice of current voice translation is too mechanical and stiff to effectively restore the speaker's tone.
[0005] The first aspect of the present application provides a method for generating translated voice with specified tone, comprising the steps of: receiving target tone audio; extracting tone from the target tone audio according to a preset tone extraction algorithm to obtain a target tone; receiving first language voice; translating the first language voice according to a preset translation algorithm to obtain second language voice; mapping and adjusting the second language voice according to the target tone to obtain target tone translated voice.
[0006] Optionally, in the first implementation manner of the first aspect of the present application, the step of translating the first language voice according to a preset translation algorithm to obtain second language voice comprises: splitting the first language voice based on a preset discrete splitting algorithm to generate N discrete voice segments, where N is a positive integer; The N discrete speech segments are processed by language mapping based on a preset sequence-to-sequence translation network to obtain M target language speech segments, where M is a positive integer. The M target language segments are combined and encoded to generate a second language speech.
[0007] Optionally, in the second implementation manner of the first aspect of the present application, the splitting processing of the first language speech based on the preset discrete splitting algorithm to generate N discrete speech segments comprises: The first language speech is split to generate N discrete speech segments based on a preset CNN encoder.
[0008] Optionally, in the third implementation manner of the first aspect of the present application, the target timbre extraction processing of the target timbre audio based on the preset timbre extraction algorithm to obtain a target timbre comprises: The target timbre audio is processed by a voiceprint extraction based on a preset ECAPA-TDNN algorithm to generate a timbre embedding vector corresponding to the target timbre.
[0009] Optionally, in the fourth implementation manner of the first aspect of the present application, the mapping adjustment processing of the second language speech based on the target timbre to obtain a target timbre translation speech comprises: The second language speech is processed by content recognition based on a preset Transformer multi-layer perception network to obtain a text embedding vector; The timbre embedding vector and the text embedding vector are processed by audio duration prediction based on a preset duration prediction network to obtain a duration prediction vector; The text embedding vector is processed by linear mapping based on a preset linear transformation network to obtain a text mapping vector; The duration prediction vector and the text mapping vector are processed by feature alignment based on the text timbre alignment network to obtain an alignment vector; The alignment vector and the mapping embedding vector are processed by residual learning based on a preset residual network to obtain a residual alignment vector; The residual alignment vector and the timbre embedding vector are processed by timbre constraint based on a preset GAN timbre adversarial network to generate a target timbre translation speech.
[0010] Optionally, in the fifth implementation manner of the first aspect of the present application, the mapping adjustment processing of the second language speech based on the target timbre to obtain a target timbre translation speech comprises: The second language speech is processed by content, duration and timbre mapping adjustment based on the target timbre to obtain a target timbre translation speech.
[0011] Optionally, in a sixth implementation form of the first aspect of the present application, the receiving the first language speech comprises: collecting an external language speech; performing noise reduction processing on the external language speech according to a preset noise reduction algorithm to generate the first language speech.
[0012] Optionally, in a seventh implementation form of the first aspect of the present application, the receiving the target timbre audio comprises: collecting a user timbre audio; performing noise reduction processing on the user timbre audio according to a preset noise reduction algorithm to generate the target timbre audio.
[0013] The second aspect of the present application provides a specified timbre translation speech generation device, comprising a memory and at least one processor, the memory storing instructions, and the memory and the at least one processor being interconnected by a circuit; the at least one processor invokes the instructions in the memory to enable the specified timbre translation speech generation device to perform the specified timbre translation speech generation method described above.
[0014] The third aspect of the present application provides a computer readable storage medium, which stores instructions, when running on a computer, enabling the computer to perform the specified timbre translation speech generation method described above.
[0015] In the embodiments of the present application, the voiceprint features of the target audio are used to control the timbre of the second language speech for subsequent voice-to-speech translation, so that the timbre of the finally generated target timbre translation speech is close to the timbre of the target timbre audio, ensuring that the target timbre translation speech is consistent with the original speaker or the desired timbre, overcoming the mechanical stiffness and monotony of the existing translation audio, and the output speech of the voice translation is close to the personalized features of human voice, solving the technical problem that the output speech of the current voice translation is too mechanical and rigid to effectively restore the speaker's timbre. BRIEF DESCRIPTION OF DRAWINGS
[0016] Figure 1 An embodiment schematic diagram of the specified timbre translation speech generation method in the embodiments of the present application; Figure 2 An embodiment schematic diagram of the 104 step of the specified timbre translation speech generation method in the embodiments of the present application; Figure 3 A neural network framework schematic diagram of the 104 step of the specified timbre translation speech generation method in the embodiments of the present application; Figure 4 A neural network framework schematic diagram of the 105 step of the specified timbre translation speech generation method in the embodiments of the present application; Figure 5 An embodiment of a designated voice translation voice generation device in the embodiments of the present application. DETAILED DESCRIPTION
[0017] The embodiments of the present application provide a designated voice translation voice generation method, device and storage medium.
[0018] The embodiments of the present application will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present application are shown in the drawings, it should be understood that the present application can be implemented in various forms and should not be interpreted as being limited to the embodiments set forth herein, but rather the embodiments are provided to more thoroughly and completely understand the present application. It should be understood that the drawings and embodiments of the present application are only for exemplary purposes and are not intended to limit the scope of protection of the present application.
[0019] In the description of the embodiments of the present application, the term "comprising" and similar terms are to be interpreted as open-ended, i.e. "including but not limited to". The term "based on" is to be interpreted as "based, at least in part, on". The term "one embodiment" or "the embodiment" is to be interpreted as "at least one embodiment". The terms "first", "second", etc. can refer to different or same objects. Other explicit and implicit definitions can also be included below.
[0020] For ease of understanding, the specific process of the embodiments of the present application is described below. Please refer to Figure 1 One embodiment of a designated voice translation voice generation method in the embodiments of the present application includes the following steps: 101, receiving target voice audio; In this embodiment, the user can upload the voice audio that needs to be converted, such as the voice of a character that the user likes, or the user can upload the voice audio of the user's own voice.
[0021] Specifically, the following specific embodiments are included in step 101: 1011, collecting user voice audio; 1012, performing noise reduction processing on the user voice audio according to a preset noise reduction algorithm to generate target voice audio.
[0022] In steps 1011-1012, the user's voice audio is collected in real time by a microphone, and after noise reduction of the collected user voice audio by a pre-set noise reduction algorithm, the target voice audio is generated.
[0023] 102, performing voice extraction processing on the target voice audio according to a preset voice extraction algorithm to obtain a target voice; In the embodiment, the target timbre audio is subjected to timbre extraction processing according to a timbre extraction algorithm related to voiceprint recognition, and a target timbre corresponding to a subsequent translation audio generation style is obtained.
[0024] Further, the following specific embodiments are included in the step 102: 1021, based on the preset ECAPA-TDNN algorithm, the target timbre audio is subjected to voiceprint extraction processing, and a timbre embedding vector corresponding to the target timbre is generated.
[0025] In the step 1021, for the audio data of the target female timbre or the target male timbre, the ECAPA-TDNN algorithm is used to extract the voiceprint of the target timbre, and a timbre embedding vector corresponding to the target timbre is generated. The timbre embedding vector contains all the features of voiceprint recognition.
[0026] 103, receiving a first language voice; In the embodiment, the Chinese language voice to be translated is received, and the subsequent target is to translate the Chinese language voice into English language voice.
[0027] Further, the following specific embodiments are included in the step 103: 1031, collecting external language voice; 1032, according to the preset noise reduction algorithm, the external language voice is subjected to noise reduction processing to generate a first language voice.
[0028] In the steps 1031-1032, the real-time Chinese language voice issued by the Chinese user is collected through the microphone. Then, the preset noise reduction algorithm is used to reduce the noise of the collected Chinese language voice, and the Chinese language voice to be translated is generated.
[0029] 104, according to the preset translation algorithm, the first language voice is subjected to translation processing to obtain a second language voice; In the embodiment, the translation algorithm uses a sound-to-sound translation method to convert the Chinese language voice into a cross-language timbre preserving voice. It is preset to convert Chinese into English. Based on the conversion setting, the Chinese language voice is subjected to neural network mapping translation, and a second language voice is generated through sound-to-sound conversion. In order to preserve the original timbre, the text is not used as an intermediate for translation conversion, but the mapping relationship between sound and sound is used.
[0030] Please refer to Figure 2 , Figure 2 For a specific embodiment of the 104 step of the specified timbre translation voice generation method in the embodiment of the application, the following specific embodiments are included in the step 104: 1041、based on the preset discrete splitting algorithm, the first language voice is split to generate N discrete voice segments, wherein N is a positive integer; 1042, based on the preset sequence to sequence translation network, the N discrete voice segments are processed by language mapping to obtain M target language voice segments, wherein M is a positive integer; 1043, the M target language segments are combined and encoded to generate a second language voice.
[0031] In steps 1041-1043, please refer to Figure 3 , Figure 3 The neural network framework diagram of the 104 step of the specified voice translation voice generation method in the embodiment of the application. The first language voice is input as an audio signal into the discrete splitting algorithm, and is split into multiple and discrete voice segment vectors through several layers of CNN. In the figure, 6 segments are generated, 3 of which are masked and input into the Transformer. The input voice is extracted through the self-supervised kernel comparison module, the MFCC feature is processed through K-means, and the output segment is compared with the codebook z1-z6 to find the closest distance, for example: the first frame output discrete voice segment feature is close to z2, the second frame output discrete voice segment feature is close to z3, and the third frame output discrete voice segment feature is close to z4, and all discrete voice segments are looped. Then, the distance between the vector in the mask and the target codebook is taken as the objective function, and the objective function is minimized. After minimizing the objective function, the new codebook is taken as the pseudo label to train the second round of model, and the above process is repeated to obtain an effective voice-to-voice model.
[0032] By processing the N discrete voice segments by language mapping, M target language voice segments are obtained, and finally based on the set decoder, the M target language segments are combined and encoded to generate a second language voice.
[0033] Specifically, in step 1041, the following specific embodiments are included: 10411, based on the preset CNN encoder, the first language voice is split to generate N discrete voice segments.
[0034] In step 10411, assuming that a voice signal has a sampling rate of 16000Hz, through 5 layers of convolutional CNN, the step length of each layer is 5, 4, 2, 2, 2, and through these five layers, the signal is sampled by 5x4x2x2x2=160. The output discrete voice segment is 1 second, 16000 / 160=100 points per second, and finally the first language voice is split to generate N discrete voice segments.
[0035] 105. mapping and adjusting the second language speech according to the target timbre to obtain target timbre translated speech.
[0036] In the embodiment, based on the preset GPT-SoVITS algorithm, the extracted voiceprint features of the target timbre are used to perform scaling and mapping adjustment on the spectrum and amplitude of the second language speech, and finally the target timbre translated speech is obtained.
[0037] Specifically, the following specific implementation is included in step 105: 1051. mapping and adjusting the content, duration and timbre of the second language speech according to the target timbre to obtain target timbre translated speech.
[0038] In step 1051, the content, duration and timbre of the second language speech are mapped and adjusted according to the target timbre to obtain the target timbre translated speech.
[0039] More specifically, please refer to Figure 4 , Figure 4 The neural network framework diagram for the 105 step of the specified timbre translated speech generation method in the embodiment of the present application is shown in the figure. In the embodiment of step 1021, the 105 step includes the following specific implementation: 1052. performing content recognition processing on the second language speech according to the preset Transformer multi-layer perception network to obtain a text embedding vector; 1053. performing audio duration prediction processing on the timbre embedding vector and the text embedding vector according to the preset duration prediction network to obtain a duration prediction vector; 1054. performing linear mapping processing on the text embedding vector according to the preset linear transformation network to obtain a text mapping vector; 1055. performing feature alignment processing on the duration prediction vector and the text mapping vector according to the text timbre alignment network to obtain an alignment vector; 1056. performing residual learning processing on the alignment vector and the mapping embedding vector according to the preset residual network to obtain a residual alignment vector; 1057. performing timbre constraint processing on the residual alignment vector and the timbre embedding vector according to the preset GAN timbre adversarial network to generate target timbre translated speech.
[0040] In steps 1052-1057, based on the neural network framework, the Transformer multi-layer perception network is first used to perform perception input on the second language speech to generate a text embedding vector. In the duration prediction network, the text embedding vector and the timbre embedding vector are used to predict the normal human speech speed to generate a duration prediction vector.
[0041] In the linear transformation network, the text embedding vector is linearly mapped, so that the network can learn the features of the second language voice in a deeper manner, and a text mapping vector is generated.
[0042] In the text timbre alignment network, the duration prediction vector is aligned with the text mapping vector, so that the output speed of the text content is the duration constrained by the duration prediction vector, and an alignment vector is generated.
[0043] In the deep learning network, the network is prevented from being too deep to cause feature overfitting and loss, the timbre feature vector is added into the alignment vector through the residual network, and a residual alignment vector is generated.
[0044] Finally, in the GAN timbre adversarial network, the timbre features of the residual alignment vector are compared with the original timbre embedding vector for adversarial learning, so that the output audio timbre is closer to the original timbre embedding vector. In the GAN timbre adversarial network, an audio decoder is integrated, the target timbre translation voice is output, and the timbre consistency between the target timbre translation voice and the target timbre audio is realized.
[0045] In the embodiment of the present application, the timbre of the second language voice for subsequent voice-to-voice translation is controlled by using the voiceprint features of the target audio, so that the timbre of the finally generated target timbre translation voice is close to the timbre of the target timbre audio, and the target timbre translation voice and the original speaker or the desired timbre are consistent, overcoming the mechanical stiffness and monotony of the existing translation audio. The output voice of the voice translation is close to the personalized features of the human voice, and the technical problem that the output voice of the current voice translation is too mechanical and rigid to effectively restore the timbre of the speaker is solved.
[0046] Figure 5 It is a structure schematic diagram of a specified timbre translation voice generation device provided by the embodiment of the present application. The specified timbre translation voice generation device 500 can have great differences due to different configurations or performances, and can include one or more central processing units (CPUs) 510 (for example, one or more processors) and a memory 520, and one or more storage media 530 (for example, one or more mass storage devices) storing application programs 533 or data 532. The memory 520 and the storage medium 530 can be temporary storage or persistent storage. The programs stored in the storage medium 530 can include one or more modules (not shown in the figure), and each module can include a series of instruction operations in the specified timbre translation voice generation device 500. Further, the processor 510 can be configured to communicate with the storage medium 530 and execute a series of instruction operations in the storage medium 530 on the specified timbre translation voice generation device 500.
[0047] The specified-pitch translated speech generation device 500 can also include one or more power supplies 540, one or more wired or wireless network interfaces 550, one or more input / output interfaces 560, and / or one or more operating systems 531, such as Windows Server, Mac OS X, Unix, Linux, FreeBSD, etc. Those skilled in the art will appreciate that Figure 5 The illustrated specified-pitch translated speech generation device architecture is not meant to imply architectural limitations with respect to the specified-pitch translated speech generation device, and that the device can include more or fewer components than shown, or combine some components, or have a different arrangement of components.
[0048] The present disclosure also provides a computer readable storage medium, which can be a non-volatile computer readable storage medium, or a volatile computer readable storage medium, and the computer readable storage medium has instructions stored therein, and the instructions, when executed on a computer, cause the computer to perform the steps of the specified-pitch translated speech generation method.
[0049] In the context of this disclosure, a machine-readable medium can be a tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include but is not limited to an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium will include one or more lines of electrical connections, portable computer disks, hard disk drives, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or Flash memory), optical fibers, portable compact disc read-only memories (CD-ROMs), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0050] Moreover, while operations can be depicted in a particular, serial order, this should not be understood as requiring or implying that such operations be performed in the particular order shown, or in sequential order, or that all illustrated operations be performed, to achieve desirable results. In certain circumstances, multitasking and parallel processing can be advantageous. Likewise, while specific implementations are discussed herein, this should not be understood as implying that these are the only implementations that can be employed to carry out the present disclosure. Rather, other implementations can be employed, without departing from the scope of the present disclosure. Similarly, while operations are depicted as occurring in certain components, this should not be understood as requiring or implying that such operations are carried out only in those components, and that other components cannot incur such operations for the specified purpose. Conversely, operations depicted as occurring in a single component can be performed across multiple components, or components can be combined in a single component, without departing from the scope of the present disclosure.
[0051] Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are disclosed as example forms of implementing the claims.
Claims
1. A method for generating translation speech with a specified timbre, characterized in that, Including the following steps: Receive the target tone audio; According to the preset timbre extraction algorithm, the target timbre audio is processed to extract the timbre and obtain the target timbre; Receives first-language speech; According to a preset translation algorithm, the first language speech is translated to obtain the second language speech; Based on the target timbre, the second language speech is mapped and adjusted to obtain the target timbre translated speech.
2. The method for generating translation speech with a specified timbre according to claim 1, characterized in that, The step of translating the first language speech according to a preset translation algorithm to obtain the second language speech includes: Based on a preset discrete segmentation algorithm, the first language speech is segmented to generate N discrete speech segments, where N is a positive integer; Based on a pre-defined sequence-to-sequence translation network, N discrete speech segments are processed by language mapping to obtain M target language speech segments, where M is a positive integer; The M target language segments are combined and encoded to generate second language speech.
3. The method for generating translation speech with a specified timbre according to claim 2, characterized in that, The step of splitting the first language speech into N discrete speech segments based on a preset discrete segmentation algorithm includes: Based on a pre-set CNN encoder, the first language speech is split into N discrete speech segments.
4. The method for generating translation speech with a specified timbre according to claim 3, characterized in that, The step of performing timbre extraction processing on the target timbre audio according to a preset timbre extraction algorithm to obtain the target timbre includes: Based on the preset ECAPA-TDNN algorithm, the target timbre audio is processed to extract voiceprints and generate a timbre embedding vector corresponding to the target timbre.
5. The method for generating translation speech with a specified timbre according to claim 4, characterized in that, The step of mapping and adjusting the second language speech based on the target timbre to obtain the target timbre translated speech includes: Based on a preset Transformer multilayer perceptron, the second language speech is processed for content recognition to obtain a text embedding vector. Based on a preset duration prediction network, audio duration prediction processing is performed on the timbre embedding vector and the text embedding vector to obtain a duration prediction vector. Based on a preset linear transformation network, the text embedding vector is linearly mapped to obtain a text mapping vector. Based on the text timbre alignment network, feature alignment processing is performed on the duration prediction vector and the text mapping vector to obtain an alignment vector; Based on a preset residual network, residual learning processing is performed on the alignment vector and the mapping embedding vector to obtain a residual alignment vector. Based on a preset GAN timbre adversarial network, timbre constraint processing is performed on the residual alignment vector and the timbre embedding vector to generate target timbre translated speech.
6. The translation speech generation method with specified timbre according to claim 1, characterized in that, The step of mapping and adjusting the second language speech based on the target timbre to obtain the target timbre translated speech includes: Based on the target timbre, the second language speech is adjusted in terms of content, duration, and timbre mapping to obtain the target timbre translated speech.
7. The method for generating translation speech with a specified timbre according to claim 1, characterized in that, The receiving of the first language voice includes: Collect external language speech; According to a preset noise reduction algorithm, the external language speech is processed to reduce noise and generate the first language speech.
8. The method for generating translation speech with a specified timbre according to claim 1, characterized in that, The received target timbre audio includes: Collect user voice audio; According to the preset noise reduction algorithm, the user's audio timbre is processed to reduce noise and generate the target audio timbre.
9. A translation speech generation device with a specified timbre, characterized in that, The translation speech generation device with a specified timbre includes: a memory and at least one processor, wherein the memory stores instructions, and the memory and the at least one processor are interconnected via a line; The at least one processor invokes the instructions in the memory to cause the translation speech generation device with the specified timbre to execute the translation speech generation method with the specified timbre as described in any one of claims 1-8.
10. A computer-readable storage medium storing a computer program thereon, characterized in that, When the computer program is executed by the processor, it implements the translation speech generation method with a specified timbre as described in any one of claims 1-8.