Mediator timbre cloning method and system, electronic equipment and storage medium
Through tone cloning technology, the autoregressive model and acoustic model are used to generate output audio consistent with the mediator's tone, which solves the problems of high pressure and low efficiency of mediators in traditional mediation methods, and achieves efficient and good quality mediation services.
Patent Information
- Application Number
- CN202510203986.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-24
- Publication Date
- 2025-06-03
AI Technical Summary
Traditional mediation methods rely on manual mediators, resulting in high work pressure, low efficiency, and difficulty in quickly adjusting tone and expression in different situations, affecting the mediation effect.
Through tone cloning technology, the autoregressive model and acoustic model are used to generate output audio consistent with the mediator's tone, achieving flexible adjustment and automatic output of the mediator's tone.
It reduces the work pressure of mediators, improves mediation efficiency and service quality, ensures tone consistency and fairness, and improves user experience and mediation effect.
Smart Images

Figure CN120089124A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of timbre cloning, and more particularly to a mediator timbre cloning method, system, electronic device, and storage medium. Background Art
[0002] In current mediation practices, traditional methods often rely on mediators to intervene in person on-site. Mediators communicate with parties through means such as phone calls or video conferences. However, with the continuous increase in the number of mediation cases and the higher requirements for mediation efficiency, human mediators are facing increasing work pressure. At the same time, given the complex diversity of cases, mediators must be able to quickly adjust their tone, intonation, and expression in various different situations, which undoubtedly increases the complexity of the mediation process and consumes a large amount of time.
[0003] In view of this, the present invention is specifically proposed. Summary of the Invention
[0004] The present invention is proposed in consideration of the above problems. According to one aspect of the present invention, there is provided a mediator timbre cloning method, including: Selecting a mediator timbre in response to a selection instruction input by a user; Obtaining input text; Predicting a target audio feature vector of the input text using an autoregressive model; Searching for a target cluster center in an audio dictionary that matches the target audio feature vector, where the audio dictionary includes multiple clusters, and each cluster includes multiple audio feature vectors; Inputting the input text phonemes of the input text, the frequency features of a reference audio, and the target cluster center into a trained acoustic model to obtain a target output audio, where the reference audio is a mediator audio corresponding to the mediator timbre.
[0005] Exemplarily, the trained acoustic model is trained in the following manner: Determining a reference audio feature vector of the reference audio; Searching for a reference cluster center in the audio dictionary that matches the target audio feature vector; Inputting the reference text phonemes of a reference text, the frequency features of the reference audio, and the reference cluster center into an initial acoustic model to obtain a reference output audio, where the reference text is the text form of the reference audio; Optimizing the initial acoustic model based on the difference between the reference audio and the reference output audio to obtain the trained acoustic model.
[0006] Exemplarily, determining the reference audio feature vector of the reference audio includes: Processing the reference audio using a Hubert model to obtain the reference audio feature vector.
[0007] Exemplarily, using an autoregressive model to predict the target audio feature vector of the input text includes: Inputting the input text feature vector of the input text, the input text phonemes, the reference text phonemes of the reference text, the reference text feature vector of the reference text, and the reference audio feature vector of the reference audio into the autoregressive model to obtain the target audio feature vector.
[0008] Exemplarily, before using the autoregressive model to predict the target audio feature vector of the input text, the method further includes: Performing phoneme conversion on the input text to obtain the input text phonemes; And / or, Inputting the input text into a BERT model to obtain the input text feature vector.
[0009] Exemplarily, the audio dictionary is obtained by the following method: Determining the feature vector of each training audio in the training audio set; Using a clustering algorithm to cluster the feature vectors of each training audio in the training audio set to obtain a plurality of clustering clusters; Constructing the audio dictionary based on the plurality of clustering clusters.
[0010] Exemplarily, obtaining the input text includes: Receiving the input text input by the user; Or, Receiving the input audio input by the user; Converting the input audio into text to obtain the input text.
[0011] According to another aspect of the present invention, there is provided a mediator voice cloning system, including: A selection module for selecting a mediator voice in response to a selection instruction input by the user; An acquisition module for acquiring the input text; An autoregressive module for using an autoregressive model to predict the target audio feature vector of the input text; A vector quantization module for finding a target clustering center in the audio dictionary that matches the target audio feature vector, where the audio dictionary includes a plurality of clustering clusters, and each clustering cluster includes a plurality of audio feature vectors; An acoustic module for inputting the input text phonemes of the input text, the frequency characteristics of the reference audio, and the target clustering center into a trained acoustic model to obtain a target output audio, where the reference audio is a mediator audio corresponding to the mediator's voice.
[0012] According to another aspect of the present invention, there is provided an electronic device including a processor and a memory, where a computer program is stored in the memory, and the processor is configured to execute the computer program to implement the method as described above.
[0013] According to still another aspect of the present invention, there is provided a computer-readable storage medium storing computer programs / instructions, and when the computer programs / instructions are executed by a processor, the method as described above is implemented.
[0014] In the above solution of the present invention, by introducing voice cloning into the field of dispute mediation technology, the deficiencies of the traditional pure artificial mediator system can be overcome. Compared with the prior art, the present solution can show excellent effects in multiple aspects, especially in terms of voice diversity, interaction naturalness, work efficiency, and service quality, and can bring revolutionary improvements to the mediation industry. Specifically: (1) This solution breaks the limitation of single voice: Through voice cloning technology, this solution can select a suitable mediator voice according to needs, which can achieve customized voice tones, emotions, and expression methods. In other words, this solution can pre-set the voices of multiple mediators. When in use, the user only needs to select a suitable voice and use their own voice or text as input to output the corresponding audio of the selected voice. This flexible way of adjusting the mediator's voice enables the mediator to switch voices according to needs, thereby providing a more diversified and mediation-demand-compliant voice effect; (2) This solution reduces the usage threshold of voice cloning: Traditional voice cloning technology often requires a large number of audio samples, and the process of voice replication is relatively complex. In contrast, the present invention only requires the mediator to provide a short audio sample (such as a 1-minute recording) to quickly generate an output audio that is consistent with the mediator's voice. This simplified process greatly reduces the usage threshold of voice cloning, enabling this solution to be widely applied to different mediation scenarios; and because only a short period of audio data is required, mediators and other relevant personnel can quickly participate in the construction and application of the system, avoiding the large-scale audio data collection and complex model training required by traditional methods; (3) Improve mediation efficiency and reduce dependence on human mediators: By applying this solution to the dispute mediation process, mediators can automatically output audio through the above solution. Especially during the first call, case information can be automatically entered and introduced preliminarily through the above solution, saving the time and energy of mediators. When dealing with a large number of cases, this solution can effectively share the work pressure of human mediators, improve mediation efficiency, and the output audio will not be affected by factors such as the fatigue and mood swings of human mediators, ensuring the consistency of the voice tone during the mediation process, which helps to guarantee the mediation effect and the mediation experience of the parties.
[0015] (4) Provide efficient and consistent services, enhancing fairness and compliance: This solution can ensure the consistency of the voice style, making the mediation process more standardized and regulated. Whether mediators are involved or not, the accuracy of the mediation content and the service quality can be guaranteed through this solution, avoiding inconsistent services caused by factors such as emotional fluctuations of human mediators, thus enhancing fairness and compliance.
[0016] (5) Enhance the popularity and scalability of the mediation system: Traditional human mediator systems require a large amount of manpower and time costs to train mediators, and the quantity and quality of mediators are often restricted by factors such as region, language, and culture, resulting in a significant limitation in the popularity of services. However, through this solution, mediators can provide services with the same voice tone in multiple scenarios. For mediator / party groups with dialects, the mediation service can break through geographical restrictions and be widely applied in different cultural backgrounds and language environments, enhancing the popularity and scalability of the mediation system.
[0017] (6) Improve the user experience and enhance customer satisfaction: In traditional human mediation systems, the customer experience is often affected by factors such as the tone and mood of mediators. When communicating with human mediators, customers may feel a rigid tone, cold attitude, etc., thus affecting their trust and satisfaction with the mediation process. This solution can output the input text according to the voice tone of the mediator, making the mediator audio in the entire mediation process always gentle, kind, and natural. When communicating with the parties, it can let the parties feel a more real, warm, and considerate voice experience, thus significantly enhancing customer satisfaction and trust.
[0018] In summary, the above technical solution effectively improves the existing pure human mediation system, solves key problems in traditional human systems such as single voice tone, unnatural interaction, and low efficiency. This solution can reduce the work pressure of mediators, simplify the mediation process, improve mediation efficiency, ensure service consistency, enhance the user experience, and has good practicality and market competitiveness.
[0019] The above description is only an overview of the technical solution of the present invention. In order to understand the technical means of the present invention more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features and advantages of the present invention more obvious and understandable, the specific embodiments of the present invention are specifically exemplified below. Brief Description of the Drawings
[0020] By describing the embodiments of the present invention in more detail in conjunction with the accompanying drawings, the above and other purposes, features and advantages of the present invention will become more obvious. The drawings are used to provide a further understanding of the embodiments of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation to the present invention. In the drawings, the same reference numerals generally represent the same components or steps.
[0021] Figure 1 A schematic flowchart showing a mediator voice cloning method according to an embodiment of the present invention; Figure 2 A schematic block diagram showing a mediator voice cloning system according to an embodiment of the present invention; Figure 3 A schematic block diagram showing an electronic device according to an embodiment of the present invention. Detailed Description of the Embodiments
[0022] In order to make the purpose, technical solution and advantages of the present invention more obvious, the exemplary embodiments according to the present invention will be described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments of the present invention. It should be understood that the present invention is not limited by the exemplary embodiments described herein. Based on the embodiments of the present invention described in the present invention, all other embodiments obtained by those skilled in the art without creative efforts shall fall within the protection scope of the present invention.
[0023] As described above, human mediators are facing increasing work pressure. At the same time, given the complex diversity of cases, mediators must be able to quickly adjust their tone, intonation, and expression in various different situations, which undoubtedly increases the complexity of the mediation process and consumes a large amount of time. Moreover, in the traditional manual mediation system, mediators usually have only a fixed voice, and in scenarios where a large number of people are involved in mediation, the monotony and fatigue of the voice may affect the mediator's mood and work efficiency, thereby affecting the mediation effect. In some related technologies, a voice interaction system is used to assist mediators in the mediation work. However, although the current voice interaction system can simulate communication in some situations, it cannot accurately reproduce the personalized voice style of mediators, nor can it flexibly adjust the voice content and expression according to different needs of cases. All these affect the mediation efficiency and effect. In view of this, the present invention provides a method, system, electronic device, and storage medium for cloning the voice of a mediator. This method can accurately simulate the voice and expectations of the mediator, ensuring the consistency between the simulated voice and the real voice of the mediator. Applying this method to the dispute mediation process helps to improve the mediation efficiency and effect. This method, system, electronic device, and storage medium will be described in detail below.
[0024] According to one aspect of an embodiment of the present invention, a method for cloning the voice of a mediator is provided. Figure 1 A schematic flowchart showing a method for cloning the voice of a mediator according to an embodiment of the present invention is shown. As Figure 1 shown, the method may include the following steps S110, step S120, step S130, step S140, and step S150.
[0025] In step S110, in response to a selection instruction input by a user, a mediator voice is selected.
[0026] In the solution of this example, the system for implementing the method for cloning the voice of a mediator may have a display interface. At least one mediator voice may be displayed on the display interface. The user may select the desired mediator voice according to needs. The selection instruction may be a click instruction, a text instruction, or a voice instruction. For example, multiple mediator voices may be displayed on the display interface, and the user may click on the desired mediator voice with a mouse to select the mediator voice. Again, for example, the user may select the mediator voice by means of text or voice input. Specifically, for example, the user may input "Mediator 1" by text or voice, so as to select "Mediator 1" as the desired mediator voice among multiple mediator voices.
[0027] In step S120, input text is obtained.
[0028] In the solution of this example, the user can input what they want to say during the mediation process in the form of text or voice. In some embodiments, obtaining the input text includes: receiving the input text input by the user. In the solution of this embodiment, the user can directly input text through a text input device such as a keyboard or a graphics tablet, or search for a suitable reply statement in the statement library and paste it into the input box on the display interface. In other embodiments, obtaining the input text includes: receiving the input audio input by the user; converting the input audio into text to obtain the input text. In the solution of this embodiment, the user can input voice into the system through a voice input device such as a microphone. After receiving the input audio, the system can convert the input voice into the input text through a speech-to-text model. This method can facilitate the user's input and improve the mediation efficiency.
[0029] In step S130, use an autoregressive model to predict the target audio feature vector of the input text.
[0030] After obtaining the input text, an autoregressive (AR) model can be used to predict information such as the timbre, tone, and speech rate that the input text may correspond to, that is, to predict the target audio feature vector of the input text. This information can provide basic data for the generation of the target output audio in subsequent steps, making the target output audio more natural.
[0031] In step S140, search for the target cluster center in the audio dictionary that matches the target audio feature vector, where the audio dictionary includes multiple clusters, and each cluster includes multiple audio feature vectors.
[0032] Optionally, searching for the target cluster center in the audio dictionary that matches the target audio feature vector may include: calculating the distance between the target audio feature vector and the cluster center of each cluster in the audio dictionary through a nearest neighbor search algorithm; determining the cluster center with the smallest distance from the target audio feature vector as the target cluster center.
[0033] As described above, the audio dictionary includes multiple clusters. Each cluster can be regarded as a set composed of similar audio feature vectors. The audio feature vector can reflect information such as timbre, tone, and speech rate, and each cluster can represent a specific audio feature pattern. By searching for the target cluster center in the audio dictionary that matches the target audio feature vector, the audio feature pattern closest to the target audio feature vector can be found among the multiple audio feature patterns in the audio dictionary, and the most representative audio feature vector (i.e., the target cluster center) in this audio feature pattern can be selected as the basic data for generating the target output audio in subsequent steps, which helps to generate the target output audio more accurately.
[0034] In step S150, the input text phonemes of the input text, the frequency features of the reference audio, and the target clustering center are input into the trained acoustic model to obtain the target output audio, where the reference audio is the mediator audio corresponding to the mediator's voice.
[0035] Optionally, the acoustic model (which can be called the SoVits model) used in this article can be any existing or future-developed acoustic model, and the present invention does not limit this.
[0036] It can be understood that the reference audio is the audio pre-input by the mediator (which can be called the target mediator) corresponding to the required mediator's voice. In some embodiments, the target mediator can pre-record a standard audio of a preset duration as the reference audio. In other embodiments, the recordings of the target mediator during the mediation process can be collected, and an audio segment of a preset duration emitted by the target mediator can be selected from the recordings. The preset duration can be selected according to actual needs. For example, it can be 1 minute. Since this reference audio is the audio of the target mediator during actual mediation work, the target output audio obtained based on this reference audio can be closer to the voice of the target mediator during actual work, which helps to improve the mediation experience of the parties and thus helps to improve the mediation effect.
[0037] In this article, the frequency features of the reference audio can be the Mel spectrogram features of the reference audio. In one embodiment, the reference audio can be input into the Hubert model to obtain the Hubert representation of the reference audio (i.e., the reference audio feature vector of the reference audio). Then, the Hubert representation of the reference audio is converted into the Mel spectrogram features of the reference audio. The Mel spectrogram features can effectively represent the frequency features of the audio and thus can be used as the basis for generating the target output audio.
[0038] In this example, the inventor considered outputting the words to be expressed by the current user in the voice of the mediator through voice cloning. Voice cloning technology, also known as speech synthesis technology, is a voice simulation technology based on deep learning that can accurately replicate the voice of a specific person and has been applied in multiple fields, such as virtual assistants, customer service, entertainment, education, medical care, and other industries. For example, virtual assistant products like Xiling Digital Human provide users with a highly realistic dialogue experience through voice cloning technology, supporting natural and personalized voice interaction. In the entertainment field, voice cloning technology can be used for the voice restoration of deceased actors or singers, or to provide customized learning voices for students. In the medical and special education fields, speech therapists use this technology to help patients practice speaking, or to provide assistive communication tools for students with speech disorders. These technologies mainly rely on deep learning models. Through the training of a large amount of speech data, they can synthesize speech content close to reality and even reproduce the speech characteristics of specific people, including pitch, pronunciation style, etc.
[0039] Although significant progress has been made in these fields with voice cloning technology, its application in legal and social affairs, especially in the field of virtual mediators, is blank. Mediators play important roles in resolving disputes, providing legal advice, generating mediation reports, etc. Applying voice cloning technology to the field of dispute mediation can use virtual mediators to simulate real mediators to conduct mediation work, so as to improve the mediation efficiency and effect.
[0040] However, through research, the inventors found that the following problems need to be solved when applying voice cloning technology to the field of dispute mediation: 1. Limitations of speech synthesis technology: Current voice cloning technology is mostly applied to commercial voice assistants and the entertainment field, and its application in the professional scenarios of mediators is still blank. During the mediation process, the emotional color and tone changes of the voice are crucial to the mediation effect, and the existing technology is difficult to achieve personalized customization for complex mediation scenarios and cannot accurately capture the voice details and emotional fluctuations of mediators, which may affect the mediation experience of the parties and thus the mediation effect; 2. Poor personalization and multi-scenario adaptability: Most existing speech synthesis technologies use general voice tones and speech models and cannot flexibly adjust the voice style according to the characteristics of different mediation cases, resulting in the inability to meet the diverse mediation needs in actual use; 3. Limitations of mediator participation: Existing mediation work usually relies on the real-time participation of mediators, which not only consumes a large amount of time and energy, but also when mediators are unable to participate, the progress of case handling will be delayed, affecting the mediation efficiency; 4. Technical limitations: The current mediation model is usually "one mediator for one case". When switching between voice cloning and real mediators, the difference in voice may lead to inconsistent experiences and effects during the case handling process.
[0041] In the above solution of the present invention, by introducing voice cloning into the field of dispute mediation technology, the deficiencies of the traditional pure manual mediator system can be overcome. Compared with the existing technology, this solution can show excellent effects in many aspects, especially in terms of voice diversity, interaction naturalness, work efficiency and service quality, and can bring revolutionary improvements to the mediation industry. Specifically: (1) This solution breaks the limitation of a single timbre: Through timbre cloning technology, this solution can select a suitable mediator's timbre according to requirements, which can achieve customized timbre tones, emotions, and expression methods. In other words, this solution can preset the timbres of multiple mediators. When using it, users only need to select a suitable timbre and use their own voice or text as input to output the corresponding audio of the selected timbre. This flexible way of adjusting the mediator's timbre enables the mediator to switch timbres according to needs, thus providing a more diversified and mediation-demand-compliant sound effect; (2) This solution reduces the usage threshold of timbre cloning: Traditional timbre cloning technology often requires a large number of audio samples, and the process of timbre replication is relatively complex. In contrast, this invention only requires the mediator to provide a short audio sample (such as a 1-minute recording) to quickly generate output audio consistent with the mediator's timbre. This simplified process greatly reduces the usage threshold of timbre cloning, enabling this solution to be widely applied to different mediation scenarios; and because only a short period of audio data is required, mediators and other relevant personnel can quickly participate in the construction and application of the system, avoiding the large-scale audio data collection and complex model training required by traditional methods; (3) Improve mediation efficiency and reduce dependence on human mediators: By applying this solution to the dispute mediation process, mediators can automatically output audio through the above solution. Especially during the first call, the case information entry and preliminary introduction can be automatically completed through the above solution, saving the time and energy of mediators. When dealing with a large number of cases, this solution can effectively share the work pressure of human mediators, improve mediation efficiency, and the output audio will not be affected by factors such as the fatigue and emotional fluctuations of human mediators, which can ensure the timbre consistency during the mediation process, helping to guarantee the mediation effect and the mediation experience of the parties.
[0042] (4) Provide efficient and consistent services, enhancing fairness and compliance: This solution can ensure the consistency of the voice style, making the mediation process more standardized and regularized. Whether mediators are involved or not, through this solution, the accuracy of mediation content and service quality can be guaranteed, avoiding service inconsistencies caused by factors such as emotional fluctuations of human mediators, thus enhancing fairness and compliance.
[0043] (5) Enhance the popularity and scalability of the mediation system: The traditional manual mediator system requires a large amount of human and time costs to train mediators, and the quantity and quality of mediators are often restricted by factors such as region, language, and culture, resulting in a significant limitation in the popularity of the service. However, through this solution, mediators can provide services with the same voice tone in multiple scenarios. For mediator / client groups with dialects, the mediation service can break through geographical restrictions and be widely applied in different cultural backgrounds and language environments, enhancing the popularity and scalability of the mediation system.
[0044] (6) Improve the user experience and enhance customer satisfaction: In the traditional manual mediation system, the customer experience is often affected by factors such as the mediator's tone and mood. When communicating with a manual mediator, customers may feel a harsh tone, cold attitude, etc., which affects their trust and satisfaction with the mediation process. Through this solution, by outputting the input text according to the mediator's voice tone, the mediator's audio throughout the mediation process can always remain gentle, kind, and natural. When communicating with the parties, it can allow the parties to feel a more real, warm, and considerate voice experience, thus significantly enhancing customer satisfaction and trust.
[0045] In summary, the above technical solution effectively improves the existing pure manual mediation system, solves the key problems in traditional manual systems such as single voice tone, unnatural interaction, and low efficiency. This solution can reduce the work pressure of mediators, simplify the mediation process, improve the mediation efficiency, ensure service consistency, enhance the user experience, and has good practicality and market competitiveness.
[0046] Exemplarily, the trained acoustic model is obtained through the following steps: determining the reference audio feature vector of the reference audio; searching for the reference clustering center in the audio dictionary that matches the target audio feature vector; inputting the reference text phonemes of the reference text, the frequency features of the reference audio, and the reference clustering center into the initial acoustic model to obtain the reference output audio, where the reference text is the text form of the reference audio; optimizing the initial acoustic model based on the difference between the reference audio and the reference output audio to obtain the trained acoustic model.
[0047] The method of searching for the reference clustering center in the audio dictionary that matches the target audio feature vector is similar to that of searching for the target clustering center in the audio dictionary, so it will not be elaborated.
[0048] Optionally, the number of trained acoustic models corresponds one-to-one with the mediator's voice timbre. Step S150 may include: inputting the input text phonemes, the frequency features of the reference audio, and the target clustering center into the trained acoustic model corresponding to the mediator's voice timbre to obtain the target output audio. In this embodiment, acoustic models corresponding one-to-one with different mediator voice timbres can be pre-trained. When the target output audio needs to be output, the trained acoustic model corresponding to the mediator's voice timbre can be used, which helps to further ensure that the generated audio is close to the real voice of the mediator.
[0049] In the solution of this example, the mediator audio corresponding to the mediator's voice timbre can be used to train the acoustic model, which helps to make the audio output by the acoustic model closer to the real voice of the mediator, thus helping to ensure the consistency of voice timbre during the mediation process. And only one piece of mediator audio is needed for this training process to obtain the SoVits model corresponding to the mediator's voice timbre, without occupying too much time of the mediator, which helps to improve the usage experience of the mediator.
[0050] Exemplarily, determining the reference audio feature vector of the reference audio includes: processing the reference audio with the Hubert model to obtain the reference audio feature vector. The Hubert model can convert the audio into a series of feature vectors representing the audio content, and this vector can fully reflect information such as the voice timbre, tone, and speech rate of the audio. Based on this, this solution can accurately extract information such as the voice timbre, tone, and speech rate of the reference audio, so as to provide an accurate benchmark for the generation of the target output audio in subsequent steps and ensure that the generated audio is close to the real voice of the mediator.
[0051] Exemplarily, using the autoregressive model to predict the target audio feature vector of the input text includes: inputting the input text feature vector of the input text, the input text phonemes, the reference text phonemes of the reference text, the reference text feature vector of the reference text, and the reference audio feature vector of the reference text into the autoregressive model to obtain the target audio feature vector.
[0052] The input text feature vector can be represented as text_bert, the input text phoneme can be represented as text_seq, the reference text phoneme can be represented as ref_seq, the reference text feature vector can be represented as ref_bert, and the reference audio feature vector can be represented as ref_ssl. In the solution of this example, text_bert, text_seq, ref_seq, ref_bert, and ref_ssl can be input into the AR model, so as to use the AR model to predict the target audio feature vector corresponding to the input text. This target feature vector can be used to obtain feature points (i.e., target clustering centers) from the audio dictionary (CodeBook), which helps to accurately clone the tone, intonation, and emotion of the mediator.
[0053] The above solution can infer the target feature vector more accurately through an autoregressive manner, which can provide an accurate basis for generating the target output audio in subsequent steps.
[0054] Exemplarily, before using the autoregressive model to predict the target audio feature vector of the input text, the method further includes: performing phoneme conversion on the input text to obtain the input text phoneme; and / or, inputting the input text into the BERT model to obtain the input text feature vector.
[0055] Optionally, before using the autoregressive model to predict the target audio feature vector of the input text, the method further includes: performing phoneme conversion on the input text to obtain the input text phoneme. In the solution of this example, an existing or future-developed phoneme conversion model can be used to perform factor conversion on the input text, and the present invention does not limit this. By converting the input text into a phoneme sequence (i.e., the input text phoneme), this solution can meet the model input requirements in subsequent steps and provide basic information for the generation of the target output audio.
[0056] Similarly, the reference text phoneme can also be obtained by using an existing or future-developed phoneme conversion model, which will not be elaborated.
[0057] Optionally, before using the autoregressive model to predict the target audio feature vector of the input text, the method further includes: inputting the input text into the BERT model to obtain the input text feature vector. The BERT model can extract the context information of the input text, which can provide a relatively accurate basis for the generation of the target output audio in subsequent steps.
[0058] Similarly, the reference text feature vector can also be obtained by using the BERT model, which will not be elaborated.
[0059] Exemplarily, the audio dictionary is obtained in the following manner: determining the feature vectors of each training audio in the training audio set; using a clustering algorithm to cluster the feature vectors of each training audio in the training audio set to obtain a plurality of clustering clusters; and constructing an audio dictionary based on the plurality of clustering clusters.
[0060] In the solution of this example, first, a large-scale training audio set (including multi-language audios such as Chinese, English, Japanese, etc.) can be processed to extract the ssl features (i.e., the feature vectors of the audios) of each audio in the training audio set. These ssl features can be classified by the KMeans clustering algorithm to generate clustering centers and form a CodeBook, where each clustering center corresponds to a vector. This audio dictionary can provide an accurate audio feature representation for the generation of the target output audio in the subsequent process.
[0061] According to another aspect of the embodiments of the present invention, a mediator voice cloning system is provided. Figure 2 A schematic block diagram showing a mediator voice cloning system according to an embodiment of the present invention. As Figure 2 shown, the system 200 may include a selection module 210, an acquisition module 220, an autoregressive module 230, a vector quantization module 240, and an acoustic module 250.
[0062] Among them, the selection module 210 is configured to select the mediator voice in response to a selection instruction input by the user; the acquisition module 220 is configured to acquire the input text; the autoregressive module 230 is configured to predict the target audio feature vector of the input text using an autoregressive model; the vector quantization module 240 is configured to find a target clustering center in the audio dictionary that matches the target audio feature vector, where the audio dictionary includes a plurality of clustering clusters, and each clustering cluster includes a plurality of audio feature vectors; the acoustic module 250 is configured to input the input text phonemes of the input text, the frequency features of the reference audio, and the target clustering center into a trained acoustic model to obtain the target output audio, where the reference audio is the mediator audio corresponding to the mediator voice.
[0063] This system can accurately reproduce the voice of the desired mediator and can make adaptive adjustments in terms of emotion and tone. The user only needs to input audio or text to obtain the target output audio, and the operation is simple, which can better assist the mediation work of the mediator.
[0064] According to still another aspect of the embodiments of the present invention, an electronic device is further provided. Figure 3 A schematic block diagram showing an electronic device according to an embodiment of the present invention. As Figure 3 shown, the electronic device 300 includes: a processor 310 and a memory 320. A computer program is stored in the memory 320, and the processor 310 is configured to execute the computer program to implement the above method.
[0065] According to another aspect of the embodiments of the present invention, a computer-readable storage medium is further provided. A computer program / instruction is stored in the storage medium, and when the computer program / instruction is executed by a processor, the above-mentioned method is implemented. The storage medium may include, for example, a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a portable compact disc read-only memory (CD-ROM), a USB memory, or any combination of the above storage media. The computer-readable storage medium may be any combination of one or more computer-readable storage media.
[0066] Those of ordinary skill in the art can easily understand the implementation structure, working principle, and beneficial effects of the system, electronic device, and computer-readable storage medium by reading the above method. For the sake of brevity, they will not be elaborated here.
[0067] Although example embodiments have been described herein with reference to the accompanying drawings, it should be understood that the above example embodiments are merely exemplary and are not intended to limit the scope of the present invention thereto. Those of ordinary skill in the art can make various changes and modifications therein without departing from the scope and spirit of the present invention. All such changes and modifications are intended to be included within the scope of the present invention as claimed in the appended claims.
[0068] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or in a combination of computer software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. A professional technician can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.
[0069] In several embodiments provided by the present invention, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed.
[0070] In the specification provided herein, a large number of specific details are set forth. However, it can be understood that the embodiments of the present invention can be practiced without these specific details. In some instances, well-known methods, structures, and technologies have not been shown in detail so as not to obscure the understanding of this specification.
[0071] Similarly, it should be understood that, for the purpose of streamlining the present invention and assisting in understanding one or more of the various aspects of the invention, in the description of the exemplary embodiments of the present invention, the various features of the present invention are sometimes grouped together into a single embodiment, figure, or description thereof. However, the method of the present invention should not be construed as reflecting an intention that the claimed invention requires more features than those expressly recited in each claim. Rather, as reflected in the corresponding claims, the inventive point lies in that the corresponding technical problem can be solved by features less than all the features of a single disclosed embodiment. Therefore, the claims following the specific implementation are hereby expressly incorporated into the specific implementation, where each claim itself serves as a separate embodiment of the present invention.
[0072] Those skilled in the art will appreciate that, except where features are mutually exclusive, any combination can be used to combine all the features disclosed in this specification (including the accompanying claims, abstract, and drawings), as well as all the processes or units of any method or device so disclosed. Unless otherwise expressly stated, each feature disclosed in this specification (including the accompanying claims, abstract, and drawings) can be replaced by an alternative feature that serves the same, equivalent, or similar purpose.
[0073] In addition, those skilled in the art will be able to understand that, although some of the embodiments described herein include certain features included in other embodiments rather than other features, the combination of features of different embodiments means that it is within the scope of the present invention and forms different embodiments. For example, in the claims, any one of the claimed embodiments can be used in any combination.
[0074] The various component embodiments of the present invention can be implemented in hardware, or in software modules running on one or more processors, or in a combination thereof. Those skilled in the art should understand that a microprocessor or a digital signal processor (DSP) can be used in practice to implement some or all of the functions of some of the modules in the system and electronic devices according to the embodiments of the present invention. The present invention can also be implemented as a device program (e.g., a computer program and a computer program product) for performing part or all of the methods described herein. Such a program for implementing the present invention can be stored on a computer-readable medium, or can be in the form of one or more signals. Such signals can be downloaded from an Internet website, or provided on a carrier signal, or provided in any other form.
[0075] It should be noted that the above embodiments are illustrative of the present invention rather than restrictive thereof, and those skilled in the art can design alternative embodiments without departing from the scope of the appended claims. In the claims, any reference signs placed between parentheses shall not be construed as limiting the claim. The word "comprising" does not exclude the presence of elements or steps not listed in a claim. The word "a" or "an" preceding an element does not exclude the presence of a plurality of such elements. The present invention can be implemented by means of hardware including several different elements and by means of a suitably programmed computer. In a unit claim listing several devices, several of these devices may be embodied by the same item of hardware. The use of the words first, second, and third, etc. does not denote any order. These words may be interpreted as names.
[0076] As described above, the above is only the specific implementation manner or the description of the specific implementation manner of the present invention, and the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should be covered within the protection scope of the present invention. The protection scope of the present invention shall be subject to the protection scope of the claims.
Claims
1. A method for cloning the tone of a mediator, characterized in that: include: In response to a selection instruction input by a user, selecting a mediator timbre; Get input text; Predicting a target audio feature vector of the input text using an autoregressive model; Searching for a target cluster center matching the target audio feature vector in an audio dictionary, wherein the audio dictionary includes a plurality of clusters, and each cluster includes a plurality of audio feature vectors; Input text phonemes of the input text, frequency features of reference audio, and the target cluster center are input into a trained acoustic model to obtain a target output audio, wherein the reference audio is the mediator audio corresponding to the mediator timbre.
2. The method according to claim 1, characterized in that The trained acoustic model is trained in the following manner: Determining a reference audio feature vector of the reference audio; Searching the audio dictionary for a reference cluster center that matches the target audio feature vector; Inputting reference text phonemes of a reference text, frequency features of the reference audio, and the reference cluster center into an initial acoustic model to obtain a reference output audio, wherein the reference text is a text form of the reference audio; Based on the difference between the reference audio and the reference output audio, the initial acoustic model is optimized to obtain the trained acoustic model.
3. The method according to claim 2, characterized in that The determining a reference audio feature vector of the reference audio includes: The reference audio is processed using a Hubert model to obtain the reference audio feature vector.
4. The method according to claim 1, characterized in that: The method of predicting the target audio feature vector of the input text by using an autoregressive model includes: The input text feature vector of the input text, the input text phonemes, the reference text phonemes of the reference text, the reference text feature vector of the reference text, and the reference audio feature vector of the reference text are input into the autoregressive model to obtain the target audio feature vector.
5. The method according to claim 4, characterized in that Before predicting the target audio feature vector of the input text using the autoregressive model, the method further includes: Performing phoneme conversion on the input text to obtain the input text phonemes; and / or, The input text is input into the BERT model to obtain the input text feature vector.
6. The method according to any one of claims 1 to 5, characterized in that: The audio dictionary is obtained in the following way: Determine a feature vector for each training audio in the training audio set; Clustering the feature vector of each training audio in the training audio set using a clustering algorithm to obtain a plurality of clusters; The audio dictionary is constructed based on the multiple clusters.
7. The method according to any one of claims 1 to 5, characterized in that: The obtaining of input text comprises: receiving the input text input by the user; or, Receive input audio input from the user; The input audio is converted into text to obtain the input text.
8. A mediator voice cloning system, characterized in that: include: A selection module, for selecting a mediator timbre in response to a selection instruction input by a user; The acquisition module is used to obtain the input text; An autoregressive module, used for predicting a target audio feature vector of the input text using an autoregressive model; A vector quantization module, used for searching a target cluster center matching the target audio feature vector in an audio dictionary, wherein the audio dictionary includes a plurality of clusters, and each cluster includes a plurality of audio feature vectors; An acoustic module is used to input the input text phonemes of the input text, the frequency characteristics of the reference audio, and the target cluster center into a trained acoustic model to obtain a target output audio, wherein the reference audio is the mediator audio corresponding to the mediator timbre.
9. An electronic device, characterized in that: The method comprises a processor and a memory, wherein the memory stores a computer program, and the processor is used to execute the computer program to implement the method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: A computer program / instruction is stored, and when the computer program / instruction is executed by a processor, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Cited By
Method and device for selecting reference audio for tone cloning
CN120977285A