Method for generating voice packet, voice broadcasting method and corresponding device

By generating augmented speech and utilizing distillation training technology, the problems of high data collection threshold and low efficiency in traditional personalized voice package generation are solved, and low-threshold, fast voice package generation and real-time synthesis on the user side are achieved.

CN120612919APending Publication Date: 2025-09-09BEIJING AMAP YUNXIN TECHNOLOGY CO LTD
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
CN202511101232.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-07
Publication Date
2025-09-09

AI Technical Summary

Technical Problem

Traditional personalized voice package generation solutions have the problems of high threshold for voice data collection and low generation efficiency.

Method used

By obtaining at least one speech of the target speaker, multiple augmented speech sounds are generated based on acoustic features, and a second speech synthesis model is trained on the first speech synthesis model using distillation training to generate a lightweight speech synthesis model to lower the data collection threshold and improve generation efficiency.

Benefits of technology

It achieves low-threshold voice data collection and fast voice packet generation, reduces model training time and computing resource requirements, and supports real-time speech synthesis on the user side.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120612919A_ABST
    Figure CN120612919A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a method for generating a voice packet, a voice broadcasting method and a corresponding device. According to the main technical scheme, the method comprises the steps of obtaining at least one voice of a target speaker; generating a plurality of augmented voices based on the acoustic features of the at least one voice; using the plurality of augmented voices to perform distillation training on a second voice synthesis model on the basis of the first voice synthesis model to obtain a second voice synthesis model after distillation training, the parameter scale of the second voice synthesis model being smaller than the parameter scale of the first voice synthesis model; and determining a voice packet corresponding to the target speaker based on the second voice synthesis model after distillation training. According to the method and the device, the voice packet can be generated by only needing a small number of target speaker voices, even one voice, so that the acquisition threshold of the voice data is greatly reduced; and on the basis of ensuring the model effect, the generation efficiency of the voice packet is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of voice technology, and in particular to a method for generating a voice packet, a voice broadcasting method, and a corresponding device. Background Art

[0002] With the popularization of smart devices and the continuous improvement of user needs, users are no longer satisfied with the same default voice, but hope to use voices that they are familiar with or like for voice broadcasting. Therefore, personalized voice broadcasting has become an important function to enhance user experience.

[0003] To achieve personalized voice broadcasting, users need to download a personalized voice package. However, traditional personalized voice package generation solutions generally have two major pain points: high voice data collection threshold and low voice package generation efficiency. Summary of the Invention

[0004] In view of this, the present application provides a method for generating a voice package, a voice broadcast method and a corresponding device to solve the above pain points.

[0005] This application provides the following solutions: According to a first aspect, a method for generating a voice packet is provided, the method comprising: Obtain at least one speech of the target speaker; generating a plurality of augmented speech messages based on the acoustic features of the at least one speech message; Using the plurality of augmented speech sounds, performing distillation training on a second speech synthesis model based on the first speech synthesis model to obtain a second speech synthesis model after distillation training, wherein a parameter scale of the second speech synthesis model is smaller than a parameter scale of the first speech synthesis model; Based on the second speech synthesis model after the distillation training, a speech package corresponding to the target speaker is determined.

[0006] According to a second aspect, a voice broadcast method is provided, which is applied to a client, and the method includes: Obtaining a speech package corresponding to a target speaker, the speech package including at least a second speech synthesis model, the second speech synthesis model being obtained through distillation training using at least one speech of the target speaker; In response to an event triggering a voice broadcast, determining a target broadcast voice using the voice packet; Play the target announcement voice.

[0007] According to a third aspect, a device for generating a voice packet is provided, the device comprising: a speech acquisition unit configured to acquire at least one speech of a target speaker; a data augmentation unit, configured to generate a plurality of augmented speech messages based on the acoustic features of the at least one speech message; a distillation training unit configured to perform distillation training on a second speech synthesis model based on the first speech synthesis model using the plurality of augmented speech pieces, to obtain a second speech synthesis model after distillation training, wherein a parameter scale of the second speech synthesis model is smaller than a parameter scale of the first speech synthesis model; The speech packet generation unit is configured to determine the speech packet corresponding to the target speaker based on the second speech synthesis model trained by the distillation.

[0008] According to a fourth aspect, a voice broadcasting device is provided, the device comprising: a speech packet acquisition unit configured to acquire a speech packet corresponding to a target speaker, wherein the speech packet includes at least a second speech synthesis model, wherein the second speech synthesis model is obtained through distillation training using at least one speech of the target speaker; A voice acquisition unit is configured to determine a target broadcast voice using the voice packet in response to an event triggering the voice broadcast; The voice broadcast unit is configured to play the target broadcast voice.

[0009] According to a fifth aspect, a computer program product is provided, comprising a computer program, which implements the steps of the method described in any one of the first aspects above when executed by a processor.

[0010] According to the specific embodiments provided in this application, this application discloses the following technical effects: Through this application, only one or more speech of the target speaker is required, and augmentation is performed based on the acoustic features of the speech to obtain a larger number of augmented speech to ensure the training effect of the model, which greatly reduces the threshold for collecting speech data; in addition, this application adopts the distillation training method to obtain the second speech synthesis model, and then generate a speech package. On the one hand, it reduces the amount of speech data required for model training, and on the other hand, it shortens the speed of model training while ensuring the model effect, thereby improving the efficiency of speech package generation.

[0011] Of course, any product implementing the present application does not necessarily need to achieve all of the advantages described above at the same time. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0013] Figure 1 This is a system architecture diagram applicable to the embodiments of the present application.

[0014] Figure 2 A flow chart of a method for generating a voice packet provided in an embodiment of the present application.

[0015] Figure 3 A schematic structural diagram of a second speech synthesis model provided in an embodiment of the present application.

[0016] Figure 4 A schematic diagram of the principle of distillation training provided in an embodiment of the present application.

[0017] Figure 5 This is a flowchart of the voice broadcast method provided in an embodiment of the present application.

[0018] Figure 6 A schematic block diagram of the apparatus for generating a voice packet provided in an embodiment of the present application.

[0019] Figure 7 A schematic block diagram of a voice broadcasting device provided in an embodiment of the present application.

[0020] Figure 8 A schematic block diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0021] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field are within the scope of protection of this application.

[0022] The terms used in the embodiments of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. The singular forms "a", "an", "the" and "the" used in the embodiments of the present invention and the appended claims are also intended to include plural forms unless the context clearly indicates otherwise.

[0023] It should be understood that the term "and / or" as used herein is merely a description of the relationship between associated objects, indicating that three possible relationships exist. For example, "A and / or B" can represent: A exists alone, A and B exist simultaneously, or B exists alone. Furthermore, the character " / " in this document generally indicates that the associated objects are in an "or" relationship.

[0024] The word "if," as used herein, may be interpreted as "at the time of" or "when" or "in response to determining" or "in response to detecting," depending on the context. Similarly, the phrases "if it is determined" or "if (stated condition or event) is detected" may be interpreted as "when it is determined" or "in response to the determination" or "when detecting (stated condition or event)" or "in response to detecting (stated condition or event)," depending on the context.

[0025] Traditional personalized voice package generation mainly includes the following two methods: The first method involves pre-recording all or most of the sentences likely to appear in a specific broadcast scenario. These sentences are then packaged into a voice package corresponding to the speaker and made available for download by the user. During the actual broadcast process, the text to be broadcast is matched with the sentences in the voice package, and the matched sentences are used for broadcasting.

[0026] The second method involves pre-collecting some of the speaker's sentences as training data. This training data is then used to train a customized speech synthesis model for that speaker. The trained speech synthesis model is then packaged as a speech package specific to that speaker for users to download. However, this method also places high demands on the quantity and quality of the speaker's sentences collected, requiring the speaker to provide tens of minutes or even hours of high-quality speech. Furthermore, training the speech synthesis model takes a long time, leading to lengthy wait times after users submit their recorded speech.

[0027] In view of this, the present application provides a new approach. To facilitate understanding of the present application, the system architecture on which the present application is based is first described. Figure 1 An exemplary system architecture to which the embodiments of the present application can be applied is shown. Figure 1 As shown in , the system architecture may include: a user end and a server end.

[0028] The server and client are the two main components of an application service. The server uses the server as the primary hardware infrastructure and can include one or more software-based service modules. The server and client form a collaborative front-end and back-end.

[0029] The user terminal can be set in the terminal device. The user terminal involved in the embodiment of the present application can be a local application, a small program running on the terminal device, or a web application running through a browser.

[0030] Terminal devices may include, but are not limited to, smart mobile terminals, wearable devices, PCs (Personal Computers), and smart home devices. Smart mobile devices include mobile phones, tablets, laptops, PDAs (Personal Digital Assistants), and internet-connected car terminals. Wearable devices include smart watches, smart glasses, smart bracelets, VR (Virtual Reality) devices, AR (Augmented Reality) devices, and mixed reality devices (i.e., devices that support both VR and AR). Smart home devices include smart TVs, smart refrigerators with voice announcements, and smart speakers.

[0031] A server can be a single server, a server cluster consisting of multiple servers, or a cloud server. A cloud server, also known as a cloud computing server or cloud host, is a hosting product within the cloud computing service ecosystem. It addresses the management difficulties and limited scalability of traditional physical hosting and virtual private server (VPS) services.

[0032] In one application scenario, a user sends the target speaker's voice to a server via a user terminal on a terminal device. The server generates a voice packet corresponding to the target speaker based on the target speaker's voice using the methods provided in the embodiments of the present application and sends the voice packet to the user terminal. The user terminal stores the voice packet locally and uses the locally stored voice packet for voice broadcasting.

[0033] It should be understood that Figure 1 The number of the client terminals and the server terminals in the embodiment is only for illustration. According to the implementation requirements, there can be any number of the client terminals and the server terminals.

[0034] Figure 2 A flow chart of a method for generating a voice packet provided in an embodiment of the present application, which can be performed by Figure 1 The server side execution in the system shown. Figure 2 As shown in , the method may include the following steps: Step 201: Obtain at least one speech of a target speaker.

[0035] Step 203: Generate multiple augmented speech messages based on the acoustic features of at least one speech message.

[0036] Step 205: Using multiple augmented speech, distillation training is performed on a second speech synthesis model based on the first speech synthesis model to obtain a second speech synthesis model after distillation training. The parameter scale of the second speech synthesis model is smaller than the parameter scale of the first speech synthesis model.

[0037] Step 207: Determine the speech package corresponding to the target speaker based on the second speech synthesis model trained by distillation.

[0038] It can be seen from the above process that this application only requires one or more speech of the target speaker, and augments the speech based on the acoustic features of the speech to obtain a larger number of augmented speech to ensure the training effect of the model, which greatly reduces the threshold for collecting speech data; in addition, this application adopts the distillation training method to obtain the second speech synthesis model, and then generates a speech package. On the one hand, it reduces the amount of speech data required for model training, and on the other hand, it shortens the speed of model training while ensuring the model effect, thereby improving the efficiency of speech package generation.

[0039] The following describes in detail each step in the above process and the effects it can produce, in conjunction with the examples. It should be noted that the terms "first" and "second" in this disclosure do not restrict size, order, or quantity, but are merely used to distinguish between them. For example, "first speech synthesis model" and "second speech synthesis model" are used to distinguish between two speech synthesis models.

[0040] First, the above step 201, namely "obtaining at least one speech of the target speaker", is described in detail with reference to the embodiment.

[0041] The purpose of this embodiment of the application is to generate a voice package with the timbre characteristics of a specific "person," referred to as a target speaker. The target speaker can be any user using the client, such as the user themselves; a natural person related to the user using the client, such as the user's family or friends; or any natural person unrelated to the user, such as a public figure the user likes, or professionals such as announcers and voice actors; or even a virtual person, such as an anime character or a digital human.

[0042] The at least one speech of the target speaker may be uploaded to the server by the user using the user terminal, may be stored locally on the server, or may be obtained from other servers or service platforms, and so on.

[0043] The above step 203, namely "generating multiple augmented speech messages based on the acoustic features of at least one speech message", is described in detail below with reference to an embodiment.

[0044] Before generating augmented speech, the at least one speech of the target speaker can be preprocessed, such as by performing content review and noise reduction, to ensure speech quality. Furthermore, speech recognition can be performed on the at least one speech of the target speaker to obtain the corresponding text, and the "speech-text" pair can be used as subsequent training data.

[0045] Since in the embodiment of the present application, the number of speech of the target speaker obtained is relatively small, or even just one speech. When a small number of speech is used to train the speech synthesis model, the training effect will be very poor. Therefore, in order to improve the training effect of the speech synthesis model, it is necessary to increase the amount of training data. The purpose of this step is to generate a larger number of speech based on the acoustic features of at least one speech of the target speaker obtained. Since the user intends to increase the number of speech used to train the speech synthesis model (the second speech synthesis model involved in the following embodiments), the generated larger number of speech is called "augmented speech".

[0046] The augmented speech needs to have the same timbre characteristics as the target speaker's speech. In this embodiment of the application, a large model with speech synthesis capabilities can be used to generate augmented speech. When the large model with speech synthesis capabilities (hereinafter referred to as the speech synthesis large model) is input with reference audio (Reference Audio), reference text (ReferenceText, i.e., the text content corresponding to the reference audio), and target text (Target Text, i.e., the text for which speech synthesis is required), the speech synthesis large model can achieve speech cloning and generate speech with the same timbre characteristics as the reference audio, but with the content of the target text, i.e., the target audio. In other words, the timbre characteristics of the reference audio are learned and transferred to the target text.

[0047] In an embodiment of the present application, the above-mentioned large speech synthesis model can be used to generate multiple augmented speech based on the acoustic features of at least one speech of the target speaker. As one of the feasible ways, multiple texts can be first obtained as augmented texts; then the large speech synthesis model can be used to generate speech corresponding to the multiple augmented texts based on the acoustic features of at least one speech as augmented speech. Specifically, the large speech synthesis model uses the speech of the target speaker as the reference speech, the text corresponding to the speech of the target speaker as the reference text, and the augmented text as the target text to generate the target speech, and uses the generated target speech as the augmented speech. This method of performing speech cloning based on augmented text to obtain augmented speech can obtain more speech efficiently and at low cost, thereby achieving the augmentation of training data and providing a basis for the subsequent training of a high-performance second speech synthesis model.

[0048] Since the generated voice packages are usually applied to specific services in specific fields, such as navigation services in the map field, when obtaining augmented text, some high-frequency texts in the specific services in the field where the voice package is applied can be used as augmented text. Taking the navigation service in the map field as an example, high-frequency texts such as "Turn left at the intersection ahead", "About to reach the destination, pull over to the right side of the road", "Red light ahead, please pay attention to your speed", "About to reach the tunnel, please pay attention to driving safety" can be used as augmented text, and augmented voices can be generated for each augmented text.

[0049] Large speech synthesis models are typically trained with massive amounts of data and feature complex neural network architectures. They are complex systems that integrate language understanding, acoustic modeling, and waveform generation. This application does not restrict the structure or specific implementation of large speech synthesis models; any large speech synthesis model can be used, such as CosyVoice, OpenVoice, Speech-2, and others. Furthermore, in addition to large speech synthesis models, other speech synthesis models with voice cloning capabilities can also be used to generate augmented speech.

[0050] Furthermore, since the augmented speech and augmented text are used for subsequent training of the second speech synthesis model, the greater the amount of augmented speech, the higher the training quality of the second speech synthesis model, but the longer it takes. Therefore, a balance must be struck, using as little augmented speech as possible while still meeting the required model quality. However, the spectra of speech vary from speaker to speaker. The more uneven the speech spectrum, the higher the corresponding complexity. The speech synthesis model typically requires more training data to learn the patterns of speech and perform detailed modeling to understand the acoustic characteristics of speech. For example, for high-frequency speech, the corresponding spectrum is primarily concentrated in the high-frequency band and is unevenly distributed. Therefore, more training data (i.e., speech-text pairs consisting of augmented speech and augmented text) is required for the speech synthesis model to learn the characteristics of this high-frequency speech.

[0051] In view of this, in an embodiment of the present application, after acquiring at least one speech of the target speaker, the spectrum of the at least one speech can be obtained; based on the uniformity of the spectrum of the at least one speech, multiple augmented texts corresponding to the uniformity are obtained to generate the corresponding augmented speech. The uniformity is negatively correlated with the number of augmented texts corresponding to the uniformity. That is, the higher the uniformity of the spectrum, the lower the complexity, and the fewer augmented texts required; conversely, the lower the uniformity of the spectrum, the higher the complexity, and the more augmented texts required. The uniformity of the spectrum can be measured using methods such as SFM (Spectral Flatness Measure) and spectral entropy. In addition, the fundamental frequency of the speech also affects the uniformity of the spectrum to a certain extent. For example, when the fundamental frequency is high (such as female voices and children's voices), the harmonics are more spaced and sparsely distributed, resulting in a more obvious "peak sense" of energy concentration and lower uniformity.

[0052] By determining the number of augmented texts and their corresponding augmented speech based on spectral uniformity, the second speech synthesis model can have more training data to learn the characteristics and rules of speech with uneven spectrum, thereby effectively ensuring the quality of training the second speech synthesis model.

[0053] The following describes in detail step 205, i.e., "using multiple augmented speech pieces to perform distillation training on a second speech synthesis model based on the first speech synthesis model to obtain a second speech synthesis model after distillation training, wherein the parameter scale of the second speech synthesis model is smaller than the parameter scale of the first speech synthesis model," in conjunction with an embodiment.

[0054] In an embodiment of the present application, the first speech synthesis model can adopt a speech synthesis model with a large parameter scale and high quality, such as the large speech synthesis model described in the previous embodiment. However, if a speech synthesis model with a large parameter scale and high quality is trained for each target speaker, for example, a large speech synthesis model is trained for each target speaker, on the one hand, the training time is very long and the user needs to wait for a long time, often requiring a waiting time of tens of minutes or even hours; on the other hand, the model is very large and difficult to deploy on the user side, or even if it can be deployed on the user side, it consumes a lot of network resources. In view of this, the embodiment of the present application adopts a distillation training method, using the first speech synthesis model with a larger parameter scale as the "teacher model" and the second speech synthesis model with a smaller parameter scale as the "student model", and performing distillation training to transfer the knowledge of the first speech synthesis model to the lighter-weight second speech synthesis model.

[0055] It should be noted that the speech synthesis model used when generating the augmented speech in step 203 and the "teacher model" used when performing distillation training in step 205 can be the same model or different models.

[0056] Figure 3 A schematic diagram of the structure of a second speech synthesis model provided in an embodiment of the present application is shown in FIG. Figure 3 As shown in , the second speech synthesis model includes a duration prediction module, a semantic encoder and an acoustic decoder.

[0057] The duration prediction module is used to predict the pronunciation duration of each phoneme in the input text. A phoneme is the smallest unit of speech in human language that can distinguish meaning. The duration prediction module is a key component for achieving natural speech prosody and can be implemented using either a recurrent neural network or a Transformer network.

[0058] Furthermore, the input of the duration prediction module may further include a speaker identification. In this case, the pronunciation duration of each phoneme in the input text may be predicted based on the input text and the speaker identification.

[0059] The semantic encoder encodes the input text and the pronunciation duration of each phoneme in the input text to obtain semantic features corresponding to the input text. Essentially, a speech encoder converts discrete text into continuous vector features that reflect semantic information. Semantic features can be used, for example, in the Hubert semantic feature. These features can be extracted using a text encoder such as a variational autoencoder (VAE).

[0060] Furthermore, the input of the semantic encoder may further include a speaker identifier. In this case, encoding may be performed based on the speaker identifier, the input text, and the pronunciation duration of each phoneme in the input text to obtain semantic features corresponding to the input text.

[0061] The acoustic decoder decodes the input text based on its semantic features to obtain the corresponding acoustic features. The acoustic decoder is the core component that converts text features into acoustic features. The output acoustic features can be spectral features, such as mel-spectrograms. Mel-spectrograms, which reflect the frequency distribution and energy variations of speech, are the most commonly used acoustic features. However, other acoustic features can also be used.

[0062] Furthermore, the input of the acoustic decoder may further include a speaker identifier. In this case, the acoustic decoder may perform decoding based on the semantic features corresponding to the input text and the speaker identifier to obtain the acoustic features corresponding to the input text.

[0063] Compared with the large speech synthesis model, the structure of the above-mentioned second speech synthesis model is simpler, with a smaller parameter scale, and is easier to package and send to the user end and deploy on the user end, thereby realizing speech synthesis on the user end.

[0064] During distillation training, the input text includes augmented text corresponding to the augmented speech and text corresponding to at least one speech. The training data consists of multiple training samples, each of which is a text-speech pair, including: the text corresponding to the target speaker's speech and the text-speech pair formed by the speech, and the text-speech pair formed by the augmented text and the augmented speech.

[0065] In each round of distillation training, the same input text is fed into both the first and second speech synthesis models, and the predicted speech output by the first and second speech synthesis models and the acoustic features output by the second speech synthesis models are obtained. A loss function value is then determined, and the model parameters of the second speech synthesis model are updated using a method such as gradient descent (the model parameters of the first speech synthesis model remain unchanged throughout the training process) until a preset training termination condition is met. The training termination condition may include, for example, the loss function value being less than or equal to a preset loss function threshold, or the number of iterations reaching a preset threshold.

[0066] As one of the possible implementation methods, the value of the loss function in the embodiment of the present application can be determined by a first loss term, a second loss term, and a third loss term. The first loss term is obtained by the difference between the semantic features output by the second speech synthesis model for the input text and the semantic features extracted from the predicted speech. The second loss term is obtained by the difference between the acoustic features output by the second speech synthesis model for the augmented text and the acoustic features extracted from the predicted speech. The third loss term is obtained by the difference between the acoustic features output by the second speech synthesis model for the above at least one speech and the acoustic features extracted from the above at least one speech. For example, the loss function The following formula can be used: in, The first loss term reflects the difference between the semantic features output by the second speech synthesis model for the input text and the semantic features extracted from the predicted speech of the first speech synthesis model. Since the first speech synthesis model performs inference on the input text, it can obtain the predicted speech, which is the result of the speech synthesis performed by the first speech synthesis model on the input text. The Hubert semantic features can be extracted from the predicted speech using a model such as HuBERT, such as Figure 4 The training goal of the first loss term is to minimize the distance between the true semantic features (actually, the predicted speech obtained by the first speech synthesis model is used as the true speech value, and the semantic features corresponding to the true speech value are considered to be the true semantic features) and the semantic features extracted by the first speech synthesis model, aligning the two.

[0067] is the second loss term, which reflects the difference between the acoustic features output by the second speech synthesis model for the augmented text and the acoustic features extracted from the predicted speech. Acoustic features can be extracted from the predicted speech using, for example, a Mel spectrum filter bank, such as Figure 4 The training goal of the second loss term is to minimize the difference between the learning features obtained by the second speech synthesis model and the true acoustic features (actually, the predicted speech by the first speech synthesis model is used as the true speech value, and the acoustic features corresponding to the true speech value are considered to be the true acoustic features).

[0068] is the third loss term, which reflects the difference between the acoustic features output by the second speech synthesis model for the at least one speech and the acoustic features extracted from the at least one speech, such as Figure 4 This loss term is similar to the second loss term, but targets the target speaker’s speech and its text.

[0069] and is the weight coefficient, which can be a preset hyperparameter. is an identifier. When the input is augmented text, =1, when the input data is the text corresponding to the target speaker's speech, =0.

[0070] It can be seen that the above loss function can align the second speech synthesis model with the first speech synthesis model not only in terms of semantic features but also in acoustic features during the distillation training process, so that the second speech synthesis model can fully learn the understanding and synthesis capabilities of the high-performance first speech synthesis model, thereby obtaining a lighter-weight, less computationally intensive second speech synthesis model that meets quality requirements.

[0071] In addition to the above-mentioned method of determining the loss function, other methods may also be used. For example, based on the basic principle of the above-mentioned formula, a simple deformation of the formula is performed, which is within the scope of protection of this application.

[0072] The above step 207, i.e., "determining the speech package corresponding to the target speaker based on the second speech synthesis model after distillation training," is described in detail below with reference to an embodiment.

[0073] After converting the second speech synthesis model trained by distillation into a format usable by the user end, the second speech synthesis model trained by distillation can be used for packaging to obtain a speech package corresponding to the target speaker.

[0074] It's important to note that the second speech synthesis model trained after distillation also includes a vocoder. However, the vocoder doesn't participate in the training process and can be added later. The vocoder converts the abstract acoustic features (such as the Mel-spectrogram) output by the acoustic model into an audible audio waveform and is the "final vocalization unit" of speech synthesis.

[0075] Furthermore, if the augmented voice is generated based on a high-frequency file for a specific scenario or a specific service, it means that the augmented voice will also be used frequently in that specific scenario or for that specific service. Therefore, the second speech synthesis model trained by distillation can be used to package the augmented voice to obtain a voice package corresponding to the target speaker. In other words, the voice package includes not only the second speech synthesis model trained by distillation but also the augmented voice. The user end can give priority to using the augmented voice for voice broadcasting. Only when the augmented voice does not contain the target broadcast voice, the second speech synthesis model is used to generate the target broadcast voice. This approach can save user end performance and improve broadcast efficiency.

[0076] The server can include the generated voice package in the installation package of the user. After generating the voice package corresponding to the target speaker, the server can also push the voice package to the user in real time. Alternatively, the server can send the voice package to the user in response to a user request. For example, the server can include a download link for the target speaker's voice package on a page, and can also include a trial voice of the target speaker on the page. After playing the trial voice, the user can determine whether to request download of the corresponding voice package through the download link. The audiovisual voice can be selected from at least one voice or augmented voice of the target speaker, or can be generated by a second speech synthesis model.

[0077] Figure 5 This is a flow chart of the voice broadcast method provided in the embodiment of the present application. This method can be performed by Figure 1 The user side execution in the system architecture shown. Figure 5 As shown in , the method may include: Step 501: Acquire a speech package corresponding to a target speaker, where the speech package includes at least a second speech synthesis model. The second speech synthesis model is obtained through distillation training using at least one speech of the target speaker.

[0078] The user end can obtain the voice package corresponding to the target speaker from the installation package, or obtain the voice package pushed by the server end, or request and download the voice package generated by the server end.

[0079] As one of the possible implementation methods, the user end can send at least one speech of the target speaker collected by the user through the recording collection tool to the server end to request the generation of the speech package corresponding to the target speaker. Figure 2 The process in the illustrated method embodiment generates a voice package of the target speaker and sends a download link of the voice package to the user terminal. The user terminal can trigger the user terminal to download the voice package from the server terminal through the download link.

[0080] As another possible implementation, the server can use Figure 2 The process in the illustrated method embodiment generates a voice package for the target speaker and includes a download link for the voice package and a trial audio on the relevant page of the voice package service. After the user accesses the page, the user can trigger the playback of the trial audio and download the voice package corresponding to the target speaker through the download link.

[0081] Step 503: In response to the event triggering the voice broadcast, the target broadcast voice is determined using the voice package.

[0082] In an embodiment of the present application, the voice to be broadcast, i.e., the target broadcast voice, can be determined in different scenarios or different services. If the content to be broadcast matches the augmented voice, i.e., the target broadcast voice exists in the augmented voice of the voice package, then the target broadcast voice is selected from multiple augmented voices for broadcast, i.e., the augmented voice is used for broadcasting first. If the content to be broadcast does not match the augmented voice, i.e., the target broadcast voice does not exist in the augmented voice of the voice package, then the target broadcast voice is generated using the second speech synthesis model.

[0083] The above application scenarios or services may include but are not limited to the following: Interaction with intelligent voice assistants. For example, users can interact with intelligent voice assistants installed in smart speakers, smart TVs, smart watches, smart car terminals, etc. After generating a conversation text based on the context, the intelligent voice assistant can use the voice package to obtain the corresponding speech of the conversation text, which has the timbre characteristics of the target speaker.

[0084] Voice announcements in map services. For example, map services can provide users with traffic announcements, station announcements, and navigation announcements. Typically, the server generates announcement text based on actual conditions and sends it to the user end, which then uses the voice package to generate the announcement voice, which has the timbre characteristics of the target speaker.

[0085] Audio content generation in content consumption services. For example, news and information platforms offer automatic news and information broadcasting. Users can use voice packages to generate audio from text of news and information, with the audio having the timbre and voice characteristics of the target speaker. Another example is audiobook services, where audio packages can be used to convert articles into audio for listening, with the audio having the timbre and voice characteristics of the target speaker.

[0086] Regardless of the scenario or service, the process of determining the target broadcast voice is performed locally on the user side. Compared with the method of obtaining the target broadcast voice from the server side, on the one hand, it is faster and has lower latency, and on the other hand, it saves more network resources and can even be achieved without an Internet connection in some scenarios.

[0087] Step 505: Play the target announcement voice.

[0088] It can be seen that this application makes it possible to collect data of a target speaker and produce a personalized voice package through voice augmentation and distillation training, which lowers the threshold for producing personalized voice packages and allows users to complete voice collection within seconds. In addition, the distillation training method greatly shortens the model training time and obtains a lightweight "student model" with a smaller parameter scale and lower computational complexity to achieve deployment on the user side. This model can support real-time, offline speech synthesis on the user side. In addition, the size of the personalized voice package is reduced, the rapid delivery of voice is achieved, and network resources are saved.

[0089] The foregoing description of this specification describes specific embodiments. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in an order different from that described in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order shown or the sequential order to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0090] According to another embodiment, a device for generating a voice packet is provided. Figure 6 A schematic block diagram of the device for generating a voice packet provided in an embodiment of the present application, wherein the device is provided at Figure 1 The server side of the architecture shown in Figure 1. Figure 6As shown, the apparatus 600 includes: a speech acquisition unit 601, a data augmentation unit 602, a distillation training unit 603, and a speech packet generation unit 604. The main functions of each component unit are as follows: The speech acquisition unit 601 is configured to acquire at least one speech of a target speaker.

[0091] The data augmentation unit 602 is configured to generate multiple augmented speech messages based on the acoustic features of at least one speech message.

[0092] The distillation training unit 603 is configured to use multiple augmented speech to perform distillation training on the second speech synthesis model based on the first speech synthesis model to obtain the second speech synthesis model after distillation training. The parameter scale of the second speech synthesis model is smaller than the parameter scale of the first speech synthesis model.

[0093] The speech packet generating unit 604 is configured to determine a speech packet corresponding to the target speaker based on the second speech synthesis model trained by distillation.

[0094] As one of the possible implementation methods, the data augmentation unit 602 can be specifically configured to: obtain multiple augmented texts; and generate speech corresponding to the multiple augmented texts as augmented speech based on the acoustic features of at least one speech using a first speech synthesis model.

[0095] As one of the possible implementation methods, the voice package generation unit 604 may be specifically configured to package the second speech synthesis model trained by distillation and multiple augmented speech pieces to obtain a voice package corresponding to the target speaker.

[0096] As one of the possible implementation methods, when obtaining multiple augmented texts, the data augmentation unit 602 can be specifically configured to: obtain the frequency spectrum of at least one speech; and obtain a number of augmented texts corresponding to the uniformity of the frequency spectrum of the at least one speech, wherein the uniformity is negatively correlated with the number of augmented texts corresponding to the uniformity.

[0097] As one possible implementation method, the second speech synthesis model includes a duration prediction module, a semantic encoder and an acoustic decoder.

[0098] The duration prediction module is used to predict the pronunciation duration of each phoneme in the input text based on the input text.

[0099] The semantic encoder is used to encode the input text and the pronunciation duration of each phoneme in the input text to obtain the semantic features corresponding to the input text.

[0100] The acoustic decoder is used to decode the input text based on the semantic features corresponding to the input text to obtain the acoustic features corresponding to the input text.

[0101] In the distillation training process, the input text includes the augmented text corresponding to the augmented speech and the text corresponding to at least one speech.

[0102] As one of the possible implementation methods, the distillation training unit 603 can be specifically configured as follows: inputting the same input text into the first speech synthesis model and the second speech synthesis model respectively, and obtaining the predicted speech output by the first speech synthesis model and the acoustic features output by the second speech synthesis model; determining the value of the loss function, the value of the loss function is determined by the first loss term, the second loss term and the third loss term, the first loss term is obtained by the difference between the semantic features output by the second speech synthesis model for the input text and the semantic features extracted from the predicted speech, the second loss term is obtained by the difference between the acoustic features output by the second speech synthesis model for the augmented text and the acoustic features extracted from the predicted speech, and the third loss term is obtained by the difference between the acoustic features output by the second speech synthesis model for at least one speech and the acoustic features extracted from at least one speech; using the value of the loss function, the model parameters of the second speech synthesis model are updated.

[0103] Figure 7 This is a schematic block diagram of a voice broadcast device provided in an embodiment of the present application, which is arranged at Figure 1 The user side of the architecture shown in Figure 1 is shown in Figure 2. Figure 7 As shown, the device 700 includes: a voice packet acquisition unit 701, a voice acquisition unit 702 and a voice broadcast unit 703. The main functions of each component unit are as follows: The speech packet acquisition unit 701 is configured to acquire a speech packet corresponding to a target speaker, where the speech packet includes at least a second speech synthesis model, which is obtained through distillation training using at least one speech of the target speaker.

[0104] The voice acquisition unit 702 is configured to determine a target broadcast voice using a voice packet in response to an event that triggers a voice broadcast.

[0105] The voice broadcast unit 703 is configured to play the target broadcast voice.

[0106] Furthermore, the voice package also includes multiple augmented voices; when the voice acquisition unit 702 uses the voice package to determine the target broadcast voice, it can be specifically configured as follows: if the target broadcast voice exists among the multiple augmented voices, then the target broadcast voice is selected from the multiple augmented voices; if the target broadcast voice does not exist among the multiple augmented voices, then the target broadcast voice is generated using the second voice synthesis model.

[0107] Each embodiment in this specification is described in a progressive manner. The same or similar parts between the embodiments can be referred to each other. Each embodiment focuses on the differences from other embodiments. In particular, for system or device embodiments, since they are basically similar to method embodiments, the description is relatively simple. For relevant parts, refer to the partial description of the method embodiment. The system and device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without expending creative work.

[0108] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.

[0109] In addition, an embodiment of the present application further provides a computer-readable storage medium on which a computer program is stored. When the program is executed by a processor, the steps of any one of the methods in the aforementioned method embodiments are implemented.

[0110] And an electronic device comprising: one or more processors; and A memory associated with the one or more processors, the memory being used to store program instructions, wherein the program instructions, when read and executed by the one or more processors, execute the steps of the method described in any one of the aforementioned method embodiments.

[0111] The present application also provides a computer program product, comprising a computer program, which implements the steps of any one of the methods described in the aforementioned method embodiments when executed by a processor.

[0112] in, Figure 8The electronic device architecture is shown as an example, and may include a processor 810, a video display adapter 811, a disk drive 812, an input / output interface 813, a network interface 814, and a memory 820. The processor 810, the video display adapter 811, the disk drive 812, the input / output interface 813, the network interface 814, and the memory 820 may be communicatively connected via a communication bus 830.

[0113] The processor 810 may be implemented as a general-purpose CPU, a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and may be used to execute relevant programs to implement the technical solutions provided in this application.

[0114] The memory 820 can be implemented in the form of ROM (Read Only Memory), RAM (Random Access Memory), static storage device, dynamic storage device, etc. The memory 820 can store an operating system 821 for controlling the operation of the electronic device 800, and a basic input and output system (BIOS) 822 for controlling the low-level operations of the electronic device 800. In addition, a web browser 823, a data storage management system 824, and a device for generating a voice packet 600 / a voice broadcast device 700, etc. can also be stored. The above-mentioned device for generating a voice packet 600 / a voice broadcast device 700 can be an application program that specifically implements the operations of the aforementioned steps in the embodiment of the present application. In short, when the technical solution provided by the present application is implemented by software or firmware, the relevant program code is stored in the memory 820 and is called and executed by the processor 810.

[0115] The input / output interface 813 is used to connect to input / output modules to enable information input and output. The input / output modules can be configured as components within the device (not shown) or externally connected to the device to provide corresponding functions. Input devices may include a keyboard, mouse, touch screen, microphone, and various sensors, while output devices may include a display, speaker, vibrator, indicator light, and the like.

[0116] The network interface 814 is used to connect to a communication module (not shown) to enable communication between the device and other devices. The communication module can communicate via wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, Wi-Fi, Bluetooth, etc.).

[0117] The bus 830 comprises a pathway for transmitting information between the various components of the device (eg, the processor 810 , the video display adapter 811 , the disk drive 812 , the input / output interface 813 , the network interface 814 , and the memory 820 ).

[0118] It should be noted that although the above device only shows the processor 810, video display adapter 811, disk drive 812, input / output interface 813, network interface 814, memory 820, bus 830, etc., in the specific implementation process, the device may also include other components necessary for normal operation. In addition, it will be understood by those skilled in the art that the above device may also include only the components necessary to implement the solution of the present application, and does not necessarily include all the components shown in the figure.

[0119] Through the description of the above embodiments, it can be seen that those skilled in the art can clearly understand that the present application can be implemented by means of software plus the necessary general hardware platform. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a computer program product. The computer program product can be stored in a storage medium, such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in various embodiments of the present application or certain parts of the embodiments.

[0120] The above is a detailed introduction to the technical solutions provided by this application. Specific examples are used herein to illustrate the principles and implementation methods of this application. The description of the above embodiments is only intended to help understand the method and core concept of this application. At the same time, for those skilled in the art, based on the concept of this application, there may be changes in the specific implementation methods and application scope. In summary, the contents of this specification should not be understood as limiting this application.

Claims

1. A method for generating a voice packet, characterized in that: The method comprises: Obtain at least one speech of the target speaker; generating a plurality of augmented speech messages based on the acoustic features of the at least one speech message; Using the plurality of augmented speech sounds, performing distillation training on a second speech synthesis model based on the first speech synthesis model to obtain a second speech synthesis model after distillation training, wherein a parameter scale of the second speech synthesis model is smaller than a parameter scale of the first speech synthesis model; Based on the second speech synthesis model after the distillation training, a speech package corresponding to the target speaker is determined.

2. The method according to claim 1, characterized in that The generating of a plurality of augmented speech pieces based on the acoustic features of the at least one piece of speech data comprises: Get multiple augmented texts; The first speech synthesis model is used to generate speech corresponding to each of the plurality of augmented texts as the augmented speech based on the acoustic features of the at least one speech.

3. The method according to claim 2, characterized in that Determining the speech package corresponding to the target speaker based on the second speech synthesis model after distillation training includes: The second speech synthesis model trained by distillation and the plurality of augmented speech pieces are packaged to obtain a speech package corresponding to the target speaker.

4. The method according to claim 2, characterized in that The obtaining of multiple augmented texts includes: Acquiring a frequency spectrum of the at least one speech; According to the uniformity of the frequency spectrum of the at least one speech, a number of augmented texts corresponding to the uniformity is obtained, wherein the uniformity is negatively correlated with the number of augmented texts corresponding to the uniformity.

5. The method according to any one of claims 1 to 4, characterized in that The second speech synthesis model includes a duration prediction module, a semantic encoder and an acoustic decoder; The duration prediction module is used to predict the pronunciation duration of each phoneme in the input text based on the input text; The semantic encoder is used to encode based on the input text and the pronunciation duration of each phoneme in the input text to obtain the semantic features corresponding to the input text; The acoustic decoder is used to decode based on the semantic features corresponding to the input text to obtain the acoustic features corresponding to the input text; In the distillation training process, the input text includes the augmented text corresponding to the augmented speech and the text corresponding to the at least one speech.

6. The method according to claim 5, characterized in that The method of performing distillation training on the second speech synthesis model based on the first speech synthesis model by using the plurality of augmented speech pieces includes: Inputting the same input text into the first speech synthesis model and the second speech synthesis model respectively, and obtaining the predicted speech output by the first speech synthesis model and the acoustic features output by the second speech synthesis model; Determining a value of a loss function, where the value of the loss function is determined by a first loss term, a second loss term, and a third loss term, wherein the first loss term is obtained by a difference between semantic features output by the second speech synthesis model for the input text and semantic features extracted from the predicted speech, the second loss term is obtained by a difference between acoustic features output by the second speech synthesis model for the augmented text and acoustic features extracted from the predicted speech, and the third loss term is obtained by a difference between acoustic features output by the second speech synthesis model for the at least one speech and acoustic features extracted from the at least one speech; The model parameters of the second speech synthesis model are updated using the value of the loss function.

7. A voice broadcast method, applied to a client, characterized in that: The method comprises: Obtaining a speech package corresponding to a target speaker, the speech package including at least a second speech synthesis model, the second speech synthesis model being obtained through distillation training using at least one speech of the target speaker; In response to an event triggering a voice broadcast, determining a target broadcast voice using the voice packet; Play the target announcement voice.

8. The method according to claim 7, characterized in that The voice package also includes a plurality of augmented voices; and determining the target broadcast voice using the voice package includes: If the target broadcast voice exists in the multiple augmented voices, selecting the target broadcast voice from the multiple augmented voices; If the target broadcast voice does not exist in the multiple augmented voices, the target broadcast voice is generated using the second voice synthesis model.

9. A device for generating a voice packet, characterized in that: The device comprises: a speech acquisition unit configured to acquire at least one speech of a target speaker; a data augmentation unit, configured to generate a plurality of augmented speech messages based on the acoustic features of the at least one speech message; a distillation training unit configured to perform distillation training on a second speech synthesis model based on the first speech synthesis model using the plurality of augmented speech pieces, to obtain a second speech synthesis model after distillation training, wherein a parameter scale of the second speech synthesis model is smaller than a parameter scale of the first speech synthesis model; The speech packet generation unit is configured to determine the speech packet corresponding to the target speaker based on the second speech synthesis model trained by the distillation.

10. A voice broadcasting device, characterized in that: The device comprises: a speech packet acquisition unit configured to acquire a speech packet corresponding to a target speaker, wherein the speech packet includes at least a second speech synthesis model, wherein the second speech synthesis model is obtained through distillation training using at least one speech of the target speaker; A voice acquisition unit is configured to determine a target broadcast voice using the voice packet in response to an event triggering the voice broadcast; The voice broadcast unit is configured to play the target broadcast voice.

11. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.

Citation Information

Patent Citations

  • Method and device of generating voice packet, equipment and computer storage medium

    CN110751940A

  • Method and device for training voice augmentation model

    CN113314107A

  • Text processing method and device, readable medium and electronic equipment

    CN115186633A

  • Model training method and speech synthesis method based on semi-supervised knowledge distillation

    CN116092469A

  • Voice data augmentation method, electronic equipment and storage medium

    CN116504233A