Extensible method and apparatus for implementing acoustic model of a speaker

CN117690410BActive Publication Date: 2026-09-15CHINA TELECOM CORP LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311723701.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-12-14
Publication Date
2026-09-15
Estimated Expiration
2043-12-14

AI Technical Summary

Technical Problem

[0006]本申请实施例提供了一种可扩展发音人的声学模型实现方法、装置,以至少解决相关声学模型多个发音人批输入语音文本进行语音合成时,通过调用各个发音人的自适应模型参数依次对每个发音人的语音文本进行推理,导致推理处理效率较低的技术问题

Benefits of technology

[0019] In this embodiment, in response to multiple target new speaker requests for speech synthesis, the target speaker identifiers of each target new speaker in the speech synthesis request are determined; the target embedding vectors of the target new speakers corresponding to the target speaker identifiers are determined from a preset speaker index table, wherein the target embedding vectors include: speaker embedding vectors and adaptive embedding vectors. The speaker index table sequentially records the first speaker identifier and its corresponding first speaker embedding vector of the first speaker, the new speaker identifiers of the multiple new speakers and their corresponding second embedding vectors, and the second embedding vectors include: the new speaker embedding vector corresponding to the new speaker and the adaptive embedding vector adapter parameters obtained by adaptively training the target acoustic model using the second speech data of the new speaker. The target acoustic model is obtained by combining the preset adapter with the baseline acoustic model trained based on the first speech data of the first speaker; the target adapter parameters in the target acoustic model are determined based on the target embedding vectors of the multiple target new speakers, and the target adapter parameters are recombined with the target acoustic model. The combined target acoustic model is then used to respond to the speech synthesis requests of the multiple target speakers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117690410B_ABST
    Figure CN117690410B_ABST
Patent Text Reader

Abstract

The application discloses a scalable speaker acoustic model implementation method and device. The method comprises the following steps: in response to a plurality of target new speaker voice synthesis requests, determining the target speaker identifier of each target new speaker in the voice synthesis request; determining the target embedding vector of the target new speaker corresponding to the target speaker identifier in the preset speaker index table; determining the target adapter parameter in the target acoustic model based on the target embedding vector of the plurality of target new speakers, and recombining the target adapter parameter and the target acoustic model; and responding to the voice synthesis request of the plurality of target speakers through the combined target acoustic model. The application solves the technical problem of low inference processing efficiency caused by sequentially inferring the voice text of each speaker by calling the adaptive model parameters of each speaker when the related acoustic model performs voice synthesis on the batch input voice text of multiple speakers.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of speech synthesis technology, and more specifically, to a method and apparatus for implementing an acoustic model of a scalable speaker. Background Technology

[0002] Speech synthesis technology endows computers (or various terminal devices) with the ability to speak like humans. TTS (Text to Speech) technology belongs to speech synthesis; it is a technology that converts text information generated by the computer itself or input from external sources into understandable, fluent spoken output. A speech synthesis system generally consists of three main parts: a text analysis module, an acoustic model, and a vocoder.

[0003] Building upon speech synthesis systems, multi-speaker speech synthesis systems can train an acoustic model using mixed multi-speaker data, and then fine-tune some model parameters using a small amount of target speaker corpus—a process known as "pre-training + fine-tuning." In recent years, adapter-based adaptive techniques have yielded significant results in fine-tuning large-scale NLP (Natural Language Processing) models. Existing literature suggests adding only a few adapter layers to the existing speech synthesis acoustic system's base model to learn the features of new speakers. However, regardless of the approach, some layers and parameters are shared, while others are unique to each speaker. Therefore, current multi-speaker speech synthesis models, when inputting more than one speaker, require different speakers to select their own unique layers and parameters for inference during the inference process. However, since each speaker's parameters exist independently after training, this approach is problematic.

[0004] To address the aforementioned issues, a relatively simple method has been proposed by technical personnel: inputs with a batch size greater than 1 are transformed into multiple inputs with a batch size of 1, decoded separately, and then the results are concatenated into a batch and returned. However, this method severely reduces the inference efficiency of the computer's GPU.

[0005] There is currently no effective solution to the above problems. Summary of the Invention

[0006] This application provides a method and apparatus for implementing an acoustic model of a scalable speaker, which at least solves the technical problem of low inference processing efficiency when multiple speakers input speech text for speech synthesis, which is caused by sequentially inferring the speech text of each speaker by calling the adaptive model parameters of each speaker.

[0007] According to one aspect of the embodiments of this application, a method for implementing an acoustic model of a scalable speaker is provided, comprising: responding to a speech synthesis request from multiple target new speakers, determining the target speaker identifier of each target new speaker in the speech synthesis request; determining the target embedding vector of the target new speaker corresponding to the target speaker identifier from a preset speaker index table, wherein the target embedding vector includes: a speaker embedding vector and an adaptive embedding vector, the speaker index table sequentially recording the first speaker identifier and the corresponding first speaker embedding vector of a first speaker, the new speaker identifiers of multiple new speakers and the corresponding second embedding vectors, the second embedding vector including: the new speaker embedding vector corresponding to the new speaker and adaptive embedding vector adapter parameters obtained by adaptively training the target acoustic model using the second speech data of the new speaker, the target acoustic model being obtained by combining the preset adapter with a baseline acoustic model trained based on the first speech data of the first speaker; determining the target adapter parameters within the target acoustic model based on the target embedding vectors of the multiple target new speakers, and recombining the target adapter parameters with the target acoustic model, and responding to the speech synthesis requests from the multiple target speakers through the combined target acoustic model.

[0008] Optionally, the training process of the benchmark acoustic model includes: constructing a deep learning model, wherein the deep learning model includes: an encoder composed of a single-layer feedforward transformer (FFT) module, a speaker embedding module, and a decoder composed of N layers of FFT modules, where N is a positive integer greater than or equal to 1; acquiring first speech data of multiple first speakers, wherein the first speech data includes: the first phoneme sequence encoding corresponding to the first speech text of the first speaker and the first acoustic feature corresponding to the first speech audio, and the first speech text corresponds to the first speech audio; iteratively training the deep learning model based on the first speech data of multiple first speakers to obtain the benchmark acoustic model.

[0009] Optionally, acquiring first speech data from multiple first speakers includes: acquiring initial speech text for each first speaker and initial speech audio corresponding to the initial speech text, wherein the initial speech text includes at least one of the following: text information, punctuation marks, and prosodic annotations; converting the text information in the initial speech text into an initial phoneme sequence, and inserting punctuation marks and prosodic annotations into the initial phoneme sequence to obtain a first phoneme sequence; encoding the first phoneme sequence using one-hot encoding to obtain a first phoneme sequence encoding; performing preprocessing operations on the initial speech audio, wherein the preprocessing operations include at least one of the following: sampling, volume adjustment, and cropping; and extracting features from the preprocessed initial speech audio to obtain first acoustic features, wherein the first acoustic features include at least one of the following: Mel-spectral features, frame-level variable features, and phoneme-level duration features.

[0010] Optionally, a baseline acoustic model is obtained by iteratively training a deep learning model based on the first speech data of multiple first speakers. This includes: for the first speech data of each first speaker, encoding the first phoneme sequence of the first speaker and inputting it into the deep learning model, sequentially passing it through the encoder and decoder within the deep learning model to output the corresponding first streaming acoustic features and first non-streaming acoustic features; determining a target loss function based on the first acoustic features, the first streaming acoustic features, and the first non-streaming acoustic features within the first speech data of each first speaker, wherein the target loss function includes: a non-streaming mean squared error loss function, a streaming mean squared error loss function, and an adversarial loss function; calculating the minimum value of the target loss function using a gradient descent algorithm, and adjusting the model parameters of the deep learning model based on the minimum value of the target loss function to obtain the baseline acoustic model.

[0011] Optionally, the adaptive embedding vector obtained by adaptively training the target acoustic model using the second speech data of the newly added speaker includes: acquiring the second speech data of the newly added speaker, wherein the second speech data includes: the second phoneme sequence encoding corresponding to the second speech text of the newly added speaker and the second acoustic features corresponding to the second speech audio, and the second speech text corresponds to the second speech audio; acquiring a baseline acoustic model and improving the decoder in the baseline acoustic model using an adapter to obtain a target acoustic model, wherein the adapter consists of a first feedforward neural network with reduced dimensionality, an activation layer, and a second feedforward neural network with increased dimensionality; adaptively training the target acoustic model using the second speech data of the newly added speaker to obtain the adapter parameters of the newly added speaker, and concatenating the adapter parameters to obtain the adaptive embedding vector corresponding to the newly added speaker, wherein the adapter parameters include at least one of the following: the first weight matrix and the first bias vector of the first feedforward neural network of the N-layer adapter, and the second weight matrix and the second bias vector of the second feedforward neural network.

[0012] Optionally, the decoder within the baseline acoustic model is improved using an adapter, including: initializing the adapter and adding an initialized adapter after each of the N-layer FFT modules of the decoder within the baseline acoustic model to obtain an improved decoder, wherein the FFT module consists of a multi-head attention mechanism module, a block module, a conditional layer normalization module, and a causal convolution module.

[0013] Optionally, the adapter is initialized by: initializing the first weight matrix of the first feedforward neural network and the second weight matrix of the second feedforward neural network to 1, and initializing the first bias vector of the first feedforward neural network and the second bias vector of the second feedforward neural network to 0.

[0014] Optionally, the adapter parameters are concatenated to obtain the adaptive embedding vector corresponding to the new speaker, including: concatenating the first weight matrix and first bias vector of the first feedforward neural network and the second weight matrix and second bias vector of the second feedforward neural network in each adapter from the first to the Nth layer of the adapter parameters corresponding to the new speaker, to obtain the adaptive embedding vector corresponding to the new speaker.

[0015] Optionally, the target adapter parameters within the target acoustic model are determined based on the target embedding vectors of multiple newly added speakers, including: slicing and transforming the target adaptive embedding vector of each newly added speaker to obtain the target adapter parameters corresponding to the corresponding N-layer target adapter; for the target adapter parameters corresponding to each layer of target adapters, concatenating the first weight matrix, the first bias vector, the second weight matrix, and the second bias vector of the first feedforward neural network corresponding to multiple newly added speakers to obtain the first target weight matrix, the first target bias vector, the second target weight matrix, and the second target bias vector; and obtaining the target adapter parameters of the target acoustic model from the target speaker embedding vectors corresponding to multiple newly added speakers, the first target weight matrix, the first target bias vector, the second target weight matrix, and the second target bias vector of each layer of target adapters.

[0016] According to another aspect of the embodiments of this application, an scalable acoustic model implementation apparatus for speakers is also provided, comprising: a first determining module, configured to determine the target speaker identifier of each target new speaker in the speech synthesis request in response to a speech synthesis request for multiple target new speakers; and a second determining module, configured to determine the target embedding vector of the target new speaker corresponding to the target speaker identifier from a preset speaker index table, wherein the target embedding vector includes: a speaker embedding vector and an adaptive embedding vector, and the speaker index table sequentially records the first speaker identifier of the first speaker and the corresponding first speaker embedding vector, and the new speaker identifiers of multiple new speakers and the corresponding second embedding vectors. The second embedding vector includes: a new speaker embedding vector corresponding to the new speaker and adaptive embedding vector adapter parameters obtained by adaptively training the target acoustic model using the second speech data of the new speaker. The target acoustic model is obtained by combining the preset adapter with the baseline acoustic model trained based on the first speech data of the first speaker. The model generation module is used to determine the target adapter parameters in the target acoustic model based on the target embedding vectors of multiple target new speakers, and to recombine the target adapter parameters with the target acoustic model, and to respond to the speech synthesis requests of multiple target speakers through the combined target acoustic model.

[0017] According to another aspect of the embodiments of this application, a non-volatile storage medium is also provided, the non-volatile storage medium including a stored computer program, wherein the device where the non-volatile storage medium is located executes the above-described method for implementing the acoustic model of a scalable speaker by running the computer program.

[0018] According to another aspect of the embodiments of this application, an electronic device is also provided, the electronic device comprising: a memory and a processor, wherein the memory stores a computer program, and the processor is configured to execute the above-described method for implementing the acoustic model of a scalable speaker through the computer program.

[0019] In this embodiment, in response to multiple target new speaker requests for speech synthesis, the target speaker identifiers of each target new speaker in the speech synthesis request are determined; the target embedding vectors of the target new speakers corresponding to the target speaker identifiers are determined from a preset speaker index table, wherein the target embedding vectors include: speaker embedding vectors and adaptive embedding vectors. The speaker index table sequentially records the first speaker identifier and its corresponding first speaker embedding vector of the first speaker, the new speaker identifiers of the multiple new speakers and their corresponding second embedding vectors, and the second embedding vectors include: the new speaker embedding vector corresponding to the new speaker and the adaptive embedding vector adapter parameters obtained by adaptively training the target acoustic model using the second speech data of the new speaker. The target acoustic model is obtained by combining the preset adapter with the baseline acoustic model trained based on the first speech data of the first speaker; the target adapter parameters in the target acoustic model are determined based on the target embedding vectors of the multiple target new speakers, and the target adapter parameters are recombined with the target acoustic model. The combined target acoustic model is then used to respond to the speech synthesis requests of the multiple target speakers.

[0020] In this embodiment, by using the speech data of a newly added speaker to adaptively train the target acoustic model after adding an adapter, layers and parameters unique to that newly added speaker are obtained, and a speaker index table is established based on this. When responding to speech synthesis requests from multiple newly added speakers, the corresponding speaker embedding vector can be found based on the speaker identifier of the newly added speaker, and the unique layers and parameters in the embedding vectors of multiple newly added speakers are combined to obtain a multi-speaker acoustic model that can handle the speech synthesis requests of each newly added speaker. This eliminates the need to call the unique layers and parameters of each newly added speaker for inference every time speech synthesis is performed for different newly added speakers, thus avoiding looping operations, improving the inference efficiency of the computer, and solving the technical problem of low inference processing efficiency when multiple speakers' batch input speech texts are used for speech synthesis in related acoustic models by calling the adaptive model parameters of each speaker to infer the speech text of each speaker sequentially. Attached Figure Description

[0021] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments of this application and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:

[0022] Figure 1 This is a hardware structure block diagram of a computer terminal for implementing an acoustic model of a scalable speaker, according to an embodiment of this application.

[0023] Figure 2 This is a flowchart illustrating an optional method for implementing an acoustic model of a scalable speaker according to an embodiment of this application.

[0024] Figure 3 This is a schematic diagram of the structure of an optional deep learning model according to an embodiment of this application;

[0025] Figure 4 This is a schematic diagram of the structure of an optional FFT module according to an embodiment of this application;

[0026] Figure 5 This is a flowchart illustrating an optional phoneme sequence encoding generation method according to an embodiment of this application.

[0027] Figure 6 This is a schematic diagram of the structure of an optional target acoustic model according to an embodiment of this application;

[0028] Figure 7 This is a schematic diagram of the structure of an optional adapter according to an embodiment of this application;

[0029] Figure 8This is a flowchart illustrating the calculation of an optional adapter supporting batch input from multiple speakers according to an embodiment of this application.

[0030] Figure 9 This is a schematic diagram of an optional apparatus for implementing an acoustic model of a scalable speaker, according to an embodiment of this application. Detailed Implementation

[0031] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.

[0032] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0033] Furthermore, all information and data (including but not limited to user device information, user personal information, etc.) involved in this application are information and data authorized by the user or fully authorized by all parties. For example, this system has an interface with the relevant user or organization. Before obtaining relevant information, it needs to send an acquisition request to the aforementioned user or organization through the interface, and obtain the relevant information after receiving consent from the aforementioned user or organization.

[0034] Example 1

[0035] According to an embodiment of this application, an embodiment of a method for implementing an acoustic model of a scalable speaker is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.

[0036] The method embodiments provided in this application can be executed on a mobile terminal, computer terminal, or similar computing device. Figure 1 A hardware block diagram of a computer terminal for implementing an acoustic model of a scalable speaker is shown. Figure 1 As shown, the computer terminal 10 may include one or more processors 102 (shown as 102a, 102b, ..., 102n in the figure) 102 (processor 102 may include, but is not limited to, a microprocessor MCU or a programmable logic device FPGA, etc.), a memory 104 for storing data, and a transmission device 106 for communication functions. In addition, it may also include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of a BUS bus), a network interface, a power supply, and / or a camera. Those skilled in the art will understand that... Figure 1 The structure shown is for illustrative purposes only and does not limit the structure of the aforementioned electronic device. For example, computer terminal 10 may also include... Figure 1 The more or fewer components shown, or having the same Figure 1 The different configurations shown.

[0037] It should be noted that the aforementioned one or more processors 102 and / or other data processing circuits are generally referred to herein as "data processing circuits". These data processing circuits may be embodied, in whole or in part, in software, hardware, firmware, or any other combination thereof. Furthermore, the data processing circuits may be a single, independent processing module, or may be integrated, in whole or in part, into any other element within the computer terminal 10. As involved in the embodiments of this application, the data processing circuits serve as a form of processor control (e.g., selection of a variable resistor termination path connected to an interface).

[0038] The memory 104 can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the method for implementing an acoustic model of a scalable speaker in the embodiments of this application. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, that is, implementing the above-mentioned application method for implementing an acoustic model of a scalable speaker. The memory 104 may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include memory remotely located relative to the processor 102, and these remote memories can be connected to the computer terminal 10 via a network. Examples of the above-mentioned networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.

[0039] The transmission device 106 is used to receive or send data via a network. Specific examples of the network described above may include a wireless network provided by the communication provider of the computer terminal 10. In one example, the transmission device 106 includes a Network Interface Controller (NIC), which can connect to other network devices via a base station to communicate with the Internet. In another example, the transmission device 106 may be a Radio Frequency (RF) module, used for wireless communication with the Internet.

[0040] The display may be, for example, a touchscreen liquid crystal display (LCD) that allows the user to interact with the user interface of the computer terminal 10.

[0041] Under the above operating environment, Figure 2 This is a schematic diagram of an optional method for implementing an acoustic model of a scalable speaker according to an embodiment of this application, such as... Figure 2 As shown, the method includes at least steps S202-S206, wherein:

[0042] Step S202: In response to multiple requests for adding a speaker from multiple targets, determine the target speaker identifier for each target in the speech synthesis request.

[0043] The target speaker identifier mentioned above can be the target speaker ID, denoted as speaker_ids.

[0044] Step S204: Determine the target embedding vector of the target new speaker corresponding to the target speaker identifier from the preset speaker index table.

[0045] Specifically, after obtaining the target speakers for each new speaker, a pre-established speaker index table can be used to obtain the target embedding vector corresponding to each target speaker identifier. This speaker index table sequentially records the first speaker identifier and its corresponding first speaker embedding vector for the first speaker, and the new speaker identifiers and their corresponding second embedding vectors for multiple new speakers. That is, the order of speakers in the speaker index table is: first, the speaker information of the first speaker, and then the speaker information of each new speaker is ordered accordingly. Furthermore, the aforementioned second embedding vector includes: the new speaker embedding vector corresponding to the new speaker and adaptive embedding vector adapter parameters obtained by adaptively training the target acoustic model using the second speech data of the new speakers. The target acoustic model is obtained by combining the preset adapter with the baseline acoustic model trained based on the first speech data of the first speaker.

[0046] As an optional implementation, the training process of the above-mentioned benchmark acoustic model includes the following steps S1-S3, wherein:

[0047] Step S1: Construct a deep learning model.

[0048] The structure of a deep learning model is as follows: Figure 3 As shown, it includes: an encoder consisting of a single-layer FFT (Feed-Forward Transformer) module, a speaker embedding module, and a decoder consisting of N layers of FFT modules, where N is a positive integer greater than or equal to 1.

[0049] in addition, Figure 4 A schematic diagram of an optional FFT module is shown, such as... Figure 4 As shown, the FFT module in this embodiment consists of a multi-head self-attention mechanism module, a block mask module, a conditional layer normalization (CLN) module, a causal convolution module, and a cascaded conditional normalization layer (CLN). During model training, the block mask is used for parallel computation after the multi-head self-attention mechanism module. During inference, it can be decoded in packets and combined with cached keys and values ​​from previous time steps to calculate the current output, achieving a result equivalent to that of parallel inference using the block mask. Furthermore, for streaming inference, the convolutional layers of all FFT modules are replaced with causal convolution modules for parallel training. During inference, the results are combined with cached receptive field results to obtain a result equivalent to that of parallel inference. It should be noted that... Figure 4 The dashed lines represent parts that only participate in training, while the solid lines connect parts that participate in both training and inference.

[0050] Step S2: Obtain the first speech data of multiple first speakers.

[0051] The first speech data includes: the first phoneme sequence encoding corresponding to the first speech text of the first speaker and the first acoustic feature corresponding to the first speech audio, and the first speech text corresponds to the first speech audio.

[0052] Specifically, the technical solution provided in step S2 above allows the first speech data of each first speaker to be acquired in the following manner:

[0053] Step 1: Obtain the initial speech text of each first speaker, and the initial speech audio corresponding to the initial speech text, wherein the initial speech text includes at least one of the following: text information, punctuation marks, and prosodic annotations.

[0054] The above steps can be understood as collecting the initial audio recording of the first speaker and the corresponding initial audio text. The audio text may include, but is not limited to, Chinese, English, sentence-ending punctuation, intra-sentence rhythmic pauses, and pronunciation annotations. For example, the initial audio recording of the first speaker can be recorded in a professional recording studio or in a quiet environment using a mobile phone. Higher audio quality results in better model performance.

[0055] Step 2: Convert the text information in the initial speech text into an initial phoneme sequence, and insert punctuation marks and prosodic annotations into the initial phoneme sequence to obtain the first phoneme sequence.

[0056] The above steps can be understood as follows: converting the text information in the initial speech text into phonemes (such as the initials and finals of Chinese, and the phonetic symbols of English) to form an initial phoneme sequence; then inserting punctuation marks and prosodic pause tags into the corresponding positions between the phonemes in the original initial speech text to form the first phoneme sequence.

[0057] Step 3: Encode the first phoneme sequence using one-hot encoding to obtain the first phoneme sequence code.

[0058] The above steps can be understood as encoding the first phoneme sequence using one-hot encoding to obtain the first phoneme sequence encoding, such as... Figure 5 As shown. One-hot encoding is a single-bit encoding method that uses an N-bit state register to encode N states. Each state has its own independent register bit, and only one bit is active at any given time.

[0059] Step 4: Preprocess the initial audio.

[0060] The preprocessing operations include at least one of the following: sampling, volume adjustment, and cropping.

[0061] The above steps can be understood as processing the initial speech audio, specifically including: converting the initial speech audio into single-channel, 16-bit encoded audio; unifying the volume of the adjusted initial speech audio to -6dB; converting the volume-adjusted initial speech audio into target sampling rate audio, such as 16k sampling rate audio; and trimming the beginning and end silence segments of the initial speech audio to retain a maximum of 100 milliseconds of silence.

[0062] Step 5: Extract features from the preprocessed initial speech audio to obtain the first acoustic features.

[0063] The first acoustic feature includes at least one of the following: Mel-spectral features, frame-level variable features, and phoneme-level duration features. For example, the extracted Mel-spectral features can be 80-dimensional, and the frame-level variable features can be fundamental frequency, energy, etc.

[0064] Step S3: Iteratively train the deep learning model based on the first speech data of multiple first speakers to obtain the baseline acoustic model.

[0065] Specifically, the technical solution provided in step S3 above may further include the following steps:

[0066] Step 1: For the first speech data of each first speaker, the first phoneme sequence of the first speaker is encoded and input into the deep learning model, and then passed through the encoder and decoder in the deep learning model to output the corresponding first streaming acoustic feature and first non-streaming acoustic feature.

[0067] Step 2: Determine a target loss function based on the first acoustic features, the first streaming acoustic features, and the first non-streaming acoustic features within the first speech data of each first speaker, wherein the target loss function includes: non-streaming mean squared error loss function, streaming mean squared error loss function, and adversarial loss function;

[0068] Step 3: Calculate the minimum value of the target loss function using the gradient descent algorithm, and adjust the model parameters of the deep learning model based on the minimum value of the target loss function to obtain the baseline acoustic model.

[0069] In this embodiment, the first phoneme sequence encoding of the first speaker is used as a feature and input into the deep learning model. The encoder encodes the input first phoneme sequence encoding, and the decoder outputs the corresponding first streaming acoustic features and first non-streaming acoustic features. Then, the streaming mean squared error loss function is determined using the acoustic features corresponding to each phoneme encoding and the first streaming acoustic features. The non-streaming mean squared error loss function is determined using the acoustic features corresponding to the entire phoneme sequence encoding and the first non-streaming acoustic features. Furthermore, to improve the sound quality of the streaming synthesized Mel spectrum, an adversarial loss function is added. The details of the adversarial loss function have been disclosed in relevant literature and will not be elaborated further. Therefore, the final target loss function includes the non-streaming mean squared error loss function, the streaming mean squared error loss function, and the adversarial loss function. Finally, the gradient is calculated through backpropagation, and the minimum value of the target loss function is calculated using the gradient descent algorithm. This minimum value is used to update the model parameters of the deep learning model to obtain the baseline acoustic model. In simple terms, the baseline acoustic model is a base model trained using a multi-speaker voice database, and its overall architecture is based on the FastSpeech architecture.

[0070] Furthermore, after obtaining the baseline acoustic model, a target acoustic model can be obtained by combining the baseline acoustic model with a preset adapter, and an adaptive embedding vector can be obtained by adaptively training the target acoustic model using the second speech data of the newly added speaker. The specific implementation is as follows:

[0071] First, the second speech data of the newly added speaker is acquired. This second speech data includes: the second phoneme sequence encoding corresponding to the second speech text of the newly added speaker and the second acoustic features corresponding to the second speech audio, with the second speech text and the second speech audio corresponding to each other. The process of acquiring the second speech data can refer to the steps for acquiring the first speech data described above; therefore, this process will not be elaborated upon here.

[0072] Then, a baseline acoustic model is obtained, and the decoder within the baseline acoustic model is improved using an adapter to obtain the target acoustic model, such as... Figure 6 As shown.

[0073] The structure of the aforementioned adapter is as follows: Figure 7 As shown, the adapter includes: a dimensionality-reduced first feedforward neural network (Linear) down ), activation layer (ReLU), and second feedforward neural network of increased dimensionality (Linear) up The Dropout layer and the Dropout layer are cascaded together.

[0074] In the construction of the target acoustic model, to accelerate its convergence, the baseline acoustic model and adapters can be initialized first. For the baseline acoustic model, the speaker embedding module can be initialized using a set of first speaker embedding vectors, and the conditional layer normalization (CLN) can be initialized using the layer parameters trained on the baseline acoustic model. For the adapters, the first weight matrix of the first feedforward neural network and the second weight matrix of the second feedforward neural network can be initialized to 1, and the first bias vector of the first feedforward neural network and the second bias vector of the second feedforward neural network can be initialized to 0. After initialization, an initialized adapter is added after each of the N-layer FFT modules of the decoder within the baseline acoustic model, resulting in an improved decoder. Each FFT module consists of a multi-head attention mechanism module, a block masking module, a causal convolution module, and a conditional layer normalization module.

[0075] Finally, the target acoustic model is trained using the second speech data of the newly added speaker to obtain adaptive parameters for the new speaker. These adaptive parameters are then concatenated to obtain the adaptive embedding vector corresponding to the new speaker. The adaptive parameters include at least one of the following: the first weight matrix of the first feedforward neural network of the N-layer adapter (denoted as...). ) and the first bias vector (denoted as The second weight matrix of the second feedforward neural network (denoted as...) ) and the second bias vector (denoted as ).

[0076] In the above process, the adaptive parameters are concatenated to obtain the adaptive embedding vector corresponding to the new speaker. This can be done according to the principle of "from layer 1 to N, first Linear..." down Linear up The order of "weight matrix first, then bias vector" is flattened sequentially. That is, in the adapter parameters corresponding to the new speaker, from the first adapter to the Nth adapter, the first weight matrix and the first bias vector of the first feedforward neural network in each adapter, and the second weight matrix and the second bias vector of the second feedforward neural network are horizontally concatenated to obtain the adaptive embedding vector corresponding to the new speaker. Therefore, the final adaptive embedding vector is a one-dimensional vector.

[0077] Through the above three steps, adaptive training of the target acoustic model can be completed, obtaining the adaptive embedding vector for each new speaker. The adaptive embedding vector for each new speaker, along with the AdaptableParams composed of the new speaker's embedding vector, are saved in a specific format and recorded in the speaker index table. Subsequently, by inputting the speaker's identifier for each speaker, their corresponding AdaptableParams can be obtained.

[0078] Then, the AdaptableParams of each newly added speaker are concatenated according to the order of the speech synthesis request to obtain a target embedding vector group containing adaptive embedding vectors of multiple target newly added speakers and speaker embedding vectors. In this embodiment, the target newly added speaker can be any one of the newly added speakers recorded in the speaker index table.

[0079] Step S206: Based on the target embedding vectors of multiple target speakers, determine the target adapter parameters in the target acoustic model, and recombine the target adapter parameters with the target acoustic model. Then, respond to the speech synthesis requests of multiple target speakers through the combined target acoustic model.

[0080] The above steps can be understood as combining the target embedding vectors of multiple newly added speakers to obtain the target adapter parameters of the target acoustic model that can perform speech synthesis for multiple newly added speakers. By recombinating the target adapter parameters with the target acoustic model, the recombined target acoustic model can respond to the speech synthesis requests of multiple newly added speakers without having to use the unique layer and parameters of each newly added speaker for adaptive model adjustment during speech synthesis inference, thus effectively improving inference efficiency.

[0081] As an optional implementation, in the technical solution provided in step S206 above, the target adapter parameters within the target acoustic model can be obtained through the following steps S2061-S2063, wherein:

[0082] Step S2061: Slice and transform the target adaptive embedding vector of the new speaker for each target to obtain the target adapter parameters corresponding to the N-layer target adapter.

[0083] The purpose of the above slicing is to obtain the adapter parameters corresponding to a certain layer of the adapter, for example, the first feedforward neural network Linear. down The first weight matrix, transformation refers to shape transformation, the purpose of which is to convert the weight matrix in the adapter parameters from one dimension to matrix form.

[0084] Step S2062: For the target adapter parameters corresponding to each layer of target adapter, the first weight matrix, the first bias vector, the second weight matrix, and the second bias vector of the first feedforward neural network corresponding to the multiple target speakers are concatenated to obtain the first target weight matrix, the first target bias vector, the second target weight matrix, and the second target bias vector.

[0085] Step S2063: The target adapter parameters of the target acoustic model are obtained from the target speaker embedding vectors corresponding to the multiple target speakers, the first target weight matrix, the first target bias vector, the second target weight matrix, and the second target bias vector corresponding to each layer of target adapters.

[0086] For example, taking any one of the N layers in the decoder's adapter that supports batch input from multiple speakers as an example, the calculation process for encoding the input sequence by each layer in the adapter is explained, such as... Figure 8 As shown, with Figure 8 The first feedforward neural network of the adapter on the left (Linear) down Taking the calculation process of (dimensions 384, 16) as an example, the calculation process is as follows: Figure 8 The right side. The process consists of the following steps:

[0087] First, based on the B targets, add the speaker IDs from the adapter embedding vector. (i.e., the adapter embedding vector, which is formed by horizontally concatenating the adaptive embedding vectors of the M newly added speakers in the speaker index table), obtain the adapter vector of each target new speaker (which includes the adapter-related parameters corresponding to each new speaker), and obtain the speaker embedding vector corresponding to the IDs of the target new speakers.

[0088] Then, based on the adapter's Iid, the corresponding Linear is obtained from the adapter vectors of the B target new speakers. down The first weight matrix and the first bias vector are concatenated to obtain the first target weight matrix. and the first target bias vector The first target weight matrix and the first target bias vector mentioned above are the standard parameters of the adapter;

[0089] Finally, based on the B targets, add speaker IDs from the speaker embedding vector. Find the speaker embedding vector corresponding to each newly added speaker in the target, where S represents the number of times the BaseModel is the first speaker. Then, use B speaker embedding vectors... and the first target weight matrix The product of the first target bias vector and the first target bias vector The results are added together and their shapes are transformed to obtain the output of the adapter.

[0090] Based on the scheme defined in steps S202 to S206 above, it can be understood that, in the embodiment, in response to the speech synthesis requests of multiple target new speakers, the target speaker identifier of each target new speaker in the speech synthesis request is determined; the target embedding vector of the target new speaker corresponding to the target speaker identifier is determined from the preset speaker index table, wherein the target embedding vector includes: speaker embedding vector and adaptive embedding vector, the speaker index table sequentially records the first speaker identifier and the corresponding first speaker embedding vector of the first speaker, the new speaker identifiers of multiple new speakers and the corresponding second embedding vectors, the second embedding vector includes: the new speaker embedding vector corresponding to the new speaker and the adaptive embedding vector adapter parameters obtained by adaptively training the target acoustic model using the second speech data of the new speaker, the target acoustic model is obtained by combining the preset adapter with the baseline acoustic model trained based on the first speech data of the first speaker; the target adapter parameters in the target acoustic model are determined based on the target embedding vectors of multiple target new speakers, and the target adapter parameters are recombined with the target acoustic model, and the combined target acoustic model is used to respond to the speech synthesis requests of multiple target speakers.

[0091] Therefore, the streaming acoustic model (i.e., the target acoustic model) supporting speaker expansion in this application embodiment possesses a streamlined timbre adaptation module and an optimized streaming training method. Specifically, by employing a joint streaming and non-streaming training approach, the synthesized audio exhibits more natural prosody compared to training a streaming model alone. Simultaneously, with the aid of an adversarial loss function, the sound quality of the streaming synthesized audio is clearer, especially with a significant improvement in high-frequency clarity. Furthermore, the streamlined speaker adaptation module used in the model can replicate the voice of a new speaker using a small number of parameters. Moreover, in multi-speaker batch input scenarios, the solution in this application embodiment avoids introducing loop operations by merging parameters and constructing a multi-speaker adaptive structure. This solves the technical problem of low inference processing efficiency when synthesizing speech text from multiple speakers in batches using related acoustic models, where the adaptive model parameters of each speaker are called sequentially for inference of each speaker's speech text.

[0092] Example 2

[0093] Based on Embodiment 1 of this application, an embodiment of a scalable speaker acoustic model implementation device is also provided. This device, when running, executes the scalable speaker acoustic model implementation method of the above embodiment. Wherein, Figure 9 This is a schematic diagram of the structure of an optional scalable acoustic model implementation device for a speaker according to an embodiment of this application, as shown below. Figure 9As shown, the scalable speaker acoustic model implementation device includes at least: a first determining module 91, a second determining module 93, and a model generation module 95, wherein:

[0094] The first module 91 is used to respond to multiple target new speaker requests for speech synthesis and determine the target speaker identifier of each target new speaker in the speech synthesis request;

[0095] The second determining module 93 is used to determine the target embedding vector of the target new speaker corresponding to the target speaker identifier from a preset speaker index table. The target embedding vector includes a speaker embedding vector and an adaptive embedding vector. The speaker index table records the first speaker identifier of the first speaker and the corresponding first speaker embedding vector, the new speaker identifiers of multiple new speakers and the corresponding second embedding vectors in sequence. The second embedding vector includes the new speaker embedding vector corresponding to the new speaker and the adaptive embedding vector adapter parameters obtained by adaptively training the target acoustic model using the second speech data of the new speaker. The target acoustic model is obtained by combining the preset adapter with the baseline acoustic model trained based on the first speech data of the first speaker.

[0096] The model generation module 95 is used to determine the target adapter parameters in the target acoustic model based on the target embedding vectors of multiple target speakers, and to recombine the target adapter parameters with the target acoustic model. The combined target acoustic model is then used to respond to the speech synthesis requests of multiple target speakers.

[0097] It should be noted that the modules in the above-mentioned scalable speaker acoustic model implementation device can be program modules (e.g., a set of program instructions to implement a certain function) or hardware modules. For the latter, they can be in the following forms, but are not limited to these: each of the above modules is in the form of a processor, or the functions of each of the above modules are implemented by a processor.

[0098] Example 3

[0099] According to an embodiment of this application, a non-volatile storage medium is also provided, which stores a program, wherein when the program runs, it controls the device where the non-volatile storage medium is located to execute the acoustic model implementation method of the scalable speaker in Embodiment 1.

[0100] Optionally, the device containing the non-volatile storage medium performs the following steps by running this program:

[0101] Step S202: In response to multiple target addition of speech synthesis requests, determine the target speech synthesizer identifier for each target addition of speech synthesis request;

[0102] Step S204: Determine the target embedding vector of the target new speaker corresponding to the target speaker identifier from the preset speaker index table. The target embedding vector includes: speaker embedding vector and adaptive embedding vector. The speaker index table records the first speaker identifier and the corresponding first speaker embedding vector of the first speaker, the new speaker identifiers of multiple new speakers and the corresponding second embedding vectors in sequence. The second embedding vector includes: the new speaker embedding vector corresponding to the new speaker and the adaptive embedding vector adapter parameters obtained by adaptively training the target acoustic model using the second speech data of the new speaker. The target acoustic model is obtained by combining the preset adapter with the baseline acoustic model trained based on the first speech data of the first speaker.

[0103] Step S206: Based on the target embedding vectors of multiple target speakers, determine the target adapter parameters in the target acoustic model, and recombine the target adapter parameters with the target acoustic model. Then, respond to the speech synthesis requests of multiple target speakers through the combined target acoustic model.

[0104] According to an embodiment of this application, a processor is also provided for running a program, wherein the program executes the scalable speaker acoustic model implementation method of embodiment 1 during runtime.

[0105] Optionally, the program executes the following steps during runtime:

[0106] Step S202: In response to multiple target addition of speech synthesis requests, determine the target speech synthesizer identifier for each target addition of speech synthesis request;

[0107] Step S204: Determine the target embedding vector of the target new speaker corresponding to the target speaker identifier from the preset speaker index table. The target embedding vector includes: speaker embedding vector and adaptive embedding vector. The speaker index table records the first speaker identifier and the corresponding first speaker embedding vector of the first speaker, the new speaker identifiers of multiple new speakers and the corresponding second embedding vectors in sequence. The second embedding vector includes: the new speaker embedding vector corresponding to the new speaker and the adaptive embedding vector adapter parameters obtained by adaptively training the target acoustic model using the second speech data of the new speaker. The target acoustic model is obtained by combining the preset adapter with the baseline acoustic model trained based on the first speech data of the first speaker.

[0108] Step S206: Based on the target embedding vectors of multiple target speakers, determine the target adapter parameters in the target acoustic model, and recombine the target adapter parameters with the target acoustic model. Then, respond to the speech synthesis requests of multiple target speakers through the combined target acoustic model.

[0109] According to an embodiment of this application, an electronic device is also provided, wherein the electronic device includes one or more processors; a memory for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors are configured to run the programs, wherein the programs are configured to execute the scalable speaker acoustic model implementation method in Embodiment 1 above.

[0110] Optionally, the processor is configured to execute the following steps via a computer program:

[0111] Step S202: In response to multiple target addition of speech synthesis requests, determine the target speech synthesizer identifier for each target addition of speech synthesis request;

[0112] Step S204: Determine the target embedding vector of the target new speaker corresponding to the target speaker identifier from the preset speaker index table. The target embedding vector includes: speaker embedding vector and adaptive embedding vector. The speaker index table records the first speaker identifier and the corresponding first speaker embedding vector of the first speaker, the new speaker identifiers of multiple new speakers and the corresponding second embedding vectors in sequence. The second embedding vector includes: the new speaker embedding vector corresponding to the new speaker and the adaptive embedding vector adapter parameters obtained by adaptively training the target acoustic model using the second speech data of the new speaker. The target acoustic model is obtained by combining the preset adapter with the baseline acoustic model trained based on the first speech data of the first speaker.

[0113] Step S206: Based on the target embedding vectors of multiple target speakers, determine the target adapter parameters in the target acoustic model, and recombine the target adapter parameters with the target acoustic model. Then, respond to the speech synthesis requests of multiple target speakers through the combined target acoustic model.

[0114] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.

[0115] In the above embodiments of this application, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0116] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units can be a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual couplings, direct couplings, or communication connections may be through some interfaces; indirect couplings or communication connections between units or modules may be electrical or other forms.

[0117] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0118] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0119] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to related technologies, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard drive, magnetic disk, or optical disk.

[0120] The above are merely preferred embodiments of this application. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of this application, and these improvements and modifications should also be considered within the scope of protection of this application.

Claims

1. A method for implementing an acoustic model of a speaker that is extensible, the method comprising: include: In response to multiple speech synthesis requests for adding a target speaker, the target speaker identifier for each of the target added speakers in the speech synthesis request is determined; The target embedding vector of the new target speaker corresponding to the target speaker identifier is determined from a preset speaker index table. The target embedding vector includes a speaker embedding vector and an adaptive embedding vector. The speaker index table sequentially records the first speaker identifier and the corresponding first speaker embedding vector of the first speaker, the new speaker identifiers of multiple new speakers and their corresponding second embedding vectors. The second embedding vector includes the new speaker embedding vector corresponding to the new speaker and the adaptive embedding vector adapter parameters obtained by adaptively training the target acoustic model using the second speech data of the new speaker. The target acoustic model is obtained by combining the preset adapter with the baseline acoustic model trained based on the first speech data of the first speaker. Based on the target embedding vectors of multiple target speakers, the target adapter parameters within the target acoustic model are determined, and the target adapter parameters are recombined with the target acoustic model. The combined target acoustic model then responds to the speech synthesis requests of multiple target speakers.

2. The method according to claim 1, characterized in that, The training process of the baseline acoustic model includes: Construct a deep learning model, wherein the deep learning model includes: an encoder composed of a single-layer feedforward transformer (FFT) module, a speaker embedding module, and a decoder composed of N layers of the FFT module, where N is a positive integer greater than or equal to 1; Acquire first speech data of multiple first speakers, wherein the first speech data includes: a first phoneme sequence encoding corresponding to the first speech text of the first speaker and a first acoustic feature corresponding to the first speech audio, and the first speech text corresponds to the first speech audio; The deep learning model is iteratively trained based on the first speech data of multiple first speakers to obtain the baseline acoustic model.

3. The method according to claim 2, characterized in that, Acquire first speech data from multiple first speakers, including: Obtain the initial speech text of each of the first speakers, and the initial speech audio corresponding to the initial speech text, wherein the initial speech text includes at least one of the following: text information, punctuation marks, and prosodic annotations; The text information in the initial speech text is converted into an initial phoneme sequence, and the punctuation marks and prosodic annotations are inserted into the initial phoneme sequence to obtain the first phoneme sequence; The first phoneme sequence is encoded using one-hot encoding to obtain the first phoneme sequence encoding; The initial speech audio is preprocessed, wherein the preprocessing operation includes at least one of the following: sampling, volume adjustment, and trimming; Feature extraction is performed on the preprocessed initial speech audio to obtain the first acoustic feature, wherein the first acoustic feature includes at least one of the following: Mel spectrum feature, frame-level variable feature, and phoneme-level duration feature.

4. The method according to claim 2, characterized in that, The deep learning model is iteratively trained based on first speech data from multiple first speakers to obtain the baseline acoustic model, including: For each first speech data of the first speaker, the first phoneme sequence of the first speaker is encoded and input into the deep learning model, and then passed through the encoder and decoder in the deep learning model to output the first streaming acoustic feature and the first non-streaming acoustic feature. A target loss function is determined based on the first acoustic features, the first streaming acoustic features, and the first non-streaming acoustic features within the first speech data of each first speaker. The target loss function includes: a non-streaming mean squared error loss function, a streaming mean squared error loss function, and an adversarial loss function. The minimum value of the target loss function is calculated using the gradient descent algorithm, and the model parameters of the deep learning model are adjusted based on the minimum value of the target loss function to obtain the baseline acoustic model.

5. The method according to claim 2, characterized in that, The adaptive embedding vector obtained by adaptively training the target acoustic model using the second speech data of the newly added speaker includes: The second speech data of the newly added speaker is obtained, wherein the second speech data includes: the second phoneme sequence encoding corresponding to the second speech text of the newly added speaker and the second acoustic feature corresponding to the second speech audio, and the second speech text corresponds to the second speech audio; The reference acoustic model is obtained, and the decoder within the reference acoustic model is improved using the adapter to obtain the target acoustic model. The adapter consists of a first feedforward neural network with reduced dimensionality, an activation layer, and a second feedforward neural network with increased dimensionality. The target acoustic model is trained using the second speech data of the newly added speaker to obtain the adapter parameters of the newly added speaker. The adapter parameters are then concatenated to obtain the adaptive embedding vector corresponding to the newly added speaker. The adapter parameters include at least one of the following: the first weight matrix and the first bias vector of the first feedforward neural network of the N-layer adapter, and the second weight matrix and the second bias vector of the second feedforward neural network.

6. The method according to claim 5, characterized in that, Improving the decoder within the reference acoustic model using the adapter includes: The adapter is initialized, and an initialized adapter is added after the N-layer FFT module of the decoder in the reference acoustic model to obtain the improved decoder. The FFT module consists of a multi-head attention mechanism module, a block module, a conditional layer normalization module, and a causal convolution module.

7. The method according to claim 6, characterized in that, Initializing the adapter includes: The first weight matrix of the first feedforward neural network and the second weight matrix of the second feedforward neural network are initialized to 1, and the first bias vector of the first feedforward neural network and the second bias vector of the second feedforward neural network are initialized to 0.

8. The method according to claim 5, characterized in that, The adapter parameters are concatenated to obtain the adaptive embedding vector corresponding to the newly added speaker, including: The first layer of the adapters to the Nth layer of the adapters corresponding to the new speaker are sequentially concatenated horizontally with the first weight matrix and the first bias vector of the first feedforward neural network and the second weight matrix and the second bias vector of the second feedforward neural network in each adapter to obtain the adaptive embedding vector corresponding to the new speaker.

9. The method according to claim 5, characterized in that, Determining the target adapter parameters within the target acoustic model based on the target embedding vectors of multiple newly added target speakers includes: The target adaptive embedding vector of each newly added speaker for the target is sliced ​​and transformed to obtain the target adapter parameters corresponding to the N-layer target adapter; For the target adapter parameters corresponding to each layer of the target adapter, the first weight matrix, the first bias vector, the second weight matrix, and the second bias vector of the first feedforward neural network corresponding to the multiple target new speakers are concatenated to obtain the first target weight matrix, the first target bias vector, the second target weight matrix, and the second target bias vector. The target adapter parameters of the target acoustic model are obtained from the target speaker embedding vectors corresponding to the multiple target newly added speakers, the first target weight matrix, the first target bias vector, the second target weight matrix, and the second target bias vector corresponding to each layer of the target adapter.

10. A device for realizing an acoustic model of a scalable speaker, characterized in that, include: The first determining module is used to determine the target speaker identifier of each of the target new speakers in the speech synthesis request in response to multiple target new speaker requests; The second determining module is used to determine the target embedding vector of the target new speaker corresponding to the target speaker identifier from a preset speaker index table. The target embedding vector includes a speaker embedding vector and an adaptive embedding vector. The speaker index table sequentially records the first speaker identifier and the corresponding first speaker embedding vector of the first speaker, the new speaker identifiers of multiple new speakers and their corresponding second embedding vectors. The second embedding vector includes the new speaker embedding vector corresponding to the new speaker and the adaptive embedding vector adapter parameters obtained by adaptively training the target acoustic model using the second speech data of the new speaker. The target acoustic model is obtained by combining the preset adapter with the baseline acoustic model trained based on the first speech data of the first speaker. The model generation module is used to determine the target adapter parameters in the target acoustic model based on the target embedding vectors of multiple target new speakers, and to recombine the target adapter parameters with the target acoustic model, and to respond to the speech synthesis requests of multiple target speakers through the combined target acoustic model.

11. A non-volatile storage medium, characterized in that, The non-volatile storage medium stores a computer program, wherein the device containing the non-volatile storage medium executes the acoustic model implementation method of any one of claims 1 to 9 by running the computer program.

12. An electronic device, characterized in that, include: A memory and a processor, the processor being configured to run a program stored in the memory, wherein the program, when running, executes the method for implementing the acoustic model of a scalable speaker as described in any one of claims 1 to 9.

Citation Information

Patent Citations

  • Speech translation method and system using multilingual text-to-speech synthesis model

    CN111566656A

  • Speech synthesis method and device based on pronunciator vector

    CN116884388A