Method and apparatus for implementing speaker-adaptive acoustic model
By constructing a deep learning model and performing adaptive training, adaptive embedding vectors are generated and a pronunciation index table is established, which solves the problem of low inference efficiency of multi-pronunciation speech synthesis model and achieves more efficient speech synthesis processing.
Patent Information
- Application Number
- PCT/CN2024/123878
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-14
- Filing Date
- 2024-10-10
- Publication Date
- 2025-08-07
AI Technical Summary
The existing multipronunciation voice synthesis model is less efficient in inference, especially when multiple pronunciation man input batches, it is necessary to call their own unique model parameters separately, resulting in a decrease in the inference efficiency of computer GPU.
By constructing a deep learning model, including an encoder and a decoder, the target acoustic model is adaptively trained using the speech data of the new pronunciator, an adaptive embedding vector is generated, and a pronunciator index table is established. In response to a speech synthesis request, the embedding vector and adapter parameters of multiple pronunciators are directly combined to form an extensible acoustic model.
It improves the inference efficiency of the speech synthesis model, avoids repeated calls to layers and parameters unique to each pronunciator, and improves the processing efficiency of the computer.
Smart Images

Figure CN2024123878_07082025_PF_FP_ABST
Abstract
Description
Method and device for implementing an acoustic model of an extensible speaker
[0001] This application claims priority to the Chinese patent application filed with the China Patent Office on December 14, 2023, with application number 2023117237013, and application name “Method and device for implementing an acoustic model of an extensible speaker”, the entire contents of which are incorporated by reference into this application. Technical Field
[0002] The present application relates to the field of speech synthesis technology, and more specifically, to a method and device for implementing an acoustic model of an extensible speaker. Background Art
[0003] Speech synthesis technology enables computers (or various terminal devices) to speak like humans. Text-to-speech (TTS) technology is a type of speech synthesis technology that converts computer-generated or externally input text into understandable, fluent spoken output. Speech synthesis systems generally consist of three main components: a text analysis module, an acoustic model, and a vocoder.
[0004] Based on a speech synthesis system, a multi-speaker speech synthesis system can be built by training an acoustic model using mixed data from multiple speakers, then fine-tuning some of the model parameters of the pre-trained acoustic model using a small amount of target speaker data—a process known as "pre-training + fine-tuning." In recent years, adapter-based adaptive technology has achieved significant results in fine-tuning large NLP (Natural Language Processing) models. Existing literature suggests that the characteristics of a new speaker can be learned by simply adding a small number of adapter layers to the base model of the existing speech synthesis acoustic system. However, either approach results in either shared layers and parameters or speaker-specific layers and parameters. Therefore, when the number of speakers in the inference engine of a current multi-speaker speech synthesis model can be greater than one, each speaker must select their own unique layers and parameters for inference. However, since each speaker's parameters remain independent after training, they are still independent.
[0005] To address the above issues, relevant technicians have provided a relatively simple method, which is to convert inputs with a batch size greater than 1 into multiple inputs with a batch size of 1, decode them separately, and then piece the results together into a batch and return them. However, this method will seriously reduce the inference efficiency of the computer GPU.
[0006] To address the above-mentioned problems, no effective solutions have been proposed so far.
[0007] Summary of the Invention
[0008] The embodiments of the present application provide a method and device for implementing an acoustic model that is scalable for speakers, so as to at least solve the technical problem that when multiple speakers of the relevant acoustic model batch input speech text for speech synthesis, the speech text of each speaker is inferred in turn by calling the adaptive model parameters of each speaker, resulting in low inference processing efficiency.
[0009] According to one aspect of an embodiment of the present application, a method for implementing an acoustic model of an extensible speaker is provided, comprising: determining, in response to speech synthesis requests of multiple target new speakers, a target speaker identifier of each target new speaker in the speech synthesis request; determining a target embedding vector of the target new speaker corresponding to the target speaker identifier from a preset speaker index table, wherein the target embedding vector comprises: a speaker embedding vector and an adaptive embedding vector, the speaker index table records in sequence the first speaker identifier and the corresponding first speaker embedding vector of the first speaker, the new speaker identifiers and the corresponding second embedding vectors of multiple new speakers, the second embedding vector comprises: the new speaker embedding vector corresponding to the new speaker and an adaptive embedding vector adapter parameter obtained by adaptively training the target acoustic model using the second speech data of the new speaker, the target acoustic model being obtained by combining a preset adapter with a baseline acoustic model trained based on the first speech data of the first speaker; determining target adapter parameters in the target acoustic model based on the target embedding vectors of the multiple target new speakers, and recombining the target adapter parameters with the target acoustic model, and responding to the speech synthesis requests of the multiple target speakers through the combined target acoustic model.
[0010] In some embodiments, the training process of the baseline acoustic model includes: constructing a deep learning model, wherein the deep learning model includes: an encoder composed of a single-layer feedforward transformer FFT module, a speaker embedding module, and a decoder composed of N layers of FFT modules, where N is a positive integer greater than or equal to 1; obtaining first speech data of multiple first speakers, wherein the first speech data includes: a first phoneme sequence encoding corresponding to the first speech text of the first speaker and a first acoustic feature corresponding to the first speech audio, and the first speech text corresponds to the first speech audio; iteratively training the deep learning model based on the first speech data of multiple first speakers to obtain a baseline acoustic model.
[0011] In some embodiments, obtaining first speech data of multiple first speakers includes: obtaining an initial speech text of each first speaker, and an initial speech audio corresponding to the initial speech text, wherein the initial speech text includes at least one of the following: text information, punctuation marks, and prosodic annotations; converting the text information in the initial speech text into an initial phoneme sequence, and inserting the punctuation marks and prosodic annotations into the initial phoneme sequence to obtain a first phoneme sequence; encoding the first phoneme sequence using one-hot encoding to obtain a first phoneme sequence encoding; performing a preprocessing operation on the initial speech audio, wherein the preprocessing operation includes at least one of the following: sampling, volume adjustment, and cropping; performing feature extraction on the preprocessed initial speech audio to obtain a first acoustic feature, wherein the first acoustic feature includes at least one of the following: Mel spectrum feature, frame-level variable feature, and phoneme-level duration feature.
[0012] In some embodiments, a deep learning model is iteratively trained based on the first speech data of multiple first speakers to obtain a baseline acoustic model, including: for the first speech data of each first speaker, the first phoneme sequence of the first speaker is encoded and input into the deep learning model, and the corresponding first streaming acoustic features and first non-streaming acoustic features are output in sequence through the encoder and decoder in the deep learning model; a target loss function is determined based on the first acoustic features in the first speech data of each first speaker and the first streaming acoustic features and the first non-streaming acoustic features, wherein the target loss function includes: a non-streaming mean square error loss function, a streaming mean square error loss function, and an adversarial loss function; the minimum value of the target loss function is calculated using a gradient descent algorithm, and the model parameters of the deep learning model are adjusted based on the minimum value of the target loss function to obtain a baseline acoustic model.
[0013] In some embodiments, an adaptive embedding vector obtained by adaptively training a target acoustic model using the second speech data of a newly added speaker includes: obtaining the second speech data of the newly added speaker, wherein the second speech data includes: a second phoneme sequence encoding corresponding to the second speech text of the newly added speaker and a second acoustic feature corresponding to the second speech audio, and the second speech text corresponds to the second speech audio; obtaining a baseline acoustic model, and using an adapter to improve the decoder in the baseline acoustic model to obtain a target acoustic model, wherein the adapter consists of a first feedforward neural network with reduced dimensionality, an activation layer, and a second feedforward neural network with increased dimensionality; using the second speech data of the newly added speaker to adaptively train the target acoustic model to obtain adapter parameters of the newly added speaker, and splicing the adapter parameters to obtain an adaptive embedding vector corresponding to the newly added speaker, wherein the adapter parameters include at least one of the following: a first weight matrix and a first bias vector of the first feedforward neural network of the N-layer adapter, and a second weight matrix and a second bias vector of the second feedforward neural network.
[0014] In some embodiments, an adapter is used to improve the decoder in a baseline acoustic model, including: initializing the adapter and adding an initialized adapter after each of the N-layer FFT modules of the decoder in the baseline acoustic model to obtain an improved decoder, wherein the FFT module consists of a multi-head attention mechanism module, a block module, a conditional layer normalization module, and a causal convolution module.
[0015] In some embodiments, initializing the adapter includes initializing a first weight matrix of the first feedforward neural network and a second weight matrix of the second feedforward neural network to 1, and initializing a first bias vector of the first feedforward neural network and a second bias vector of the second feedforward neural network to 0.
[0016] In some embodiments, the adapter parameters are spliced to obtain an adaptive embedding vector corresponding to the newly added speaker, including: connecting the first layer adapter to the Nth layer adapter in the adapter parameters corresponding to the newly added speaker, and horizontally splicing the first weight matrix and the first bias vector of the first feedforward neural network and the second weight matrix and the second bias vector of the second feedforward neural network in each adapter in turn to obtain an adaptive embedding vector corresponding to the newly added speaker.
[0017] In some embodiments, target adapter parameters within a target acoustic model are determined based on target embedding vectors of multiple target new speakers, including: slicing and converting the target adaptive embedding vector of each target new speaker to obtain target adapter parameters corresponding to the corresponding N layers of target adapters; for the target adapter parameters corresponding to each layer of target adapters, the first weight matrix of the first feedforward neural network, the first bias vector of the first feedforward neural network, the second weight matrix of the second feedforward neural network, and the second bias vector of the second feedforward neural network corresponding to the multiple target new speakers are respectively spliced to obtain a first target weight matrix, a first target bias vector, a second target weight matrix, and a second target bias vector; the target adapter parameters of the target acoustic model are obtained from the target speaker embedding vectors corresponding to the multiple target new speakers, the first target weight matrix, the first target bias vector, the second target weight matrix, and the second target bias vector corresponding to each layer of target adapter.
[0018] According to another aspect of the embodiment of the present application, an acoustic model implementation device for an extensible speaker is also provided, including: a first determination module, used to respond to speech synthesis requests of multiple target new speakers, and determine the target speaker identifier of each target new speaker in the speech synthesis request; a second determination module, used to determine the target embedding vector of the target new speaker corresponding to the target speaker identifier from a preset speaker index table, wherein the target embedding vector includes: a speaker embedding vector and an adaptive embedding vector, and the speaker index table records the first speaker identifier and the corresponding first speaker embedding vector of the first speaker, the new speaker identifiers and the corresponding second embedding vector of multiple new speakers in sequence. The input vector, the second embedding vector includes: the new speaker embedding vector corresponding to the new speaker and the adaptive embedding vector adapter parameters obtained by adaptively training the target acoustic model using the second speech data of the new speaker, and the target acoustic model is obtained by combining a preset adapter with a baseline acoustic model trained based on the first speech data of the first speaker; a model generation module is used to determine the target adapter parameters in the target acoustic model based on the target embedding vectors of multiple target new speakers, and recombine the target adapter parameters with the target acoustic model, and respond to the speech synthesis requests of multiple target speakers through the combined target acoustic model.
[0019] According to another aspect of an embodiment of the present application, a non-volatile storage medium is provided, which stores a computer program, wherein the device where the non-volatile storage medium is located executes the above-mentioned method for implementing the acoustic model of an extensible speaker by running the computer program.
[0020] According to another aspect of an embodiment of the present application, an electronic device is provided, which includes: a memory and a processor, wherein a computer program is stored in the memory, and the processor is configured to execute the above-mentioned method for implementing an extensible speaker's acoustic model through the computer program.
[0021] In an embodiment of the present application, in response to speech synthesis requests from multiple target new speakers, the target speaker identifier of each target new speaker in the speech synthesis request is determined; the target embedding vector of the target new speaker corresponding to the target speaker identifier is determined from a preset speaker index table, wherein the target embedding vector includes: a speaker embedding vector and an adaptive embedding vector, and the speaker index table records the first speaker identifier of the first speaker and the corresponding first speaker embedding vector, the new speaker identifiers of multiple new speakers and the corresponding second embedding vectors in sequence, and the second embedding vector includes: the new speaker embedding vector corresponding to the new speaker and the adaptive embedding vector adapter parameters obtained by adaptively training the target acoustic model using the second speech data of the new speaker, and the target acoustic model is obtained by combining a preset adapter with a baseline acoustic model trained based on the first speech data of the first speaker; based on the target embedding vectors of the multiple target new speakers, the target adapter parameters in the target acoustic model are determined, and the target adapter parameters are recombined with the target acoustic model, and the speech synthesis requests of the multiple target speakers are responded to by the combined target acoustic model.
[0022] In an embodiment of the present application, the target acoustic model after adding the adapter is adaptively trained by utilizing the voice data of the newly added speaker to obtain the layers and parameters unique to the newly added speaker, and a speaker index table is established based on this. When subsequently responding to speech synthesis requests from multiple newly added speakers, the corresponding speaker embedding vector can be found based on the speaker identifier of the newly added speaker, and the unique layers and parameters in the multiple newly added speaker embedding vectors are combined to obtain a multi-speaker acoustic model that can handle the speech synthesis requests of each newly added speaker. This eliminates the need to call the layers and parameters unique to the newly added speaker for inference each time speech synthesis is performed for different newly added speakers, thereby eliminating the need for loop operations and improving the computer's inference efficiency. This solves the technical problem of low inference processing efficiency when multiple speakers of the relevant acoustic model batch input speech text for speech synthesis, and the speech text of each speaker is inferred in turn by calling the adaptive model parameters of each speaker. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:
[0024] FIG1 is a hardware structure block diagram of a computer terminal for implementing a method for implementing an acoustic model of an extensible speaker according to an embodiment of the present application;
[0025] FIG2 is a flow chart of an optional method for implementing an acoustic model of an extensible speaker according to an embodiment of the present application;
[0026] FIG3 is a schematic diagram of the structure of an optional deep learning model according to an embodiment of the present application;
[0027] FIG4 is a schematic structural diagram of an optional FFT module according to an embodiment of the present application;
[0028] FIG5 is a flow chart of an optional generation of phoneme sequence encoding according to an embodiment of the present application;
[0029] FIG6 is a schematic diagram of the structure of an optional target acoustic model according to an embodiment of the present application;
[0030] FIG7 is a schematic structural diagram of an optional adapter according to an embodiment of the present application;
[0031] FIG8 is a calculation flow chart of an optional adapter supporting batch input of multiple speakers according to an embodiment of the present application;
[0032] FIG9 is a schematic structural diagram of an optional apparatus for implementing an acoustic model of an extensible speaker according to an embodiment of the present application. DETAILED DESCRIPTION
[0033] In order to enable those skilled in the art to better understand the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of this application.
[0034] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in a sequence other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0035] In addition, the relevant information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for display, data for analysis, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. For example, an interface is set up between this system and the relevant user or organization. Before obtaining relevant information, it is necessary to send an acquisition request to the aforementioned user or organization through the interface, and obtain the relevant information after receiving the consent information fed back by the aforementioned user or organization.
[0036] Example 1
[0037] According to an embodiment of the present application, an embodiment of a method for implementing an acoustic model of an extensible speaker is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0038] The method embodiment provided in the embodiment of the present application can be executed in a mobile terminal, a computer terminal or a similar computing device. Figure 1 shows a hardware structure block diagram of a computer terminal for realizing an acoustic model implementation method for an extensible speaker. As shown in Figure 1, the computer terminal 10 may include one or more (102a, 102b, ..., 102n are used to illustrate) processors 102 (the processor 102 may include but is not limited to a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 104 for storing data, and a transmission device 106 for a communication function. In addition, it may also include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of the BUS bus), a network interface, a power supply and / or a camera. Those skilled in the art will appreciate that the structure shown in Figure 1 is only illustrative and does not limit the structure of the above-mentioned electronic device. For example, the computer terminal 10 may also include more or fewer components than shown in Figure 1, or have a configuration different from that shown in Figure 1.
[0039] It should be noted that the one or more processors 102 and / or other data processing circuits described above may generally be referred to herein as "data processing circuitry." The data processing circuitry may be embodied in whole or in part as software, hardware, firmware, or any other combination thereof. Furthermore, the data processing circuitry may be a single, independent processing module, or may be incorporated in whole or in part into any of the other components of the computer terminal 10. As described in the embodiments of the present application, the data processing circuitry serves as a processor control (e.g., selection of a variable resistor terminal path connected to an interface).
[0040] Memory 104 can be used for storing software programs and modules of application software, such as the program instruction / data storage device corresponding to the acoustic model implementation method for realizing extensible pronunciation people in the embodiment of the present application, processor 102 performs various functional applications and data processing by running the software programs and modules stored in the memory 104, namely realizes the acoustic model implementation method for realizing extensible pronunciation people of above-mentioned application program. Memory 104 can include high-speed random access memory, and can also include non-volatile memory, such as one or more magnetic storage devices, flash memory or other non-volatile solid-state memories. In some instances, memory 104 can further include a memory remotely arranged relative to processor 102, and these remote memories can be connected to computer terminal 10 through a network. The example of above-mentioned network includes but is not limited to the Internet, intranet, local area network, mobile communication network and combination thereof.
[0041] The transmission device 106 is configured to receive or transmit data via a network. A specific example of the aforementioned network may include a wireless network provided by the communications provider of the computer terminal 10. In one embodiment, the transmission device 106 includes a network interface controller (NIC), which can be connected to other network devices via a base station to enable communication with the Internet. In another embodiment, the transmission device 106 may be a radio frequency (RF) module, which is configured to communicate with the Internet wirelessly.
[0042] The display may be, for example, a touch screen liquid crystal display (LCD) that enables a user to interact with a user interface of the computer terminal 10 .
[0043] In the above operating environment, FIG2 is a schematic diagram of an optional method for implementing an acoustic model of an extensible speaker according to an embodiment of the present application. As shown in FIG2 , the method includes at least steps S202-S206, wherein:
[0044] Step S202 : In response to speech synthesis requests from multiple target new speakers, determine a target speaker identifier of each target new speaker in the speech synthesis request.
[0045] The target speaker identifier may be a target speaker ID, which is denoted as speaker_ids.
[0046] Step S204 : determining a target embedding vector of a target newly added speaker corresponding to the target speaker identifier from a preset speaker index table.
[0047] In some embodiments, after obtaining the target speaker of each target newly added speaker, a pre-established speaker index table can be called to obtain the target embedding vector corresponding to each target speaker identifier, wherein the speaker index table records the first speaker identifier and the corresponding first speaker embedding vector of the first speaker, the newly added speaker identifiers and the corresponding second embedding vectors of multiple newly added speakers in sequence, that is, the order of each speaker in the speaker index table is the speaker information of the first speaker first, and then the speaker information of each newly added speaker is sorted in this way. In addition, the above-mentioned second embedding vector includes: the newly added speaker embedding vector corresponding to the newly added speaker and the adaptive embedding vector adapter parameter obtained by adaptively training the target acoustic model using the second speech data of the newly added speaker, wherein the target acoustic model is obtained by combining the preset adapter with the baseline acoustic model trained based on the first speech data of the first speaker.
[0048] As an optional implementation, the training process of the above-mentioned reference acoustic model includes the following steps S1-S3, wherein:
[0049] Step S1: Build a deep learning model.
[0050] Among them, the structure of the deep learning model is shown in Figure 3, which includes: an encoder composed of a single-layer FFT (Feed-Forward Transformer) module, a speaker embedding module, and a decoder composed of N-layer FFT modules, where N is a positive integer greater than or equal to 1.
[0051] In addition, Figure 4 shows a structural diagram of an optional FFT module. As shown in Figure 4, the FFT module in the embodiment of the present application is composed of a multi-head self-attention mechanism module, a block mask module, a conditional layer normalization (CLN), a causal convolution module, and a conditional normalization layer (CLN) cascade. In model training, block mask parallel calculation is used after the multi-head self-attention mechanism module. During reasoning, the decoding can be packetized and combined with the cached keys and values at the previous moment to calculate the current output, and the result equivalent to the block mask parallel reasoning can be obtained. In addition, for streaming reasoning, the convolution layers of all FFT modules are replaced with causal convolution modules, and trained in parallel. During reasoning, the calculation is combined with the cached receptive field results to obtain a result equivalent to parallel reasoning. It should be noted that the dotted part in Figure 4 represents only participation in training, and the solid line connection part participates in training and reasoning.
[0052] Step S2: Acquire first speech data of multiple first speakers.
[0053] The first voice data includes: a first phoneme sequence code corresponding to a first voice text of a first speaker and a first acoustic feature corresponding to a first voice audio, and the first voice text corresponds to the first voice audio.
[0054] In some embodiments, in the technical solution provided in step S2 above, the first voice data of each first speaker can be obtained in the following manner:
[0055] Step 1: Acquire the initial speech text of each first speaker and the initial speech audio corresponding to the initial speech text, wherein the initial speech text includes at least one of the following: text information, punctuation marks, and prosody annotations.
[0056] The above steps can be understood as collecting the initial speech audio of the first speaker and the initial speech text corresponding to the initial speech audio, where the speech text may include, but is not limited to, Chinese, English, sentence-end punctuation, prosodic pauses, and pronunciation annotation information. For example, the initial speech audio of the first speaker can be recorded in a professional recording studio or with a mobile phone in a quiet environment. The higher the quality of the recorded audio, the better the model effect.
[0057] Step 2: Convert the text information in the initial speech text into an initial phoneme sequence, and insert punctuation marks and prosody annotations into the initial phoneme sequence to obtain a first phoneme sequence.
[0058] The above steps can be understood as converting the text information in the initial speech text into phonemes (such as Chinese vowels and English phonetic symbols) to form an initial phoneme sequence, and then inserting punctuation marks and prosodic pause labels into the corresponding positions between the phonemes corresponding to the original initial speech text to form a first phoneme sequence.
[0059] Step 3: Encode the first phoneme sequence using one-hot encoding to obtain the first phoneme sequence encoding.
[0060] The above steps can be understood as encoding the first phoneme sequence through one-hot encoding to obtain the first phoneme sequence encoding, as shown in Figure 5. Among them, one-hot encoding is a single-bit effective encoding, which mainly uses an N-bit state register to encode N states. Each state has its own independent register bit, and only one bit is valid at any time.
[0061] Step 4: Preprocess the initial speech audio.
[0062] The pre-processing operation includes at least one of the following: sampling, volume adjustment, and clipping.
[0063] The above steps can be understood as processing the initial voice audio, specifically including: converting the initial voice audio into single-channel, 16-bit encoded audio; unifying the volume of the adjusted initial voice audio to -6db; converting the initial voice audio after volume adjustment to target sampling rate audio, such as 16k sampling rate audio; cropping the silence segments at the beginning and end of the initial voice audio so that a maximum of 100 milliseconds of silence is retained.
[0064] Step 5: Extract features from the preprocessed initial speech audio to obtain the first acoustic feature.
[0065] The first acoustic feature includes at least one of the following: a mel-spectrogram feature, a frame-level variable feature, and a phoneme-level duration feature. For example, the extracted mel-spectrogram feature may be 80-dimensional, and the frame-level variable feature may be fundamental frequency, energy, etc.
[0066] Step S3: Iteratively train the deep learning model based on the first speech data of multiple first speakers to obtain a benchmark acoustic model.
[0067] In some embodiments, in the technical solution provided in step S3 above, the method may further include the following steps:
[0068] Step 1: For each first speaker's first speech data, encode the first speaker's first phoneme sequence and input it into the deep learning model. The encoder and decoder in the deep learning model sequentially output the corresponding first streaming acoustic features and first non-streaming acoustic features.
[0069] Step 2: determining a target loss function based on the first acoustic feature in the first speech data of each first speaker, the first streaming acoustic feature, and the first non-streaming acoustic feature, wherein the target loss function includes: a non-streaming mean square error loss function, a streaming mean square error loss function, and an adversarial loss function;
[0070] Step 3: Use the gradient descent algorithm to calculate the minimum value of the target loss function, and adjust the model parameters of the deep learning model based on the minimum value of the target loss function to obtain the baseline acoustic model.
[0071] In this embodiment, the first phoneme sequence code of the first speaker is used as a feature and input into the deep learning model. The input first phoneme sequence code is encoded by the encoder, and then the decoder outputs the corresponding first streaming acoustic feature and the first non-streaming acoustic feature. Then, the acoustic features corresponding to each phoneme code and the first streaming acoustic feature are used to determine the streaming mean square error loss function, and the acoustic features corresponding to the entire phoneme sequence code and the first non-streaming acoustic feature are used to determine the non-streaming mean square error loss function. In addition, in order to improve the sound quality of the streaming synthesized mel spectrum, an adversarial loss function is added. The content of the adversarial loss function has been disclosed in relevant literature and will not be explained in detail. Therefore, the final target loss function includes three parts: the non-streaming mean square error loss function, the streaming mean square error loss function, and the adversarial loss function. Finally, the gradient is calculated by backpropagation, and the gradient descent algorithm is used to calculate the minimum value of the target loss function. The minimum value is used to update the model parameters of the deep learning model to obtain a baseline acoustic model. Simply put, the baseline acoustic model is a base model trained using a multi-speaker voice library, and its overall architecture is based on the FastSpeech architecture.
[0072] Furthermore, after obtaining the baseline acoustic model, the target acoustic model can be obtained by combining the baseline acoustic model with a preset adapter, and the adaptive embedding vector obtained by adaptively training the target acoustic model using the second speech data of the newly added speaker. The specific implementation method is as follows:
[0073] First, second voice data of a newly added speaker is obtained, where the second voice data includes: a second phoneme sequence code corresponding to a second voice text of the newly added speaker and a second acoustic feature corresponding to a second voice audio, and the second voice text corresponds to the second voice audio. The process of obtaining the second voice data can refer to the steps for obtaining the first voice data described above, and therefore, this process is not further described here.
[0074] Then, a baseline acoustic model is obtained, and the adapter is used to improve the decoder in the baseline acoustic model to obtain the target acoustic model, as shown in FIG6 .
[0075] The structure of the adapter is shown in FIG7 , which includes: a first feedforward neural network (Linear down ), activation layer (ReLU), the second feedforward neural network with increased dimension (Linear up ) and Dropout layers are cascaded in sequence.
[0076] In the process of constructing the above-mentioned target acoustic model, in order to accelerate the rapid convergence of the target acoustic model, the baseline acoustic model and the adapter can be initialized first, wherein, for the initialization of the baseline acoustic model, the speaker embedding module can be initialized using a set of first speaker embedding vectors of the first speaker, and the initialization of the conditional layer normalization (CLN) can use the layer parameters trained by the baseline acoustic model. For initialization of the adapter, the first weight matrix of the first feedforward neural network and the second weight matrix of the second feedforward neural network can be initialized to 1, and the first bias vector of the first feedforward neural network and the second bias vector of the second feedforward neural network can be initialized to 0. After the initialization is completed, an initialized adapter is added after each of the N-layer FFT modules of the decoder in the baseline acoustic model to obtain an improved decoder, wherein each FFT module is composed of a multi-head attention mechanism module, a block mask module, a causal convolution module, and a conditional layer normalization module.
[0077] Finally, the target acoustic model is adaptively trained using the second speech data of the newly added speaker to obtain the adaptive parameters of the newly added speaker, and the adaptive parameters are concatenated to obtain the adaptive embedding vector corresponding to the newly added speaker, wherein the adaptive parameters include at least one of the following: the first weight matrix of the first feedforward neural network of the N-layer adapter (denoted as ) and the first bias vector (denoted as ), the second weight matrix of the second feedforward neural network (denoted as ) and the second bias vector (denoted as ).
[0078] In the above process, the adaptive parameters are concatenated to obtain the adaptive embedding vector corresponding to the newly added speaker. The adaptive embedding vector can be obtained by “from 1 to N layers, first linear down Linear up , first the weight matrix and then the bias vector. That is, from the first layer adapter to the Nth layer adapter in the adapter parameters corresponding to the newly added speaker, the first weight matrix and the first bias vector of the first feedforward neural network, and the second weight matrix and the second bias vector of the second feedforward neural network in each adapter are horizontally spliced in turn to obtain the adaptive embedding vector corresponding to the newly added speaker. Therefore, the final adaptive embedding vector is a one-dimensional vector.
[0079] Through these three steps, adaptive training of the target acoustic model is completed, and adaptive embedding vectors for each newly added speaker are obtained. The AdaptableParams consisting of the adaptive embedding vectors for each newly added speaker and the newly added speaker embedding vector are saved as a file in a specific format and recorded in the speaker index table. Subsequently, by entering the speaker ID for each speaker, the corresponding AdaptableParams can be obtained.
[0080] Then, the AdaptableParams of each newly added speaker are concatenated in the order of the speech synthesis request to obtain a target embedding vector group containing adaptive embedding vectors and speaker embedding vectors for multiple target newly added speakers. In the embodiment of the present application, the target newly added speaker can be any one of the newly added speakers recorded in the speaker index table.
[0081] Step S206, determining target adapter parameters within the target acoustic model based on the target embedding vectors of the multiple target newly added speakers, and recombining the target adapter parameters with the target acoustic model, and responding to the speech synthesis requests of the multiple target speakers through the combined target acoustic model.
[0082] The above steps can be understood as combining the target embedding vectors of multiple target new speakers to obtain the target adapter parameters of the target acoustic model that can realize speech synthesis for multiple target new speakers. By recombining the target adapter parameters with the target acoustic model, the recombined target acoustic model can respond to the speech synthesis requests of multiple target new speakers. There is no need to use the unique layers and parameters of the target new speaker for model adaptive adjustment during speech synthesis inference for each target new speaker, thereby effectively improving the inference efficiency.
[0083] As an optional implementation, in the technical solution provided in the above step S206, the target adapter parameters in the target acoustic model can be obtained through the following steps S2061-S2063, where:
[0084] Step S2061 , slicing and transforming the target adaptive embedding vector of each target newly added speaker to obtain target adapter parameters corresponding to the corresponding N-layer target adapter.
[0085] The purpose of the above slicing is to obtain the adapter parameters corresponding to a certain layer of adapters, for example, the first feedforward neural network Linear down The first weight matrix, conversion refers to the shape conversion, which aims to convert the weight matrix within the adapter parameters from one-dimensional to matrix form.
[0086] Step S2062, for the target adapter parameters corresponding to each layer of the target adapter, the first weight matrix of the first feedforward neural network, the first bias vector of the first feedforward neural network, the second weight matrix of the second feedforward neural network, and the second bias vector of the second feedforward neural network corresponding to multiple target newly added speakers are spliced to obtain the first target weight matrix, the first target bias vector, the second target weight matrix, and the second target bias vector.
[0087] Step S2063, obtain the target adapter parameters of the target acoustic model from the target speaker embedding vectors corresponding to multiple target newly added speakers, the first target weight matrix, the first target bias vector, the second target weight matrix, and the second target bias vector corresponding to each layer of the target adapter.
[0088] For example, taking any one of the N layers of adapters in the decoder that supports batch input of multiple speakers as an example, the computational process of encoding the input sequence in each layer of the adapter is described. As shown in FIG8 , the first feedforward neural network Linear down (dimensions are 384, 16) as an example, the calculation process is shown on the right side of Figure 8. The process is divided into the following steps:
[0089] First, the IDs of the new speakers are added from the adapter embedding vector according to the B targets Get the adapter vector of each target new speaker (which includes the adapter-related parameters corresponding to each new speaker) from the adaptive embedding vector of the M new speakers in the speaker index table (i.e., the adapter embedding vector composed of the horizontal splicing of the adaptive embedding vectors of the M new speakers in the speaker index table), and get the speaker embedding vector corresponding to the IDs of the target new speakers
[0090] Then, according to the Iid of the adapter, the corresponding Linear down The first weight matrix and the first bias vector of B target newly added speakers are concatenated to obtain the first target weight matrix and the first target bias vector The first target weight matrix and the first target bias vector are the standard parameters of the adapter;
[0091] Finally, according to the IDs of the B target speakers, the speaker embedding vector is added. Find the speaker embedding vector corresponding to each target new speaker, where S represents the number of BaseModels that are the first speaker. And through B speaker embedding vectors and the first target weight matrix The product of the first target bias vector The output of the adapter is obtained by adding and changing the shape.
[0092] Based on the scheme defined by the above steps S202 to S206, it can be known that in an embodiment, in response to speech synthesis requests of multiple target new speakers, the target speaker identifier of each target new speaker in the speech synthesis request is determined; the target embedding vector of the target new speaker corresponding to the target speaker identifier is determined from a preset speaker index table, wherein the target embedding vector includes: a speaker embedding vector and an adaptive embedding vector, and the speaker index table records the first speaker identifier of the first speaker and the corresponding first speaker embedding vector, the new speaker identifiers of multiple new speakers and the corresponding second embedding vectors in sequence, and the second embedding vector includes: the new speaker embedding vector corresponding to the new speaker and the adaptive embedding vector adapter parameters obtained by adaptively training the target acoustic model using the second speech data of the new speaker, and the target acoustic model is obtained by combining the preset adapter with the baseline acoustic model trained based on the first speech data of the first speaker; the target adapter parameters in the target acoustic model are determined based on the target embedding vectors of the multiple target new speakers, and the target adapter parameters are recombined with the target acoustic model, and the speech synthesis requests of the multiple target speakers are responded to by the combined target acoustic model.
[0093] It can be seen that in the technical scheme of the embodiment of the present application, the streaming acoustic model (i.e., the target acoustic model) that supports speaker expansion has a streamlined timbre adaptation module and an optimized streaming training method. Among them, by adopting a streaming and non-streaming joint training method, the rhythm of the synthesized audio is more natural compared to training the streaming model alone. At the same time, with the help of the adversarial loss function, the sound quality of the streaming synthesized audio is clearer, and the clarity of the especially high frequencies is significantly improved. In addition, the streamlined speaker adaptation module used by the model can achieve the sound reproduction of the new speaker using a small amount of parameters. In addition, by the embodiment of the present application, in a multi-speaker batch input scenario, by merging parameters and constructing a multi-speaker adaptive structure, it is possible to avoid introducing a loop operation, thereby solving the problem that when multiple speakers of the relevant acoustic model batch input speech text for speech synthesis, the speech text of each speaker is inferred in turn by calling the adaptive model parameters of each speaker, resulting in a technical problem with low reasoning processing efficiency.
[0094] Example 2
[0095] Based on Example 1 of the present application, an embodiment of an acoustic model implementation device for an extensible speaker is also provided, which executes the acoustic model implementation method of the extensible speaker of the above embodiment when running. Figure 9 is a structural diagram of an optional acoustic model implementation device for an extensible speaker according to an embodiment of the present application. As shown in Figure 9, the acoustic model implementation device for an extensible speaker includes at least: a first determination module 91, a second determination module 93 and a model generation module 95, wherein:
[0096] The first determination module 91 is used to respond to speech synthesis requests of multiple target new speakers and determine the target speaker identifier of each target new speaker in the speech synthesis request;
[0097] The second determination module 93 is used to determine the target embedding vector of the target newly added speaker corresponding to the target speaker identifier from a preset speaker index table, wherein the target embedding vector includes: a speaker embedding vector and an adaptive embedding vector, the speaker index table sequentially records the first speaker identifier of the first speaker and the corresponding first speaker embedding vector, the newly added speaker identifiers of multiple newly added speakers and the corresponding second embedding vectors, the second embedding vector includes: the newly added speaker embedding vector corresponding to the newly added speaker and the adaptive embedding vector adapter parameter obtained by adaptively training the target acoustic model using the second speech data of the newly added speaker, and the target acoustic model is obtained by combining the preset adapter with the baseline acoustic model trained based on the first speech data of the first speaker;
[0098] The model generation module 95 is used to determine the target adapter parameters in the target acoustic model based on the target embedding vectors of multiple target new speakers, and recombine the target adapter parameters with the target acoustic model, and respond to the speech synthesis requests of multiple target speakers through the combined target acoustic model.
[0099] It should be noted that the various modules in the above-mentioned extensible speaker acoustic model implementation device can be program modules (for example, a set of program instructions that implement a certain specific function) or hardware modules. For the latter, it can be expressed in the following form, but is not limited to this: the expression form of each of the above-mentioned modules is a processor, or the functions of each of the above-mentioned modules are implemented by a processor.
[0100] Example 3
[0101] According to an embodiment of the present application, a non-volatile storage medium is also provided, in which a program is stored. When the program is running, the device where the non-volatile storage medium is located is controlled to execute the method for implementing the acoustic model of the scalable speaker in Example 1.
[0102] In some embodiments, the device where the non-volatile storage medium is located implements the following steps by running the program:
[0103] Step S202, in response to speech synthesis requests from multiple target new speakers, determining a target speaker identifier for each target new speaker in the speech synthesis request;
[0104] Step S204, determining a target embedding vector of a target newly added speaker corresponding to the target speaker identifier from a preset speaker index table, wherein the target embedding vector includes: a speaker embedding vector and an adaptive embedding vector, the speaker index table sequentially records a first speaker identifier of a first speaker and a corresponding first speaker embedding vector, and newly added speaker identifiers of multiple newly added speakers and corresponding second embedding vectors, the second embedding vector includes: a newly added speaker embedding vector corresponding to the newly added speaker and an adaptive embedding vector adapter parameter obtained by adaptively training a target acoustic model using the second speech data of the newly added speaker, and the target acoustic model is obtained by combining a preset adapter with a reference acoustic model trained based on the first speech data of the first speaker;
[0105] Step S206, determining target adapter parameters within the target acoustic model based on the target embedding vectors of the multiple target newly added speakers, and recombining the target adapter parameters with the target acoustic model, and responding to the speech synthesis requests of the multiple target speakers through the combined target acoustic model.
[0106] According to an embodiment of the present application, a processor is also provided, which is used to run a program, wherein the method for implementing the acoustic model of an extensible speaker in Example 1 is executed when the program is running.
[0107] In some embodiments, the program is executed to implement the following steps:
[0108] Step S202, in response to speech synthesis requests from multiple target new speakers, determining a target speaker identifier for each target new speaker in the speech synthesis request;
[0109] Step S204, determining a target embedding vector of a target newly added speaker corresponding to the target speaker identifier from a preset speaker index table, wherein the target embedding vector includes: a speaker embedding vector and an adaptive embedding vector, the speaker index table sequentially records a first speaker identifier of a first speaker and a corresponding first speaker embedding vector, and newly added speaker identifiers of multiple newly added speakers and corresponding second embedding vectors, the second embedding vector includes: a newly added speaker embedding vector corresponding to the newly added speaker and an adaptive embedding vector adapter parameter obtained by adaptively training a target acoustic model using the second speech data of the newly added speaker, and the target acoustic model is obtained by combining a preset adapter with a reference acoustic model trained based on the first speech data of the first speaker;
[0110] Step S206, determining target adapter parameters within the target acoustic model based on the target embedding vectors of the multiple target newly added speakers, and recombining the target adapter parameters with the target acoustic model, and responding to the speech synthesis requests of the multiple target speakers through the combined target acoustic model.
[0111] According to an embodiment of the present application, an electronic device is also provided, wherein the electronic device includes one or more processors; a memory for storing one or more programs, which, when the one or more programs are executed by one or more processors, enables the one or more processors to run the programs, wherein the program is configured to execute the acoustic model implementation method of the scalable speaker in the above-mentioned embodiment 1 when running.
[0112] In some embodiments, the processor is configured to implement the following steps by executing a computer program:
[0113] Step S202, in response to speech synthesis requests from multiple target new speakers, determining a target speaker identifier for each target new speaker in the speech synthesis request;
[0114] Step S204, determining a target embedding vector of a target newly added speaker corresponding to the target speaker identifier from a preset speaker index table, wherein the target embedding vector includes: a speaker embedding vector and an adaptive embedding vector, the speaker index table sequentially records a first speaker identifier of a first speaker and a corresponding first speaker embedding vector, and newly added speaker identifiers of multiple newly added speakers and corresponding second embedding vectors, the second embedding vector includes: a newly added speaker embedding vector corresponding to the newly added speaker and an adaptive embedding vector adapter parameter obtained by adaptively training a target acoustic model using the second speech data of the newly added speaker, and the target acoustic model is obtained by combining a preset adapter with a reference acoustic model trained based on the first speech data of the first speaker;
[0115] Step S206, determining target adapter parameters within the target acoustic model based on the target embedding vectors of the multiple target newly added speakers, and recombining the target adapter parameters with the target acoustic model, and responding to the speech synthesis requests of the multiple target speakers through the combined target acoustic model.
[0116] The serial numbers of the above embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.
[0117] In the above embodiments of the present application, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, please refer to the relevant description of other embodiments.
[0118] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only exemplary. For example, the division of units can be a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.
[0119] Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple units. Some or all of the units may be selected to achieve the purpose of the present embodiment according to actual needs.
[0120] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0121] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the relevant technology, or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server or network device, etc.) to execute all or part of the steps of the various embodiments of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk.
[0122] The above is only a preferred embodiment of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.
Claims
1. A method for implementing an acoustic model that can extend a speaker, comprising: In response to speech synthesis requests from multiple target new speakers, determining a target speaker identifier for each of the target new speakers in the speech synthesis requests; Determine the target embedding vector of the target newly added speaker corresponding to the target speaker identifier from a preset speaker index table, wherein the target embedding vector includes: a speaker embedding vector and an adaptive embedding vector, the speaker index table sequentially records the first speaker identifier of the first speaker and the corresponding first speaker embedding vector, the newly added speaker identifiers of multiple newly added speakers and the corresponding second embedding vectors, the second embedding vector includes: the newly added speaker embedding vector corresponding to the newly added speaker and an adaptive embedding vector adapter parameter obtained by adaptively training a target acoustic model using the second speech data of the newly added speaker, and the target acoustic model is obtained by combining a preset adapter with a baseline acoustic model trained based on the first speech data of the first speaker; Based on the target embedding vectors of the multiple target new speakers, the target adapter parameters in the target acoustic model are determined, and the target adapter parameters are recombined with the target acoustic model, and the speech synthesis requests of the multiple target speakers are responded to through the combined target acoustic model.
2. The method according to claim 1, wherein The training process of the benchmark acoustic model includes: Constructing a deep learning model, wherein the deep learning model includes: an encoder composed of a single-layer feedforward transformer FFT module, a speaker embedding module, and a decoder composed of N layers of the FFT module, where N is a positive integer greater than or equal to 1; Acquire first speech data of a plurality of first speakers, wherein the first speech data includes: a first phoneme sequence code corresponding to a first speech text of the first speaker and a first acoustic feature corresponding to a first speech audio, and the first speech text corresponds to the first speech audio; The deep learning model is iteratively trained based on the first speech data of the plurality of first speakers to obtain the benchmark acoustic model.
3. The method according to claim 2, wherein: Acquiring first voice data of the plurality of first speakers includes: Acquire an initial speech text of each first speaker and an initial speech audio corresponding to the initial speech text, wherein the initial speech text includes at least one of the following: text information, punctuation marks, and prosodic annotations; Converting text information in the initial speech text into an initial phoneme sequence, and inserting the punctuation marks and the prosodic annotations into the initial phoneme sequence to obtain a first phoneme sequence; Encoding the first phoneme sequence using one-hot encoding to obtain the first phoneme sequence encoding; Performing a preprocessing operation on the initial speech audio, wherein the preprocessing operation includes at least one of the following: Sampling, volume adjustment, and trimming; Feature extraction is performed on the preprocessed initial speech audio to obtain the first acoustic feature, wherein the first acoustic feature includes at least one of the following: a Mel spectrum feature, a frame-level variable feature, and a phoneme-level duration feature.
4. The method according to claim 2, wherein: Iteratively training the deep learning model based on the first speech data of the plurality of first speakers to obtain the benchmark acoustic model includes: For each first speaker's first speech data, encoding the first speaker's first phoneme sequence and inputting it into the deep learning model, and sequentially passing through an encoder and a decoder within the deep learning model to output a first streaming acoustic feature and a first non-streaming acoustic feature; Determining a target loss function based on the first acoustic feature in the first speech data of each first speaker, the first streaming acoustic feature, and the first non-streaming acoustic feature, wherein the target loss function includes: a non-streaming mean square error loss function, a streaming mean square error loss function, and an adversarial loss function; The minimum value of the target loss function is calculated using a gradient descent algorithm, and the model parameters of the deep learning model are adjusted based on the minimum value of the target loss function to obtain the baseline acoustic model.
5. The method according to claim 2, wherein: The adaptive embedding vector obtained by adaptively training the target acoustic model using the second speech data of the newly added speaker includes: Acquire second voice data of the newly added speaker, wherein the second voice data includes: a second phoneme sequence code corresponding to a second voice text of the newly added speaker and a second acoustic feature corresponding to a second voice audio, and the second voice text corresponds to the second voice audio; Obtaining the baseline acoustic model and improving a decoder within the baseline acoustic model using the adapter to obtain the target acoustic model, wherein the adapter comprises a first feedforward neural network for dimensionality reduction, an activation layer, and a second feedforward neural network for dimensionality increase; The target acoustic model is adaptively trained using the second speech data of the newly added speaker to obtain adapter parameters of the newly added speaker, and the adapter parameters are concatenated to obtain an adaptive embedding vector corresponding to the newly added speaker, wherein the adapter parameters include at least one of the following: the first weight matrix and the first bias vector of the first feedforward neural network of the N-layer adapter, and the second weight matrix and the second bias vector of the second feedforward neural network.
6. The method according to claim 5, wherein: Improving a decoder within the baseline acoustic model using the adapter includes: The adapter is initialized, and one of the initialized adapters is added after each of the N layers of the FFT modules of the decoder in the baseline acoustic model to obtain the improved decoder, wherein the FFT module consists of a multi-head attention mechanism module, a block module, a conditional layer normalization module, and a causal convolution module.
7. The method according to claim 5, wherein: Initializing the adapter includes: Initializing a first weight matrix of the first feedforward neural network and a second weight matrix of the second feedforward neural network to 1, and initializing a first bias vector of the first feedforward neural network and a second bias vector of the second feedforward neural network to 0.
8. The method according to claim 5, wherein The adapter parameters are concatenated to obtain an adaptive embedding vector corresponding to the newly added speaker, including: The first layer of adapters to the Nth layer of adapters in the adapter parameters corresponding to the newly added speaker are horizontally spliced with the first weight matrix and the first bias vector of the first feedforward neural network and the second weight matrix and the second bias vector of the second feedforward neural network in each of the adapters in turn to obtain an adaptive embedding vector corresponding to the newly added speaker.
9. The method according to claim 5, wherein: Determining target adapter parameters within the target acoustic model based on target embedding vectors of the plurality of target newly added speakers includes: Slicing and transforming the target adaptive embedding vector of each target newly added speaker to obtain target adapter parameters corresponding to the corresponding N-layer target adapter; For the target adapter parameters corresponding to each layer of the target adapter, the first weight matrix of the first feedforward neural network, the first bias vector of the first feedforward neural network, the second weight matrix of the second feedforward neural network, and the second bias vector of the second feedforward neural network corresponding to the multiple target newly added speakers are respectively spliced to obtain a first target weight matrix, a first target bias vector, a second target weight matrix, and a second target bias vector; The target adapter parameters of the target acoustic model are obtained by the target speaker embedding vectors corresponding to the multiple target newly added speakers, the first target weight matrix corresponding to each layer of the target adapter, the first target bias vector, the second target weight matrix, and the second target bias vector.
10. The method according to claim 3, wherein: The initial speech audio is preprocessed, including: Converting the initial speech audio into single-channel, 16-bit encoded audio; Unifying the volume of the converted initial voice audio to -6db; Converting the volume-adjusted initial voice audio to target sampling rate audio; The silence segments at the beginning and end of the initial speech audio are trimmed so that a maximum of 100 milliseconds of silence segment is retained.
11. A device for implementing an acoustic model capable of scalably expressing a speaker, comprising: A first determining module is configured to determine a target speaker identifier of each target new speaker in the speech synthesis request in response to speech synthesis requests of multiple target new speakers; The second determining module is used to determine the target speaker identifier from the preset speaker index table. The target embedding vector of the target newly added speaker, wherein the target embedding vector includes: a speaker embedding vector and an adaptive embedding vector, the speaker index table records in sequence the first speaker identifier of the first speaker and the corresponding first speaker embedding vector, the newly added speaker identifiers of multiple newly added speakers and the corresponding second embedding vectors, the second embedding vector includes: the newly added speaker embedding vector corresponding to the newly added speaker and the adaptive embedding vector adapter parameters obtained by adaptively training the target acoustic model using the second speech data of the newly added speaker, and the target acoustic model is obtained by combining a preset adapter with a baseline acoustic model trained based on the first speech data of the first speaker; A model generation module is used to determine the target adapter parameters in the target acoustic model based on the target embedding vectors of multiple target new speakers, and recombine the target adapter parameters with the target acoustic model, and respond to the speech synthesis requests of multiple target speakers through the combined target acoustic model.
12. A non-volatile storage medium storing a computer program, wherein: The device where the non-volatile storage medium is located executes the method for implementing the acoustic model of an expandable speaker as described in any one of claims 1 to 10 by running the computer program.
13. An electronic device comprising: A memory and a processor, wherein the memory stores a computer program, and the processor is configured to execute the method for implementing an extensible speaker acoustic model according to any one of claims 1 to 10 through the computer program.