Streaming voice synthesis method and apparatus, electronic device and storage medium
By dynamically adjusting the feature block size, combining the size information and real-time rate of the current processing cycle, the problem of difficult to balance the block size in streaming speech synthesis is solved, achieving more efficient inference speed and better speech synthesis effect, and improving user experience.
Patent Information
- Application Number
- PCT/CN2024/116765
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-13
- Filing Date
- 2024-09-04
- Publication Date
- 2025-05-22
AI Technical Summary
The existing streaming speech synthesis technology is difficult to balance in the chunking size, resulting in the inability to balance the first frame delay, overall inference speed and synthesis effect, affecting the user's perception experience.
By dynamically adjusting the size of the feature block, the next size information is determined based on the size information and real-time rate of the current processing cycle to achieve reasonable feature allocation and avoiding the delay and inference speed problems caused by fixed block size.
Without affecting the delay of the first frame, the overall inference speed is improved and the voice synthesis effect is improved, improving the user's perceived experience.
Smart Images

Figure CN2024116765_22052025_PF_FP_ABST
Abstract
Description
Streaming speech synthesis method, device, electronic device and storage medium
[0001] This application claims priority to the Chinese patent application filed with the China Patent Office on November 13, 2023, with application number 202311508652.1, the entire contents of which are incorporated by reference into this application. Technical Field
[0002] The present application relates to the field of information processing technology, for example, to a streaming speech synthesis method, device, electronic device and storage medium. Background Art
[0003] With the successful application of deep neural networks in the field of speech synthesis, speech synthesis technology has made great breakthroughs in recent years and has been widely used in many fields.
[0004] In related technologies, speech synthesis generally includes three main modules: a text front-end, an acoustic model, and a vocoder. The acoustic model further includes two main modules: an encoder and a decoder. Streaming speech synthesis generally refers to streaming the encoder output features, expanded to the frame level, into the decoder and vocoder in sequence to obtain the corresponding speech output. Streaming speech synthesis generally divides the encoder output features into blocks of fixed size, which are then fed into the decoder and vocoder for speech synthesis. However, larger blocks will directly increase the first-frame latency and reduce system response speed; smaller blocks will reduce the overall inference speed and affect the synthesis effect. Therefore, the block division method is often difficult to balance, which inevitably affects the user's perceptual experience.
[0005] Summary of the Invention
[0006] The present application provides a streaming speech synthesis method, device, electronic device and storage medium to solve the problem of the inability to balance the first frame delay with the overall inference speed and synthesis effect due to a fixed block size. It achieves the goal of improving the overall inference speed and synthesis effect without affecting the first frame delay, thereby improving the user's perceptual experience.
[0007] The present invention provides a streaming speech synthesis method, which is applied to a speech synthesis model. The method includes:
[0008] Determining current size information used in a current processing cycle, and determining a current feature block based on the current size information, wherein the current feature block is composed of features of the current size information intercepted from current input features, and the current input features remaining after interception are current remaining features. The current size information is the number of frames from which features should be intercepted from the current input features. The input features are features obtained by encoding the text information to be synthesized and expanding them to the frame level according to phoneme duration information. The phoneme duration information is the number of frames in the speech audio corresponding to one phoneme.
[0009] Performing speech synthesis inference on the current feature block, outputting current speech audio corresponding to the current feature block, and determining a current inference consumption time;
[0010] Determining current duration information of the current feature block based on the current size information, and determining a current real-time rate of the speech synthesis model based on the current inference consumption time and the current duration information, where the current duration information is duration information of the speech audio output corresponding to the current feature block;
[0011] Determining next size information based on the current size information and the current real-time rate for loading in the next processing cycle;
[0012] When the next size information is greater than or equal to the number of frames of the current remaining features, all the current remaining features are sent to the speech synthesis module for inference to obtain the remaining speech audio corresponding to the current remaining features. When the next size information is less than the number of frames of the current remaining features, the current remaining features are used as the current input features, the next size information is used as the current size information and the current size information used to determine the current processing cycle is returned, and the current feature block is determined based on the current size information.
[0013] The present invention provides a streaming speech synthesis device for use in a speech synthesis model. The device includes:
[0014] a feature block determination module configured to determine current size information used in a current processing cycle and determine a current feature block based on the current size information, wherein the current feature block is composed of features of the current size information intercepted from current input features, the current input features remaining after interception being the current remaining features, the current size information being the number of frames from which features should be intercepted from the current input features, the input features being features obtained by encoding the text information to be synthesized and expanding it to the frame level according to phoneme duration information, the phoneme duration information being the number of frames in the speech audio corresponding to one phoneme;
[0015] a time determination module configured to perform speech synthesis inference on the current feature block, output the current speech audio corresponding to the current feature block, and determine the current inference consumption time;
[0016] a first information determination module configured to determine current duration information of the current feature block based on the current size information, and determine a current real-time rate of the speech synthesis model based on the current inference consumption time and the current duration information, wherein the current duration information is duration information of the speech audio output corresponding to the current feature block;
[0017] a second information determination module configured to determine next size information based on the current size information and the current real-time rate, for loading in the next processing cycle;
[0018] The judgment module is configured to send all the current remaining features to the speech synthesis module for inference to obtain the remaining speech audio corresponding to the current remaining features when the next size information is greater than or equal to the number of frames of the current remaining features; and when the next size information is less than the number of frames of the current remaining features, use the current remaining features as the current input features, use the next size information as the current size information and return the current size information used to determine the current processing cycle, and determine the current feature block based on the current size information.
[0019] An embodiment of the present application provides an electronic device, the electronic device comprising:
[0020] at least one processor; and
[0021] a memory communicatively connected to the at least one processor; wherein,
[0022] The memory stores a computer program that can be executed by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the streaming speech synthesis method described in any embodiment of the present application.
[0023] An embodiment of the present application provides a computer-readable storage medium storing computer instructions, which are used to enable a processor to implement the streaming speech synthesis method described in any embodiment of the present application when executed. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0025] FIG1 is a flow chart of a streaming speech synthesis method provided according to an embodiment of the present application;
[0026] FIG2 is a schematic diagram of the structure of a speech synthesis model applicable to an embodiment of the present application;
[0027] FIG3 is a schematic structural diagram of a streaming speech synthesis device according to an embodiment of the present application;
[0028] FIG4 is a schematic structural diagram of an electronic device for implementing the streaming speech synthesis method according to an embodiment of the present application. DETAILED DESCRIPTION
[0029] The following will describe the technical solutions in the embodiments of this application in conjunction with the drawings in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of this application.
[0030] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in a sequence other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0031] Example 1
[0032] Figure 1 is a flow chart of a streaming speech synthesis method provided in an embodiment of the present application. This embodiment is applicable to situations where streaming speech synthesis is performed on input text information. The method can be performed by a streaming speech synthesis device, which can be implemented in hardware and / or software and can be configured in any electronic device with network communication capabilities. As shown in Figure 1, the method includes the following steps.
[0033] S110: Determine current size information used in the current processing cycle, and determine a current feature block based on the current size information.
[0034] The current feature block is composed of the features of the current size information cut out from the current input features. The current input features remaining after cutting out are the current remaining features. The current size information is the number of frames from which the features should be cut out from the current input features. The input features are the features obtained by encoding the text information to be synthesized and expanding it to the frame level according to the phoneme duration information. The phoneme duration information is the number of frames in the speech audio corresponding to a phoneme.
[0035] Figure 2 is a structural diagram of the speech synthesis model applicable to the embodiment of the present application, wherein the input may include: text information for controlling the synthesis content, speaker information for controlling the synthesis tone, style information for controlling the synthesis style, etc.
[0036] After obtaining text information for speech synthesis, the text information is converted into phoneme information, and the phoneme information is encoded to obtain coding features; the phoneme duration information corresponding to each phoneme information is determined; based on the phoneme duration information, the coding features of the phoneme information are expanded into frame-level feature information to obtain input features.
[0037] Optionally, determine the current size information used in the current processing cycle, including:
[0038] If the current processing cycle is the first processing cycle, the current size information is determined by the receptive field of the speech synthesis model. The receptive field is the number of frames of features that should be intercepted from the current input features to output one frame of speech audio. For example, the receptive field is recorded as rf, and the current size information is
[0039] If the current processing cycle is not the first processing cycle, the current size information is the size information determined in the previous processing cycle.
[0040] S120: Perform speech synthesis inference on the current feature block, output the current speech audio corresponding to the current feature block, and determine the current inference consumption time.
[0041] The current speech audio is the audio output after the current feature block is input into the speech synthesis model for speech synthesis inference.
[0042] S130. Determine the current duration information of the current feature block based on the current size information, and determine the current real-time rate of the speech synthesis model based on the current inference consumption time and the current duration information.
[0043] The current duration information is the duration of the speech audio output corresponding to the current feature block, and is in a fixed ratio with the current size information. This ratio, which is the audio length corresponding to one frame of features, is uniquely determined by the speech synthesis model algorithm. The current real-time rate is the unit time required for the speech synthesis system to generate one unit of speech. The current size information, the current inference time, and the time length corresponding to the feature input information are obtained. The product of the current size information and the duration of one frame of speech audio output is determined as the current duration information. The current real-time rate is determined as the ratio of the current inference time to the current duration information.
[0044] For example, the current size information is N frames, the current inference time is T, and the audio length corresponding to one frame of input features is t. Then the current duration information is N*t, and the current real-time rate R=T / (N*t).
[0045] S140: Determine next size information based on the current size information and the current real-time rate for loading in the next processing cycle.
[0046] The principle for determining the next size information is to complete the synthesis before the audio synthesized by the current feature block finishes playing, so as not to affect continuous broadcasting. Therefore, the current duration information and the current real-time rate can be obtained, and the ratio of the two can be used to determine the next duration information. The ratio of the next duration information to the length of a single audio frame can be used to determine the next size information. This can be equivalently done by obtaining the current size information and the current real-time rate, and then using the ratio of the current size information to the current real-time rate as the next size information.
[0047] For example, if the current size is N frames, then the current duration is N*t. If the current real-time rate is R, then the next duration is calculated as N*t / R based on the current duration, and the next size is calculated as N*t / R / t = N / R. This is equivalent to deriving the next size from the current size. The next size should be rounded down.
[0048] S150. When the next size information is greater than or equal to the number of frames of the current remaining features, all the current remaining features are sent to the speech synthesis module for inference to obtain the remaining speech audio corresponding to the current remaining features. When the next size information is less than the number of frames of the current remaining features, the current remaining features are used as the current input features, the next size information is used as the current size information and the process returns to execute S110.
[0049] The current remaining feature is the feature remaining after the current input feature is cut off from the current input feature according to the current size information.
[0050] Determine whether the next size information is greater than or equal to the number of frames of the current remaining features. If the next size information is greater than or equal to the number of frames of the current remaining features, it means that the output audio playback time of the previous feature block is sufficient to complete the synthetic reasoning of the current remaining features. Therefore, the current remaining features are combined into the last feature block and input into the speech synthesis module to obtain the corresponding speech output, completing the streaming speech synthesis;
[0051] If the next size information is less than the number of frames of the current remaining features, it means that the output audio playback time of the previous feature block is insufficient to complete the synthesis and inference of the current remaining features. Further segmentation is required to obtain continuous voice playback without affecting the user's perceptual experience. The next size information just calculated is used as the current size information, and the current remaining features are used as the current input features. The steps of segmentation, inference, and obtaining the next size information are continued, i.e., steps S110, S120, S130, S140, and S150 described above.
[0052] Optionally, the block operation of this application can convert multiple single-frame inference processes into a single multi-frame inference process, and convert multiple small-frame inference processes into a single large-frame inference process, thus achieving matrix input feature transformation and fully utilizing the parallel characteristics of neural network operators to improve inference speed. In other words, the time consumption of N times X frames inference is greater than the time consumption of one N*X frame inference.
[0053] The technical solution of the embodiment of the present application determines the current size information used in the current processing cycle, and determines the current feature block based on the current size information, determines the current inference consumption time for speech synthesis inference on the current feature block, determines the current duration information of the current feature block based on the current size information, and determines the current real-time rate of the speech synthesis model based on the current inference consumption time and the current duration information, and determines the next size information based on the current size information and the current real-time rate; when the next size information is greater than or equal to the number of frames of the current remaining features, all the current remaining features are sent to the speech synthesis module to obtain the remaining speech audio, and when the next size information is less than the number of frames of the current remaining features, the next size information is continued to be determined and the block operation is performed. The technical solution of the present application accurately determines the next size information based on the current size information and the current real-time rate, thereby achieving a reasonable allocation of input features, solving the problem of the inability to balance the first frame delay with the overall inference speed and synthesis effect caused by the fixed block size, and achieving an improvement in the overall inference speed and synthesis effect without affecting the first frame delay, thereby improving the user's perceptual experience.
[0054] Example 2
[0055] FIG3 is a schematic diagram of the structure of a streaming speech synthesis device provided in an embodiment of the present application. This embodiment is applicable to the case of streaming speech synthesis of input text information. The streaming speech synthesis device can be implemented in the form of hardware and / or software. The streaming speech synthesis device can be configured in any electronic device with network communication function. As shown in FIG3, the device includes:
[0056] The feature block determination module 210 is configured to determine current size information used in the current processing cycle and determine a current feature block based on the current size information, wherein the current feature block is composed of features of the current size information intercepted from the current input features, and the current input features remaining after interception are the current remaining features. The current size information is the number of frames from which features should be intercepted from the current input features. The input features are features obtained by encoding the text information to be synthesized and expanding them to the frame level according to the phoneme duration information, and the phoneme duration information is the number of frames in the speech audio corresponding to one phoneme.
[0057] The time determination module 220 is configured to perform speech synthesis inference on the current feature block, output the current speech audio corresponding to the current feature block, and determine the current inference consumption time;
[0058] A first information determination module 230 is configured to determine current duration information of the current feature block based on the current size information, and determine a current real-time rate of the speech synthesis model based on the current inference consumption time and the current duration information, where the current duration information is duration information of the output speech audio corresponding to the current feature block;
[0059] A second information determination module 240 is configured to determine next size information based on the current size information and the current real-time rate, for loading in the next processing cycle;
[0060] The judgment module 250 is configured to send all the current remaining features to the speech synthesis module for inference to obtain the remaining speech audio corresponding to the current remaining features when the next size information is greater than or equal to the number of frames of the current remaining features; otherwise, the current remaining features are used as the current input features, the next size information is used as the current size information and the current size information used to determine the current processing cycle is returned, and the current feature block is determined based on the current size information.
[0061] Optionally, the feature block determination module includes an input feature determination unit configured to:
[0062] Acquiring text information for speech synthesis, converting the text information into phoneme information, and encoding the phoneme information to obtain encoding features;
[0063] Determining the phoneme duration information corresponding to each of the phoneme information;
[0064] The encoding features of the phoneme information are expanded into frame-level feature information based on the phoneme duration information to obtain the input features.
[0065] Optionally, the feature block determination module includes a size information determination unit configured to:
[0066] If the current processing cycle is the first processing cycle, the current size information is determined by the receptive field of the speech synthesis model, where the receptive field is the number of frames of features that should be intercepted from the current input features to output one frame of speech audio;
[0067] If the current processing cycle is not the first processing cycle, the current size information is the size information determined in the previous processing cycle.
[0068] Optionally, the first information determining module 230 is configured to:
[0069] Determine the current duration information by multiplying the current size information by the duration of one frame of output speech audio;
[0070] The ratio of the current inference consumption time to the current duration information is determined as the current real-time rate.
[0071] Optionally, the second information determining module 240 is configured to:
[0072] The ratio of the current size information to the current real-time rate is determined as the next size information.
[0073] The streaming speech synthesis device provided in the embodiments of the present application can execute the streaming speech synthesis method provided in any embodiment of the present application.
[0074] Example 3
[0075] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0076] FIG4 shows a schematic diagram of the structure of an electronic device that can be used to implement the streaming speech synthesis method of an embodiment of the present application. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present application described and / or required herein.
[0077] As shown in FIG4 , the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 and a random access memory (RAM) 13, that is communicatively connected to the at least one processor 11. The memory stores a computer program that can be executed by the at least one processor, and the processor 11 can perform various appropriate actions and processes according to the computer program stored in the ROM 12 or the computer program loaded from the storage unit 18 into the RAM 13. The RAM 13 can also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are connected to each other via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0078] Multiple components in the electronic device 10 are connected to the I / O interface 15, including an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a magnetic disk, an optical disk, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0079] The processor 11 can be a variety of general-purpose and / or specialized processing components with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The processor 11 performs the various methods and processes described above, such as the streaming speech synthesis method.
[0080] In some embodiments, the streaming speech synthesis method can be implemented as a computer program that is tangibly contained in a computer-readable storage medium, such as the storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the streaming speech synthesis method described above can be performed. Alternatively, in other embodiments, the processor 11 can be configured to perform the streaming speech synthesis method in any other appropriate manner (e.g., by means of firmware).
[0081] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard parts (ASSPs), system on chips (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0082] Computer programs for implementing the methods of the present application may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the computer program is executed by the processor, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The computer program may be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0083] In the context of the present application, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by an instruction execution system, device or equipment or used in combination with an instruction execution system, device or equipment. A computer-readable storage medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared or semiconductor systems, devices or equipment, or any suitable combination of the foregoing. Alternatively, a computer-readable storage medium can be a machine-readable signal medium. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a RAM, a ROM, an erasable programmable read-only memory (EPROM) or a flash memory, an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0084] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a cathode ray tube (CRT) or a liquid crystal display (LCD) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0085] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with embodiments of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.
[0086] A computing system may include a client and a server. The client and server are generally remote from each other and typically interact via a communication network. The client-server relationship arises through computer programs running on the respective computers and establishing a client-server relationship with each other. The server may be a cloud server, also known as a cloud computing server or cloud host, a host product within a cloud computing service ecosystem that addresses the management difficulties and limited scalability of traditional physical hosts and virtual private server (VPS) services.
[0087] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the multiple steps described in this application can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of this application can be achieved. This is not limited herein.
Claims
1. A streaming speech synthesis method, applied to a speech synthesis model, the method comprising: Determine the current size information used in the current processing cycle, and determine the current feature block based on the current size information, wherein the current feature block is composed of features of the current size information intercepted from the current input features, and the current input features remaining after interception are the current remaining features, and the current size information is the number of frames from which the features should be intercepted from the current input features, and the input features are the features obtained by encoding the text information to be synthesized and expanding it to the frame level according to the phoneme duration information, and the phoneme duration information is the number of frames in the speech audio corresponding to one phoneme; Performing speech synthesis reasoning on the current feature block, outputting current speech audio corresponding to the current feature block, and determining current reasoning consumption time; Determine the current duration information of the current feature block based on the current size information, and determine the current real-time rate of the speech synthesis model based on the current inference consumption time and the current duration information, wherein the current duration information is the duration information of the speech audio output corresponding to the current feature block; Determine next size information based on the current size information and the current real-time rate for loading and use in the next processing cycle; When the next size information is greater than or equal to the number of frames of the current remaining features, all the current remaining features are sent to the speech synthesis module for inference to obtain the remaining speech audio corresponding to the current remaining features. When the next size information is less than the number of frames of the current remaining features, the current remaining features are used as the current input features, the next size information is used as the current size information and the current size information used to determine the current processing cycle is returned, and the current feature block is determined based on the current size information.
2. The method according to claim 1, wherein: The process of determining the input features includes: Acquiring text information for speech synthesis, converting the text information into phoneme information, and encoding the phoneme information to obtain encoding features; Determine the phoneme duration information corresponding to each of the phoneme information; The encoded features of the phoneme information are expanded into frame-level feature information based on the phoneme duration information to obtain the input features.
3. The method according to claim 1, wherein: The determining of the current size information used in the current processing cycle includes: In response to the current processing cycle being the first processing cycle, the current size information is determined by a receptive field of the speech synthesis model, the receptive field being the number of frames of features that should be intercepted from the current input features required to output one frame of speech audio; In response to the current processing cycle not being the first processing cycle, the current size information is the previous Processing cycle determined sizing information.
4. The method according to claim 1, wherein: The determining the current duration information of the current feature block based on the current size information, and determining the current real-time rate of the speech synthesis model based on the current reasoning consumption time and the current duration information, includes: The current duration information is determined by multiplying the current size information by the duration of one frame of output speech audio; The ratio of the current reasoning consumption time to the current duration information is determined as the current real-time rate.
5. The method according to claim 1, wherein: The determining the next size information based on the current size information and the current real-time rate includes: The ratio of the current size information to the current real-time rate is determined as the next size information.
6. A streaming speech synthesis device, applied to a speech synthesis model, the device comprising: A feature block determination module is configured to determine current size information used in a current processing cycle, and determine a current feature block based on the current size information, wherein the current feature block is composed of features of the current size information intercepted from current input features, and the current input features remaining after interception are current remaining features, and the current size information is the number of frames from which features should be intercepted from the current input features, and the input features are features obtained by encoding the text information to be synthesized and expanding it to a frame level according to the phoneme duration information, and the phoneme duration information is the number of frames in the speech audio corresponding to a phoneme; A time determination module, configured to perform speech synthesis reasoning on the current feature block, output the current speech audio corresponding to the current feature block, and determine the current reasoning consumption time; A first information determination module is configured to determine the current duration information of the current feature block based on the current size information, and determine the current real-time rate of the speech synthesis model based on the current reasoning consumption time and the current duration information, wherein the current duration information is the duration information of the speech audio output corresponding to the current feature block; A second information determination module is configured to determine next size information based on the current size information and the current real-time rate for loading and use in the next processing cycle; The judgment module is configured to, when the next size information is greater than or equal to the number of frames of the current remaining features, send all the current remaining features to the speech synthesis module for inference to obtain the remaining speech audio corresponding to the current remaining features; when the next size information is less than the number of frames of the current remaining features, use the current remaining features as the current input features, use the next size information as the current size information and repeatedly determine the current size information used in the current processing cycle, and determine the current feature block based on the current size information; perform speech synthesis inference on the current feature block, output the current speech audio corresponding to the current feature block, and determine the current inference consumption time; Determine the current duration information of the current feature block based on the current size information, and determine the current real-time rate of the speech synthesis model based on the current reasoning consumption time and the current duration information; Next size information is determined based on the current size information and the current real-time rate for loading and use in the next processing cycle.
7. The device according to claim 6, wherein: The feature block determination module includes an input feature determination unit, which is configured as follows: Acquiring text information for speech synthesis, converting the text information into phoneme information, and encoding the phoneme information to obtain encoding features; Determine the phoneme duration information corresponding to each of the phoneme information; The encoding feature corresponding to the phoneme information is expanded into frame-level feature information based on the phoneme duration information to obtain the input feature.
8. The device according to claim 6, wherein: The feature block determination module includes a size information determination unit, which is configured as follows: In response to the current processing cycle being the first processing cycle, the current size information is determined by a receptive field of the speech synthesis model, the receptive field being the number of frames of features that should be intercepted from the current input features required to output one frame of speech audio; In response to the current processing cycle not being the first processing cycle, the current size information is size information determined in a previous processing cycle.
9. An electronic device, comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can perform the streaming speech synthesis method according to any one of claims 1 to 5.
10. A computer-readable storage medium storing computer instructions, wherein the computer instructions are used to enable a processor to implement the streaming speech synthesis method according to any one of claims 1 to 5 when executed.
Citation Information
Patent Citations
Voice synthesis method and device, storage medium and electronic equipment
CN111402855A
Speech synthesis method based on stream generation model
CN113299268A
Speech synthesis method and device thereof, electronic equipment and readable storage medium
CN113781995A
Speech synthesis method and device, equipment and storage medium
CN113903326A
Method for synthesizing streaming speech into vocoder based on NPU (Network Processing Unit) and related product
CN115440235A