Audio generation method and device, equipment and medium
Through the combination of the voice conversion model and the fusion model, target audio that meets user needs is generated, which solves the problem that audio generation in the prior art cannot adapt to diverse user preferences and improves the user experience.
Patent Information
- Application Number
- CN202510245348.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-03
- Publication Date
- 2025-07-04
AI Technical Summary
In the prior art, the text input by the user is converted into a fixed voice line or a replica authorized voice line, which cannot meet the personalized needs of different users, resulting in the inability to adapt to the diverse user preferences.
By obtaining the first reference audio and reference text, the pronunciation mark and text mark are extracted using the speech conversion model, and the target audio is generated in the fusion model with the set tone data, so as to realize personalized synthesis of the audio.
It realizes the generation of audio with target tones based on user needs, improves the user experience and meets the personalized needs of different users.
Smart Images

Figure CN120260541A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical field of audio processing, and more specifically, to an audio generation method, apparatus, device, and medium. Background Art
[0002] With the rapid development of artificial intelligence technology, artificial intelligence technology can be applied to song generation, dubbing, voice imitation, etc. Currently, in order to convert the text that a user needs to dub into audio, the existing conversion methods usually convert the text input by the user into a fixed voice line or reproduce an authorized voice line, and then obtain an audio with a single rhythm and timbre. However, since different users have different preferences, it is difficult for the audio with a single rhythm and timbre to meet the needs of users. Summary of the Invention
[0003] An object of an embodiment of the present disclosure is to provide a new technical solution for audio generation.
[0004] According to a first aspect of the present disclosure, there is provided an audio generation method, the method including:
[0005] Obtain a first reference audio and a reference text;
[0006] Input a prosody identifier of the first reference audio and a text identifier of the reference text into a preset voice conversion model to obtain a voice identifier in which the prosody of the first reference audio matches the text content of the reference text;
[0007] Input the voice identifier and set timbre data into a preset fusion model to obtain a target audio with a target timbre; wherein, the timbre reflected in the timbre data is the target timbre.
[0008] In a possible implementation manner, the method further includes:
[0009] Input the first reference audio into a preset audio encoding model to obtain a first prosody identifier corresponding to the first reference audio;
[0010] Input the reference text into a preset text encoding model to obtain a first text identifier corresponding to the reference text.
[0011] In a possible implementation manner, the method further includes:
[0012] Obtain a second reference audio; wherein, the second reference audio is used to represent the target timbre;
[0013] Input the second reference audio into a preset timbre encoder to obtain set timbre data.
[0014] In a possible implementation, the fusion model is a conditional flow matching model.
[0015] In a possible implementation, the method further includes:
[0016] Obtain noise data and conditional information associated with the noise data;
[0017] Use the noise data and the conditional information as inputs to train a pre-set fusion model to obtain a trained fusion model.
[0018] In a possible implementation, the noise data includes a Mel spectrogram; the step of using the noise data and the conditional information as inputs to train a pre-set fusion model to obtain a trained fusion model includes:
[0019] Input the Mel spectrogram into a pre-set audio encoding model to obtain a second prosody identifier in the Mel spectrogram;
[0020] Input the Mel spectrogram into a pre-set timbre encoder to obtain a timbre identifier in the Mel spectrogram;
[0021] Use the conditional information as a conditional input to the pre-set fusion model, and use the second prosody identifier and the timbre identifier as sample inputs to the fusion model to train the fusion model to obtain a trained fusion model.
[0022] In a possible implementation, the Mel spectrogram has a masked segment; the step of using the conditional information as a conditional input to the pre-set fusion model, and using the second prosody identifier and the timbre identifier as sample inputs to the fusion model to train the fusion model to obtain a trained fusion model includes:
[0023] Construct a context condition of the pre-set fusion model through the second prosody identifier, the timbre identifier, and the masked segment;
[0024] Use the conditional information as a conditional input to the fusion model, and use the context condition as a sample input to the fusion model to train the fusion model to obtain a trained fusion model.
[0025] According to a second aspect of the present disclosure, there is also provided an audio generation device, including:
[0026] An acquisition module, configured to acquire a first reference audio and a reference text;
[0027] A first obtaining module, configured to input the prosody identifier of the first reference audio and the text identifier of the reference text into a preset speech conversion model, and obtain a speech identifier in which the prosody of the first reference audio matches the text content of the reference text;
[0028] A second obtaining module, configured to input the speech identifier and set timbre data into a preset fusion model, and obtain a target audio with a target timbre; wherein, the timbre reflected in the timbre data is the target timbre.
[0029] According to a third aspect of the present disclosure, there is also provided a computer system, which includes a processor. When the processor executes program instructions or code, the computer system implements the audio generation method in the first aspect. Exemplarily, the computer system further includes a memory for storing the program instructions or code.
[0030] According to a fourth aspect of the present disclosure, there is also provided a computer-readable storage medium, in which a computer program is stored. The computer program is configured to execute the above audio generation method when running.
[0031] According to a fifth aspect of the present disclosure, there is also provided a computer program product, including a game program. When the game program is executed, the computer is caused to execute the steps of the above audio generation method.
[0032] According to a sixth aspect of the present disclosure, there is also provided an electronic device, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to execute the above audio generation method through the computer program.
[0033] One beneficial effect of the embodiments of the present disclosure is that the audio generation method provided by the embodiments of the present disclosure can, after the user provides the first reference audio and the reference text, extract the prosody of the first reference audio through a speech conversion model, and obtain a speech identifier in which the prosody matches the text content of the reference text. Then, through a fusion model, the target timbre reflected by the set timbre data is fused into the speech identifier, and further a target audio with the target timbre is obtained, so as to realize synthesizing audio without using a fixed voice line or a replicated authorized voice line, to meet the different needs of different users, and effectively improve the user experience.
[0034] Through the following detailed description of the exemplary embodiments of the present specification with reference to the accompanying drawings, the features and advantages of the embodiments of the present specification will become clear. Description of the Drawings
[0035] The accompanying drawings incorporated in and constituting a part of this specification illustrate embodiments of the specification and, together with the description, serve to explain the principles of the embodiments of the specification.
[0036] Figure 1 FIG. shows a schematic hardware structure diagram of an electronic device that can be used to implement the audio generation method according to an embodiment of the present disclosure;
[0037] Figure 2 FIG. shows a schematic flowchart of an audio generation method according to some embodiments;
[0038] Figure 3 FIG. shows a topology diagram embodying a voice conversion model and a fusion model according to some embodiments;
[0039] Figure 4 FIG. shows a schematic flowchart of the training of a fusion model according to some embodiments;
[0040] Figure 5 FIG. shows a schematic structural diagram of an audio generation device according to some embodiments;
[0041] Figure 6 FIG. shows a schematic hardware structure diagram of an electronic device according to some embodiments. Detailed Embodiments
[0042] Various exemplary embodiments of the present specification will now be described in detail with reference to the accompanying drawings.
[0043] The following description of at least one exemplary embodiment is merely illustrative in nature and is in no way intended to limit the embodiments of the present specification, their applications, or uses.
[0044] It should be noted that like reference numerals and letters denote like items in the following drawings, and thus, once an item is defined in one drawing, it need not be further discussed in subsequent drawings.
[0045] It should be noted that all actions of obtaining signals, information, or data in the embodiments of the present disclosure are carried out on the premise of complying with the corresponding data protection regulations and policies of the country where the location is located and obtaining the authorization given by the corresponding device owner.
[0046] The embodiments of the present disclosure provide a new audio generation solution, which allows a user to provide an audio and text, extract the prosody of the audio and the text content of the text, and in a voice conversion model, fuse the prosody of the audio into the text content to obtain a voice identifier. Then, inputting the timbre reflected by a timbre data and the voice identifier into a fusion model, an audio with the timbre reflected by the timbre data and conforming to the text content can be obtained.
[0047] Figure 1 FIG. 1 shows a schematic diagram of the hardware structure of an electronic device that can be used to implement the audio generation method according to an embodiment of the present disclosure.
[0048] The electronic device 1000 is a device capable of running audio processing software, which can be a local application installed on the electronic device, or a web application, a lightweight application, or a small program, etc., which is not limited herein. The electronic device 1000 can be a mobile phone, a tablet computer, a PC, etc., which is not limited herein.
[0049] As Figure 1 shown, the electronic device 1000 may include a processor 1101, a memory 1102, an interface device 1103, a communication device 1104, an output device 1105, an input device 1106, and so on. Figure 1 The hardware configuration shown is merely illustrative and is in no way intended to limit the present disclosure, its application, or its use.
[0050] The processor 1101 is used to execute a computer program, which can be written in an instruction set such as x86, Arm, RISC, MIPS, SSE, etc. The memory 1102 includes, for example, a ROM (read-only memory), a RAM (random access memory), a non-volatile memory such as a hard disk, etc. The interface device 1103 includes, for example, a USB interface, a network cable interface, a headphone interface, etc. The communication device 1104 can perform wired or wireless communication, and the communication device 1104 may include at least one short-range communication module, for example, any module that performs short-range wireless communication based on short-range wireless communication protocols such as the Hilink protocol, WiFi (IEEE 802.11 protocol), Mesh, Bluetooth, ZigBee, Thread, Z-Wave, NFC, UWB, LiFi, etc. The communication device 1104 may also include a remote communication module, for example, any module that performs WLAN, GPRS, 2G / 3G / 4G / 5G remote communication. The output device 1105 may include, for example, a liquid crystal display screen or a touch display screen, a speaker, etc. The input device 1106 may include, for example, a touch screen, a keyboard, a microphone, various sensors, etc.
[0051] In this embodiment, the memory 1102 of the electronic device 1000 is used to store a computer program, which is used to control the processor 1101 to operate to execute the audio generation method according to any embodiment of the present disclosure.
[0052] Next, taking the electronic device 1000 as shown in Figure 1 as an implementation subject, various embodiments of the audio generation method will be described.
[0053] <First Embodiment>
[0054] Figure 2 The figure shows an audio generation method according to some embodiments. The audio generation method may include the following steps S210 to S230:
[0055] Step S210: Obtain a first reference audio and a reference text.
[0056] In this embodiment, the first reference audio may be any piece of audio, which may be the audio in a video file, a given Mel spectrogram, or a string representing audio, and is not limited herein.
[0057] In this embodiment, the reference text includes text and a string reflecting the text, that is, the reference text may be any piece of text or a string representing a piece of text, and is not limited herein.
[0058] Step S220: Input the prosody identifier of the first reference audio and the text identifier of the reference text into a pre-set speech conversion model to obtain a speech identifier in which the prosody of the first reference audio matches the text content of the reference text.
[0059] In this embodiment, the prosody identifier of the first reference audio reflects audio features such as the pitch, rhythm, stress, and speech rate of the first reference audio, and further expresses information such as emotional mood, intention, semantic structure, and grammar.
[0060] In this embodiment, the text identifier of the reference text reflects text features such as sentence segmentation and pauses of the text content.
[0061] In this embodiment, as Figure 3 shown, the speech conversion model may be an autoregressive conversion model (LLM-Autoregressive Transformer). When the speech conversion model is an autoregressive conversion model, through self-supervised learning of a large-scale text, high-quality speech generation, understanding, and reasoning capabilities are achieved, and further, the prosody identifier of the first reference audio and the text identifier of the reference text are converted into speech identifiers (Speech Tokens) that meet the requirements.
[0062] Step S230: Input the speech identifier and set timbre data into a pre-set fusion model to obtain a target audio with a target timbre; wherein, the timbre reflected in the timbre data is the target timbre.
[0063] In some examples, if the timbre data is, for example, a segment of audio data extracted from Jay Chou's "Nunchucks", then the target timbre is the timbre imitating "Jay Chou".
[0064] In some embodiments, the fusion model can be a Condition Flow Matching model, which can seamlessly integrate the set timbre data as conditional information into the flow matching framework to enable the target timbre to be fused into the voice identifier, thereby obtaining a target audio with the target timbre, conforming to the rhythm of the first reference audio, and conforming to the text content of the reference text.
[0065] According to the audio generation method of the first embodiment of the present invention, it solves the problem in the prior art that converting the text input by the user into a fixed voice line or replicating an authorized voice line only obtains an audio with a single rhythm and timbre, which cannot meet the different needs of different users. Based on this method, after the user provides the first reference audio and the reference text, through the voice conversion model, the rhythm of the first reference audio is extracted, and a voice identifier that matches the text content of the rhythm and the reference text is obtained. Then, through the fusion model, the target timbre reflected by the set timbre data is fused into the voice identifier, thereby obtaining a target audio with the target timbre, so as to realize the synthesis of audio without using a fixed voice line or replicating an authorized voice line, which is applicable to the different needs of different users and effectively improves the user experience.
[0066] <Second Embodiment>
[0067] In this embodiment, in order to extract the rhythm identifier of the first reference audio and the text identifier of the reference text to obtain a voice identifier that matches the text content of the rhythm of the first reference audio, the rhythm identifier of the first reference audio can be obtained through an audio encoding model, and the text identifier of the reference text can be obtained through a text encoding model.
[0068] In these embodiments, compared with the above first embodiment, before step S220, the method further includes the following steps S310 and S320:
[0069] Step S310: Input the first reference audio into a preset audio encoding model to obtain a first rhythm identifier corresponding to the first reference audio.
[0070] In this embodiment, as Figure 3 shown, the audio encoding model can be a multilingual encoder (SpeechTokenizer), which is fine-tuned and optimized based on the Whisper Automatic Speech Recognition model (Whisper ASR). The audio encoding model can encode the input first reference audio to obtain a first rhythm identifier, that is, rhythm tokens. Here, the rhythm tokens are the numbers corresponding to the basic units of the feature vectors of the rhythm expression.
[0071] Step S320: Input the reference text into a pre-set text encoding model to obtain a first text identifier corresponding to the reference text.
[0072] In this embodiment, as Figure 3 shown, the text encoding model can be a text encoder (TextTokenizer). This text encoding model can encode the input reference text to obtain a first text identifier, that is, text tokens. Here, the text tokens are the numbers corresponding to the basic units of the feature vectors of the text expression.
[0073] <Third Embodiment>
[0074] In this embodiment, to meet the user's needs for different dubbing timbres, the user is supported to input a second reference audio. Through a timbre encoder, the timbre data of the second reference audio is extracted to determine the target timbre of the timbre data.
[0075] In these embodiments, compared with the above first embodiment, before step S230, the method further includes the following steps S410 and S420:
[0076] Step S410: Obtain a second reference audio; wherein, the second reference audio is used to represent the target timbre.
[0077] In this embodiment, if the second reference audio is, for example, "Double Section Staff" by Jay Chou, then the target timbre is the timbre imitating "Jay Chou".
[0078] Step S420: Input the second reference audio into a pre-set timbre encoder to obtain set timbre data.
[0079] In this embodiment, the timbre encoder has a multi-layer neural network with multi-head attention. The second reference audio is encoded and compressed into a 512-dimensional feature vector by the multi-layer neural network with multi-head attention, and then through the bottleneck architecture of the timbre encoder, the information representing the text content and rhythm in the second reference audio is eliminated, so that the obtained feature vector, that is, the timbre data, can represent the target timbre. In other words, by setting the timbre encoder, the timbre data of the second reference audio is extracted to determine the target timbre of the timbre data, thereby realizing timbres suitable for different user needs.
[0080] <Fourth Embodiment>
[0081] In this embodiment, in order to improve the accuracy of the target audio output by the fusion model, the fusion model needs to be trained.
[0082] In these embodiments, compared with the first embodiment above, before step S210, the method further includes the following steps S510 and S520:
[0083] Step S510, obtain noise data and condition information associated with the noise data.
[0084] In this embodiment, the noise data may be a Mel spectrogram, and the condition information may be a constraint condition or a training objective. For example, the training objective of the Mel spectrogram is target prosody Tokens and target timbre Tokens.
[0085] Step S520, use the noise data and the condition information as inputs to train a preset fusion model to obtain a trained fusion model.
[0086] In this embodiment, by using the noise data and the condition information as inputs, it is possible to achieve precise control over the training process of the fusion model, so that the subsequent fusion model can efficiently generate target audio that meets complex condition constraints.
[0087] <Fifth Embodiment>
[0088] In this embodiment, in order to further improve the accuracy of the target audio output by the fusion model, the noise data may include a Mel spectrogram, and by using the Mel spectrogram as an input.
[0089] In these embodiments, compared with the fourth embodiment above, step S520 may include the following steps S610 to S630:
[0090] Step S610, input the Mel spectrogram into a preset audio encoding model to obtain a second prosody identifier in the Mel spectrogram.
[0091] In this embodiment, as Figure 4 shown, the Mel spectrogram has a noise segment Noise, a prosody segment Token, and a timbre segment Spk. Through the audio encoding model, the second prosody identifier of the prosody segment Token of the input Mel spectrogram is extracted.
[0092] Step S620, input the Mel spectrogram into a preset timbre encoder to obtain a timbre identifier in the Mel spectrogram.
[0093] In this embodiment, as Figure 4 shown, through the timbre encoder, the timbre identifier of the timbre segment Spk of the input Mel spectrogram is extracted.
[0094] Step S630: Use the condition information as the condition input of a pre-set fusion model, and use the second prosody identifier and the timbre identifier as the sample input of the fusion model to train the fusion model and obtain a trained fusion model.
[0095] In this embodiment, as Figure 4 shown, the condition information may include target prosody Tokens output by an audio encoding model and target timbre Tokens output by a timbre encoder. The target prosody Tokens output by the audio encoding model, the target timbre Tokens output by the timbre encoder, and the Mel spectrogram are used as context conditions (In-Context Conditioning) to achieve the purpose of training a condition flow matching model.
[0096] <Sixth Embodiment>
[0097] In this embodiment, in order to further improve the accuracy of the target audio output by the fusion model, the Mel spectrogram also has a masked segment, which forces the fusion model to learn the rules of audio generation from the condition information, effectively improving the sensitivity of the fusion model to context conditions and effectively improving the accuracy of the target audio generated by the fusion model.
[0098] In these embodiments, compared with the above-mentioned first embodiment, step S630 may include the following steps S710 and S720:
[0099] Step S710: Construct the context conditions of a pre-set fusion model through the second prosody identifier, the timbre identifier, and the masked segment.
[0100] In this embodiment, as Figure 4 shown, the second prosody identifier Token, the timbre identifier Spk, and the masked segment Masked Mel are used to construct the context condition Condition of the fusion model.
[0101] Step S720: Use the condition information as the condition input of the fusion model, and use the context condition as the sample input of the fusion model to train the fusion model and obtain a trained fusion model.
[0102] In this embodiment, by using the Mel spectrogram with a masked segment as the context condition to restore the complete Mel spectrogram, the speech-infilling function can be realized to complement and correct the audio when the audio is damaged or noisy, effectively enhancing the robustness of the fusion model.
[0103] <Device Embodiment>
[0104] Figure 5 Shows a schematic structural diagram of an audio generation device according to an embodiment of the present disclosure. As Figure 5 shown, the audio generation device 500 includes an acquisition module 510, a first obtaining module 520, and a second obtaining module 530.
[0105] The acquisition module 510 is configured to acquire a first reference audio and a reference text;
[0106] The first obtaining module 520 is configured to input the prosody identifier of the first reference audio and the text identifier of the reference text into a preset speech conversion model to obtain a speech identifier in which the prosody of the first reference audio matches the text content of the reference text;
[0107] The second obtaining module 530 is configured to input the speech identifier and set timbre data into a preset fusion model to obtain a target audio with a target timbre; wherein, the timbre reflected in the timbre data is the target timbre.
[0108] In some embodiments, the audio generation device 500 further includes an identifier obtaining module, which is configured to input the first reference audio into a preset audio encoding model to obtain a first prosody identifier corresponding to the first reference audio; input the reference text into a preset text encoding model to obtain a first text identifier corresponding to the reference text.
[0109] In some embodiments, the audio generation device 500 further includes a timbre obtaining module, which is configured to acquire a second reference audio; wherein, the second reference audio is used to characterize the target timbre; input the second reference audio into a preset timbre encoder to obtain set timbre data.
[0110] In some embodiments, the audio generation device 500 further includes a training module, which is configured to acquire noise data and condition information associated with the noise data; use the noise data and the condition information as inputs to train a preset fusion model to obtain a trained fusion model.
[0111] In some embodiments, the training module is further configured to input the Mel spectrogram into a preset audio encoding model to obtain a second prosody identifier in the Mel spectrogram; input the Mel spectrogram into a preset timbre encoder to obtain a timbre identifier in the Mel spectrogram; use the condition information as a condition input of a preset fusion model, and use the second prosody identifier and the timbre identifier as sample inputs of the fusion model to train the fusion model to obtain a trained fusion model.
[0112] In some embodiments, the training module is further configured to construct context conditions of a preset fusion model by using the second prosody identifier, the timbre identifier, and the masking segment; use the condition information as the conditional input of the fusion model, use the context conditions as the sample input of the fusion model, and train the fusion model to obtain a trained fusion model.
[0113] <Device Embodiment>
[0114] Figure 6 FIG. shows a schematic hardware structure diagram of an electronic device according to some other embodiments. As Figure 6 shown, the electronic device 600 includes a processor 610 and a memory 620. The memory 620 is configured to store a computer program, and the computer program is configured to control the processor 610 to operate so as to control the electronic device 600 to execute the audio generation method according to any embodiment of the present disclosure.
[0115] Embodiments of the present disclosure also provide a computer-readable storage medium storing a computer program which, when executed by a processor, implements the audio generation method according to any embodiment of the present disclosure.
[0116] Embodiments of the present disclosure also provide a computer program product including a computer program or instructions which, when executed by a processor, implement the audio generation method according to any embodiment of the present disclosure.
[0117] Each embodiment in this specification is described in a progressive manner. Similar parts among the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for device and equipment embodiments, since they are basically similar to method embodiments, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiments.
[0118] The specific embodiments of this specification have been described above. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in a different order than in the embodiments and still achieve the desired results. Additionally, the processes depicted in the figures do not necessarily require the particular order or sequential order shown to achieve the desired results. In certain implementations, multitasking and parallel processing are also possible or may be advantageous.
[0119] Embodiments of this specification can be devices, methods, and / or computer program products. The computer program products can include a computer-readable storage medium having computer-readable program instructions thereon for causing a processor to implement various aspects of the embodiments of this specification.
[0120] A computer-readable storage medium can be a tangible device that can retain and store instructions for use by an instruction execution device. A computer-readable storage medium may be, for example—but not limited to—an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer-readable storage medium include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), a memory stick, a floppy disk, a mechanically encoded device such as a punched card or raised structures in grooves having instructions stored thereon, and any suitable combination of the foregoing. The computer-readable storage medium as used herein is not construed to be a transitory signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., an optical pulse through an optical fiber cable), or an electrical signal transmitted through a wire.
[0121] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to respective computing / processing devices, or can be downloaded to an external computer or an external storage device through a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network may include a copper transmission cable, an optical fiber transmission, a wireless transmission, a router, a firewall, a switch, a gateway computer, and / or an edge server. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium in each computing / processing device.
[0122] The computer program instructions for performing the operations of the embodiments of this specification may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-related instructions, microcode, firmware instructions, state setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk, C++, etc., and conventional procedural programming languages such as the "C" language or similar programming languages. The computer-readable program instructions may be executed entirely on the first user computer, partially on the first user computer, executed as a stand-alone software package, partially on the first user computer and partially on a remote computer, or entirely on the remote computer or server. In the case of a remote computer, the remote computer may be connected to the first user computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., using an Internet service provider to connect through the Internet). In some embodiments, by using the state information of the computer-readable program instructions to customize an electronic circuit, such as a programmable logic circuit, a field-programmable gate array (FPGA), or a programmable logic array (PLA), the electronic circuit can execute the computer-readable program instructions to implement various aspects of the embodiments of this specification.
[0123] Aspects of the embodiments of this specification are described herein with reference to the flowcharts and / or block diagrams of methods, apparatuses (devices), and computer program products according to the embodiments of this specification. It should be understood that each block of the flowcharts and / or block diagrams, and the combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.
[0124] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine such that when these instructions are executed by the processor of the computer or other programmable data processing device, a device is produced that implements the functions / acts specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium, and these instructions cause the computer, the programmable data processing device, and / or other devices to work in a specific manner. Thus, the computer-readable medium storing the instructions includes a manufacture, which includes instructions for implementing various aspects of the functions / acts specified in one or more blocks of the flowchart and / or block diagram.
[0125] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device, causing a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process such that the instructions executed on the computer, other programmable data processing apparatus, or other device implement the functions / acts specified in one or more boxes of the flowchart and / or block diagram.
[0126] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present specification. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two consecutive blocks may in fact be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending upon the functionality involved. It should also be noted that each block of the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented by special purpose hardware-based systems that perform the specified functions or acts, or by combinations of special purpose hardware and computer instructions. As will be apparent to those of ordinary skill in the art, implementations in hardware, in software, and in combinations of software and hardware are all equivalent.
[0127] The embodiments of the present specification have been described above. The above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope of the described embodiments. The choice of terms used herein is intended to best explain the principles of the embodiments, the practical application, or improvements made to the technology in the market, or to enable other ordinary skill in the art to understand the embodiments disclosed herein.
Claims
1. An audio generation method, the method comprising: Obtaining a first reference audio and a reference text; Inputting the prosody identifier of the first reference audio and the text identifier of the reference text into a preset voice conversion model to obtain a voice identifier in which the prosody of the first reference audio matches the text content of the reference text; Inputting the voice identifier and set timbre data into a preset fusion model to obtain a target audio with a target timbre; wherein, the timbre reflected in the timbre data is the target timbre.
2. The method according to claim 1, wherein, The method further comprises: Inputting the first reference audio into a preset audio encoding model to obtain a first prosody identifier corresponding to the first reference audio; Inputting the reference text into a preset text encoding model to obtain a first text identifier corresponding to the reference text.
3. The method according to claim 1, wherein, The method further comprises: Obtaining a second reference audio; wherein, the second reference audio is used to characterize the target timbre; Inputting the second reference audio into a preset timbre encoder to obtain set timbre data.
4. The method according to claim 1, wherein The fusion model is a conditional flow matching model.
5. The method according to any one of claims 1 to 4, wherein, The method further comprises: Obtaining noise data and conditional information associated with the noise data; Using the noise data and the conditional information as inputs to train a preset fusion model to obtain a trained fusion model.
6. The method according to claim 5, wherein, The noise data includes a Mel spectrogram; the using the noise data and the conditional information as inputs to train a preset fusion model to obtain a trained fusion model comprises: Inputting the Mel spectrogram into a preset audio encoding model to obtain a second prosody identifier in the Mel spectrogram; Inputting the Mel spectrogram into a preset timbre encoder to obtain a timbre identifier in the Mel spectrogram; Using the conditional information as a conditional input to a preset fusion model, and using the second prosody identifier and the timbre identifier as sample inputs to the fusion model to train the fusion model to obtain a trained fusion model.
7. The method according to claim 6, wherein, The Mel spectrogram has a masked segment; the using the conditional information as a conditional input to a preset fusion model, and using the second prosody identifier and the timbre identifier as sample inputs to the fusion model to train the fusion model to obtain a trained fusion model comprises: Constructing a context condition of a preset fusion model through the second prosody identifier, timbre identifier, and the masked segment; Using the conditional information as a conditional input to the fusion model, and using the context condition as a sample input to the fusion model to train the fusion model to obtain a trained fusion model.
8. An audio generation device, wherein, Comprising: An obtaining module, configured to obtain a first reference audio and a reference text; A first obtaining module, configured to input the prosody identifier of the first reference audio and the text identifier of the reference text into a preset voice conversion model to obtain a voice identifier in which the prosody of the first reference audio matches the text content of the reference text; A second obtaining module, configured to input the voice identifier and the set timbre data into a pre-set fusion model to obtain a target audio with a target timbre; wherein, the timbre reflected in the timbre data is the target timbre.
9. An electronic device, wherein, It includes a memory and a processor, the memory is used to store a computer program, and the processor is used to execute the method steps according to any one of claims 1 to 7 under the control of the computer program.
10. A computer-readable storage medium, the computer-readable storage medium including a stored computer program, wherein, When the computer program is running, it executes the method steps according to any one of claims 1 to 7.
Citation Information
Cited By
Voice data generation method based on large model and method for training large model
CN121393416A