Method and apparatus for generating audio, and electronic device and medium

By setting up pre-trained models on the server side and using cloud processing technology, the problems of high model complexity and high computing volume in the existing tone conversion technology are solved, and fast and convenient audio generation is achieved, improving the user experience.

WO2025118859A1PCT designated stage expired Publication Date: 2025-06-12BEIJING ZITIAO NETWORK TECH CO LTD

Patent Information

Application Number
PCT/CN2024/127073
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-04
Filing Date
2024-10-24
Publication Date
2025-06-12

AI Technical Summary

Technical Problem

The existing tone conversion technology has problems such as high model complexity, large computing volume, long training time, and hardware performance limitations, resulting in low audio generation efficiency and poor user experience.

Method used

Setting up a pre-trained model on the server side, users only need to upload a small amount of audio to train, generate a tone conversion model, and realize the cloud processing of audio generation, avoiding complex calculations on user devices.

Benefits of technology

It realizes that users can upload audio conveniently, without being restricted by time, venue and equipment, quickly generate audio of user tone, and improve audio generation efficiency and user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024127073_12062025_PF_FP_ABST
    Figure CN2024127073_12062025_PF_FP_ABST
Patent Text Reader

Abstract

A method and apparatus for generating audio, and an electronic device and a medium. The method (200) comprises: receiving user audio of a user reading reference text, wherein the user audio is used for determining the tone of the user (202); determining target text to be played (204); and generating target audio of reading the target text by using the tone of the user (206).
Need to check novelty before this filing date? Find Prior Art

Description

Method, device, electronic device and medium for generating audio

[0001] This application claims priority to Chinese patent application number 202311649940.9, filed on December 4, 2023, and entitled “Methods, devices, electronic devices and media for generating audio”. The entire contents of this application are incorporated by reference into this application. Technical Field

[0002] Embodiments of the present disclosure relate to the field of computers, and more particularly, to a method, apparatus, electronic device, and medium for generating audio. Background Art

[0003] With the continuous development of information technology, the widespread application of computer devices such as smart phones, tablets and laptops, computer devices are developing in the direction of diversification and personalization. Computer devices can now synthesize voices comparable to real people, enriching the experience of human-computer interaction.

[0004] Common speech processing technologies currently include speech synthesis and speech conversion. Speech synthesis involves extracting timbre information from user-provided speech and synthesizing it using the user's timbre. Speech synthesis technology not only enables text-to-speech conversion for a fixed speaker but also allows for further specification of the speaker's timbre.

[0005] Summary of the Invention

[0006] Embodiments of the present disclosure provide a method, apparatus, electronic device, and medium for generating audio.

[0007] According to a first aspect of the present disclosure, a method for generating audio is provided. The method includes receiving user audio of a user reading a reference text, the user audio being used to determine the user's timbre; determining a target text to be played; and generating target audio reading the target text in the user's timbre.

[0008] In a second aspect of the present disclosure, a device for generating audio is provided. The device includes a user audio receiving module configured to receive user audio of a user reading a reference text, the user audio used to determine the user's timbre; a target text determination module configured to determine a target text to be played; and a target audio generation module configured to generate target audio of the target text read aloud in the user's timbre.

[0009] According to a third aspect of the present disclosure, an electronic device is provided, comprising a processor and a memory coupled to the processor, wherein the memory has instructions stored therein, and when the instructions are executed by the processor, the electronic device executes the method according to the first aspect.

[0010] In a fourth aspect of the present disclosure, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores computer-executable instructions, wherein the computer-executable instructions are executed by a processor to implement the method according to the first aspect.

[0011] This summary is provided to introduce a selection of concepts in a simplified form that are further described below in the detailed description. It is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to limit the scope of the claimed subject matter. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] The above and other features, advantages and aspects of the embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. In the accompanying drawings, the same or similar reference numerals represent the same or similar elements, wherein:

[0013] FIG1 shows a schematic diagram of an example environment in which a method for generating audio according to some embodiments of the present disclosure may be implemented;

[0014] FIG2 shows a flowchart of a method for generating audio according to some embodiments of the present disclosure;

[0015] FIG3 shows a schematic diagram of an interface for generating audio according to some embodiments of the present disclosure;

[0016] FIG4A is a schematic diagram showing an interface when a user reads a reference text according to some embodiments of the present disclosure;

[0017] FIG4B shows a schematic diagram of an interface during the generation of a timbre conversion model according to some embodiments of the present disclosure;

[0018] FIG5 shows a framework diagram of a process for generating audio according to some embodiments of the present disclosure;

[0019] FIG6 shows a flowchart of training a pre-trained model according to some embodiments of the present disclosure;

[0020] FIG7 shows a flowchart of generating target audio according to some embodiments of the present disclosure;

[0021] FIG8 shows a block diagram of an apparatus for generating audio according to some embodiments of the present disclosure; and

[0022] FIG9 shows a block diagram of an electronic device according to some embodiments of the present disclosure.

[0023] Throughout the drawings, the same or similar reference numbers denote the same or similar elements. DETAILED DESCRIPTION

[0024] It is understandable that before using the technical solutions disclosed in the various embodiments of this disclosure, the type, scope of use, usage scenarios, etc. of the personal information involved in this disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.

[0025] For example, upon receiving a user's active request, a prompt message is sent to the user to clearly inform the user that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the electronic device, application, server, storage medium, or other software or hardware that performs the operations of the disclosed technical solution based on the prompt message.

[0026] As an optional but non-limiting implementation, in response to receiving a user's active request, the prompt information may be sent to the user in the form of a pop-up window, in which the prompt information may be presented in text form. Furthermore, the pop-up window may also contain a selection control for the user to select "agree" or "disagree" to provide personal information to the electronic device.

[0027] It is understandable that the above notification and user authorization process are merely illustrative and do not limit the implementation of the present disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of the present disclosure.

[0028] It is understandable that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) must comply with the requirements of relevant laws, regulations and relevant provisions.

[0029] The following describes embodiments of the present disclosure in more detail with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments described herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.

[0030] In the description of the embodiments of the present disclosure, the term "including" and similar terms should be understood as open inclusion, that is, "including but not limited to". The term "based on" should be understood as "based at least in part on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The terms "first", "second", etc. can refer to different or the same objects, unless explicitly stated otherwise. Other explicit and implicit definitions may also be included below.

[0031] Traditional voice conversion methods primarily convert user-uploaded voices based on statistical models. Existing statistical models include Gaussian mixture models, hidden Markov models, and dynamic time warping. These models all learn from voice data to build a voice conversion model, which is then used to predict the input voice signal to produce the desired voice.

[0032] It should be understood that once a voice conversion model is established, the user can reuse it multiple times. When the user subsequently uses the model for voice conversion, the model can synthesize speech using the user's voice without storing learned voice data or undergoing secondary training. It should be understood that the voice of the synthesized speech is a voice already in the model's sound library or a voice authorized by the user.

[0033] Traditional timbre conversion uses highly complex models and requires extensive computation, often requiring hours of training to generate audio that meets the desired timbre requirements. Furthermore, local model execution is limited by its size, requiring significant CPU or GPU resources for training. This complexity exceeds the performance capabilities of the user's device hardware. Due to the limitations of hardware performance, timbre conversion cannot capture audio of a specified timbre in all scenarios. Furthermore, audio acquisition takes a long time, resulting in low timbre conversion efficiency, which degrades the user experience.

[0034] In order to solve the above problems, the embodiments of the present disclosure provide a solution for generating audio. The solution constructs a recording interface on the web page, and users can record audio and select target text in the cloud. The target text and the recorded audio are transmitted to the server. The server generates a target audio that reads the target text in the user's timbre based on the audio and target text recorded by the user, and transmits the target audio to the user. Users can perform timbre conversion conveniently and flexibly without being restricted by time, place, and equipment. The solution sets a pre-trained model on the server side. Users only need to upload a small amount of audio to train the pre-trained model in a short time and obtain a timbre conversion model for generating target audio. The use of this solution enables users to upload audio conveniently without being restricted by time, place, and equipment, and can quickly generate audio with the user's timbre, thereby improving the efficiency of audio generation and user experience.

[0035] Figure 1 shows a schematic diagram of an example environment 100 in which the method for generating audio according to some embodiments of the present disclosure can be implemented. As shown in Figure 1, the example environment 100 may include a user device 101, which may be any device with computing hardware, such as a computer device such as a smartphone, a tablet computer, and a laptop computer. The user device 101 may include a user interface, which may be a touch screen display, and the user may input commands to the client by touching and / or gesturing on the display screen. In some embodiments, a first control 102 may be displayed on the user interface, and the user may implement cloud recording of the user audio by touch operation on the first control 102. In some embodiments, text content 102 may be displayed on the user interface, and the user may select the desired text content as the target text to generate target audio read aloud by the user's voice.

[0036] 1 , the example environment 100 may further include a server 104, which may be a single server or a centralized or distributed server cluster (e.g., a cloud). In the embodiments of the present disclosure, the recorded user audio may be obtained by the server. It should be understood that the user audio may also be obtained by any other device with computing capabilities. The architecture and functionality of the server 104 are described herein for exemplary purposes only and do not imply any limitation on the scope of the present disclosure.

[0037] In some embodiments, the server 104 may be configured with a pre-trained model 105 for generating a target video. The process of generating the target video includes the server receiving user audio recorded in the cloud and inputting it into the pre-trained model 105. The pre-trained model 105 is trained based on the recorded user audio to generate a voice-to-sound conversion model. The voice-to-sound conversion model generates a target audio based on the text content selected by the user, i.e., the target text, and transmits the target audio to the user device 101.

[0038] From the above description, it can be seen that the solution disclosed in the present invention is to build a recording interface on the web page. Users can record audio in the cloud, set a pre-trained model on the server side, train the pre-trained model according to the user audio, and obtain a timbre conversion model for generating the target audio. The adoption of this solution enables users to upload audio conveniently without being restricted by time, place and equipment, and can quickly generate audio with the user's timbre, thereby improving the efficiency of audio generation and user experience.

[0039] It should be understood that the architecture and functions in the example environment 100 are described for exemplary purposes only and do not imply any limitation on the scope of the present disclosure. The embodiments of the present disclosure may also be applied to other environments with different structures and / or functions.

[0040] The following describes the process of the embodiment of the present disclosure in detail with reference to Figures 2 to 9. For ease of understanding, the specific data mentioned in the following description are exemplary and are not intended to limit the scope of protection of the present disclosure. It is understood that the embodiments described below may also include additional actions not shown and / or may omit the actions shown, and the scope of the present disclosure is not limited in this respect.

[0041] FIG2 shows a flow chart of a method 200 for generating audio according to some embodiments of the present disclosure. At block 202, user audio of a user reading a reference text is received, and the user audio is used to determine the user's timbre. For example, referring to FIG1 , the user audio can be recorded on a web page, and the microphone on the web page can be used to sample the user's input audio in real time. The user only needs to perform a touch operation on the first control 102 to record the sound. Based on the timbre of the recorded user audio, the timbre of the target audio to be converted can be determined to generate audio with a personal timbre.

[0042] At block 204 , a target text to be played is determined. For example, as shown in FIG1 , a user interface may display multiple text contents, and a user may select from the multiple text contents. The selected text content is the target text, and the target text may be converted into a target audio having the user's timbre by the server 104 .

[0043] At block 206, target audio is generated that reads the target text in the user's voice. For example, as shown in FIG1 , the pre-trained model 105 can be generated by training a speech model based on a multi-person voice dataset, which includes multiple audio files with different voices. After pre-training, the pre-trained model 105 only requires a small amount of user audio input to complete training. During the training process, the parameters of the pre-trained model 105 can be adjusted, and training can be completed in a short period of time to generate a voice conversion model. The voice conversion model generates target audio that reads the target text in the user's voice based on the target text selected by the user.

[0044] Therefore, according to the method 200 of the embodiment of the present disclosure, the method constructs a recording interface on the web page side, and the user can record audio and select target text in the cloud. The target text and the recorded audio are transmitted to the server, and the server generates a target audio that reads the target text in the user's timbre based on the audio and target text recorded by the user, and transmits the target audio to the user. The user can perform timbre conversion conveniently and flexibly without being restricted by time, place and equipment. The method sets a pre-trained model on the server side, and the user only needs to upload a small amount of audio to train the pre-trained model and obtain a timbre conversion model for generating the target audio. The use of this solution enables users to upload audio conveniently without being restricted by time, place and equipment, and can quickly generate audio with the user's timbre, thereby improving the efficiency of audio generation and user experience.

[0045] The following will specifically illustrate the process of target audio generation with reference to Figures 3 to 7. In the embodiment of the present disclosure, the explanation is given in the order of the user audio recording interface, the timbre conversion model generation interface, the target audio generation framework, the timbre conversion model training process, and the target audio generation process. The specific data mentioned in the following description are all exemplary and are not intended to limit the scope of protection of the present disclosure. It will be understood that the embodiments described below may also include additional actions not shown and / or the actions shown may be omitted, and the scope of the present disclosure is not limited in this respect.

[0046] Figure 3 shows a schematic diagram of an interface 300 for generating audio in some embodiments of the present disclosure. The interface 300 may include a user interface 301. Text content 302 is displayed on the user interface 301. The text content 302 may be a sample text provided by a web page or may be content that the user searches for and needs to be converted into a target audio. The user may select a target text from the text content 302 to generate the target audio. The user interface 301 also displays a first control 304 and a waveform display area 303. The user may record the user audio by touching the first control 304. During the recording process, the size of the audio waveform will be displayed in the waveform display area 303 to prompt the user whether the sound input is appropriate. The user interface 301 also displays a second control 305. The user may audition the recorded audio by touching the second control 305.

[0047] Figure 4A illustrates a schematic diagram of an interface 400A for a user reading reference text according to some embodiments of the present disclosure. When a user touches the first control 304 in Figure 3 , the interface shown in Figure 4A is displayed on the user interface 401. The user can follow the prompts to read the reference text on the user interface 401. The cloud-recorded audio of the user reading the reference text serves as the user audio. The user interface 401 also displays a third control 402, a fourth control 403, and a fifth control 404. The user can touch the third control 402 to re-record, touch the fourth control 403 to pause recording, and touch the fifth control 404 to complete recording.

[0048] FIG4B shows a schematic diagram of an interface 400B generated by a timbre conversion model according to some embodiments of the present disclosure. When a user touches the fifth control 404 in FIG4A , an interface as shown in FIG4B is displayed on the user interface 405. After the user audio is recorded and transmitted to the pre-trained model of the server, while the pre-trained model is being trained based on the user audio, the user interface 405 is shown in FIG4B . The user interface 405 displays a progress bar for auditioning the recorded user audio and a sixth control 406. The user can audition the audio and jump to playback. The user can also touch the sixth control 406 to pause the auditioned user audio.

[0049] Figure 5 shows a schematic diagram of a framework diagram 500 of the process of generating audio according to some embodiments of the present disclosure. User audio 501 may include a first user audio 502, a second user audio 503, and an Nth user audio 504. It should be understood that user audio 501 may include multiple different audio segments, and the number of audio segments may be selected according to actual needs, specifically based on meeting the purpose of model training. The user audio 501 recorded in the cloud is input into a pre-trained model 505, and the pre-trained model 505 is trained based on the recorded user audio to generate a timbre conversion model 506. In specific implementation, multiple text contents may be displayed on the user interface, and the user may select a target text 507 from the multiple text contents. After the target text 507 is input into the server, it will be converted into speech. The text-to-speech 508 method may adopt an existing commonly used text-to-speech method, such as an end-to-end architecture based on a variational autoencoder (VAE) or a generative adversarial network (GAN) and an end-to-end architecture based on a speech synthesis model, etc., which can be selected according to actual needs. The target text 508 converted into speech is input into the timbre conversion model 506. The timbre conversion model 506 performs timbre conversion on the input audio to generate target audio 509 in which the user reads the target text in the timbre. After the target audio 509 is generated, it can be pushed to the front end 510, that is, the user can accept the generated target audio 509, or it can be stored as a static resource 511 in the cloud.

[0050] Figure 6 shows a flowchart of training a pre-trained model according to some embodiments of the present disclosure. At box 601, user audio is recorded. The user can record user audio on the web page. The audio here can be audio with a low sampling frequency. The audio recording of the present disclosure does not require professional equipment. It can be recorded on the web page, which improves the convenience of user recording. At box 602, the spectrum information is extracted. After the user audio recorded in the cloud is transmitted to the server, the server maps the user audio with the spectrum and extracts the spectrum information in the user audio. At box 603, the pre-trained model is input. The pre-trained model can be trained based on the spectrum information in the user audio.

[0051] At block 604, sound features and semantic features are extracted. In the embodiment of the present disclosure, sound features may include pitch features, waveform features, and tone features, among others. The present disclosure also extracts semantic features from the user audio, sets parameters corresponding to the semantics in the pre-trained model, and during training, adjusts the semantic parameters to enable the trained timbre conversion model to have semantic understanding capabilities. Specifically, the timbre conversion model trained based on semantic features can understand the semantic information in the speech signal, and the tone of the converted audio can match the semantics of the target text, conforming to the characteristics of human speech.

[0052] At box 605, synthesized audio is generated based on the sound features and semantic features. After a certain number of rounds of training based on the sound features and semantic features, the pre-trained model synthesizes the extracted sound features and semantic features to generate synthesized audio. The synthesized audio can use a speech synthesis model, such as an end-to-end architecture model. In some embodiments of the present disclosure, the process of training the pre-trained model can be to pre-set training rounds, and the training is completed when the number of trained words reaches a preset value. The process of training the pre-trained model can also be to train a certain number of words, generate synthesized audio, and then the user selects the synthesized audio that is closest to his or her own timbre. The pre-trained model corresponding to the selected synthesized audio is the timbre conversion model.

[0053] FIG7 illustrates a flowchart for generating target audio according to some embodiments of the present disclosure. At block 701, audio to be converted from target text is input. During the personalized audio generation process, multiple text segments may be displayed on the user interface, and the user may select a target text from the multiple segments. After the target text is input into the server, it will be converted into speech. The text-to-speech method may employ existing commonly used text-to-speech methods. The target text converted into speech is input into a timbre conversion model. At block 702, sound features and semantic features are extracted. The timbre conversion model extracts the sound features and semantic features from the audio of the standard timbre converted from the target text and performs timbre conversion. The timbre conversion model can understand the semantic information in the speech signal, and the converted audio tone can match the semantics of the target text and conform to the characteristics of human speech. At block 703, target audio is generated based on the sound features and semantic features. The speech synthesis model synthesizes the sound features and semantic features after timbre conversion to generate the target audio.

[0054] FIG8 shows a block diagram of an apparatus 800 for generating target audio according to some embodiments of the present disclosure. As shown in FIG8 , the apparatus 800 includes a user audio receiving module 802, a target text determination module 804, and a target audio generation module 806. The user audio receiving module 802 is configured to receive user audio of a user reading a reference text, wherein the user audio is used to determine the user's timbre. The target text determination module 804 is configured to determine the target text to be played. The target audio generation module 806 is configured to generate target audio of the target text read aloud in the user's timbre.

[0055] Figure 9 shows a block diagram of an electronic device 900 according to some embodiments of the present disclosure, and device 900 can be the device or apparatus described in the embodiments of the present disclosure. As shown in Figure 9, device 900 includes a central processing unit (CPU) and / or a graphics processing unit (GPU) 901, which can perform various appropriate actions and processes according to computer program instructions stored in a read-only memory (ROM) 902 or computer program instructions loaded from a storage unit 908 into a random access memory (RAM) 903. In RAM 903, various programs and data required for the operation of device 900 can also be stored. CPU and / or GPU 901, ROM 902 and RAM 903 are connected to each other via a bus 904. Input and / or output (I and / or O) interface 905 is also connected to bus 904. Although not shown in Figure 9, device 900 can also include a coprocessor.

[0056] Various components in the device 900 are connected to the I and / or O interface 905, including an input unit 906, such as a keyboard, a mouse, etc.; an output unit 907, such as various types of displays, speakers, etc.; a storage unit 908, such as a magnetic disk, an optical disk, etc.; and a communication unit 909, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 909 allows the device 900 to exchange information and / or data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0057] The various methods or processes described above may be performed by the CPU and / or GPU 901. For example, in some embodiments, the methods may be implemented as a computer software program that is tangibly embodied in a machine-readable medium, such as the storage unit 908. In some embodiments, part or all of the computer program may be loaded and / or installed onto the device 900 via the ROM 902 and / or the communication unit 909. When the computer program is loaded into the RAM 903 and executed by the CPU and / or GPU 901, one or more steps or actions in the methods or processes described above may be performed.

[0058] In some embodiments, the methods and processes described above may be implemented as a computer program product. The computer program product may include a computer-readable storage medium carrying computer-readable program instructions for executing various aspects of the present disclosure.

[0059] Computer-readable storage medium can be a tangible device that can keep and store the instructions used by the instruction execution device.Computer-readable storage medium can be, for example, but not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device or any suitable combination thereof.More specific examples (non-exhaustive list) of computer-readable storage medium include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanical encoding device, for example, a punch card or a convex structure in a groove having instructions stored thereon, and any suitable combination thereof.Computer-readable storage medium used herein is not interpreted as a transient signal itself, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagated by waveguides or other transmission media (for example, light pulses by fiber optic cables), or electrical signals transmitted by wires.

[0060] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to each computing and / or processing device, or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. The network adapter card or network interface in each computing and / or processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions to be stored in the computer-readable storage medium in each computing and / or processing device.

[0061] The computer program instructions for performing the disclosed operation can be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-related instructions, microcode, firmware instructions, state setting data or source code or the object code written in any combination of one or more programming languages, programming languages ​​include object-oriented programming languages, and conventional procedural programming languages.Computer-readable program instructions can be performed completely on a user's computer, partially on a user's computer, performed as an independent software package, partly on a user's computer and partly on a remote computer, or performed completely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer by any type of network-including local area network (LAN) or wide area network (WAN), or can be connected to an external computer (such as utilizing an internet service provider to connect by the internet). In certain embodiments, by utilizing the state information of computer-readable program instructions to carry out personalized customization electronic circuits, such as programmable logic circuits, field programmable gate arrays (FPGAs) or programmable logic arrays (PLA), this electronic circuit can perform computer-readable program instructions, thereby realizing various aspects of the present disclosure.

[0062] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine, so that when these instructions are executed by the processing unit of the computer or other programmable data processing device, a device is generated that implements the functions and / or actions specified in one or more blocks in the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium, where these instructions cause the computer, programmable data processing device, and / or other device to operate in a specific manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing various aspects of the functions and / or actions specified in one or more blocks in the flowchart and / or block diagram.

[0063] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device, so that a series of operational steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to implement the functions and / or actions specified in one or more boxes in the flowchart and / or block diagram.

[0064] The flow charts and block diagrams in the accompanying drawings show the possible architecture, functions and operations of the devices, methods and computer program products according to multiple embodiments of the present disclosure. In this regard, each box in the flow chart or block diagram can represent a part of a module, program segment or instruction, and the part of the module, program segment or instruction contains one or more executable instructions for realizing the prescribed logical function. In some alternative implementations, the functions marked in the box can also occur in a sequence different from that marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart, can be implemented by a special hardware-based system that performs the prescribed function or action, or can be implemented by a combination of special hardware and computer instructions.

[0065] The embodiments of the present disclosure have been described above. The above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is selected to best explain the principles of the embodiments, practical applications, or technical improvements to existing technologies, or to enable other persons skilled in the art to understand the embodiments disclosed herein.

[0066] Some example implementations of the present disclosure are listed below.

[0067] Example 1. A method for generating audio, comprising:

[0068] receiving a user audio of a user reading a reference text, wherein the user audio is used to determine the user's timbre;

[0069] determining a target text to be played; and

[0070] A target audio is generated in which the target text is read aloud in the voice of the user.

[0071] Example 2. The method of Example 1, wherein receiving user audio of the user reading the reference text comprises:

[0072] displaying the reference text and the first control; and

[0073] In response to a user touch operation on the first control, audio input by the user is recorded.

[0074] Example 3. The method of any of Examples 1-2, further comprising:

[0075] While the user is reading the reference text, a waveform of the recorded audio is displayed.

[0076] Example 4. The method of any one of Examples 1-3, wherein determining the target text to be played comprises:

[0077] Displaying multiple text contents on the user interface; and

[0078] In response to a user selecting a text content from the plurality of text contents, the target text to be played is determined.

[0079] Example 5. The method of any one of Examples 1-4, wherein determining the user's timbre comprises:

[0080] Based on the user audio, training a pre-trained model to generate a timbre conversion model for the user; and

[0081] Based on the timbre conversion model, the timbre of the target audio is determined.

[0082] Example 6. The method of any one of Examples 1-5, wherein generating the timbre conversion model comprises:

[0083] Determining spectrum information of the user audio;

[0084] Inputting the spectrum information into the pre-trained model;

[0085] Extracting sound features and semantic features from the sound spectrum information; and

[0086] Based on the sound features and the semantic features, the parameters of the pre-trained model are adjusted to generate the timbre conversion model.

[0087] Example 7. The method of any one of Examples 1-6, wherein adjusting parameters of the pre-trained model to generate a timbre conversion model comprises:

[0088] Based on the sound features and the semantic features, the parameters of the preset rounds are adjusted to generate a timbre conversion model.

[0089] Example 8. The method of any one of Examples 1-7, wherein adjusting parameters of the pre-trained model to generate a timbre conversion model comprises:

[0090] Adjusting parameters of a multi-round pre-training model based on the sound features and the semantic features;

[0091] For the speech model after each round of parameter adjustment, synthesizing the sound features and the semantic features to obtain a plurality of synthesized audios; and

[0092] In response to a user selecting one of the plurality of synthesized audios, the timbre conversion model is generated, wherein the timbre conversion model corresponds to the selected synthesized audio.

[0093] Example 9. The method of any one of Examples 1-8, wherein the sound features include pitch features, waveform features, and tone features.

[0094] Example 10. A method according to any one of Examples 1-9, wherein the pre-trained model is generated by training a speech model based on a multi-person timbre dataset, and the multi-person timbre dataset includes audios of multiple different timbres.

[0095] Example 11. The method of any one of Examples 1-10, wherein generating target audio for reading the target text in the user's voice comprises:

[0096] Inputting the speech corresponding to the target text into the timbre conversion model;

[0097] Extracting the sound features and semantic features of the speech corresponding to the target text; and

[0098] The extracted sound features and the semantic features are synthesized to generate the target audio in which the target text is read aloud in the voice of the user.

[0099] Example 12. An apparatus for generating audio, comprising:

[0100] A user audio receiving module is configured to receive user audio of a user reading a reference text, wherein the user audio is used to determine the timbre of the user;

[0101] a target text determination module, configured to determine a target text to be played; and

[0102] The target audio generation module is configured to generate target audio for reading the target text in the user's voice.

[0103] Example 13. The apparatus of Example 12, wherein the user audio receiving module comprises:

[0104] A first control display module configured to display the reference text and a first control; and

[0105] The audio recording module is configured to record audio input by the user in response to a user touch operation on the first control.

[0106] Example 14. The apparatus of any of Examples 12-13, further comprising:

[0107] The waveform display module is configured to display the waveform of the recorded audio while the user is reading the reference text.

[0108] Example 15. The apparatus of any one of Examples 12-14, wherein the target text determination module comprises:

[0109] a text content display module configured to display a plurality of text contents on a user interface; and

[0110] The first target text determination module is configured to determine the target text to be played in response to a user selecting a text content from the multiple text contents.

[0111] Example 16 The apparatus of any one of Examples 12-15, further comprising:

[0112] a timbre conversion model generation module, configured to train a pre-trained model based on the user audio to generate a timbre conversion model for the user; and

[0113] The timbre determination module is configured to determine the timbre of the target audio based on the timbre conversion model.

[0114] Example 17. The apparatus of any one of Examples 12-16, wherein the timbre conversion model generation module comprises:

[0115] a sound spectrum information determination module, configured to determine the sound spectrum information of the user audio;

[0116] A sound spectrum information input module, configured to input the sound spectrum information into the pre-trained model;

[0117] A first feature extraction module is configured to extract sound features and semantic features from the sound spectrum information; and

[0118] The parameter adjustment module is configured to adjust the parameters of the pre-trained model based on the sound features and the semantic features to generate the timbre conversion model.

[0119] Example 18. The apparatus of any one of Examples 12-17, wherein the parameter adjustment module comprises:

[0120] The first parameter adjustment module is configured to adjust the parameters of the preset rounds based on the sound features and the semantic features to generate a timbre conversion model.

[0121] Example 19. The apparatus of any one of Examples 12-18, wherein the parameter adjustment module further comprises:

[0122] A second parameter adjustment module is configured to adjust parameters of the multi-round pre-training model based on the sound features and the semantic features;

[0123] A first audio synthesis module is configured to synthesize the sound features and the semantic features for the speech model after each round of parameter adjustment to obtain a plurality of synthesized audios; and

[0124] The audio selection module is configured to generate the timbre conversion model in response to a user selecting one of the plurality of synthesized audios, wherein the timbre conversion model corresponds to the selected synthesized audio.

[0125] Example 20. The apparatus of any one of Examples 12-19, wherein the sound features include pitch features, waveform features, and tone features.

[0126] Example 21. An apparatus according to any one of Examples 12-20, wherein the pre-trained model is generated by training a speech model based on a multi-person timbre dataset, and the multi-person timbre dataset includes audios of multiple different timbres.

[0127] Example 22. The apparatus of any of Examples 12-21, wherein the target audio generation module comprises:

[0128] A speech input module, configured to input the speech corresponding to the target text into the timbre conversion model;

[0129] A second feature extraction module is configured to extract the sound features and semantic features of the speech corresponding to the target text; and

[0130] The first target audio generation module is configured to synthesize the extracted sound features and the semantic features to generate the target audio for reading the target text in the timbre of the user.

[0131] Example 23. An electronic device comprising:

[0132] processor; and

[0133] a memory coupled to the processor, the memory having instructions stored therein, the instructions causing the electronic device to perform actions when executed by the processor, the actions comprising:

[0134] receiving a user audio of a user reading a reference text, wherein the user audio is used to determine the user's timbre;

[0135] determining a target text to be played; and

[0136] A target audio is generated in which the target text is read aloud in the voice of the user.

[0137] Example 24. The electronic device of Example 23, wherein receiving user audio of the user reading the reference text comprises:

[0138] displaying the reference text and the first control; and

[0139] In response to a user touch operation on the first control, audio input by the user is recorded.

[0140] Example 25. The electronic device of any of Examples 23-24, further comprising:

[0141] While the user is reading the reference text, a waveform of the recorded audio is displayed.

[0142] Example 26. The electronic device of any one of Examples 23-25, wherein determining the target text to be played comprises:

[0143] Displaying multiple text contents on the user interface; and

[0144] In response to a user selecting a text content from the plurality of text contents, the target text to be played is determined.

[0145] Example 27. The electronic device of any of Examples 23-26, wherein determining the user's timbre comprises:

[0146] Based on the user audio, training a pre-trained model to generate a timbre conversion model for the user; and

[0147] Based on the timbre conversion model, the timbre of the target audio is determined.

[0148] Example 28. The electronic device of any of Examples 23-27, wherein generating the timbre conversion model comprises:

[0149] Determining spectrum information of the user audio;

[0150] Inputting the spectrum information into the pre-trained model;

[0151] Extracting sound features and semantic features from the sound spectrum information; and

[0152] Based on the sound features and the semantic features, the parameters of the pre-trained model are adjusted to generate the timbre conversion model.

[0153] Example 29. The electronic device of any of Examples 23-28, wherein adjusting parameters of the pre-trained model to generate a timbre conversion model comprises:

[0154] Based on the sound features and the semantic features, the parameters of the preset rounds are adjusted to generate a timbre conversion model.

[0155] Example 30. The electronic device of any of Examples 23-29, wherein adjusting parameters of the pre-trained model to generate a timbre conversion model comprises:

[0156] Adjusting parameters of a multi-round pre-training model based on the sound features and the semantic features;

[0157] For the speech model after each round of parameter adjustment, synthesizing the sound features and the semantic features to obtain a plurality of synthesized audios; and

[0158] In response to a user selecting one of the plurality of synthesized audios, the timbre conversion model is generated, wherein the timbre conversion model corresponds to the selected synthesized audio.

[0159] Example 31. An electronic device according to any one of Examples 23-30, wherein the sound features include pitch features, waveform features, and tone features.

[0160] Example 32. An electronic device according to any one of Examples 23-31, wherein the pre-trained model is generated by training a speech model based on a multi-person timbre dataset, and the multi-person timbre dataset includes audios of multiple different timbres.

[0161] Example 33. The electronic device of any one of Examples 23-32, wherein generating target audio that reads the target text in the user's voice comprises:

[0162] Inputting the speech corresponding to the target text into the timbre conversion model;

[0163] Extracting the sound features and semantic features of the speech corresponding to the target text; and

[0164] The extracted sound features and the semantic features are synthesized to generate the target audio in which the target text is read aloud in the voice of the user.

[0165] Example 34. A computer-readable storage medium having computer-executable instructions stored thereon, wherein the computer-executable instructions are executed by a processor to implement the method according to any one of claims 1-11.

[0166] Example 35. A computer program product tangibly stored on a computer-readable medium and comprising computer-executable instructions that, when executed by a device, cause the device to perform the method of any one of Examples 1-11.

[0167] Although the present disclosure has been described in language specific to structural features and / or methodological logical acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms of implementing the claims.

Claims

1. A method for generating audio, comprising: receiving a user audio of a user reading a reference text, wherein the user audio is used to determine the timbre of the user; Determine the target text to be played; as well as A target audio is generated in which the target text is read aloud in the voice of the user.

2. The method according to claim 1, wherein receiving user audio of the user reading the reference text comprises: Display the reference text and the first control; as well as In response to a user's touch operation on the first control, audio input by the user is recorded.

3. The method according to claim 2, further comprising: While the user is reading the reference text, a waveform of the recorded audio is displayed.

4. The method according to claim 1, wherein determining the target text to be played comprises: Display multiple text contents on the user interface; as well as In response to a user selecting a text content from the plurality of text contents, the target text to be played is determined.

5. The method of claim 1, wherein determining the user's timbre comprises: Based on the user audio, training the pre-trained model to generate a timbre conversion model for the user; as well as Based on the timbre conversion model, the timbre of the target audio is determined.

6. The method according to claim 5, wherein generating the timbre conversion model comprises: Determining spectrum information of the user audio; Inputting the spectrum information into the pre-trained model; Extracting sound features and semantic features from the spectrum information; as well as Based on the sound features and the semantic features, the parameters of the pre-trained model are adjusted to generate the timbre conversion model.

7. The method according to claim 6, wherein adjusting the parameters of the pre-trained model to generate a timbre conversion model comprises: Based on the sound features and the semantic features, the parameters of the preset rounds are adjusted to generate a timbre conversion model.

8. The method according to claim 6, wherein adjusting the parameters of the pre-trained model to generate a timbre conversion model comprises: Adjusting parameters of multiple rounds of pre-training models based on the sound features and the semantic features; For each round of the speech model after adjusting the parameters, synthesizing the sound features and the semantic features to obtain a plurality of synthesized audios; and In response to a user selecting one of the plurality of synthesized audios, the timbre conversion model is generated, wherein the timbre conversion model corresponds to the selected synthesized audio.

9. The method according to claim 6, wherein the sound features include pitch features, waveform features, and tone features.

10. The method according to claim 6, wherein the pre-trained model is generated by training a speech model based on a multi-person timbre dataset, and the multi-person timbre dataset includes audios of multiple different timbres.

11. The method according to claim 5, wherein generating a target audio for reading the target text in the voice of the user comprises: Inputting the speech corresponding to the target text into the timbre conversion model; Extracting the sound features and semantic features of the speech corresponding to the target text; as well as The extracted sound feature and the semantic feature are synthesized to generate the target audio in which the target text is read aloud in the voice of the user.

12. An apparatus for generating audio, comprising: A user audio receiving module is configured to receive user audio of a user reading a reference text, wherein the user audio is used to determine the timbre of the user; A target text determination module is configured to determine a target text to be played; as well as The target audio generation module is configured to generate a target audio for reading the target text in the voice of the user.

13. An electronic device comprising: processor; as well as A memory coupled to the processor, the memory having instructions stored therein, wherein when the instructions are executed by the processor, the electronic device executes the method according to any one of claims 1-11.

14. A computer-readable storage medium having computer-executable instructions stored thereon, wherein the computer-executable instructions are executed by a processor to implement the method according to any one of claims 1-11.

Citation Information

Patent Citations

  • Audio generation method and device, storage medium and electronic equipment

    CN113205793A

  • Audio synthesis method and device, electronic equipment and readable storage medium

    CN113870828A

  • Speech synthesis method, speech synthesis device, electronic equipment and storage medium

    CN116343747A

  • Speech synthesis method and device, electronic equipment and computer readable storage medium

    CN116645955A

  • Artificial intelligence-based audio processing method and apparatus, device, storage medium, and computer program product

    WO2022252904A1

Cited By

  • Story audio timbre processing method and related device

    CN121148402A