Method and device for generating audio, electronic equipment and medium
By setting up pre-trained models on the server side and recording audio in the cloud, the existing tone conversion technology is solved and the problems of low efficiency and hardware performance limitations are achieved, convenient tone conversion for users and rapid generation of user tone audio, improving the audio generation efficiency and user experience.
Patent Information
- Application Number
- CN202311649940.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-04
- Publication Date
- 2025-06-06
AI Technical Summary
The existing tone conversion technology has problems such as high model complexity, large computing volume, long training time, and hardware performance limitations, resulting in low tone conversion efficiency and inability to obtain audio of the specified tone in any scene, affecting the user experience.
The recording interface is built on the web page side. The user can record audio and select target text in the cloud. The target text and recorded audio are transmitted to the server. The server generates target audio that reads the target text with the user's tone based on the audio and target text recorded by the user and the target text, and transmits the target audio to the user. The server-side sets up a pretrained model, and users only need to upload a small amount of audio to train the pretrained model in a short time, and obtain a tone conversion model used to generate the target audio.
It realizes user convenient tone conversion without being restricted by time, venue and equipment, and can quickly generate audio of user tone, improving audio generation efficiency and user experience.
Smart Images

Figure CN120108376A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure relate to the field of computers, and more particularly, to a method, apparatus, electronic device, and medium for generating audio. Background Art
[0002] With the continuous development of information technology, the widespread application of computer devices such as smart phones, tablets and laptops, computer devices are developing in the direction of diversification and personalization. Computer devices can already synthesize voices comparable to real people, enriching the experience of human-computer interaction.
[0003] Common speech processing technologies currently include speech synthesis and speech conversion. Speech synthesis refers to the technology in which a machine extracts timbre information from the speech provided by the user and synthesizes speech using the user's timbre. Speech synthesis technology can not only achieve text-to-speech conversion on a fixed speaker, but also further specify the speaker's timbre. Summary of the invention
[0004] Embodiments of the present disclosure provide a method, an apparatus, an electronic device, and a medium for generating audio.
[0005] According to a first aspect of the present disclosure, a method for generating audio is provided, which includes receiving user audio of a user reading a reference text, the user audio being used to determine the user's timbre; determining a target text to be played; and generating target audio of the target text being read aloud in the user's timbre.
[0006] In a second aspect of the present disclosure, a device for generating audio is provided. The device includes a user audio receiving module configured to receive user audio of a user reading a reference text, the user audio being used to determine the user's timbre; a target text determining module configured to determine a target text to be played; and a target audio generating module configured to generate target audio of reading the target text in the user's timbre.
[0007] According to a third aspect of the present disclosure, an electronic device is provided, comprising a processor and a memory coupled to the processor, wherein the memory has instructions stored therein, and when the instructions are executed by the processor, the electronic device executes the method according to the first aspect.
[0008] In a fourth aspect of the present disclosure, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores computer-executable instructions, wherein the computer-executable instructions are executed by a processor to implement the method according to the first aspect.
[0009] This Summary is provided to introduce a selection of concepts in a simplified form that are further described below in the Detailed Description. This Summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to limit the scope of the claimed subject matter. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] The above and other features, advantages and aspects of the embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. In the accompanying drawings, the same or similar reference numerals represent the same or similar elements, wherein:
[0011] Figure 1 A schematic diagram showing an example environment in which a method for generating audio according to some embodiments of the present disclosure may be implemented;
[0012] Figure 2 A flowchart of a method for generating audio according to some embodiments of the present disclosure is shown;
[0013] Figure 3 A schematic diagram of an interface for generating audio according to some embodiments of the present disclosure is shown;
[0014] Figure 4A A schematic diagram showing an interface when a user reads a reference text according to some embodiments of the present disclosure;
[0015] Figure 4B A schematic diagram showing an interface in a process of generating a timbre conversion model according to some embodiments of the present disclosure;
[0016] Figure 5 A framework diagram showing a process of generating audio according to some embodiments of the present disclosure;
[0017] Figure 6 A flowchart of training a pre-trained model according to some embodiments of the present disclosure is shown;
[0018] Figure 7 A flowchart of generating target audio according to some embodiments of the present disclosure is shown;
[0019] Figure 8 A block diagram showing an apparatus for generating audio according to some embodiments of the present disclosure; and
[0020] Fig. 9 A block diagram of an electronic device according to some embodiments of the present disclosure is shown.
[0021] Throughout the drawings, the same or similar reference numbers denote the same or similar elements. DETAILED DESCRIPTION
[0022] It is understandable that before using the technical solutions disclosed in the embodiments of the present disclosure, the types, scope of use, usage scenarios, etc. of the personal information involved in the present disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.
[0023] For example, when receiving an active request from a user, a prompt message is sent to the user to clearly prompt the user that the operation requested to be performed will require obtaining and using the user's personal information. Thus, the user can autonomously choose whether to provide personal information to software or hardware such as an electronic device, application, server, or storage medium that performs the operation of the technical solution of the present disclosure according to the prompt message.
[0024] As an optional but non-limiting implementation, in response to receiving an active request from the user, the prompt information may be sent to the user in the form of a pop-up window, in which the prompt information may be presented in text form. In addition, the pop-up window may also carry a selection control for the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0025] It is understandable that the above notification and the process of obtaining user authorization are merely illustrative and do not constitute a limitation on the implementation of the present disclosure. Other methods that meet the relevant laws and regulations may also be applied to the implementation of the present disclosure.
[0026] It is understandable that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and relevant provisions.
[0027] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as being limited to the embodiments described herein, which are instead provided for a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not intended to limit the scope of protection of the present disclosure.
[0028] In the description of the embodiments of the present disclosure, the term "including" and similar terms should be understood as open inclusion, that is, "including but not limited to". The term "based on" should be understood as "based at least in part on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The terms "first", "second", etc. may refer to different or the same objects, unless explicitly stated. Other explicit and implicit definitions may also be included below.
[0029] Traditional timbre conversion methods mainly convert user-uploaded voices based on statistical models. Existing statistical models include Gaussian mixture models, hidden Markov models, dynamic time warping, etc. These models perform timbre conversion by learning voice data to establish a timbre conversion model, and then use the timbre conversion model to predict the input voice signal to obtain the desired timbre conversion.
[0030] It should be understood that after the timbre conversion model is established, the user can use it repeatedly. When the user subsequently uses the model for timbre conversion, the model can use the user's timbre to synthesize speech without storing the learned voice data or performing secondary training. It is understood that the timbre of the synthesized speech is the timbre already in the model sound library or the timbre authorized by the user.
[0031] The models used in traditional timbre conversion are highly complex and require a lot of computation. It often takes several hours of training to generate audio that meets the required timbre conditions. The local model running method is limited by the size of the model. The model training calculations require a lot of CPU or GPU resources to support the model complexity, and the user device hardware performance tolerance. Tone conversion is limited by the performance of the hardware device. It is used to obtain audio of a specified timbre in any scenario. It takes a long time to obtain audio, and the timbre conversion efficiency is low, which will reduce the user experience.
[0032] In order to solve the above problems, an embodiment of the present disclosure provides a solution for generating audio. The solution constructs a recording interface on the web page, and the user can record audio and select target text in the cloud. The target text and the recorded audio are transmitted to the server. The server generates a target audio that reads the target text in the user's timbre based on the audio and target text recorded by the user, and transmits the target audio to the user. The user can perform timbre conversion conveniently and flexibly without being restricted by time, place and equipment. The solution sets a pre-trained model on the server side. The user only needs to upload a small amount of audio to train the pre-trained model in a short time and obtain a timbre conversion model for generating the target audio. The use of this solution enables users to upload audio conveniently without being restricted by time, place and equipment, and can quickly generate audio with the user's timbre, thereby improving the efficiency of audio generation and user experience.
[0033] Figure 1 1 is a schematic diagram of an example environment 100 in which a method for generating audio according to some embodiments of the present disclosure may be implemented. Figure 1As shown, the example environment 100 may include a user device 101, which may be any device with computing hardware, such as a computer device such as a smart phone, a tablet computer, and a laptop computer. The user device 101 may include a user interface, which may be a touch screen display, and the user may input commands to the client by touching and / or gesturing on the display screen. In some embodiments, a first control 102 may be displayed on the user interface, and the user may implement cloud recording of the user audio by touch operation on the first control 102. In some embodiments, the text content 102 may be displayed on the user interface, and the user may select the desired text content as the target text to generate the target audio read aloud by the user's voice.
[0034] refer to Figure 1 , the example environment 100 may further include a server 104, which may be a single server or a server cluster (e.g., cloud) in a centralized or distributed manner. In the embodiment of the present disclosure, the recorded user audio may be obtained by the server. It should be understood that the user audio may also be obtained by any other device with computing capabilities. The architecture and functions of the server 104 are described here only for exemplary purposes, and do not imply any limitation on the scope of the present disclosure.
[0035] In some embodiments, the server 104 may be provided with a pre-trained model 105 for generating a target video. The process of generating the target video includes the server receiving the user audio recorded in the cloud and inputting it into the pre-trained model 105, the pre-trained model 105 is trained based on the recorded user audio, and a voice conversion model is generated, and the voice conversion model generates a target audio according to the text content selected by the user, that is, the target text, and transmits the target audio to the user device 101.
[0036] From the above description, it can be seen that the solution of the present invention is to build a recording interface on the web page. Users can record audio in the cloud, set a pre-trained model on the server side, train the pre-trained model according to the user audio, and obtain a timbre conversion model for generating the target audio. The adoption of this solution enables users to upload audio conveniently without being restricted by time, place and equipment, and can quickly generate audio of the user's timbre, thereby improving the efficiency of audio generation and user experience.
[0037] It should be understood that the architecture and functions in the example environment 100 are described for exemplary purposes only and do not imply any limitation on the scope of the present disclosure. Embodiments of the present disclosure may also be applied to other environments with different structures and / or functions.
[0038] The following will combine Figures 2 to 9The process of the embodiment of the present disclosure is described in detail. For ease of understanding, the specific data mentioned in the following description are exemplary and are not intended to limit the scope of protection of the present disclosure. It is understood that the embodiments described below may also include additional actions not shown and / or may omit the actions shown, and the scope of the present disclosure is not limited in this respect.
[0039] Figure 2 A flow chart of a method 200 for generating audio according to some embodiments of the present disclosure is shown. At block 202, user audio of a user reading a reference text is received, and the user audio is used to determine the user's timbre. For example, the reference text Figure 1 , user audio can be recorded on the web page, and the microphone on the web page can be used to sample the audio input by the user in real time. The user only needs to touch the first control 102 to record the sound. The timbre of the target audio that needs to be converted can be determined based on the timbre of the recorded user audio to generate audio with personal timbre.
[0040] At block 204, the target text to be played is determined. Figure 1 As shown, multiple text contents can be displayed on the user interface, and the user can select from the multiple text contents. The selected text content is the target text, and the target text can be converted into a target audio with the user's timbre through the server 104.
[0041] At block 206, target audio is generated that reads the target text in the user's voice. Figure 1 As shown, the pre-trained model 105 can be generated by training a speech model based on a multi-person timbre data set, and the multi-person timbre data set includes multiple audios with different timbres. After the pre-trained model 105 is pre-trained, only a small amount of user audio is required to input the pre-trained model 105 to implement training. During the training process, the parameters of the pre-trained model 105 can be adjusted, and the training can be completed in a short time to generate a timbre conversion model. The timbre conversion model generates a target audio that reads the target text in the user's timbre according to the target text selected by the user.
[0042] Thus, according to method 200 of an embodiment of the present disclosure, the method constructs a recording interface on the web page, and the user can record audio and select target text in the cloud. The target text and the recorded audio are transmitted to the server, and the server generates a target audio that reads the target text in the user's timbre based on the audio and target text recorded by the user, and transmits the target audio to the user. The user can perform timbre conversion conveniently and flexibly without being restricted by time, place, and equipment. The method sets a pre-trained model on the server side, and the user only needs to upload a small amount of audio to train the pre-trained model and obtain a timbre conversion model for generating the target audio. The use of this solution enables the user to conveniently upload audio without being restricted by time, place, and equipment, and can quickly generate audio with the user's timbre, thereby improving the efficiency of audio generation and user experience.
[0043] The following will combine Figures 3 to 7 The process of generating the target audio is described in detail. In the embodiment of the present disclosure, the explanation is given in the order of the user audio recording interface, the timbre conversion model generation interface, the target audio generation framework, the timbre conversion model training process, and the target audio generation process. The specific data mentioned in the following description are exemplary and are not intended to limit the scope of protection of the present disclosure. It can be understood that the embodiments described below may also include additional actions not shown and / or the actions shown may be omitted, and the scope of the present disclosure is not limited in this respect.
[0044] Figure 3 A schematic diagram of an interface 300 for generating audio in some embodiments of the present disclosure is shown. The interface 300 may include a user interface 301. Text content 302 is displayed on the user interface 301. The text content 302 may be a sample text provided by a web page or may be content that is searched by a user and needs to be converted into a target audio. The user may select a target text from the text content 302 to generate a target audio. The user interface 301 also displays a first control 304 and a waveform display area 303. The user may record the user audio by touching the first control 304. During the recording process, the size of the audio waveform may be displayed in the waveform display area 303 to prompt the user whether the sound input is appropriate. The user interface 301 also displays a second control 305. The user may audition the recorded audio by touching the second control 305.
[0045] Figure 4A A schematic diagram of an interface 400A for a user to read a reference text according to some embodiments of the present disclosure is shown. Figure 3 After the first control 304 in the drawing is touched, the following will be displayed on the user interface 401: Figure 4AThe user can read the reference text on the user interface 401 according to the prompt, and the sound of the user reading the reference text recorded in the cloud is the user audio. The user interface 401 displays a third control 402, a fourth control 403, and a fifth control 404. The user can touch the third control 402 to re-record, touch the fourth control 403 to pause recording, and the user can touch the fifth control 404 to complete recording.
[0046] Figure 4B A schematic diagram of an interface 400B generated by a timbre conversion model according to some embodiments of the present disclosure is shown. Figure 4A After the fifth control 404 is touched, the user interface 405 will display the following Figure 4B After the user audio is recorded, it is transmitted to the pre-trained model of the server. During the training process of the pre-trained model according to the user audio, the user interface 405 is as shown in FIG. Figure 4B As shown, the user interface 405 displays a progress bar for auditioning the recorded user audio and a sixth control 406, and the user can audition and jump to play the audio, and can also touch the sixth control 406 to pause the auditioning user audio.
[0047] Figure 5 A schematic diagram of a framework diagram 500 of a process of generating audio according to some embodiments of the present disclosure is shown. User audio 501 may include a first user audio 502, a second user audio 503, and an Nth user audio 504. It should be understood that user audio 501 may include multiple different audios, and the number of audios may be selected according to actual needs, specifically to meet the purpose of model training. The user audio 501 recorded in the cloud is input into a pre-trained model 505, and the pre-trained model 505 is trained based on the recorded user audio to generate a timbre conversion model 506. In specific implementation, multiple text contents may be displayed on the user interface, and the user may select a target text 507 from the multiple text contents. After the target text 507 is input into the server, it will be converted into speech. The method of converting text to speech 508 may adopt an existing commonly used text-to-speech method, such as an end-to-end architecture based on a variational autoencoder (VAE) or a generative adversarial network (GAN) and an end-to-end architecture based on a speech synthesis model, etc., which may be selected according to actual needs. The target text 508 converted into speech is input into the timbre conversion model 506. The timbre conversion model 506 performs timbre conversion on the input audio to generate a target audio 509 in which the user reads the target text in the timbre. After the target audio 509 is generated, it can be streamed to the front end 510, that is, the user can accept the generated target audio 509, or it can be stored as a static resource 511 in the cloud.
[0048] Figure 6A flow chart of training a pre-trained model according to some embodiments of the present disclosure is shown. At box 601, record user audio. The user can record user audio on the web page. The audio here can be audio with a low sampling frequency. The audio recording of the present disclosure does not require professional equipment and can be recorded on the web page, which improves the convenience of user recording. At box 602, extract spectral information. After the user audio recorded in the cloud is transmitted to the server, the server maps the user audio with the spectrum and extracts the spectral information in the user audio. At box 603, input the pre-trained model. The pre-trained model can be trained according to the spectral information in the user audio.
[0049] At box 604, sound features and semantic features are extracted. In the embodiment of the present disclosure, the sound features may include pitch features, waveform features, and tone features, etc. The present disclosure also extracts semantic features from the user audio, sets parameters corresponding to the semantics in the pre-trained model, and during training, adjusts the semantic parameters so that the trained timbre conversion model has a semantic understanding function. Specifically, the timbre conversion model trained based on semantic features can understand the semantic information in the speech signal, and the tone of the converted audio can match the semantics of the target text, which conforms to the characteristics of human speech.
[0050] At box 605, synthesized audio is generated based on the sound features and semantic features. After a certain number of rounds of training based on the sound features and semantic features, the pre-trained model synthesizes the extracted sound features and semantic features to generate synthesized audio. The synthesized audio can use a speech synthesis model, such as an end-to-end architecture model. In some embodiments of the present disclosure, the process of pre-trained model training can be to pre-set training rounds, and the training is completed when the number of trained words reaches a preset value. The process of pre-trained model training can also be to train a certain number of words, and after generating the synthesized audio, the user selects the synthesized audio that is closest to his or her own timbre. The pre-trained model corresponding to the selected synthesized audio is the timbre conversion model.
[0051] Figure 7A flow chart for generating target audio according to some embodiments of the present disclosure is shown. At box 701, the audio converted from the target text is input. In the process of generating personalized audio, multiple text contents may be displayed on the user interface, and the user may select the target text from the multiple text contents. After the target text is input into the server, it will be converted into speech, and the text-to-speech method may adopt the existing commonly used text-to-speech method. The target text converted into speech is input into the timbre conversion model. At box 702, the sound features and semantic features are extracted. The timbre conversion model extracts the sound features and semantic features in the audio of the standard timbre converted from the target text and performs timbre conversion. The timbre conversion model can understand the semantic information in the speech signal, and the tone of the converted audio can match the semantics of the target text and conform to the characteristics of human speech. At 703, the target audio is generated based on the sound features and semantic features. The speech synthesis model synthesizes the sound features and semantic features after the timbre conversion to generate the target audio.
[0052] Figure 8 8 is a block diagram of an apparatus 800 for generating target audio according to some embodiments of the present disclosure. Figure 8 As shown, the device 800 includes a user audio receiving module 802, a target text determining module 804 and a target audio generating module 806. The user audio receiving module 802 is configured to receive the user audio of the user reading a reference text, wherein the user audio is used to determine the user's timbre. The target text determining module 804 is configured to determine the target text to be played. The target audio generating module 806 is configured to generate the target audio of reading the target text with the user's timbre.
[0053] Fig. 9 1 shows a block diagram of an electronic device 900 according to some embodiments of the present disclosure. The device 900 may be a device or apparatus described in an embodiment of the present disclosure. Fig. 9 As shown, the device 900 includes a central processing unit (CPU) and / or a graphics processing unit (GPU) 901, which can perform various appropriate actions and processes according to computer program instructions stored in a read-only memory (ROM) 902 or computer program instructions loaded from a storage unit 908 to a random access memory (RAM) 903. In the RAM 903, various programs and data required for the operation of the device 900 can also be stored. The CPU and / or GPU 901, the ROM 902, and the RAM 903 are connected to each other via a bus 904. An input and / or output (I and / or O) interface 905 is also connected to the bus 904. Although not shown in FIG. Fig. 9 As shown in FIG. 9 , device 900 may further include a co-processor.
[0054] A number of components in the device 900 are connected to the I and / or O interface 905, including: an input unit 906, such as a keyboard, a mouse, etc.; an output unit 907, such as various types of displays, speakers, etc.; a storage unit 908, such as a disk, an optical disk, etc.; and a communication unit 909, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 909 allows the device 900 to exchange information and / or data with other devices through a computer network such as the Internet and / or various telecommunication networks.
[0055] The various methods or processes described above may be performed by the CPU and / or GPU 901. For example, in some embodiments, the methods may be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as a storage unit 908. In some embodiments, part or all of the computer program may be loaded and / or installed on the device 900 via the ROM 902 and / or the communication unit 909. When the computer program is loaded into the RAM 903 and executed by the CPU and / or GPU 901, one or more steps or actions in the methods or processes described above may be performed.
[0056] In some embodiments, the methods and processes described above may be implemented as a computer program product. The computer program product may include a computer-readable storage medium carrying computer-readable program instructions for executing various aspects of the present disclosure.
[0057] Computer readable storage medium can be a tangible device that can hold and store instructions used by an instruction execution device. Computer readable storage medium can be, for example, but not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination thereof. More specific examples (non-exhaustive list) of computer readable storage medium include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disk read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanical encoding device, for example, a punch card or a convex structure in a groove on which instructions are stored, and any suitable combination thereof. The computer readable storage medium used here is not interpreted as a transient signal itself, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagated by a waveguide or other transmission medium (for example, a light pulse by an optical fiber cable), or an electrical signal transmitted by a wire.
[0058] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to each computing and / or processing device, or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include copper transmission cables, optical fiber transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. The network adapter card or network interface in each computing and / or processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in the computer-readable storage medium in each computing and / or processing device.
[0059] The computer program instructions for performing the disclosed operation may be assembler instructions, instruction set architecture (ISA) instructions, machine instructions, machine-related instructions, microcode, firmware instructions, state setting data, or source code or object code written in any combination of one or more programming languages, programming languages including object-oriented programming languages, and conventional procedural programming languages. Computer-readable program instructions may be executed completely on a user's computer, partially on a user's computer, executed as an independent software package, partially on a user's computer, partially on a remote computer, or completely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., utilizing an Internet service provider to connect via the Internet). In certain embodiments, by utilizing the state information of a computer-readable program instruction to customize an electronic circuit, such as a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA), the electronic circuit may execute a computer-readable program instruction, thereby realizing various aspects of the present disclosure.
[0060] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine, so that when these instructions are executed by the processing unit of the computer or other programmable data processing device, a device that implements the functions and / or actions specified in one or more boxes in the flowchart and / or block diagram is generated. These computer-readable program instructions can also be stored in a computer-readable storage medium, and these instructions cause the computer, programmable data processing device, and / or other equipment to work in a specific manner, so that the computer-readable medium storing the instructions includes a manufactured product, which includes instructions for implementing various aspects of the functions and / or actions specified in one or more boxes in the flowchart and / or block diagram.
[0061] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device so that a series of operating steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to implement the functions and / or actions specified in one or more boxes in the flowchart and / or block diagram.
[0062] The flow chart and block diagram in the accompanying drawings show the possible architecture, function and operation of the equipment, method and computer program product according to multiple embodiments of the present disclosure. In this regard, each frame in the flow chart or block diagram can represent a part of a module, program segment or instruction, and a part of the module, program segment or instruction includes one or more executable instructions for realizing the specified logical function. In some alternative implementations, the function marked in the frame can also occur in a sequence different from that marked in the accompanying drawings. For example, two continuous frames can actually be executed substantially in parallel, and they can also be executed in the opposite order sometimes, depending on the functions involved. It should also be noted that each frame in the block diagram and / or flow chart, and the combination of frames in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or action, or can be implemented with a combination of dedicated hardware and computer instructions.
[0063] The embodiments of the present disclosure have been described above, and the above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and changes will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The choice of terms used herein is intended to best explain the principles of the embodiments, practical applications, or technical improvements to the technology in the market, or to enable other persons of ordinary skill in the art to understand the embodiments disclosed herein.
[0064] Some example implementations of the present disclosure are listed below.
[0065] Example 1. A method for generating audio, comprising:
[0066] receiving a user audio of a user reading a reference text, wherein the user audio is used to determine the timbre of the user;
[0067] determining a target text to be played; and
[0068] A target audio is generated in which the target text is read aloud in the voice of the user.
[0069] Example 2. The method of Example 1, wherein receiving user audio of a user reading a reference text comprises:
[0070] displaying the reference text and the first control; and
[0071] In response to a user's touch operation on the first control, audio input by the user is recorded.
[0072] Example 3. The method according to any one of Examples 1-2, further comprising:
[0073] While the user is reading the reference text, a waveform of the recorded audio is displayed.
[0074] Example 4. The method according to any one of Examples 1-3, wherein determining the target text to be played comprises:
[0075] Displaying multiple text contents on the user interface; and
[0076] In response to a user selecting a text content from the plurality of text contents, the target text to be played is determined.
[0077] Example 5. The method of any one of Examples 1-4, wherein determining the timbre of the user comprises:
[0078] Based on the user audio, training the pre-trained model to generate a timbre conversion model for the user; and
[0079] Based on the timbre conversion model, the timbre of the target audio is determined.
[0080] Example 6. The method of any one of Examples 1-5, wherein generating a timbre conversion model comprises:
[0081] Determining spectrum information of the user audio;
[0082] Inputting the spectrum information into the pre-trained model;
[0083] Extracting sound features and semantic features from the spectrum information; and
[0084] Based on the sound features and the semantic features, the parameters of the pre-trained model are adjusted to generate the timbre conversion model.
[0085] Example 7. The method of any one of Examples 1-6, wherein adjusting the parameters of the pre-trained model to generate a timbre conversion model comprises:
[0086] Based on the sound features and the semantic features, the parameters of the preset rounds are adjusted to generate a timbre conversion model.
[0087] Example 8. The method of any one of Examples 1-7, wherein adjusting the parameters of the pre-trained model to generate a timbre conversion model comprises:
[0088] Adjusting parameters of multiple rounds of pre-training models based on the sound features and the semantic features;
[0089] For each round of the speech model after adjusting the parameters, synthesizing the sound features and the semantic features to obtain a plurality of synthesized audios; and
[0090] In response to a user selecting one of the plurality of synthesized audios, the timbre conversion model is generated, wherein the timbre conversion model corresponds to the selected synthesized audio.
[0091] Example 9. A method according to any one of Examples 1-8, wherein the sound features include pitch features, waveform features, and tone features.
[0092] Example 10. A method according to any one of Examples 1-9, wherein the pre-trained model is generated by training a speech model based on a multi-person timbre dataset, and the multi-person timbre dataset includes audios of multiple different timbres.
[0093] Example 11. The method according to any one of Examples 1-10, wherein generating a target audio for reading the target text in the voice of the user comprises:
[0094] Inputting the speech corresponding to the target text into the timbre conversion model;
[0095] Extracting the sound features and semantic features of the speech corresponding to the target text; and
[0096] The extracted sound feature and the semantic feature are synthesized to generate the target audio in which the target text is read aloud in the voice of the user.
[0097] Example 12. An apparatus for generating audio, comprising:
[0098] A user audio receiving module is configured to receive user audio of a user reading a reference text, wherein the user audio is used to determine the timbre of the user;
[0099] a target text determination module, configured to determine a target text to be played; and
[0100] The target audio generation module is configured to generate a target audio for reading the target text in the voice of the user.
[0101] Example 13. The apparatus of Example 12, wherein the user audio receiving module comprises:
[0102] A first control display module, configured to display the reference text and a first control; and
[0103] The audio recording module is configured to record the audio input by the user in response to the user's touch operation on the first control.
[0104] Example 14. The apparatus of any of Examples 12-13, further comprising:
[0105] The waveform display module is configured to display the waveform of the recorded audio while the user is reading the reference text.
[0106] Example 15. The apparatus according to any one of Examples 12-14, wherein the target text determination module comprises:
[0107] A text content display module, configured to display a plurality of text contents on a user interface; and
[0108] The first target text determination module is configured to determine the target text to be played in response to a user selecting a text content from the multiple text contents.
[0109] Example 16 The apparatus according to any one of Examples 12-15, further comprising:
[0110] A timbre conversion model generation module is configured to train a pre-trained model based on the user audio to generate a timbre conversion model for the user; and
[0111] The timbre determination module is configured to determine the timbre of the target audio based on the timbre conversion model.
[0112] Example 17. The apparatus of any one of Examples 12-16, wherein the timbre conversion model generation module comprises:
[0113] A sound spectrum information determination module, configured to determine the sound spectrum information of the user audio;
[0114] A sound spectrum information input module, configured to input the sound spectrum information into the pre-trained model;
[0115] A first feature extraction module is configured to extract sound features and semantic features from the spectrum information; and
[0116] The parameter adjustment module is configured to adjust the parameters of the pre-trained model based on the sound features and the semantic features to generate the timbre conversion model.
[0117] Example 18. The apparatus of any one of Examples 12-17, wherein the parameter adjustment module comprises:
[0118] The first parameter adjustment module is configured to generate a timbre conversion model after adjusting the parameters of a preset round based on the sound features and the semantic features.
[0119] Example 19. The apparatus of any one of Examples 12-18, wherein the parameter adjustment module further comprises:
[0120] A second parameter adjustment module is configured to adjust the parameters of the multi-round pre-training model based on the sound feature and the semantic feature;
[0121] A first audio synthesis module is configured to synthesize the sound features and the semantic features to obtain a plurality of synthesized audios for the speech model after each round of parameter adjustment; and
[0122] The audio selection module is configured to generate the timbre conversion model in response to a user selecting one of the multiple synthesized audios, wherein the timbre conversion model corresponds to the selected synthesized audio.
[0123] Example 20. An apparatus according to any one of Examples 12-19, wherein the sound features include pitch features, waveform features, and tone features.
[0124] Example 21. An apparatus according to any one of Examples 12-20, wherein the pre-trained model is generated by training a speech model based on a multi-person timbre dataset, and the multi-person timbre dataset includes audios of multiple different timbres.
[0125] Example 22. The apparatus of any of Examples 12-21, wherein the target audio generation module comprises:
[0126] A speech input module, configured to input the speech corresponding to the target text into the timbre conversion model;
[0127] A second feature extraction module is configured to extract the sound features and semantic features of the speech corresponding to the target text; and
[0128] The first target audio generation module is configured to synthesize the extracted sound features and the semantic features to generate the target audio for reading the target text in the tone of the user.
[0129] Example 23. An electronic device comprising:
[0130] Processor; and
[0131] A memory coupled to the processor, the memory having instructions stored therein, the instructions causing the electronic device to perform actions when executed by the processor, the actions comprising:
[0132] receiving a user audio of a user reading a reference text, wherein the user audio is used to determine the timbre of the user;
[0133] determining a target text to be played; and
[0134] A target audio is generated in which the target text is read aloud in the voice of the user.
[0135] Example 24. The electronic device of Example 23, wherein receiving user audio of the user reading the reference text comprises:
[0136] displaying the reference text and the first control; and
[0137] In response to a user's touch operation on the first control, audio input by the user is recorded.
[0138] Example 25. The electronic device of any of Examples 23-24, further comprising:
[0139] While the user is reading the reference text, a waveform of the recorded audio is displayed.
[0140] Example 26. An electronic device according to any one of Examples 23-25, wherein determining the target text to be played comprises:
[0141] Displaying multiple text contents on the user interface; and
[0142] In response to a user selecting a text content from the plurality of text contents, the target text to be played is determined.
[0143] Example 27. The electronic device of any of Examples 23-26, wherein determining the user's timbre comprises:
[0144] Based on the user audio, training the pre-trained model to generate a timbre conversion model for the user; and
[0145] Based on the timbre conversion model, the timbre of the target audio is determined.
[0146] Example 28. An electronic device according to any of Examples 23-27, wherein generating a timbre conversion model comprises:
[0147] Determining spectrum information of the user audio;
[0148] Inputting the spectrum information into the pre-trained model;
[0149] Extracting sound features and semantic features from the spectrum information; and
[0150] Based on the sound features and the semantic features, the parameters of the pre-trained model are adjusted to generate the timbre conversion model.
[0151] Example 29. An electronic device according to any one of Examples 23-28, wherein adjusting the parameters of the pre-trained model to generate a timbre conversion model comprises:
[0152] Based on the sound features and the semantic features, the parameters of the preset rounds are adjusted to generate a timbre conversion model.
[0153] Example 30. An electronic device according to any one of Examples 23-29, wherein adjusting the parameters of the pre-trained model to generate a timbre conversion model comprises:
[0154] Adjusting parameters of multiple rounds of pre-training models based on the sound features and the semantic features;
[0155] For each round of the speech model after adjusting the parameters, synthesizing the sound features and the semantic features to obtain a plurality of synthesized audios; and
[0156] In response to a user selecting one of the plurality of synthesized audios, the timbre conversion model is generated, wherein the timbre conversion model corresponds to the selected synthesized audio.
[0157] Example 31. An electronic device according to any one of Examples 23-30, wherein the sound features include pitch features, waveform features, and tone features.
[0158] Example 32. An electronic device according to any one of Examples 23-31, wherein the pre-trained model is generated by training a speech model based on a multi-person timbre dataset, and the multi-person timbre dataset includes audios of multiple different timbres.
[0159] Example 33. The electronic device according to any one of Examples 23-32, wherein generating a target audio for reading the target text in the voice of the user comprises:
[0160] Inputting the speech corresponding to the target text into the timbre conversion model;
[0161] Extracting the sound features and semantic features of the speech corresponding to the target text; and
[0162] The extracted sound feature and the semantic feature are synthesized to generate the target audio in which the target text is read aloud in the voice of the user.
[0163] Example 34. A computer-readable storage medium having computer-executable instructions stored thereon, wherein the computer-executable instructions are executed by a processor to implement the method according to any one of claims 1-11.
[0164] Example 35. A computer program product tangibly stored on a computer-readable medium and comprising computer-executable instructions that, when executed by a device, cause the device to perform a method according to any one of Examples 1-11.
[0165] Although the disclosure has been described in language specific to structural features and / or methodological logical actions, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Instead, the specific features and actions described above are merely example forms of implementing the claims.
Claims
1. A method for generating audio, include: receiving a user audio of a user reading a reference text, wherein the user audio is used to determine the timbre of the user; Determine the target text to be played; as well as A target audio is generated in which the target text is read aloud in the voice of the user.
2. The method according to claim 1, wherein receiving user audio of the user reading the reference text include: Display the reference text and the first control; as well as In response to a user's touch operation on the first control, audio input by the user is recorded.
3. The method according to claim 2, further comprising: include: While the user is reading the reference text, a waveform of the recorded audio is displayed.
4. The method according to claim 1, wherein determining the target text to be played include: Display multiple text contents on the user interface; as well as In response to a user selecting a text content from the plurality of text contents, the target text to be played is determined.
5. The method according to claim 1, wherein determining the user's timbre include: Based on the user audio, training the pre-trained model to generate a timbre conversion model for the user; as well as Based on the timbre conversion model, the timbre of the target audio is determined.
6. The method according to claim 5, wherein generating a timbre conversion model include: Determining spectrum information of the user audio; Inputting the spectrum information into the pre-trained model; Extracting sound features and semantic features from the spectrum information; as well as Based on the sound features and the semantic features, the parameters of the pre-trained model are adjusted to generate the timbre conversion model.
7. The method according to claim 6, wherein the parameters of the pre-trained model are adjusted to generate a timbre conversion model include: Based on the sound features and the semantic features, the parameters of the preset rounds are adjusted to generate a timbre conversion model.
8. The method according to claim 6, wherein the parameters of the pre-trained model are adjusted to generate a timbre conversion model include: Adjusting parameters of multiple rounds of pre-training models based on the sound features and the semantic features; For each round of the speech model after adjusting the parameters, synthesizing the sound features and the semantic features to obtain a plurality of synthesized audios; and In response to a user selecting one of the plurality of synthesized audios, the timbre conversion model is generated, wherein the timbre conversion model corresponds to the selected synthesized audio.
9. The method according to claim 6, wherein the sound features include pitch features, waveform features, and tone features.
10. The method according to claim 6, wherein the pre-trained model is generated by training a speech model based on a multi-person timbre dataset, and the multi-person timbre dataset includes audios of multiple different timbres.
11. The method according to claim 5, wherein a target audio is generated to read the target text in the voice of the user. include: Inputting the speech corresponding to the target text into the timbre conversion model; Extracting the sound features and semantic features of the speech corresponding to the target text; as well as The extracted sound feature and the semantic feature are synthesized to generate the target audio in which the target text is read aloud in the voice of the user.
12. A device for generating audio, include: A user audio receiving module is configured to receive user audio of a user reading a reference text, wherein the user audio is used to determine the timbre of the user; A target text determination module is configured to determine a target text to be played; as well as The target audio generation module is configured to generate a target audio for reading the target text in the voice of the user.
13. An electronic device, include: processor; as well as A memory coupled to the processor, the memory having instructions stored therein, wherein when the instructions are executed by the processor, the electronic device executes the method according to any one of claims 1-11.
14. A computer-readable storage medium having computer-executable instructions stored thereon, wherein the computer-executable instructions are executed by a processor to implement the method according to any one of claims 1-11.