Method and apparatus for creating and using patch, device, and storage medium
By displaying reference text and recording controls on the tone creation page, receiving user audio and generating audition audio, the problem of high and time-consuming tone creation is solved, low-cost and fast personalized tone generation is achieved, and the user experience is improved.
Patent Information
- Application Number
- PCT/CN2024/135297
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-02-29
- Filing Date
- 2024-11-28
- Publication Date
- 2025-09-04
AI Technical Summary
In the prior art, the creation of tone is expensive and time-consuming, and the user experience is poor, making it difficult to achieve that everyone can generate their own voice based on tone.
By presenting a tone creation page in response to a tone creation request, displaying reference text and recording controls, receiving audio input from users, creating target tone based on audio and reference text, and generating audition audio in the tone confirmation page, and storing target tone in response to user confirmation.
Reduces the cost and time of tone creation, improves the user experience, ensures the naturalness and personalization of tone generation, and avoids the problems of tone imitation or forgery.
Smart Images

Figure CN2024135297_04092025_PF_FP_ABST
Abstract
Description
Method, apparatus, device and storage medium for creating and using timbre
[0001] This application claims priority to the Chinese invention patent application entitled “Methods, devices, apparatus and storage media for creating and using timbre” filed on February 29, 2024, with application number 202410233361.4, the entire contents of which are incorporated herein by reference. Technical Field
[0002] Example embodiments of the present disclosure relate generally to the field of computers, and more particularly to methods, apparatuses, devices, and computer-readable storage media for creating and using timbres. Background Art
[0003] Currently, an increasing number of applications are designed to provide users with various services. Many applications support messaging. As people interact with each other online, various types of audio have become important media for social expression and information exchange. Therefore, there is a desire to generate audio with high-quality sound quality. Summary of the Invention
[0004] In a first aspect of the present disclosure, a method for creating and using a timbre is provided. The method includes: in response to a timbre creation request, presenting a timbre creation page, the timbre creation page displaying reference text and a recording control; in response to detecting a triggering of the recording control, receiving audio input from a user; creating a target timbre for the user based on the received audio and the reference text; presenting a timbre confirmation page, the timbre confirmation page including at least audition audio generated based on the target timbre; and in response to receiving confirmation of the target timbre on the timbre confirmation page, storing the target timbre for use in audio.
[0005] In a second aspect of the present disclosure, a device for creating and using timbre is provided. The device includes: a creation page presentation module configured to present a timbre creation page in response to a timbre creation request, the timbre creation page displaying reference text and a recording control; an audio receiving module configured to receive audio input by a user in response to detecting a triggering of the recording control; a timbre creation module configured to create a target timbre for the user based on the received audio and reference text; a confirmation page presentation module configured to present a timbre confirmation page, the timbre confirmation page including at least audition audio generated based on the target timbre; and a timbre storage module configured to store the target timbre for use in audio in response to receiving confirmation of the target timbre on the timbre confirmation page.
[0006] In a third aspect of the present disclosure, an electronic device is provided. The device includes at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit. When executed by the at least one processing unit, the instructions cause the device to perform the method of the first aspect.
[0007] In a fourth aspect of the present disclosure, a computer-readable storage medium is provided, wherein a computer program is stored on the computer-readable storage medium, and the computer program can be executed by a processor to implement the method of the first aspect.
[0008] In a fifth aspect of the present disclosure, a computer program product is provided, which is tangibly stored in a computer storage medium and includes computer-executable instructions, which, when executed by a device, cause the device to perform the method of the first aspect.
[0009] It should be understood that the content described in this summary section is not intended to limit the key features or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] The above and other features, advantages and aspects of the embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. In the accompanying drawings, the same or similar reference numerals represent the same or similar elements, wherein:
[0011] FIG1 shows a schematic diagram of an example environment in which embodiments of the present disclosure can be implemented;
[0012] FIG2 shows a schematic diagram of an example architecture for creating timbres according to some embodiments of the present disclosure;
[0013] 3A to 3L are schematic diagrams illustrating example interfaces for creating timbres according to some embodiments of the present disclosure;
[0014] 4A to 4C are schematic diagrams illustrating example interfaces for using a target timbre according to some embodiments of the present disclosure;
[0015] 5A to 5J are schematic diagrams showing example interfaces for a user to discover a sound creation portal during content browsing according to some embodiments of the present disclosure;
[0016] FIG6 illustrates a flow chart of a process for timbre creation and use according to some embodiments of the present disclosure;
[0017] FIG7 illustrates a flow chart of a process for timbre creation and use according to some embodiments of the present disclosure;
[0018] FIG8 shows a block diagram of an apparatus for creating and using timbre according to some embodiments of the present disclosure; and
[0019] FIG9 illustrates a block diagram of an electronic device capable of implementing one or more embodiments of the present disclosure. DETAILED DESCRIPTION
[0020] The following describes embodiments of the present disclosure in more detail with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments described herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.
[0021] In the description of the embodiments of the present disclosure, the term "including" and similar terms should be understood as open inclusion, i.e., "including but not limited to". The term "based on" should be understood as "based at least in part on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". Other explicit and implicit definitions may be included below.
[0022] Herein, unless explicitly stated otherwise, executing a step “in response to A” does not mean executing the step immediately after “A” but may include one or more intermediate steps.
[0023] It is understandable that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) must comply with the requirements of relevant laws, regulations and relevant provisions.
[0024] It is understandable that before using the technical solutions disclosed in the various embodiments of this disclosure, the type, scope of use, usage scenarios, etc. of the personal information involved in this disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.
[0025] For example, in response to receiving a user's active request, a prompt message is sent to the user to clearly remind the user that the operation requested to be performed will require obtaining and using the user's personal information, so that the user can independently choose whether to provide personal information to the electronic device, application, server or storage medium and other software or hardware that performs the operation of the technical solution of the present disclosure based on the prompt message.
[0026] As an optional but non-limiting implementation, in response to receiving a user's active request, a prompt message may be sent to the user, for example, in the form of a pop-up window, in which the prompt message may be presented in text form. Furthermore, the pop-up window may also include a selection control for the user to select "agree" or "disagree" to provide personal information to the electronic device.
[0027] It is understandable that the above notification and the process of obtaining user authorization are merely illustrative and do not constitute a limitation on the implementation of the present disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of the present disclosure.
[0028] As used herein, the term "model" can learn the association between corresponding inputs and outputs from training data, so that after training is completed, corresponding outputs can be generated for given inputs. The generation of the model can be based on machine learning technology. Deep learning is a machine learning algorithm that processes inputs and provides corresponding outputs by using multiple layers of processing units. A neural network model is an example of a model based on deep learning. In this article, "model" may also be referred to as "machine learning model", "learning model", "machine learning network" or "learning network", and these terms are used interchangeably in this article.
[0029] As briefly mentioned above, as people interact with each other online, various types of audio have become an important medium for social expression and information exchange. Currently, creating timbre is expensive. For example, if a user wants to replicate their own voice, they often need to record a whole day's worth of sound data. Extracting the timbre and reproducing the voice takes a long time, and the results are often mediocre, with unnatural pronunciation.
[0030] This method is time-consuming, provides a poor user experience, and due to the high cost, it is difficult to achieve a situation where everyone can generate their own voice based on timbre.
[0031] In view of this, an embodiment of the present disclosure provides an improved method for creating and using timbres. In this solution, in response to a timbres creation request, a timbres creation page is presented, which displays reference text and recording controls. If a triggering of the recording controls is detected, audio input by the user is received. Then, based on the received audio and reference text, a target timbres for the user are created. Accordingly, a timbres confirmation page is presented, which includes at least audition audio generated based on the target timbres. Furthermore, in response to receiving confirmation of the target timbres on the timbres confirmation page, the target timbres for the user are stored for use in audio. In this way, users can quickly and easily create their own timbres for generating personalized audio, enhancing the user experience. During the timbres creation process, users can better understand the timbres creation method. In addition, by requiring users to extract timbres by reading reference text, rather than having them upload free audio, the problem of timbres being misused due to misused audio can be avoided to a large extent.
[0032] The term "work" in this disclosure refers to any type of media content or media work, which includes one or more types of content, including but not limited to audio files, video files, image files, text files, etc. Specifically, a work may include a short video, music, images, image compilations, multimedia clips, audiovisual materials, etc. This disclosure is not limited in this respect.
[0033] FIG1 shows a schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented. Environment 100 includes one or more users 110-1, 110-2, 110-3, ..., 110-N, which can send and receive messages through their respective associated terminal devices 120-1, 120-2, 120-3, ..., 120-N. For ease of discussion, users 110-1, 110-2, 110-3, ..., 110-N may be collectively or individually referred to as users 110, and terminal devices 120-1, 120-2, 120-3, ..., 120-N may be collectively or individually referred to as terminal devices 120. In some scenarios, user 110 may publish and comment on works in a target platform through associated terminal devices 120. In some scenarios, user 110 is also referred to as the publisher of the work.
[0034] The terminal device 120 may be installed with an application 125 that supports message interaction (i.e., application 125-1 is installed in terminal device 120-1, application 125-2 is installed in terminal device 120-2, application 125-3 is installed in terminal device 120-3, ..., application 125-N is installed in terminal device 120-N). It should be noted that the applications 125 installed in different terminal devices 120 may be exactly the same applications or different applications (e.g., different versions). The application 125 may be any appropriate application with message sending and receiving functions, such as a dedicated chat application, a social application, a content sharing application, an office support application, and the like.
[0035] In the environment 100 of FIG. 1 , if application 125 is active, terminal device 120 may present the user interface of application 125. This user interface may include various interfaces provided by application 125, such as a user interface supporting message interaction, a user interface supporting content browsing, a message sending and receiving interface, and so on. Application 125 may provide different content to user 110 via different user interfaces. Application 125 may also provide user 110 with the ability to select and switch presentation methods for related content via appropriate means, such as by clicking or selecting any appropriate element in the user interface.
[0036] In some embodiments, different terminal devices 120 may also communicate with server 130 via network 132 to provide message interaction services. Server 130 may provide management, configuration, and maintenance functions for application 125.
[0037] The terminal device 120 can be any type of mobile terminal, fixed terminal or portable terminal, including a mobile phone, a desktop computer, a laptop computer, a notebook computer, a netbook computer, a tablet computer, a media computer, a multimedia tablet, a personal communication system (PCS) device, a personal navigation device, a personal digital assistant (PDA), an audio / video player, a digital camera / camcorder, a television receiver, a radio broadcast receiver, an e-book device, a gaming device or any combination thereof, including accessories and peripherals of these devices or any combination thereof. In some embodiments, the terminal device 120 can also support any type of interface for the user (such as a "wearable" circuit, etc.). The server 130 can be various types of computing systems / servers that can provide computing power, including but not limited to mainframes, edge computing nodes, computing devices in cloud environments, and the like.
[0038] It should be understood that the structure and function of the various elements in the environment 100 are described for illustrative purposes only and do not imply any limitation on the scope of the present disclosure.
[0039] Various example implementations of the present disclosure are described in detail below.
[0040] FIG2 illustrates a schematic diagram of an example architecture 200 for creating timbre according to some embodiments of the present disclosure. As shown in FIG2 , architecture 200 involves a terminal device 120, a server 130, and a timbre extraction model 210. The timbre extraction model 210 is configured to extract timbre feature information from input audio for reuse in new audio. The timbre extraction model 210 can be implemented based on any suitable model, and the embodiments of the present disclosure are not limited in this respect.
[0041] The timbre extraction model 210 can run locally on the terminal device 120 or be deployed remotely, for example, on a server (such as server 130), in a cloud environment, etc. In the case of local operation, the terminal device 120 can directly provide the model input to the locally installed timbre extraction model 210 and obtain the model output generated by the timbre extraction model 210. In the case of remote operation, the terminal device 120 provides data to other electronic devices through a communication connection with other electronic devices. The other electronic devices determine the model input based on the acquired data and provide the model input to the timbre extraction model 210. After the other electronic devices obtain the model output of the timbre extraction model 210, they provide the model output to the terminal device 120 through the communication connection between them so that the terminal device 120 obtains the model output.
[0042] For ease of description, the following description uses the example of the timbre extraction model 210 running on the server 130. That is, the terminal device 120 can provide model input to the server 130, so that the server 130 can call the timbre extraction model 210 to perform timbre extraction. It should be understood that, for ease of discussion, the various operations involved in timbre creation and use are described from the perspective of the terminal device. However, it should be understood that some or all of these operations can also be executed on the server side, with the execution results transmitted to the terminal device, or the server can control the terminal device to execute them.
[0043] When the application 125 is active, the terminal device 120 can present a user interface (e.g., a timbre creation page) provided by the application 125. In an embodiment of the present disclosure, the terminal device 120 presents a timbre creation page in response to a timbre creation request. The timbre creation page displays reference text and recording controls. In the application 125, an appropriate interface and timbre creation entry can be configured to allow the user 110 to initiate a timbre creation request. Some examples of user 110 initiating a timbre creation request are provided below.
[0044] When the timbre creation page is presented to the user, if user 110 clicks the recording control displayed on the terminal device 120 on the timbre creation page, user 110 begins recording. It should be understood that before or after the user clicks the recording control, the user is prompted with the purpose of obtaining user 110's recording, and recording will only begin after user 110 agrees. The timbre creation page can also provide a prompt to prompt the user to record audio by reading aloud the reference text displayed on the page. Furthermore, terminal device 120 receives audio 202 input by the user. Based on the audio 202 and the reference text, terminal device 120 creates a target timbre 212 for user 110.
[0045] In some embodiments, a timbre extraction model 210 can be used to generate a target timbre 212 based on the audio 202 and the reference text. In some embodiments, the terminal device 120 can send the audio 202 and the reference text to the server 130, which can generate the target timbre 212. The terminal device 120 can send the audio 202 and the reference text to the server 130, which can generate the target timbre 212 using the timbre extraction model 210.
[0046] The following continues to describe an example embodiment of target timbre creation with reference to the example interfaces / example pages shown in FIG. 3A to FIG. 3L .
[0047] For the sole purpose of better understanding the various embodiments of the present disclosure, the example user interfaces 301 to 309 shown in Figures 3A to 3L are referenced in the example embodiments below. Figures 3A to 3I show schematic diagrams of example interfaces for creating timbres according to some embodiments of the present disclosure. It should be understood that the user interfaces shown in Figures 3A to 3L and other figures below are merely examples, and various designs may exist in practice. For example, the various graphical elements and / or controls in the user interface may have different arrangements and different visual representations, one or more elements and / or controls therein may be omitted or replaced, and one or more other elements and / or controls may also exist. In addition, any appropriate content may be included in the user interface. The scope of the present disclosure is not limited in this respect.
[0048] In an example embodiment of the present disclosure, example user interfaces 301 to 309 may be implemented at terminal device 120. For ease of discussion, example user interfaces 301 to 309 will be described with reference to environment 100 of FIG. 1 and / or FIG.
[0049] In some embodiments, the terminal device 120 presents a timbre creation page in response to the timbre creation request. The timbre creation page displays reference text and recording controls. In some examples, the reference text displayed on the timbre creation page can be used to record audio text.
[0050] The example shown in Figures 3A to 3L involves a user initiating a voice creation request on the editing page of media content. Assume that the user is editing media content and hopes to publish the content. As shown in example interfaces 301 to 304 of Figures 3A to 3D, after user 110 clicks the "Text" control 311 presented by terminal device 120 on interface 301, terminal device 120 presents a text input box 321 on interface 302. User 110 can enter text (e.g., "Document A") in text input box 321 presented by terminal device 120. Subsequently, terminal device 120 presents user-entered text 330 (e.g., "Document A") on interface 303. During the media content editing process, assuming that the user wishes to convert the entered text into speech, the user can trigger the "Text-to-Speech" function 322. In interface 303, a voice selection panel 334 may be provided. If user 110 selects "Voice 1" 331, terminal device 120 will read "Document A" in "Voice 1" 331. It is understandable that the timbre presented to the user 110 is a timbre already in the sound library or a timbre that is authorized for use.
[0051] In some embodiments, the terminal device 120 may present an entry 332 for creating a timbre on the interface 303 (for example, the interface may be a corresponding interface of the timbre selection panel 334). If the user 110 clicks on the entry 332 for creating a timbre, the terminal device 120 will detect the timbre creation request and present a function introduction interface 340 about creating a timbre on the interface 304. The function introduction interface 340 may provide some prompt information about timbre creation. For example, on the interface 140, the terminal device 120 will display the title: "Create your timbre", and related descriptions: "Record a few words out loud to create a timbre that can be used with reference text in XX", "You can set your timbre to private and can delete it at any time", "Recording only takes 10 seconds. Make sure you are in a quiet environment and avoid using wireless headphones". Of course, this information is just an example.
[0052] After the function introduction screen, the terminal device 120 may display a timbre creation page, such as interface 305 shown in FIG3E . The timbre creation page may provide recording controls and reference text. In some embodiments, the terminal device 120 receives user input audio in response to detecting that the user triggers the recording control. Based on the audio and reference text, the terminal device 120 creates a target timbre for the user.
[0053] As shown in the example interfaces 305 to 313 of Figures 3E to 3M, the terminal device 120 presents a reference text 351 (for example, copy B) and a recording control 350 on the interface 305. The terminal device 120 starts recording based on the user 110 clicking the recording control 350 to receive the audio of the user 110. In some examples, the reference text can be pre-configured. In some embodiments, a reference text library can be set, and when each user requests to create a timbre, a reference text is randomly selected from the reference text library to be displayed to the user. In some embodiments, the length of the reference text can be set to make the obtained audio sufficient for the timbre extraction model 210 to correctly extract the timbre. Depending on the capabilities of the timbre extraction model 210, when the user is asked to record audio, only one sentence or several sentences need to be recorded, and the recording is performed with a countdown of a certain number of seconds. In this way, users can create timbre easily and quickly.
[0054] To guide the user in using the voice creation function, a prompt may be provided on the voice creation page to indicate that the user needs to read the reference text displayed on the page to record the voice, and to guide the user in using the recording controls, as shown in Figure 3E. After user 110 clicks the recording control 350, user 110 begins reading copy B. At this point, terminal device 120 begins receiving the corresponding audio of user 110 reading copy B. If user 110 clicks the "Stop" control 360 in Figure 3F, terminal device 120 will stop receiving the audio input from user 110. In some embodiments, terminal device 120 creates user 110's target voice based on the received audio and reference text. In some embodiments, given that voice creation takes a certain amount of time, a progress indicator for voice generation may be provided. As shown in Figure 3G, terminal device 120 displays a corresponding "Generating" control 370 on interface 307 to indicate the progress of voice generation.
[0055] In some embodiments, after extracting the target timbre based on the audio input by the user, the terminal device 120 also presents a timbre confirmation page. This timbre confirmation page includes at least an audition audio generated based on the target timbre, so that the user can judge whether the generated timbre is satisfactory and whether it needs to be adjusted or regenerated by listening to the audition audio. As shown in Figure 3H, the terminal device 120 displays an audition control 381 for the audition audio on the timbre confirmation page 308 and a corresponding indication 380 of "Your voice is ready". Based on the user 110 (e.g., user A) clicking the audition control 381, the terminal device 120 plays the audition audio, which now has the extracted timbre. In some embodiments, if timbre creation is performed during the media content creation process and user input text is present, the audition audio can be generated based on the user input text. For example, in the example shown in Figure 3H, the audition audio is the audio generated after the user input text 330 (i.e., "Document A") is subjected to a text-to-speech operation, and the audio is imbued with the newly extracted timbre.
[0056] In some embodiments, if there is no user input text associated with the timbre creation request, the terminal device 120 presents a timbre confirmation page 309 including at least a trial audio generated based on the target timbre. The trial audio includes audio with the target timbre generated by converting the predetermined text into a text-to-speech operation.
[0057] As shown in the timbre confirmation page 309 of FIG3I , after the terminal device 120 has loaded the target timbre, the timbre of the user 110 (e.g., user A) is automatically applied to the media content currently being created. Accordingly, the terminal device 120 displays an audition control 381 for auditioning the audio on the timbre confirmation page 309 and a corresponding indication of "Your voice is ready." Based on the user 110 (e.g., user A) clicking the audition control 381, the terminal device 120 plays the corresponding audio of the user 110 (e.g., user A) reading the "predetermined text" 390.
[0058] The above describes that the terminal device 120 generates audio with a target timbre in a scenario where the user does not input a reference text. The following will continue to describe how to re-record audio and generate audio with a target timbre with reference to FIG.
[0059] In some embodiments, after the terminal device 120 loads the target timbre, the timbre of user 110 (e.g., user A) is automatically applied to the media content currently being created. As shown in FIG3H , with the consent of user 110, the generated target timbre is identified as "User A's timbre" 388 and applied to the media content currently being edited by user 110.
[0060] The terminal device 120 is in response to receiving the confirmation of the target timbre in the timbre confirmation page, and stores the target timbre for the user for use in audio. As shown in Figure 3H, if the user 110 confirms the timbre currently generated, the generated timbre can be confirmed by triggering the "use" control 382. In some examples, if the user 110 confirms the target timbre in the timbre confirmation page, the terminal device 120 stores the target timbre for use in audio. For example, after the user 110 confirms the target timbre, the terminal device 120 can directly apply the timbre to the media content currently created. The terminal device 120 can also store the timbre so that the user can directly use it next time. In other examples, if the user 110 (for example, user A) actively sets the timbre to be public, that is, the timbre of the user 110 authorizes himself to be used, then user B can directly use it.
[0061] In this way, the cost of creating timbres for users can be reduced, and each user can be supported to freely create and use their own timbres as needed.
[0062] In some embodiments, the terminal device 120 displays another reference text and recording controls in the timbre creation page in response to detecting a regeneration request in the timbre creation page. For example, if the user 110 is not satisfied with the currently generated audio with the target timbre, the user can click the "Restart" control 383 to re-record the audio.
[0063] In other embodiments, the terminal device 120 prompts the user to re-record the audio in response to detecting that the received audio does not match the reference text. Subsequently, the terminal device 120 displays another reference text and recording controls in the timbre creation page.
[0064] In some examples, if the terminal device 120 detects that the audio input by the user 110 does not match the reference text, the user is prompted to re-record the audio. After entering the audio recording page, the terminal device 120 automatically switches to recording text.
[0065] The present disclosure can avoid voice imitation or forgery by requiring the user to create the voice by reading a specific text. The following describes re-recording audio when an error occurs during the process of generating the target voice with reference to the example interface 3010 shown in FIG3J .
[0066] In some embodiments, the terminal device 120 prompts the user to re-record the audio in response to detecting that the background noise in the audio exceeds a noise threshold or the duration of the audio is less than a preset threshold. The terminal device 120 then displays another reference text and recording controls on the timbre creation page. In some examples, if the target timbre is not successfully generated, or if the close button is clicked during recording, the terminal device 120 will exit the recording page without saving the recorded data and timbre, and a prompt message will be popped up.
[0067] As shown in example interface 3010 in FIG3J , when terminal device 120 detects an error while generating the target timbre, it will present a prompt message 3010-1 stating "Error! Please try again" on interface 3010. When the user re-enters the recording page, terminal device 120 will present another reference text 3010-2 (e.g., Text C) and recording controls 3010-3 on interface 3010.
[0068] For example, when the terminal device 120 is generating the target timbre, if it detects that the background noise exceeds a default value, a prompt message “Noise environment detected, please try again” will be displayed on the interface 3010 .
[0069] For another example, when the terminal device 120 is generating a tone, if it detects that the recording duration is less than 3 seconds, a prompt message “Please record for at least 3 seconds to ensure the best results” will be presented on the interface 3010 .
[0070] In some examples, when the terminal device 120 is generating the target timbre, an error occurs in the detection of the automatic speech recognition technology (ASR), and a prompt message "There is a pronunciation problem in this recording. Please try again" is displayed on the interface 3010.
[0071] In other examples, if the terminal device 120 detects a network error while generating the target timbre, a prompt message "Please ensure network connection. Please try again" will be displayed on the interface 3010. If the terminal device 120 detects a service connection failure while generating the target timbre, a prompt message "This service is currently unavailable. Please try again in a few minutes" will be displayed on the interface 3010.
[0072] The above describes the case where the terminal device 120 re-records the audio when an error occurs during the process of generating the target timbre. The following describes the case where the terminal device 120 applies the target timbre to the edited work after the target timbre is created, with reference to Figures 3I and 3K to 3L.
[0073] In some embodiments, the terminal device 120 detects a timbre creation request from a user on a media content editing page. In some embodiments, the terminal device 120 jumps to the editing page in response to receiving confirmation of a target timbre on a timbre confirmation page. The target timbre is selected for application to the audio of the media item being edited.
[0074] In some examples, the terminal device 120 creates a target timbre after detecting a timbre creation request from the user on the media content editing page. Then, after the target timbre is created, the user directly clicks the "Use" control to apply the created target timbre directly to the edited work.
[0075] As shown in the example interfaces 309 and 3011 of Figures 3I, 3H, and 3K, after user A's target timbre is created, user A directly clicks the "Use" control 382 to directly apply the created user A's target timbre to the edited work. At this time, the terminal device 120 presents user A's target timbre in the upper area 3011-1 of the timbre 3011-2 (e.g., timbre 1, timbre 2, ..., timbre N) pre-provided by the terminal device 120 in the interface 3011. It should be understood that the pre-provided timbres presented on the terminal device 120 are all timbres already in the sound library or timbres authorized for use.
[0076] In some embodiments, the terminal device 120 deletes the stored target timbre in response to receiving a request to delete the stored target timbre. In response to receiving a request to re-record the stored target timbre, the terminal device 120 presents a timbre creation page. The timbre creation page displays another reference text and recording controls.
[0077] As shown in the example interface 3012 of FIG3L , the user 110 clicks the “More” control 3012-1 in the area 3011-1 to manage the generated target timbres (e.g., the target timbres of user A). Based on the user clicking the control 3012-1, the terminal device 120 presents a function card 3012-2 on the interface 1032.
[0078] If the user selects to delete the target timbre of user A on function card 3012-2, the terminal device 120 deletes the stored target timbre of user A. If the user selects to re-record the target timbre of user A on function card 3012-2, the terminal device 120 presents a timbre creation page.
[0079] User A can choose to set the permissions for the target sound of user A on function card 3012-2. For example, if the permissions for the target sound of user A are set to private, the target sound of user A can only be used by user A. If the permissions for the target sound of user A are set to public, the target sound of user A can be used by other users.
[0080] The above describes the process of creating a timbre. The following describes the user using a target timbre with reference to the interface 401 shown in FIG4A. FIG4A to FIG4C are schematic diagrams of example interfaces for using a target timbre according to some embodiments of the present disclosure.
[0081] In some embodiments, terminal device 120 receives a publication request for target media. The target media includes at least target audio with a target timbre. For example, terminal device 120 creates a timbre based on reference text and audio, and stores the target timbre of user 110 (e.g., user A) for use in the target audio. User 110 (e.g., user A) publishes the target media with the target audio with the target timbre.
[0082] In some embodiments, the terminal device 120 publishes the target media content in response to the publish request, along with tag information related to the target timbre. In some examples, the target media may be audio, text plus image plus audio, or video. The tag information may be a timbre anchor or a sticker bubble.
[0083] As shown in example interface 401 in FIG4A , if user 110 (e.g., user A) publishes target media, terminal device 120 presents the target media content in interface 401. If user 110 (e.g., user A) publishes the target media, terminal device 120 also presents tag information in interface 401. For example, a control 411 (e.g., user A's timbre) corresponds to a timbre anchor point, and a control 410 corresponds to a sticker bubble.
[0084] In some embodiments, the tag information includes at least an entry to a details page corresponding to the target timbre and / or an entry to a timbre usage page corresponding to the target timbre. In some examples, if the user sets the permissions for the target timbre to private, the tag information includes an entry to the timbre usage page corresponding to the target timbre. In this case, the terminal device 120 presents the recording page in response to the user's click.
[0085] In other examples, if the user sets the permission of the target timbre to public, the timbre anchor corresponding control 411 can also be the entrance to the detail page corresponding to the target timbre. In this case, the terminal device 120 presents the recording page in response to the user clicking the control 411.
[0086] In some examples, if the user 110 (eg, user A) sets the permission of his / her target timbre to private, the terminal device 120 will not display the timbre anchor corresponding control 411 .
[0087] In some embodiments, the tag information also includes intelligently generated indication information for the target timbre. As shown in FIG4A , terminal device 120 displays intelligently generated indication information 412 in interface 401. This helps alert other users, after the media content is published, that the timbre in the media content is not a real-life recording, but rather a reproduction of the sound based on the intelligently generated timbre.
[0088] The following description will refer to example interfaces 501 to 5010 shown in Figures 5A to 5J , describing how user 110 (e.g., user B) discovers a sound creation portal while browsing content. Figures 5A to 5J illustrate example interfaces for users to discover a sound creation portal while browsing content, according to some embodiments of the present disclosure. For ease of understanding, the following description will refer to example interfaces 401 to 403 shown in Figures 4A to 4C .
[0089] In some embodiments, if the user 110 requests to create a target tone in a non-editing page, the terminal device 120 jumps to the editing page of the media content in response to receiving confirmation of the target tone in the tone confirmation page, where the target tone is selected for creating the media content.
[0090] For example, if you select the Create Target Tone function from the tone details page to enter the tone creation process, then after confirming the target tone, click "Try It" to enter the editing page, where the user can choose to publish the work.
[0091] In some embodiments, the non-editing page includes a sound details page or a sound details page for another sound. In some examples, if user 110 (e.g., user A) publishes a target media including their target sound, user B clicks on the sound anchor corresponding control 410 on page 401, and the terminal device 120 presents the sound details page 403. In some examples, if user A publishes a target media including their target sound, and the permission of the target sound is set to private, the terminal device 120 will not display the sound anchor corresponding control 410 on page 401.
[0092] In some examples, if user 110 (e.g., user A) posts a target media including their target sound, user B clicks control 410 corresponding to a sticker bubble on page 401, and terminal device 120 presents sound details page 402. For example, for the target media posted by user A, user B can choose to buy the same sticker, i.e., user B enters Markov decision process (MDP) display control 410.
[0093] In some examples, user 110 (e.g., user B) can click on the "Create Tone" control 420 presented by the terminal device 120 in interface 402, and / or click on the "Create Tone" control 430 presented by the terminal device 120 in interface 403 to start creating the target tone of user B.
[0094] As shown in example interfaces 501 to 5010 in Figures 5A to 5J , terminal device 120 presents interface 501 based on user B clicking control 430 and / or control 420. User B clicks "Continue" control 510 on interface 501 to send a request to terminal device 120 to create a target timbre. Terminal device 120 receives the request from user B and presents interface 502.
[0095] Terminal device 120 receives a message from user B clicking on the "Start Recording" control 520 to start recording audio. Terminal device 120 then stops recording based on user B clicking on the "Stop" control 530. Terminal device 120 generates user B's voice based on the audio and text (e.g., "Document D"). During the generation process, terminal device 120 displays a corresponding prompt control 540 indicating "Processing" on interface 504.
[0096] After creating user B's target timbre, terminal device 120 presents a timbre confirmation page 505, which includes a trial audio file generated based on user B's target timbre. Upon user B clicking a trial audio control 550, terminal device 120 plays an audio file reciting a predetermined text. Upon user B clicking a "Try It Out" control 551, terminal device 120 presents an editing page 506. User B clicks a "Next" control 560 on the editing interface to publish the work.
[0097] In some examples, on the edit page 506, user B can tap area 561 to view or edit text. On the edit page 507, the terminal device 120 receives user B's click on the "Edit" control 570 and presents a page 508 for user B to edit text 572 in the input box 580.
[0098] On the editing page 507 , the terminal device 120 receives user B clicking on the “text to speech” control 571 , entering the text-to-speech panel to convert the text 572 into audio.
[0099] After user B's tone is created, user B directly clicks the "Use" control to apply the created tone to the edited work. At this time, terminal device 120 displays user B's target tone in upper area 590 on interface 509. User B can click the "Preview with Volume Up" control 5010-1 on interface 5010 displayed by terminal device 120 to preview the media content.
[0100] For ease of understanding, the process of creating and using a timbre will be described below with reference to FIG6 . FIG6 shows a flowchart of a process 600 for creating and using a timbre according to some embodiments of the present disclosure. For ease of discussion, the following description will be made with reference to FIG1 .
[0101] In block 610, user 110 enters text in an input box provided by terminal device 120 and clicks a text-to-speech control on terminal device 120 to access the text-to-speech panel. In block 611, terminal device 120 adds a new sound to the text-to-speech panel. In block 612, terminal device 120 detects a sound creation request sent by user 110 (e.g., by detecting a user triggering a sound creation entry in the panel) and presents the sound creation page.
[0102] In block 613, the terminal device 120 causes the user 110 to enter a recording process based on the user 110 clicking a recording control on the tone creation page. After creating a tone based on the user's recorded audio, in block 614, the terminal device 120 determines whether the user 110 has entered text.
[0103] In box 616, if user 110 does not enter text, terminal device 120 uses the preset text to audition the reading effect. In box 617, terminal device 120 determines whether to submit the content. For example, terminal device 120 determines whether user 110 chooses to publish the media content corresponding to the audio with the specified timbre. In box 618, if user 110 does not choose to submit the content, terminal device 120 displays a prompt message on the interface to "Save the timbre." In box 619, if user 110 chooses to submit the content, terminal device 120 submits the content by presenting the text plus the text-to-speech reading.
[0104] At block 615, if user 110 has entered text, terminal device 120 then reads the text (e.g., the text entered by user 110) using the recorded user's voice for audition. At block 620, terminal device 120 contributes the media content, which is reproduced with the voice. At block 627, terminal device 120 contributes the media content corresponding to the text-to-speech conversion.
[0105] In block 621, the terminal device 120 determines whether the permission of the user 110's timbre is public. In block 622, if the user 110 sets the permission of his timbre to be private, the terminal device 120 does not display the control corresponding to the timbre anchor on the page and directly enters the recording process.
[0106] In box 623, if user 110 sets the permissions for their timbre to public, terminal device 120 displays the controls corresponding to the timbre anchor and the controls corresponding to the sticker bubble on the page. In box 624, user 110 clicks the control corresponding to the timbre anchor displayed by terminal device 120 on the interface to enter the timbre details page. In box 625, terminal device 120 supports reusing user 110's timbres. In box 626, terminal device 120 displays function entrances on the timbre details page it presents, for example, an entrance for recording audio.
[0107] At block 628, after the media content corresponding to the audio with timbre is published, the terminal device 120 displays a page entry for "Shoot the Same Style" on the interface. After the media content corresponding to the standard text-to-speech is published, the terminal device 120 displays a page entry for "Shoot the Same Style" on the interface. At block 629, the terminal device 120 directly enters the recording process based on the user 110 clicking on the page entry for "Shoot the Same Style" and / or the user 110 clicking on the displayed function entry.
[0108] Through this disclosure, the discovery of functions is improved, and the video of the consumer side (user) helps users better understand the functions. In the audio recording link, by limiting uploads (for example, only allowing current recording) and strictly limiting the recording time, it can help reduce the risk of impersonating other people's voices, thereby protecting the user's own voice copyright. The corresponding controls of the voice anchor presented by the terminal device 120 can allow the sound to be reused by others. Furthermore, further interactive gameplay of the sound is supported to create a sound community.
[0109] 7 shows a flow chart of a process 700 for creating and using a timbre according to some embodiments of the present disclosure. The process 700 may be implemented at the terminal device 120. The process 700 is described below with reference to FIG1.
[0110] In block 710 , the terminal device 120 presents a timbre creation page in response to the timbre creation request, where the timbre creation page displays reference text and recording controls.
[0111] At block 720 , the terminal device 120 receives user input audio in response to detecting triggering of the recording control.
[0112] In block 730 , the terminal device 120 creates a target voice for the user based on the received audio and the reference text.
[0113] In block 740 , the terminal device 120 presents a timbre confirmation page, which at least includes audition audio generated based on the target timbre.
[0114] In block 750 , in response to receiving confirmation of the target timbre in the timbre confirmation page, the terminal device 120 stores the target timbre for the user for use in audio.
[0115] In some embodiments, in the presence of user input text associated with a timbre creation request, the audition audio includes audio with a target timbre generated by converting the user input text through a text-to-speech operation, or if there is no user input text associated with a timbre creation request, the audition audio includes audio with a target timbre generated by converting a predetermined text through a text-to-speech operation.
[0116] In some embodiments, process 700 further includes: in response to detecting that the received audio does not match the reference text, prompting the user to re-record the audio; and displaying another reference text and recording controls in the sound creation page.
[0117] In some embodiments, process 700 further includes: prompting the user to re-record the audio in response to detecting that background noise in the audio exceeds a noise threshold or the duration of the audio is lower than a preset threshold; and displaying another reference text and recording controls in the tone creation page.
[0118] In some embodiments, process 700 further includes, in response to detecting the regenerate request in the sound creation page, displaying another reference text and recording controls in the sound creation page.
[0119] In some embodiments, the timbre creation request is detected in an editing page of the media content, and the method further includes: in response to receiving confirmation of the target timbre in the timbre confirmation page, jumping to the editing page, where the target timbre is selected for application to the audio of the media item being edited.
[0120] In some embodiments, the tone creation request is detected in a non-editing page, and the method further includes: in response to receiving confirmation of the target tone in the tone confirmation page, jumping to the editing page of the media content, where the target tone is selected for creating the media content.
[0121] In some embodiments, the non-editing page includes a sound detail page, or a sound detail page for another sound.
[0122] In some embodiments, process 700 further includes: receiving a publishing request for target media, the target media including at least target audio with a target timbre; and publishing the target media content and tag information related to the target timbre in association with the publishing request.
[0123] In some embodiments, the tag information includes at least an entry to a detail page corresponding to the target timbre and / or an entry to a timbre usage page corresponding to the target timbre.
[0124] In some embodiments, the tag information further includes information indicating intelligent generation of the target timbre.
[0125] In some embodiments, process 700 further includes: in response to receiving a request to delete the stored target timbre, deleting the stored target timbre; and in response to receiving a request to re-record the stored target timbre, presenting a timbre creation page, the timbre creation page displaying another reference text and recording controls.
[0126] 8 shows a schematic block diagram of an apparatus 800 for creating and using timbre according to certain embodiments of the present disclosure. The apparatus 800 may be implemented as or included in the terminal device 120. Each module / component in the apparatus 800 may be implemented by hardware, software, firmware, or any combination thereof.
[0127] As shown in the figure, the apparatus 800 includes a creation page presentation module 810 configured to present a timbre creation page in response to a timbre creation request, wherein the timbre creation page displays reference text and recording controls.
[0128] The apparatus 800 further includes an audio receiving module 820 configured to receive audio input by a user in response to detecting a triggering of the recording control.
[0129] The apparatus 800 further includes a timbre creation module 830 configured to create a target timbre for the user based on the received audio and reference text.
[0130] The apparatus 800 further includes a confirmation page presentation module 840 configured to present a timbre confirmation page, where the timbre confirmation page at least includes audition audio generated based on the target timbre.
[0131] The apparatus 800 further includes a timbre storage module 850 configured to store the target timbre for the user for use in audio in response to receiving confirmation of the target timbre in the timbre confirmation page.
[0132] In some embodiments, in the presence of user input text associated with a timbre creation request, the audition audio includes audio with a target timbre generated by converting the user input text through a text-to-speech operation, or if there is no user input text associated with a timbre creation request, the audition audio includes audio with a target timbre generated by converting a predetermined text through a text-to-speech operation.
[0133] In some embodiments, the device 800 further includes an audio prompt module configured to prompt the user to re-record the audio in response to detecting that the received audio does not match the reference text; and display another reference text and recording controls in the tone creation page.
[0134] In some embodiments, the audio prompt module is configured to prompt the user to re-record the audio in response to detecting that the background noise in the audio exceeds a noise threshold or the duration of the audio is lower than a preset threshold; and display another reference text and recording controls in the tone creation page.
[0135] In some embodiments, the apparatus 800 further includes a generation request module configured to display another reference text and a recording control in the timbre creation page in response to detecting a regeneration request in the timbre creation page.
[0136] In some embodiments, a timbre creation request is detected in an editing page of media content, and the device 800 further includes a jump module configured to jump to the editing page in response to receiving confirmation of a target timbre in a timbre confirmation page, wherein the target timbre is selected for application to the audio of the media item being edited.
[0137] In some embodiments, the tone creation request is detected in a non-editing page, and the jump module is further configured to jump to the editing page of the media content in response to receiving confirmation of the target tone in the tone confirmation page, where the target tone is selected for creating the media content.
[0138] In some embodiments, the non-editing page includes a sound detail page, or a sound detail page for another sound.
[0139] In some embodiments, the device 800 also includes a request receiving module, which is configured to receive a publishing request for a target media, the target media including at least a target audio with a target timbre; in response to the publishing request, the target media content and tag information related to the target timbre are associated and published.
[0140] In some embodiments, the tag information includes at least an entry to a detail page corresponding to the target timbre and / or an entry to a timbre usage page corresponding to the target timbre.
[0141] In some embodiments, the tag information further includes information indicating intelligent generation of the target timbre.
[0142] In some embodiments, the device 800 also includes a deletion request module, which is configured to delete the stored target timbre in response to receiving a deletion request for the stored target timbre; and present a timbre creation page in response to receiving a re-recording request for the stored target timbre, the timbre creation page displaying another reference text and recording controls.
[0143] FIG9 shows a block diagram of an electronic device 900 in which one or more embodiments of the present disclosure may be implemented. It should be understood that the electronic device 900 shown in FIG9 is merely exemplary and should not be construed as limiting the functionality and scope of the embodiments described herein. The electronic device 900 shown in FIG9 can be used to implement the terminal device 120 of FIG1 or the apparatus 800 of FIG8.
[0144] As shown in FIG9 , electronic device 900 is a general-purpose electronic device. Components of electronic device 900 may include, but are not limited to, one or more processors or processing units 910, memory 920, storage device 930, one or more communication units 940, one or more input devices 950, and one or more output devices 960. Processing unit 910 may be a real or virtual processor and is capable of performing various processes according to programs stored in memory 920. In a multi-processor system, multiple processing units execute computer-executable instructions in parallel to enhance the parallel processing capabilities of electronic device 900.
[0145] The electronic device 900 typically includes a plurality of computer storage media. Such media can be any accessible media that can be obtained by the electronic device 900, including but not limited to volatile and non-volatile media, removable and non-removable media. The memory 920 can be a volatile memory (e.g., a register, a cache, a random access memory (RAM)), a non-volatile memory (e.g., a read-only memory (ROM), an electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. The storage device 930 can be a removable or non-removable medium and can include a machine-readable medium, such as a flash drive, a disk, or any other medium that can be used to store information and / or data (e.g., training data for training) and can be accessed within the electronic device 900.
[0146] The electronic device 900 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in FIG. 9 , a disk drive for reading or writing from a removable, non-volatile disk (e.g., a “floppy disk”) and an optical drive for reading or writing from a removable, non-volatile optical disk may be provided. In these cases, each drive may be connected to a bus (not shown) by one or more data media interfaces. The memory 920 may include a computer program product 925 having one or more program modules configured to perform various methods or actions of various embodiments of the present disclosure.
[0147] The communication unit 940 enables communication with other electronic devices via a communication medium. Additionally, the functions of the components of the electronic device 900 can be implemented as a single computing cluster or multiple computing machines that can communicate via a communication connection. Thus, the electronic device 900 can operate in a networked environment using a logical connection with one or more other servers, a network personal computer (PC), or another network node.
[0148] The input device 950 may be one or more input devices, such as a mouse, keyboard, or trackball. The output device 960 may be one or more output devices, such as a display, a speaker, or a printer. The electronic device 900 may also communicate with one or more external devices (not shown) through the communication unit 940 as needed, such as a storage device, a display device, or the like, with one or more devices that allow a user to interact with the electronic device 900, or with any device that allows the electronic device 900 to communicate with one or more other electronic devices (e.g., a network card, a modem, etc.). Such communication may be performed via an input / output (I / O) interface (not shown).
[0149] According to an exemplary implementation of the present disclosure, a computer-readable storage medium is provided, on which computer-executable instructions are stored, wherein the computer-executable instructions are executed by a processor to implement the method described above. According to an exemplary implementation of the present disclosure, a computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, and the computer-executable instructions are executed by a processor to implement the method described above.
[0150] Various aspects of the present disclosure are described herein with reference to flowcharts and / or block diagrams of methods, apparatuses, devices, and computer program products implemented according to the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.
[0151] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine, such that when these instructions are executed by the processing unit of the computer or other programmable data processing device, a device is generated that implements the functions / actions specified in one or more blocks in the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium, where these instructions cause the computer, programmable data processing device, and / or other device to operate in a specific manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing various aspects of the functions / actions specified in one or more blocks in the flowchart and / or block diagram.
[0152] Computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device so that a series of operational steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to implement the functions / actions specified in one or more boxes in the flowchart and / or block diagram.
[0153] The flow charts and block diagrams in the accompanying drawings show the possible architecture, functions and operations of the systems, methods and computer program products according to multiple implementations of the present disclosure. In this regard, each box in the flow chart or block diagram can represent a part for a module, program segment or instruction, and a part for a module, program segment or instruction comprises one or more executable instructions for realizing the logical function of the specification. In some alternative implementations, the functions marked in the box can also occur in a sequence different from that marked in the accompanying drawings. For example, two continuous boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be realized by a special hardware-based system that performs the function or action of the specification, or can be realized by a combination of special hardware and computer instructions.
[0154] While various implementations of the present disclosure have been described above, the foregoing description is intended to be illustrative, not exhaustive, and not limited to the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is selected to best explain the principles of the implementations, their practical applications, or improvements to existing technologies, or to enable others skilled in the art to understand the various implementations disclosed herein.
Claims
1. A method for creating and using a timbre, comprising: In response to the timbre creation request, presenting a timbre creation page, the timbre creation page displaying reference text and recording controls; In response to detecting triggering of the recording control, receiving user input audio; Creating a target timbre for the user based on the received audio and the reference text; Presenting a timbre confirmation page, the timbre confirmation page at least including audition audio generated based on the target timbre; as well as In response to receiving confirmation of the target timbre in the timbre confirmation page, the target timbre for the user is stored for use in audio.
2. The method according to claim 1, wherein, in a case where there is user input text associated with the timbre creation request, the audition audio comprises audio having the target timbre generated by converting the user input text into a text-to-speech operation, or If there is no user input text associated with the timbre creation request, the audition audio includes audio having the target timbre generated by converting the predetermined text through a text-to-speech operation.
3. The method according to any one of claims 1 to 2, further comprising: In response to detecting that the received audio does not match the reference text, prompting the user to re-record the audio; as well as Another reference text and the recording controls are displayed in the sound creation page.
4. The method according to any one of claims 1 to 3, further comprising: In response to detecting that background noise in the audio exceeds a noise threshold or that the duration of the audio is less than a preset threshold, prompting the user to re-record the audio; as well as Another reference text and the recording controls are displayed in the sound creation page.
5. The method according to any one of claims 1 to 4, further comprising: In response to detecting a regeneration request in the sound creation page, another reference text and the recording control are displayed in the sound creation page.
6. The method according to any one of claims 1 to 5, wherein the timbre creation request is detected in an edit page of media content, and the method further comprises: In response to receiving confirmation of the target timbre in the timbre confirmation page, jumping to the editing page, wherein the target timbre is selected for application to the audio of the media item being edited.
7. The method according to any one of claims 1 to 6, wherein the timbre creation request is detected in a non-editing page, and the method further comprises: In response to receiving confirmation of the target timbre in the timbre confirmation page, the process jumps to an editing page of media content, where the target timbre is selected for composing media content.
8. The method according to claim 7, wherein the non-editing page comprises a sound detail page, or a timbre detail page of another timbre.
9. The method according to any one of claims 1 to 8, further comprising: receiving a publishing request for target media, wherein the target media at least includes target audio with the target timbre; In response to the publishing request, the target media content and the tag information related to the target timbre are associated with each other and published. 10 . The method according to claim 9 , wherein the tag information at least includes an entry to a detail page corresponding to the target timbre and / or an entry to a timbre usage page corresponding to the target timbre.
11. The method according to claim 9, wherein the tag information further includes instruction information for intelligent generation of the target timbre.
12. The method according to any one of claims 1 to 11, further comprising: In response to receiving a deletion request for the stored target timbre, deleting the stored target timbre; as well as In response to receiving a re-recording request for the stored target timbre, a timbre creation page is presented, which displays another reference text and recording controls.
13. A device for creating and using a timbre, comprising: a creation page presentation module configured to present a timbre creation page in response to a timbre creation request, wherein the timbre creation page displays reference text and recording controls; an audio receiving module, configured to receive audio input by a user in response to detecting a trigger on the recording control; a timbre creation module configured to create a target timbre for the user based on the received audio and the reference text; a confirmation page presentation module configured to present a timbre confirmation page, the timbre confirmation page including at least an audition audio generated based on the target timbre; as well as The timbre storage module is configured to store the target timbre for the user for use in audio in response to receiving confirmation of the target timbre in the timbre confirmation page.
14. An electronic device comprising: at least one processing unit; as well as At least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions causing the electronic device to perform the method according to any one of claims 1 to 12 when executed by the at least one processing unit.
15. A computer-readable storage medium having a computer program stored thereon, wherein the computer program can be executed by a processor to implement the method according to any one of claims 1 to 12.
16. A computer program product tangibly stored in a computer storage medium and comprising computer executable instructions which, when executed by a device, cause the device to perform the method according to any one of claims 1 to 12.
Citation Information
Patent Citations
Tone creation and use method and device, equipment and storage medium
CN120564691A
User tone based method and device for voice synthesis
CN108847215A
Voice tone customizing method and system for automobile prompt tones
CN112614481A
Timbre template customization method and device, equipment, medium and product
CN113744759A
Voice playing method and device, equipment and storage medium
CN114121028A