Audio processing method and device, equipment and storage medium

By obtaining the singing content and using the target tone model to convert it into a specified tone, the problem that traditional audio processing methods cannot restore the singing tone is solved, and efficient voice change effect in singing scenes is achieved.

CN120356478APending Publication Date: 2025-07-22BEIJING ZITIAO NETWORK TECH CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410084545.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-19
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

Traditional audio processing methods cannot effectively restore the tone in singing scenes, and the voice changer can only change the tone of the speaker and cannot output the tone of the singing.

Method used

By obtaining the singing content input by the user, converting it to a specified tone using the target tone model, retaining the original tone and rhythm, and generating the second media content.

Benefits of technology

It realizes the low cost of converting the singing voice to a specified tone in singing scenes, improving the voice change effect, and retaining the original tone characteristics.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120356478A_ABST
    Figure CN120356478A_ABST
Patent Text Reader

Abstract

The embodiment of the invention relates to an audio processing method and device, equipment and a storage medium. The method provided by the invention comprises the following steps: acquiring first media content input by a user, wherein the first media content comprises first audio content corresponding to singing content; and providing a second media content based on the user's selection of the target tone, the second media content including a second audio content corresponding to the singing content, the second audio content corresponding to the selected target tone. Through the mode, the first audio content corresponding to the singing content in the audio can be converted into the specified tone, so that the voice changing effect is improved on the basis of keeping the tone.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Example embodiments of the present disclosure generally relate to the field of computers, and in particular, to methods, apparatuses, devices, and computer-readable storage media for audio processing. Background Art

[0002] With the development of computer technology, the Internet has become an important platform for people to interact with information. In the process of people interacting with information through the Internet, various types of audio have become an important medium for people to socialize and exchange information. Therefore, there is an expectation to process audio to achieve voice conversion technology in singing scenarios. Summary of the Invention

[0003] In a first aspect of the present disclosure, a method for audio processing is provided. The method includes: obtaining first media content input by a user, the first media content including first audio content corresponding to singing content; and providing second media content based on a user's selection of a target timbre, the second media content including second audio content corresponding to the singing content, the second audio content corresponding to the selected target timbre.

[0004] In a second aspect of the present disclosure, an apparatus for audio processing is provided. The apparatus includes: an obtaining module configured to obtain first media content input by a user, the first media content including first audio content corresponding to singing content; and a providing module configured to provide second media content based on a user's selection of a target timbre, the second media content including second audio content corresponding to the singing content, the second audio content corresponding to the selected target timbre.

[0005] In a third aspect of the present disclosure, an electronic device is provided. The device includes at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit. The instructions, when executed by the at least one processing unit, cause the device to execute the method according to any one of the first aspect to the fourth aspect.

[0006] In a fourth aspect of the present disclosure, a computer-readable storage medium is provided. A computer program is stored on the computer-readable storage medium and can be executed by a processor to implement the method according to any one of the first aspect to the fourth aspect.

[0007] It should be understood that the content described in this content part is not intended to limit the key features or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. Brief Description of the Drawings

[0008] The above and other features, advantages, and aspects of the embodiments of the present disclosure will become more apparent by referring to the accompanying drawings and the following detailed description. In the drawings, the same or similar reference numerals denote the same or similar elements, where:

[0009] Figure 1 A schematic diagram showing an example environment in which embodiments according to the present disclosure can be implemented;

[0010] Figure 2 A flowchart showing an example process for audio processing according to some embodiments of the present disclosure;

[0011] Figures 3A to 3C A schematic diagram showing an example interface according to some embodiments of the present disclosure;

[0012] Figure 4 A schematic structural block diagram showing an example apparatus for audio processing according to some embodiments of the present disclosure; and

[0013] Figure 5 A block diagram of an electronic device capable of implementing multiple embodiments of the present disclosure. Detailed Embodiments

[0014] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.

[0015] It should be noted that the titles of any sections / subsections provided herein are not restrictive. Various embodiments are described throughout this document, and any type of embodiment can be included under any section / subsection. Additionally, the embodiments described in any section / subsection can be combined with any other embodiments described in the same section / subsection and / or different sections / subsections in any manner.

[0016] In the description of the embodiments of the present disclosure, the term "including" and its like shall be understood as an open inclusion, that is, "including but not limited to". The term "based on" shall be understood as "at least partially based on". The term "one embodiment" or "the embodiment" shall be understood as "at least one embodiment". The term "some embodiments" shall be understood as "at least some embodiments". There may also be other explicit and implicit definitions hereinafter. The terms "first", "second", etc. may refer to different or the same objects. There may also be other explicit and implicit definitions hereinafter.

[0017] Embodiments of the present disclosure may involve user data, data acquisition and / or use, etc. All of these aspects comply with the corresponding laws, regulations and related provisions. In the embodiments of the present disclosure, the collection, acquisition, processing, processing, forwarding, use, etc. of all data are carried out on the premise that the user is aware of and confirms. Accordingly, when implementing the embodiments of the present disclosure, the types, usage scopes, usage scenarios, etc. of the data or information that may be involved should be informed to the user and the user's authorization should be obtained in an appropriate manner according to relevant laws and regulations. The specific notification and / or authorization methods may vary according to the actual situation and application scenarios, and the scope of the present disclosure is not limited in this regard.

[0018] In the solutions described in this specification and the embodiments, if personal information processing is involved, it will be processed on the premise of having a legal basis (such as obtaining the consent of the personal information subject, or being necessary for performing a contract, etc.), and will only be processed within the specified or agreed scope. If the user refuses to process personal information other than the necessary information required for the basic functions, it will not affect the user's use of the basic functions.

[0019] In the process of people's information interaction through the Internet, people expect to use high-quality audio processing methods to conveniently achieve the desired voice-changing effect. The traditional audio processing method is to use a voice changer to change the voice. However, the voice changer can only change the speaker's tone in the voiceover scenario. If the input audio is a person singing. The voice-changing effect cannot restore the input singing pitch, and the output will still be voiceover content.

[0020] In view of this, embodiments of the present disclosure propose an audio processing solution. According to this solution, the first audio content corresponding to the singing content in the audio can be converted into a specified tone. The specified tone is a tone existing in the tone library or a tone authorized for use. Thus, on the basis of retaining the tone, the voice-changing effect is improved. For example, it can enable users to hear at low cost what the effect is when their own singing is sung in the tone of other characteristics.

[0021] The following further describes various example implementations of this solution in detail with reference to the accompanying drawings.

[0022] Example environment

[0023] Figure 1 The figure shows a schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented. As Figure 1 shown, the example environment 100 may include a terminal device 110.

[0024] In this example environment 100, the terminal device 110 may run a platform that supports audio processing. For example, performing voice-changing processing on the audio. The user 140 can interact with this platform via the terminal device 110 and / or its attached devices.

[0025] In Figure 1 environment 100, if the platform is active, the terminal device 110 can present an interface 150 for supporting interface interaction through the platform.

[0026] In some embodiments, the terminal device 110 communicates with the server 130 to implement the supply of services for the platform. The terminal device 110 can be any type of mobile terminal, fixed terminal or portable terminal, including mobile phones, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, media computers, multimedia tablets, handheld computers, portable game terminals, VR / AR devices, Personal Communication System (PCS) devices, personal navigation devices, Personal Digital Assistant (PDA), audio / video players, digital cameras / camcorders, positioning devices, television receivers, radio broadcast receivers, e-book devices, game devices, or any combination of the foregoing, including accessories and peripherals of these devices or any combination thereof. In some embodiments, the terminal device 110 can also support any type of user interface (such as a "wearable" circuit, etc.).

[0027] The server 130 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks, and big data and artificial intelligence platforms. The server 130 can include, for example, a computing system / server, such as a mainframe, an edge computing node, a computing device in a cloud environment, and so on. The server 130 can provide background services for the application 120 that supports virtual scenarios in the terminal device 110.

[0028] A communication connection can be established between the server 130 and the terminal device 110. The communication connection can be established by wired or wireless means. The communication connection can include, but is not limited to, a Bluetooth connection, a mobile network connection, a Universal Serial Bus (USB) connection, a Wireless Fidelity (WiFi) connection, etc., and the embodiments of the present disclosure are not limited in this regard. In the embodiments of the present disclosure, the server 130 and the terminal device 110 can implement signaling interaction through the communication connection therebetween.

[0029] It should be understood that the structures and functions of the various elements in the environment 100 are described only for exemplary purposes and do not imply any limitation on the scope of the present disclosure.

[0030] Some example embodiments of the present disclosure will be further described below with reference to the accompanying drawings.

[0031] Example process

[0032] Figure 2 A flowchart of an example process 200 for audio processing according to some embodiments of the present disclosure is shown. The process 200 may be implemented at the terminal device 110. The following will be described with reference to Figure 1 to describe the process 200.

[0033] As Figure 2 shown, at block 210, the terminal device 110 obtains first media content input by the user 140. The first media content includes first audio content corresponding to singing content.

[0034] In some embodiments, the first media content input by the user 140 obtained by the terminal device 110 may be first media content recorded by the user 140. For example, a video taken by the user, or a recorded voice as the first media content. The first media content includes an audio corresponding to singing content. For example, a video taken by the user includes a song sung by the user himself.

[0035] In some embodiments, the first media content input by the user 140 obtained by the terminal device 110 may be first media content uploaded by the user 140. For example, a previous video stored on the terminal device 110, or a previously recorded voice stored as the first media content.

[0036] The following will be described with reference to Figures 3A to 3C to describe the process 200. Figures 3A to 3C A schematic diagram of example interfaces 301 to 303 according to some embodiments of the present disclosure is shown. The interfaces 301 to 303 may be provided by, for example, Figure 1 the terminal device 110 shown.

[0037] As Figure 3A shown, the terminal device 110 obtains that the user 140 records a video on the shooting page, and the video includes first audio content corresponding to singing content. Subsequently, the user 140 clicks on the voice control 311 in the interface 301.

[0038] Based on the click of the user 140, the terminal device 110 will present a selection panel 320. As Figure 3BAs shown, the interface 301 can be, for example, the session interface of a session. The terminal device 110 can display a selection panel 320 in the interface 301. As an example, the selection panel 320 can display timbres of different styles, such as style one 321, style two 322, style three, and so on.

[0039] In some embodiments, the selection panel 320 displayed by the terminal device 110 can provide a first set of candidate effects for processing speech content. For example, when the terminal device 110 obtains the speech content input by the user (such as a voiceover), the selection panel 320 can provide different styles that the speech content can be converted into.

[0040] In some embodiments, the selection panel 320 displayed by the terminal device 110 can also provide a second set of candidate effects for processing singing content. For example, when the terminal device 110 obtains the singing content input by the user, the selection panel 320 can provide different styles that the input singing content can be converted into while retaining its original pitch, rhythm, and so on.

[0041] In some embodiments, the terminal device 110 receives the user's selection of a target effect from the second set of candidate effects, and the target effect corresponds to a target timbre. For example, the terminal device 110 can receive the timbres of different genders or different timbres corresponding to different ages selected by the user from the second set of candidate effects.

[0042] Continuing to refer to Figure 2 , at block 220, the terminal device 110 provides second media content based on the user 140's selection of the target timbre. In some embodiments, the second media content includes second audio content corresponding to the singing content. The second audio content corresponds to the selected target timbre.

[0043] In some embodiments, the second audio content included in the second media content provided by the terminal device 110 retains at least the target audio attributes of the first audio content. The target audio attributes include at least pitch, rhythm, and so on. For example, if the pitch of the first audio content includes key A and key B, the second audio content retains at least key A and key B.

[0044] Taking Figure 3C as an example, the user 140 selects style three 331 as the target timbre on the timbre selection panel 320. The terminal device 110 receives the selection of the user 140 and converts the first audio content in the first media content into the second audio content in the second media content. For example, it changes the voice of the singing in the video uploaded by the user.

[0045] For example, if the terminal device 110 obtains the speech content input by the user, it can change the speech content in different styles. For another example, if the terminal device 110 obtains the singing content input by the user, it can change the singing content in different styles and retain the pitch, rhythm, etc. of the original singing. Based on the user's selection of the target effect, the terminal device 110 calls the server 130 to convert the first audio content and provides the second audio content.

[0046] If the first audio content corresponding to singing included in the first media content of the terminal device 110, the terminal device 110 converts the first audio content into the second audio content by calling the server 130 based on the target timbre selected by the user and provides it to the user. In this way, the second audio content can include the pitch of the first audio input. It can be understood that in this way, the user can hear at low cost what the effect is when their own singing is sung in the timbre of other characteristics.

[0047] The generation of the second audio content is described below. In some embodiments, the terminal device 110 extracts the first audio content corresponding to the singing content from the first media content. Subsequently, the terminal device 110 inputs the first audio content into the target model to obtain the second audio content. In some embodiments, the target model can be a model trained by the server according to the sample data corresponding to the target timbre.

[0048] In some embodiments, the terminal device 110 extracts the background audio content from the first media content. The background audio content is different from the first audio content, and the background audio content corresponds to the accompaniment content. For example, for a video, the background audio content can be the background music of the video. For this video, the first audio content can be the song sung by the user himself. The terminal device 110 generates the second media content by fusing the second audio content and the background audio content.

[0049] In some embodiments, the terminal device 110 obtains the intermediate audio content by fusing the second audio content and the background audio content. The terminal device 110 adjusts the reverberation effect or the volume of the intermediate audio content. In some examples, adjusting the volume of the intermediate audio content includes performing global volume equalization on the intermediate audio content. The terminal device 110 generates the second media content according to the adjusted intermediate audio content.

[0050] In some embodiments, the terminal device 110 can first determine the reverberation parameters according to the first media content. Then, the terminal device 110 adjusts the reverberation effect of the intermediate audio content according to the reverberation parameters.

[0051] For example, the reverberation matching model and the support vector classifier (SVC) model are the models that need to be engineered throughout the entire link. The input and output of the two models are independent of each other and can be processed in parallel. The reverberation matching module is a lightweight convolutional neural network (CNN) and long short-term memory network (LSTM) module, which generally finishes running earlier than the SVC model. The reverberation matching model is only responsible for calculating the reverberation parameters from the original user's voice (the output is 3 scalars). The actual addition of audio reverberation is completed by the central processing unit (CPU) in the last step before the link output. The biggest difference between the overall link and the VC link should be the logic of reverberation matching and volume equalization, as well as the additional robust model voice pitch extractor (RMVPE f0 extractor) for polyphonic music.

[0052] Embodiments of the present disclosure obtain first media content input by a user, where the first media content includes first audio content corresponding to singing content; and provide second media content based on the user's selection of a target timbre, where the second media content includes second audio content corresponding to the singing content, and the second audio content corresponds to the selected target timbre. In this way, embodiments of the present disclosure can convert the first audio content corresponding to the singing content in the audio into a specified timbre, thereby improving the voice-changing effect while preserving the timbre.

[0053] Example devices and equipment

[0054] Embodiments of the present disclosure also provide corresponding devices for implementing the above methods or processes. Figure 4 FIG. shows a schematic structural block diagram of an example device 400 for audio processing according to certain embodiments of the present disclosure. The device 400 can be implemented as or included in a terminal device 110. Each module / component in the device 400 can be implemented by hardware, software, firmware, or any combination thereof.

[0055] As Figure 4 shown, the device 400 includes an acquisition module 410 configured to acquire first media content input by a user, where the first media content includes first audio content corresponding to singing content.

[0056] The device 400 further includes a providing module 420 configured to provide second media content based on the user's selection of a target timbre, where the second media content includes second audio content corresponding to the singing content, and the second audio content corresponds to the selected target timbre.

[0057] In some embodiments, the second audio content at least preserves the target audio attributes of the first audio content, and the target audio attributes include at least one of the following: pitch, rhythm.

[0058] In some embodiments, the providing module 420 further includes a selection module configured to display a selection panel, where the selection panel provides at least a first set of candidate effects for processing speech content and a second set of candidate effects for processing singing content; and receive a user's selection of a target effect from the second set of candidate effects, where the target effect corresponds to a target timbre.

[0059] In some embodiments, the obtaining module 410 is further configured to obtain first media content recorded by a user; obtain first media content uploaded by the user.

[0060] In some embodiments, the providing module 420 further includes a generating module configured to extract first audio content corresponding to singing content from the first media content; and process the first audio content using a target model to generate second audio content, where the target model is trained based on sample data corresponding to the target timbre.

[0061] In some embodiments, the generating module is further configured to extract background audio content from the first media content, where the background audio content is different from the first audio content; and generate second media content by fusing the second audio content and the background audio content.

[0062] In some embodiments, the background audio content corresponds to accompaniment content.

[0063] In some embodiments, the generating module is further configured to fuse the second audio content and the background audio content to obtain intermediate audio content; adjust the reverberation effect or volume of the intermediate audio content; and generate second media content based on the adjusted intermediate audio content.

[0064] In some embodiments, the providing module 420 further includes an adjusting module configured to determine reverberation parameters based on the first media content; and adjust the reverberation effect of the intermediate audio content based on the reverberation parameters.

[0065] In some embodiments, the adjusting module is further configured to perform global volume equalization on the intermediate audio content.

[0066] Figure 5 A block diagram of an electronic device 500 in which one or more embodiments of the present disclosure may be implemented is shown. It should be understood that Figure 5 The illustrated electronic device 500 is merely exemplary and should not constitute any limitation to the functions and scope of the embodiments described herein. Figure 5 The illustrated electronic device 500 may be used to implement Figure 1 the electronic device 110.

[0067] As Figure 5As shown, the electronic device 500 is in the form of a general-purpose electronic device. The components of the electronic device 500 may include, but are not limited to, one or more processors or processing units 510, a memory 520, a storage device 530, one or more communication units 540, one or more input devices 550, and one or more output devices 560. The processing unit 510 may be an actual or virtual processor and be capable of performing various processes according to the programs stored in the memory 520. In a multi-processor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing ability of the electronic device 500.

[0068] The electronic device 500 generally includes multiple computer storage media. Such media may be any accessible media that can be obtained by the electronic device 500, including but not limited to volatile and non-volatile media, removable and non-removable media. The memory 520 may be a volatile memory (such as registers, caches, random access memory (RAM)), a non-volatile memory (such as read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. The storage device 530 may be a removable or non-removable medium and may include machine-readable media, such as a flash drive, a magnetic disk, or any other medium that can be used to store information and / or data and can be accessed within the electronic device 500.

[0069] The electronic device 500 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in Figure 5 a disk drive for reading from or writing to a removable, non-volatile magnetic disk (such as a "floppy disk") and an optical disk drive for reading from or writing to a removable, non-volatile optical disk may be provided. In these cases, each drive may be connected to a bus (not shown) by one or more data media interfaces. The memory 520 may include a computer program product 525 having one or more program modules that are configured to perform various methods or actions of the various embodiments of the present disclosure.

[0070] The communication unit 540 enables communication with other electronic devices through a communication medium. Additionally, the functions of the components of the electronic device 500 may be implemented by a single computing cluster or multiple computer machines that are capable of communicating through a communication connection. Thus, the electronic device 500 may operate in a networked environment using a logical connection to one or more other servers, network personal computers (PCs), or another network node.

[0071] The input device 550 can be one or more input devices, such as a mouse, a keyboard, a trackball, etc. The output device 560 can be one or more output devices, such as a display, a speaker, a printer, etc. The electronic device 500 can also communicate with one or more external devices (not shown) as needed through the communication unit 540. The external devices such as a storage device, a display device, etc., communicate with one or more devices that enable a user to interact with the electronic device 500, or communicate with any device that enables the electronic device 500 to communicate with one or more other electronic devices (e.g., a network card, a modem, etc.). Such communication can be performed via an input / output (I / O) interface (not shown).

[0072] According to an exemplary implementation of the present disclosure, there is provided a computer-readable storage medium having computer-executable instructions stored thereon, wherein the computer-executable instructions are executed by a processor to implement the method described above. According to an exemplary implementation of the present disclosure, there is also provided a computer program product, the computer program product being tangibly stored on a non-transitory computer-readable medium and including computer-executable instructions, and the computer-executable instructions being executed by a processor to implement the method described above.

[0073] Aspects of the present disclosure are described herein with reference to the flowcharts and / or block diagrams of methods, apparatuses, devices, and computer program products according to the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.

[0074] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus, thereby producing a machine such that when these instructions are executed by the processing unit of the computer or other programmable data processing apparatus, a device is produced that implements the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium, and these instructions cause a computer, a programmable data processing apparatus, and / or other devices to work in a specific manner. Thus, the computer-readable medium storing the instructions includes a manufacture, which includes instructions for implementing various aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.

[0075] The computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other devices, such that a series of operation steps are performed on the computer, other programmable data processing apparatus, or other devices to produce a computer-implemented process, so that the instructions executed on the computer, other programmable data processing apparatus, or other devices implement the functions / actions specified in one or more blocks of the flowchart and / or block diagram.

[0076] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various implementations of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a part of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the blocks may occur in a different order than noted in the accompanying drawings. For example, two consecutive blocks may actually be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented by a dedicated hardware-based system that performs the specified functions or actions, or by a combination of dedicated hardware and computer instructions.

[0077] The various implementations of the present disclosure have been described above. The above description is exemplary, not exhaustive, and is not limited to the disclosed implementations. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described implementations. The choice of terms used herein is intended to best explain the principles of the implementations, the practical application, or the improvement of technologies in the market, or to enable other ordinary skilled persons in the art to understand the various implementation manners disclosed herein.

Claims

1. An audio processing method, comprising: Obtaining first media content input by a user, where the first media content includes first audio content corresponding to singing content; And Based on the user's selection of a target timbre, providing second media content, where the second media content includes second audio content corresponding to the singing content, and the second audio content corresponds to the selected target timbre.

2. The method according to claim 1, wherein The second audio content at least retains target audio attributes of the first audio content, and the target audio attributes include at least one of the following: pitch, rhythm.

3. The method according to claim 1, further comprising: Displaying a selection panel, where the selection panel at least provides a first group of candidate effects for processing speech content and a second group of candidate effects for processing singing content; And Receiving the user's selection of a target effect in the second group of candidate effects, where the target effect corresponds to the target timbre.

4. The method according to claim 1, where obtaining the first media content input by the user includes: Obtaining the first media content recorded by the user; Obtaining the first media content uploaded by the user.

5. The method according to claim 1, where the second audio content is generated through the following process: Extracting the first audio content corresponding to the singing content from the first media content; and Processing the first audio content using a target model to generate the second audio content, where the target model is trained based on sample data corresponding to the target timbre.

6. The method according to claim 5, where the second media content is generated through the following process: Extracting background audio content from the first media content, where the background audio content is different from the first audio content; and Generating the second media content by fusing the second audio content and the background audio content.

7. The method according to claim 6, where the background audio content corresponds to accompaniment content.

8. The method according to claim 6, where generating the second media content by fusing the second audio content and the background audio content includes: Fusing the second audio content and the background audio content to obtain intermediate audio content; Adjusting the reverberation effect or volume of the intermediate audio content; And Based on the adjusted intermediate audio content, generating the second media content.

9. The method according to claim 8, where adjusting the reverberation effect of the intermediate audio content includes: Determining reverberation parameters based on the first media content; And Based on the reverberation parameters, adjusting the reverberation effect of the intermediate audio content.

10. The method according to claim 8, where adjusting the volume of the intermediate audio content includes: Performing global volume equalization on the intermediate audio content.

11. An apparatus for audio processing, comprising: An obtaining module configured to obtain first media content input by a user, where the first media content includes first audio content corresponding to singing content; And A providing module, configured to provide second media content based on the user's selection of a target timbre, where the second media content includes second audio content corresponding to the singing content, and the second audio content corresponds to the selected target timbre.

12. An electronic device, comprising: At least one processing unit; And At least one memory, the at least one memory being coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions, when executed by the at least one processing unit, causing the electronic device to perform the method according to any one of claims 1 to 10.

13. A computer-readable storage medium having stored thereon a computer program, the computer program being executable by a processor to implement the method according to any one of claims 1 to 10.