Audio processing method and apparatus, and device and storage medium
By obtaining the singing content and using the target tone model to convert it into a specified tone, the problem that traditional audio processing methods cannot restore the singing tone is solved, and efficient voice change effect in singing scenes is achieved, and tone and rhythm are retained.
Patent Information
- Application Number
- PCT/CN2024/126105
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-19
- Filing Date
- 2024-10-21
- Publication Date
- 2025-07-24
AI Technical Summary
Traditional audio processing methods cannot effectively restore the tone in singing scenes, and the voice changer can only change the speaker's tone and cannot output the correct tone.
By obtaining the singing content input by the user, converting it to a specified tone using the target tone model, retaining the original tone and rhythm, and generating the second audio content corresponding to the singing content.
It realizes the low cost of converting audio content to specified tones in singing scenes, improving the sound change effect, and retaining the tone and rhythm of the original tones.
Smart Images

Figure CN2024126105_24072025_PF_FP_ABST
Abstract
Description
Audio processing method, device, equipment and storage medium
[0001] This application claims priority to the Chinese invention patent application entitled “Method, device, equipment and storage medium for audio processing” and application number 202410084545.9 filed on January 19, 2024, the entire contents of which are incorporated by reference into this application. Technical Field
[0002] Example embodiments of the present disclosure generally relate to the field of computers, and more particularly, to methods, devices, apparatuses, and computer-readable storage media for audio processing. Background Art
[0003] With the development of computer technology, the internet has become a vital platform for information exchange. In this process, various types of audio have become crucial media for social expression and information exchange. Therefore, there is a desire to process audio to achieve voice-changing technology for singing.
[0004] Summary of the Invention
[0005] In a first aspect of the present disclosure, a method for audio processing is provided. The method comprises: obtaining first media content input by a user, the first media content comprising first audio content corresponding to singing content; and providing second media content, based on the user's selection of a target timbre, the second media content comprising second audio content corresponding to the singing content, the second audio content corresponding to the selected target timbre.
[0006] In a second aspect of the present disclosure, a device for audio processing is provided. The device includes: an acquisition module configured to acquire first media content input by a user, the first media content including first audio content corresponding to singing content; and a provision module configured to provide second media content based on the user's selection of a target timbre, the second media content including second audio content corresponding to the singing content, the second audio content corresponding to the selected target timbre.
[0007] In a third aspect of the present disclosure, an electronic device is provided. The device includes at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit. When executed by the at least one processing unit, the instructions cause the device to perform the method of any one of the first to fourth aspects.
[0008] In a fourth aspect of the present disclosure, a computer-readable storage medium is provided, wherein a computer program is stored on the computer-readable storage medium, and the computer program can be executed by a processor to implement the method of any one of the first to fourth aspects.
[0009] It should be understood that the content described in this summary section is not intended to limit the key features or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] The above and other features, advantages and aspects of the embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. In the accompanying drawings, the same or similar reference numerals represent the same or similar elements, wherein:
[0011] FIG1 shows a schematic diagram of an example environment in which embodiments according to the present disclosure may be implemented;
[0012] FIG2 illustrates a flow chart of an example process for audio processing according to some embodiments of the present disclosure;
[0013] 3A to 3C are schematic diagrams illustrating example interfaces according to some embodiments of the present disclosure;
[0014] FIG4 shows a schematic structural block diagram of an example apparatus for audio processing according to some embodiments of the present disclosure; and
[0015] FIG5 shows a block diagram of an electronic device capable of implementing various embodiments of the present disclosure. DETAILED DESCRIPTION
[0016] The following describes embodiments of the present disclosure in more detail with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments described herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.
[0017] It should be noted that the titles of any section / subsection provided herein are not limiting. Various embodiments are described throughout this document, and any type of embodiment may be included under any section / subsection. Furthermore, the embodiments described in any section / subsection may be combined in any manner with any other embodiments described in the same section / subsection and / or in different sections / subsections.
[0018] In the description of the embodiments of the present disclosure, the term "including" and similar terms should be understood as open inclusion, that is, "including but not limited to". The term "based on" should be understood as "based at least in part on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". Other explicit and implicit definitions may be included below. The terms "first", "second", etc. may refer to different or the same objects. Other explicit and implicit definitions may be included below.
[0019] The embodiments of the present disclosure may involve user data, data acquisition and / or use, etc. These aspects shall comply with the corresponding laws, regulations and relevant provisions. In the embodiments of the present disclosure, all data collection, acquisition, processing, processing, forwarding, use, etc. are carried out on the premise that the user is aware of and confirms them. Accordingly, when implementing the various embodiments of the present disclosure, the types, scope of use, and usage scenarios of the data or information that may be involved should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with the relevant laws and regulations. The specific notification and / or authorization method may vary according to the actual situation and application scenario, and the scope of the present disclosure is not limited in this respect.
[0020] If this specification and the solutions in the examples involve the processing of personal information, such processing will be done only with a legitimate basis (such as with the consent of the subject of personal information or as necessary for the performance of a contract) and only within the prescribed or agreed scope. A user's refusal to process personal information other than that required for basic functions will not affect the user's use of basic functions.
[0021] As people interact with each other online, they expect to use high-quality audio processing methods to easily achieve the desired voice-changing effects. Traditional audio processing methods rely on voice changers. However, voice changers can only alter the speaker's timbre in spoken audio scenarios. If the input is singing audio, the voice-changing effect cannot restore the input singing pitch, and the output will still be spoken audio.
[0022] In light of this, embodiments of the present disclosure propose an audio processing solution. This solution converts the first audio content corresponding to the singing into a specified timbre. This specified timbre can be a timbre already in the sound library or a timbre that has been authorized for use. This improves the voice-changing effect while preserving the timbre. For example, users can cost-effectively hear what their own singing would sound like with a different characteristic timbre.
[0023] Various example implementations of this solution are described in detail below in conjunction with the accompanying drawings.
[0024] Sample Environment
[0025] FIG1 shows a schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented. As shown in FIG1 , the example environment 100 may include a terminal device 110 .
[0026] In this example environment 100, the terminal device 110 may run a platform that supports audio processing, such as voice modification, and the user 140 may interact with the platform via the terminal device 110 and / or its attached devices.
[0027] In the environment 100 of FIG. 1 , if the platform is in an active state, the terminal device 110 may present an interface 150 for supporting interface interaction through the platform.
[0028] In some embodiments, the terminal device 110 communicates with the server 130 to enable the provision of services to the platform. The terminal device 110 can be any type of mobile terminal, fixed terminal or portable terminal, including a mobile phone, a desktop computer, a laptop computer, a notebook computer, a netbook computer, a tablet computer, a media computer, a multimedia tablet, a handheld computer, a portable game terminal, a VR / AR device, a personal communication system (Personal Communication System, PCS) device, a personal navigation device, a personal digital assistant (Personal Digital Assistant, PDA), an audio / video player, a digital camera / camcorder, a positioning device, a television receiver, a radio broadcast receiver, an e-book device, a gaming device or any combination thereof, including accessories and peripherals of these devices or any combination thereof. In some embodiments, the terminal device 110 can also support any type of interface for the user (such as a "wearable" circuit, etc.).
[0029] Server 130 can be a standalone physical server, a server cluster or distributed system consisting of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content distribution networks, and big data and artificial intelligence platforms. Server 130 can include, for example, computing systems / servers such as mainframes, edge computing nodes, computing devices in a cloud environment, and the like. Server 130 provides backend services for application 120 that supports virtual scenarios in terminal device 110.
[0030] A communication connection may be established between the server 130 and the terminal device 110. The communication connection may be established via a wired or wireless method. The communication connection may include, but is not limited to, a Bluetooth connection, a mobile network connection, a Universal Serial Bus (USB) connection, a Wireless Fidelity (WiFi) connection, etc., and the embodiments of the present disclosure are not limited in this respect. In the embodiments of the present disclosure, the server 130 and the terminal device 110 may implement signaling interaction via the communication connection between the two.
[0031] It should be understood that the structure and function of the various elements in the environment 100 are described for illustrative purposes only and do not imply any limitation on the scope of the present disclosure.
[0032] Some example embodiments of the present disclosure will be described below with continued reference to the accompanying drawings.
[0033] Example Process
[0034] FIG2 shows a flow chart of an example process 200 for audio processing according to some embodiments of the present disclosure. The process 200 may be implemented at the terminal device 110. The process 200 is described below with reference to FIG1.
[0035] As shown in Figure 2, in block 210, the terminal device 110 obtains first media content input by the user 140. The first media content includes first audio content corresponding to singing content.
[0036] In some embodiments, the first media content input by user 140 and obtained by terminal device 110 may be first media content recorded by user 140. For example, the first media content may include a video shot by the user or a voice recording. The first media content may include an audio clip corresponding to singing content. For example, a video shot by the user may include a song sung by the user.
[0037] In some embodiments, the first media content input by the user 140 and acquired by the terminal device 110 may be the first media content uploaded by the user 140. For example, a previously stored video or a previously recorded voice on the terminal device 110 may be used as the first media content.
[0038] The process 200 will be described below with reference to Figures 3A to 3C . Figures 3A to 3C show schematic diagrams of example interfaces 301 to 303 according to some embodiments of the present disclosure. Interfaces 301 to 303 may be provided by the terminal device 110 shown in Figure 1 , for example.
[0039] As shown in FIG3A , the terminal device 110 obtains a video recorded by the user 140 on the shooting page, and the video includes first audio content corresponding to singing content. Subsequently, the user 140 clicks the voice control 311 in the interface 301 .
[0040] Upon user 140's click, terminal device 110 presents selection panel 320. As shown in FIG3B , interface 301 may be, for example, a conversation interface. Terminal device 110 may display selection panel 320 within interface 301. For example, selection panel 320 may display timbres of different styles, such as Style 1 321, Style 2 322, Style 3, and so on.
[0041] In some embodiments, the selection panel 320 displayed by the terminal device 110 may provide a first set of candidate effects for processing the spoken content. For example, if the terminal device 110 receives spoken content input by the user (e.g., spoken audio), the selection panel 320 may provide different styles that can be converted to the spoken content.
[0042] In some embodiments, the selection panel 320 displayed by the terminal device 110 may also provide a second set of candidate effects for processing the singing content. For example, if the terminal device 110 receives singing content input by the user, the selection panel 320 may provide different styles that can be converted to the input singing content while retaining its original pitch, rhythm, etc.
[0043] In some embodiments, the terminal device 110 receives a user's selection of a target effect from the second set of candidate effects, where the target effect corresponds to a target timbre. For example, the terminal device 110 may receive timbres of different genders or different timbres corresponding to different ages from the second set of candidate effects selected by the user.
[0044] 2 , at block 220 , the terminal device 110 provides second media content based on the target timbre selected by the user 140 . In some embodiments, the second media content includes second audio content corresponding to the singing content. The second audio content corresponds to the selected target timbre.
[0045] In some embodiments, the second audio content included in the second media content provided by the terminal device 110 retains at least the target audio attributes of the first audio content. The target audio attributes include at least pitch, rhythm, etc. For example, if the pitch of the first audio content includes A and B, the second audio content at least retains A and B.
[0046] Taking Figure 3C as an example, user 140 selects Style 3 331 as the target timbre on timbre selection panel 320. Upon receiving user 140's selection, terminal device 110 converts the first audio content in the first media content into the second audio content in the second media content. For example, this may involve voice-changing the singing voice in a user-uploaded video.
[0047] For example, if the terminal device 110 receives user input of spoken content, it can modify the spoken content to a different style. For another example, if the terminal device 110 receives user input of singing content, it can modify the singing content to a different style while preserving the original song's pitch, rhythm, and so on. Based on the user's selection of the target effect, the terminal device 110 calls the server 130 to convert the first audio content and provide the second audio content.
[0048] If the terminal device 110 receives first audio content corresponding to singing included in the first media content, it can convert the first audio content into second audio content based on the target timbre selected by the user by calling server 130, and provide it to the user. In this way, the second audio content can contain the tones of the input first audio. It can be understood that through this method, users can hear what their own singing voice would sound like with a different characteristic timbre at a low cost.
[0049] The following describes the generation of the second audio content. In some embodiments, the terminal device 110 extracts first audio content corresponding to the singing content from the first media content. Subsequently, the terminal device 110 inputs the first audio content into a target model to obtain the second audio content. In some embodiments, the target model can be a model trained by a server based on sample data corresponding to the target timbre.
[0050] In some embodiments, terminal device 110 extracts background audio content from the first media content. This background audio content is different from the first audio content and corresponds to accompaniment content. For example, for a video, the background audio content may be the background music of the video. For the video, the first audio content may be a song sung by the user. Terminal device 110 generates second media content by fusing the second audio content with the background audio content.
[0051] In some embodiments, terminal device 110 obtains intermediate audio content by fusing the second audio content and the background audio content. Terminal device 110 adjusts the reverberation effect or volume of the intermediate audio content. In some examples, adjusting the volume of the intermediate audio content includes performing global volume equalization on the intermediate audio content. Terminal device 110 generates second media content based on the adjusted intermediate audio content.
[0052] In some embodiments, the terminal device 110 may first determine a reverberation parameter based on the first media content, and then adjust the reverberation effect of the intermediate audio content based on the reverberation parameter.
[0053] For example, the reverberation matching model and the support vector classifier (SVC) model are models that require engineering across the entire link. Their inputs and outputs are independent and can run in parallel. The reverberation matching module, comprised of a lightweight convolutional neural network (CNN) and a long short-term memory (LSTM) network, typically completes before the support vector classifier (SVC) model. The reverberation matching model is solely responsible for calculating reverberation parameters from the original user voice (outputting three scalars). The actual addition of audio reverberation is performed by the central processing unit (CPU) as the final step before the link output. The biggest difference between the overall link and the VC link is likely the logic for reverberation matching and volume equalization, as well as the additional robust model extractor (RMVPEf0) for estimating high-pitched sounds in polyphonic music.
[0054] The embodiments of the present disclosure obtain first media content input by a user, the first media content including first audio content corresponding to singing content; and based on the user's selection of a target timbre, provide second media content including second audio content corresponding to the singing content, the second audio content corresponding to the selected target timbre. In this way, the embodiments of the present disclosure can convert the first audio content corresponding to the singing content in the audio to the specified timbre, thereby improving the voice changing effect while preserving the timbre.
[0055] Example devices and equipment
[0056] Embodiments of the present disclosure also provide corresponding apparatuses for implementing the above-described methods or processes. FIG4 shows a schematic block diagram of an example apparatus 400 for audio processing according to certain embodiments of the present disclosure. Apparatus 400 may be implemented as or included in terminal device 110. Each module / component in apparatus 400 may be implemented by hardware, software, firmware, or any combination thereof.
[0057] As shown in FIG4 , the apparatus 400 includes an acquisition module 410 configured to acquire first media content input by a user, where the first media content includes first audio content corresponding to singing content.
[0058] The apparatus 400 further includes a providing module 420 configured to provide second media content based on the user's selection of the target timbre, where the second media content includes second audio content corresponding to the singing content, and the second audio content corresponds to the selected target timbre.
[0059] In some embodiments, the second audio content at least retains the target audio attribute of the first audio content, where the target audio attribute includes at least one of the following: pitch and rhythm.
[0060] In some embodiments, the providing module 420 also includes a selection module, which is configured to display a selection panel, wherein the selection panel provides at least a first group of candidate effects for processing spoken content and a second group of candidate effects for processing singing content; and receive a user's selection of a target effect from the second group of candidate effects, the target effect corresponding to a target timbre.
[0061] In some embodiments, the acquisition module 410 is further configured to acquire the first media content recorded by the user; and acquire the first media content uploaded by the user.
[0062] In some embodiments, the providing module 420 also includes a generation module, which is configured to extract first audio content corresponding to the singing content from the first media content; and process the first audio content using a target model to generate second audio content, wherein the target model is trained based on sample data corresponding to the target timbre.
[0063] In some embodiments, the generation module is further configured to extract background audio content from the first media content, where the background audio content is different from the first audio content; and generate second media content by fusing the second audio content with the background audio content.
[0064] In some embodiments, the background audio content corresponds to accompaniment content.
[0065] In some embodiments, the generation module is further configured to fuse the second audio content and the background audio content to obtain intermediate audio content; adjust the reverberation effect or volume of the intermediate audio content; and generate the second media content based on the adjusted intermediate audio content.
[0066] In some embodiments, the providing module 420 further includes an adjusting module, which is configured to determine a reverberation parameter based on the first media content; and adjust the reverberation effect of the intermediate audio content based on the reverberation parameter.
[0067] In some embodiments, the adjustment module is further configured to perform global volume equalization on the intermediate audio content.
[0068] FIG5 shows a block diagram of an electronic device 500 in which one or more embodiments of the present disclosure may be implemented. It should be understood that the electronic device 500 shown in FIG5 is merely exemplary and should not be construed as limiting the functionality and scope of the embodiments described herein. The electronic device 500 shown in FIG5 can be used to implement the electronic device 110 of FIG1 .
[0069] As shown in FIG5 , electronic device 500 is a general-purpose electronic device. Components of electronic device 500 may include, but are not limited to, one or more processors or processing units 510, memory 520, storage device 530, one or more communication units 540, one or more input devices 550, and one or more output devices 560. Processing unit 510 may be a real or virtual processor and is capable of performing various processes according to programs stored in memory 520. In a multi-processor system, multiple processing units execute computer-executable instructions in parallel to enhance the parallel processing capabilities of electronic device 500.
[0070] The electronic device 500 typically includes a plurality of computer storage media. Such media can be any accessible media that can be obtained by the electronic device 500, including but not limited to volatile and non-volatile media, removable and non-removable media. The memory 520 can be a volatile memory (e.g., registers, cache, random access memory (RAM)), a non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. The storage device 530 can be a removable or non-removable medium and can include a machine-readable medium, such as a flash drive, a disk, or any other medium that can be used to store information and / or data and can be accessed within the electronic device 500.
[0071] The electronic device 500 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in FIG. 5 , a disk drive for reading from or writing to a removable, non-volatile disk (e.g., a “floppy disk”) and an optical drive for reading from or writing to a removable, non-volatile optical disk may be provided. In these cases, each drive may be connected to a bus (not shown) by one or more data media interfaces. The memory 520 may include a computer program product 525 having one or more program modules configured to perform various methods or actions of various embodiments of the present disclosure.
[0072] The communication unit 540 enables communication with other electronic devices via a communication medium. Additionally, the functions of the components of the electronic device 500 can be implemented in a single computing cluster or multiple computing machines that can communicate via a communication connection. Thus, the electronic device 500 can operate in a networked environment using a logical connection with one or more other servers, a network personal computer (PC), or another network node.
[0073] Input device 550 may be one or more input devices, such as a mouse, keyboard, or trackball. Output device 560 may be one or more output devices, such as a display, a speaker, or a printer. Electronic device 500 may also communicate with one or more external devices (not shown) via communication unit 540 as needed, such as a storage device, a display device, or the like, with one or more devices that allow a user to interact with electronic device 500, or with any device that allows electronic device 500 to communicate with one or more other electronic devices (e.g., a network card, a modem, etc.). Such communication may be performed via an input / output (I / O) interface (not shown).
[0074] According to an exemplary implementation of the present disclosure, a computer-readable storage medium is provided, on which computer-executable instructions are stored, wherein the computer-executable instructions are executed by a processor to implement the method described above. According to an exemplary implementation of the present disclosure, a computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, and the computer-executable instructions are executed by a processor to implement the method described above.
[0075] Various aspects of the present disclosure are described herein with reference to flowcharts and / or block diagrams of methods, apparatuses, devices, and computer program products implemented according to the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.
[0076] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine, such that when these instructions are executed by the processing unit of the computer or other programmable data processing device, a device is generated that implements the functions / actions specified in one or more blocks in the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium, where these instructions cause the computer, programmable data processing device, and / or other device to operate in a specific manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing various aspects of the functions / actions specified in one or more blocks in the flowchart and / or block diagram.
[0077] Computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device so that a series of operational steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to implement the functions / actions specified in one or more boxes in the flowchart and / or block diagram.
[0078] The flow charts and block diagrams in the accompanying drawings show the possible architecture, functions and operations of the systems, methods and computer program products according to multiple implementations of the present disclosure. In this regard, each box in the flow chart or block diagram can represent a part for a module, program segment or instruction, and a part for a module, program segment or instruction comprises one or more executable instructions for realizing the logical function of the specification. In some alternative implementations, the functions marked in the box can also occur in a sequence different from that marked in the accompanying drawings. For example, two continuous boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be realized by a special hardware-based system that performs the function or action of the specification, or can be realized by a combination of special hardware and computer instructions.
[0079] While various implementations of the present disclosure have been described above, the foregoing description is intended to be illustrative, not exhaustive, and not limited to the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is selected to best explain the principles of the implementations, their practical applications, or improvements to existing technologies, or to enable others skilled in the art to understand the various implementations disclosed herein.
Claims
1. An audio processing method, comprising: Obtaining first media content input by a user, the first media content including first audio content corresponding to singing content; And Based on the user's selection of a target timbre, providing second media content, the second media content including second audio content corresponding to the singing content, the second audio content corresponding to the selected target timbre.
2. The method according to claim 1, wherein The second audio content at least retains target audio attributes of the first audio content, and the target audio attributes include at least one of the following: pitch, rhythm.
3. The method according to claim 1 or 2, further comprising: Displaying a selection panel, wherein the selection panel at least provides a first group of candidate effects for processing speech content and a second group of candidate effects for processing singing content; And Receiving the user's selection of a target effect in the second group of candidate effects, the target effect corresponding to the target timbre.
4. The method according to claim 1 or 2, wherein obtaining the first media content input by the user includes at least one of the following: Obtaining the first media content recorded by the user, or Obtaining the first media content uploaded by the user.
5. The method according to claim 1 or 2, wherein the second audio content is generated through the following process: Extracting the first audio content corresponding to the singing content from the first media content; and Processing the first audio content by using a target model to generate the second audio content, wherein the target model is trained based on sample data corresponding to the target timbre.
6. The method according to claim 5, wherein the second media content is generated through the following process: Extracting background audio content from the first media content, the background audio content being different from the first audio content; and Generating the second media content by fusing the second audio content and the background audio content.
7. The method according to claim 6, wherein the background audio content corresponds to accompaniment content.
8. The method according to claim 6, wherein generating the second media content by fusing the second audio content and the background audio content includes: Fusing the second audio content and the background audio content to obtain intermediate audio content; Adjusting the reverberation effect or volume of the intermediate audio content; And Generating the second media content based on the adjusted intermediate audio content.
9. The method according to claim 8, wherein adjusting the reverberation effect of the intermediate audio content includes: Determining reverberation parameters based on the first media content; and Adjusting the reverberation effect of the intermediate audio content based on the reverberation parameters.
10. The method according to claim 8, wherein adjusting the volume of the intermediate audio content includes: Performing global volume equalization on the intermediate audio content.
11. An apparatus for audio processing, comprising: An obtaining module, configured to obtain first media content input by a user, the first media content including first audio content corresponding to singing content; And A providing module, configured to provide second media content based on the user's selection of a target timbre, where the second media content includes second audio content corresponding to the singing content, and the second audio content corresponds to the selected target timbre.
12. An electronic device, comprising: At least one processing unit; And At least one memory, the at least one memory being coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions, when executed by the at least one processing unit, causing the electronic device to execute the method according to any one of claims 1 to 10.
13. A computer-readable storage medium, having stored thereon a computer program, the computer program being executable by a processor to implement the method according to any one of claims 1 to 10.
Citation Information
Patent Citations
Sound conversion method and device, readable storage medium and electronic equipment
CN111798821A
Song multimedia synthesis method and device, electronic equipment and storage medium
CN112331234A
Audio production method and device, terminal, storage medium and program product
CN116229996A
Cognitive system and method to select best suited audio content based on individual's past reactions
US20190205469A1
Audio synthesis method, electronic device and readable storage medium
WO2023207472A1