Audio processing method and apparatus, device, and storage medium
Patent Information
- Application Number
- PCT/CN2026/078293
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-03-26
- Filing Date
- 2026-02-10
- Publication Date
- 2026-10-01
Smart Images

Figure CN2026078293_01102026_PF_FP_ABST
Abstract
Description
Audio processing methods, apparatus, devices and storage media
[0001] This application claims priority to Chinese Patent Application No. 202510368753.6, filed on March 26, 2025, entitled "Audio Processing Method, Apparatus, Device and Storage Medium", the entire contents of which are incorporated herein by reference. Technical Field
[0002] The exemplary embodiments disclosed herein generally relate to the field of computers, and particularly to audio processing methods, apparatus, devices, computer-readable storage media, and computer program products. Background Technology
[0003] Audio data is a common information carrier in daily life, work, and social interactions. People can disseminate information and content by generating and acquiring audio data. With the development of information technology, more and more applications are providing users with personalized, diversified, and differentiated audio processing functions to enrich their audiovisual experience. Summary of the Invention
[0004] In a first aspect of this disclosure, an audio processing method is provided. The method includes: in response to receiving a selection of multiple effects on a first interface, processing a set of source audios associated with the first interface based on the multiple effects respectively, generating multiple effect audios corresponding to the multiple effects, each effect audio containing a sound effect corresponding to the respective effect; and providing a target audio on the first interface by performing audio synthesis on the multiple effect audios, the target audio containing a target sound effect formed by mixing the sound effects corresponding to each of the multiple effects.
[0005] In a second aspect of this disclosure, an apparatus for audio processing is provided. The apparatus includes: a processing module configured to, in response to receiving a selection of multiple effects on a first interface, process a set of source audios associated with the first interface based on the multiple effects, generating multiple effect audios corresponding to the multiple effects, each effect audio containing a sound effect corresponding to the respective effect; and a synthesis module configured to, by performing audio synthesis on the multiple effect audios, provide a target audio on the first interface, the target audio containing a target sound effect formed by mixing the sound effects corresponding to each of the multiple effects.
[0006] In a third aspect of this disclosure, an electronic device is provided. The device includes at least one processor; and at least one memory coupled to the at least one processor and storing instructions for execution by the at least one processor. When executed by the at least one processor, the instructions cause the device to perform the method of the first aspect.
[0007] In a fourth aspect of this disclosure, a computer-readable storage medium is provided. The computer-readable storage medium stores computer-executable instructions that can be executed by a processor to implement the method of the first aspect.
[0008] In a fifth aspect of this disclosure, a computer program product is provided, including computer-executable instructions, wherein when executed by a processor, the computer-executable instructions implement the method according to a first aspect of this disclosure.
[0009] It should be understood that the content described in this content section is not intended to limit the key or essential features of the embodiments of this disclosure, nor is it intended to restrict the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0010] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. In the drawings, the same or similar reference numerals denote the same or similar elements, wherein:
[0011] Figure 1 shows a schematic diagram of an example environment in which embodiments of the present disclosure may be implemented;
[0012] Figure 2 shows a flowchart of an audio processing procedure according to some embodiments of the present disclosure;
[0013] Figure 3 illustrates a schematic diagram of an example architecture for audio processing according to some embodiments of the present disclosure;
[0014] Figure 4 shows a schematic diagram of an example architecture of an audio graph according to some embodiments of the present disclosure;
[0015] Figure 5 illustrates a schematic diagram of an example architecture for audio processing according to some embodiments of the present disclosure;
[0016] Figure 6 illustrates a schematic diagram of an example scene of target video generation according to some embodiments of the present disclosure;
[0017] Figure 7 shows a schematic structural block diagram of an example device for audio processing according to some embodiments of the present disclosure; and
[0018] Figure 8 shows a block diagram of an electronic device capable of implementing several embodiments of the present disclosure. Detailed Implementation
[0019] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.
[0020] In the description of embodiments of this disclosure, the term "comprising" and similar terms should be understood as open-ended inclusion, i.e., "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". Other explicit and implicit definitions may also be included below.
[0021] In this document, unless explicitly stated otherwise, performing a step in response to A does not mean that the step is performed immediately after A, but may include one or more intermediate steps.
[0022] It is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and related provisions.
[0023] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure through appropriate means in accordance with relevant laws and regulations, and user authorization should be obtained.
[0024] For example, in response to receiving a user's active request, a prompt message is sent to the user to clearly inform the user that the requested operation will require the acquisition and use of the user's personal information, thereby enabling the user to choose whether to provide personal information to the software or hardware such as electronic devices, applications, servers or storage media that perform the operation of the technical solution disclosed herein, based on the prompt message.
[0025] As an optional but non-restrictive implementation, in response to a user's active request, a prompt message can be sent to the user, such as a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0026] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.
[0027] As mentioned above, audio data is a common information carrier in daily life, work, and social interactions. People can disseminate information and content by generating and acquiring audio data. With the development of information technology, more and more applications are providing users with personalized, diversified, and differentiated audio processing functions to enrich the audiovisual experience. For example, some applications support adding special effects to audio or video, and using these effects to process audio can create unique sound effects. However, traditional technologies typically only support processing audio data using a single effect, and usually do not support adding multiple effects simultaneously. This approach, to some extent, limits the diversity and flexibility of audio processing functions.
[0028] In view of this, embodiments of this disclosure propose an improved audio processing scheme. In this scheme, if a selection of multiple effects is received on a first interface, a set of source audios associated with the first interface is processed based on each of the multiple effects to generate multiple effect audios corresponding to the multiple effects. Each effect audio contains a sound effect corresponding to the respective effect. By performing audio synthesis on the multiple effect audios, a target audio is provided on the first interface, the target audio containing a target sound effect formed by mixing the sound effects corresponding to each of the multiple effects.
[0029] In the embodiments of this disclosure, multiple special effects (also known as audio props) can be used to process audio. By mixing the sound effects corresponding to multiple special effects, a mixed special effect (i.e., the target sound effect) different from a single sound effect can be formed. Therefore, users can freely combine multiple special effects to create rich sound effects, which is beneficial for improving audio processing capabilities and enhancing the audio-based interactive experience.
[0030] The following section provides a detailed description of various example implementations of this scheme, with reference to the accompanying drawings.
[0031] Example Environment
[0032] Figure 1 illustrates a schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented. In this example environment 100, an application 120 is installed on a terminal device 110. A user 140 can interact with the application 120 via the terminal device 110 and / or an attached device of the terminal device 110. For example, the application 120 can capture the user 140's voice 145 via a voice capture device (e.g., a microphone) of the terminal device 110, and the application 120 can also capture images of the user 140 via an image capture device (e.g., a camera) of the terminal device 110. It should be understood that the collection of user-related data (e.g., voice, images, text, video, etc.) is performed after prior notification to the user and with the user's authorization.
[0033] In some embodiments of this disclosure, application 120 may be a content generation application, a content sharing application, or a social application, capable of providing user 140 with services associated with media content, including browsing, commenting, forwarding, creating (e.g., shooting and / or editing), publishing, live streaming, etc. "Media content" may include one or more types of content, such as video, images, GIFs, image sets, audio, text, etc. Application 120 may support user 140 in creating multimedia content. Such multimedia content may include image data, audio data, or video data.
[0034] In some embodiments of this disclosure, if application 120 is active, terminal device 110 can display the user interface 150 of application 120. The user interface 150 may include various interfaces provided by application 120, such as a content presentation interface, a content creation interface, a content publishing interface, a live streaming interface, etc. Application 120 can provide content browsing functionality to browse various types of content published within application 120. Application 120 can also provide live streaming functionality, allowing application 120 to push created content to terminal devices associated with the live stream, such as the terminal devices of participants or viewers, etc.
[0035] In some embodiments, application 120 can use special effects to edit media content. Specifically, application 120 can use processing strategies corresponding to special effects to process media content (e.g., audio, video, or images) to create the specific effect indicated by the special effect. For example, application 120 can use special effects for audio to process source audio (e.g., audio captured by a microphone) to create specific sound effects (e.g., spatial effects, voice alteration, background noise, etc.). In practical applications, special effects can also be referred to as props.
[0036] In some embodiments, terminal device 110 can communicate with server 130 to provide services to application 120. For example, as shown in FIG1, user 140 can select effects for editing audio. Terminal device 110 can send an editing request to server 130 in response to the selection of effects. Server 130 can edit the corresponding audio based on the selected effects and feed the edited audio back to terminal device 110. Terminal device 110 can play the audio through user interface 150, or terminal device 110 can also broadcast the edited audio live.
[0037] In some embodiments, terminal device 110 can be any type of mobile terminal, fixed terminal, or portable terminal, including mobile phones, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, media computers, multimedia tablets, personal communication system (PCS) devices, personal navigation devices, personal digital assistants (PDAs), audio / video players, digital cameras / camcorders, positioning devices, television receivers, radio receivers, e-book devices, gaming devices, or any combination thereof, including accessories and peripherals of these devices or any combination thereof. In some embodiments, terminal device 110 can also support any type of user-facing interface (such as "wearable" circuitry). Server 130 can be various types of computing systems / servers capable of providing computing power, including but not limited to mainframes, edge computing nodes, computing devices in cloud environments, etc. Server 130 can, for example, be implemented in a cloud environment.
[0038] It should be understood that the structure and function of the various elements in environment 100 are described for illustrative purposes only and do not imply any limitation on the scope of this disclosure.
[0039] Example process
[0040] The following description will continue with reference to the accompanying drawings, which will illustrate some exemplary embodiments of the present disclosure. Figure 2 shows a flowchart of an audio processing procedure 200 according to one embodiment of the present disclosure. In the following discussion, the execution of procedure 200 will be described from the perspective of server 130, but this is merely exemplary.
[0041] In box 210, server 130 may, in response to receiving a selection of multiple effects on the first interface, process a set of source audios associated with the first interface based on the multiple effects, generating multiple effect audios corresponding to the multiple effects. In some embodiments, the effects may indicate sound effects for audio, or the effects may indicate both sound effects for audio and visual effects for video frames. Since the embodiments of this disclosure primarily focus on audio processing, the following will mainly describe sound effects for audio, without elaborating on visual effects for video frames.
[0042] The multiple effects here can indicate different sound effects for audio. Each effect audio contains the sound effect corresponding to the specific effect. As an example, one of the multiple effects can indicate an ocean wave sound effect, and the effect audio generated based on this effect can contain the ocean wave sound effect. Another of the multiple effects can indicate a voice-changing effect, and the effect audio generated based on this other effect can convert a human voice into a corresponding voice-changing effect (such as a dialect). As another example, one of the multiple effects can indicate a spatial audio effect, and the corresponding effect audio generated based on this effect can contain the spatial audio effect. Another of the multiple effects can indicate a wind sound effect, and the effect audio generated based on this other effect can overlay the wind sound effect. Of course, the above effects and the sound effects indicated by the effects are merely exemplary. In actual application scenarios, application 120 can provide a variety of effects. The embodiments disclosed herein do not limit this.
[0043] In some embodiments, the first interface may be the user interface 150 of application 120, such as the recording interface, shooting interface, or live streaming interface of application 120, etc. As an example, in the recording interface, user 140 can select multiple effects for audio. As another example, in the shooting interface and live streaming interface, user 140 can select multiple effects for audio, or select effects for video (including effects for audio and effects for video frames).
[0044] In some embodiments, terminal device 110 may send a selection instruction for multiple effects to server 130 in response to user 140's selection of multiple effects on a first interface. Server 130 may determine the multiple effects selected by user 140 based on the selection instruction. Then, it may process a set of source audio associated with the first interface based on each of the multiple effects to generate multiple effect audios.
[0045] In some embodiments, the source audio may include audio captured using an audio acquisition device. As an example, FIG3 illustrates a schematic diagram of an example architecture 300 for audio processing according to some embodiments of the present disclosure. As shown in FIG3, in a video shooting or live streaming scenario, user 140 may choose to turn on microphone 302. Microphone 302 may be the microphone of terminal device 110 itself, a microphone attached to terminal device 110, or another microphone. Terminal device 110 may use the audio captured by microphone 302 as source audio 304, and terminal device 110 may transmit source audio 302 to server 130 to request server 130 to perform special effects processing on source audio 304 based on a plurality of selected special effects.
[0046] Alternatively or additionally, the source audio may also include audio contained within multiple effects. Specifically, some effects may indicate the overlay of specific sounds (e.g., ocean waves, wind, rocket launch sounds, etc.). In this case, the audio contained within the effect can be used as the source audio. As an example, as shown in Figure 3, in box 306, the effect selected by the user can be determined. If the effect contains audio, the audio contained within the effect can be used as the source audio 308. The source audio 308 can be obtained by server 130 from associated storage space or received by server 130 from terminal device 110.
[0047] Alternatively or additionally, the source audio may also include audio selected according to a selection instruction. The selected audio may be audio selected locally from the terminal device 110 or audio selected from an audio database. As an example, the terminal device 110 may select audio locally as the source audio according to a selection instruction from the user 140. The terminal device 110 may then send this source audio to the server 130.
[0048] As another example, as shown in Figure 3, user 140 can also select source audio 312 from audio database 310 via terminal device 110 or an attached device of terminal device 110. Terminal device 110 can send a selection instruction for source audio 312 to server 130, and server 130 can retrieve source audio 312 from audio database 310 according to the selection instruction. It should be noted that the above-described source audio is merely exemplary, and source audio can also include any other suitable audio, such as audio obtained from the network, or audio generated using a machine learning model, etc. The embodiments of this disclosure do not limit this.
[0049] In some embodiments, the set of source audios associated with the first interface may include multiple source audios. For a first effect among multiple effects, server 130 may process the multiple source audios based on the first effect to generate effect audio corresponding to the first effect. The first effect may be any one of the multiple effects selected by user 140.
[0050] As an example, as shown in Figure 3, a set of source audios associated with the first interface may include source audio 304, source audio 308, and source audio 312. Source audio 312 may be, for example, background music of a song selected by user 140 from audio database 310. Source audio 308 may be audio contained in at least one of a plurality of effects selected by user 140. Source audio 312 may be, for example, audio of user 140 singing a song, captured by microphone 302 associated with terminal device 110. Server 130 may process source audio 304, source audio 308, and source audio 312 based on a first effect to generate effect audio corresponding to the first effect. Of course, a set of source audios associated with the first interface may also include only one source audio; for example, the set of source audios may only include source audio 304 captured by microphone 302, or only include source audio 312 obtained from audio database 310.
[0051] In some embodiments, server 130 may deploy multiple audio graphs corresponding to multiple effects, each audio graph indicating an audio processing strategy for the corresponding effect. If a selection of multiple effects is received on the first interface, server 130 can determine multiple audio graphs corresponding to the multiple effects based on the selection. Then, server 130 can distribute this set of source audio to the multiple audio graphs corresponding to the multiple effects. Each set of source audio is processed using its respective audio graph to generate multiple effect audios corresponding to the multiple effects.
[0052] As an example, as shown in Figure 3, this set of source audios may include source audio 304, source audio 308, and source audio 312. Server 130 determines multiple audio graphs corresponding to the selected effects, including audio graph 314, audio graph 316, and audio graph 318. Server 130 can send source audios 304, 308, and 312 to audio graphs 314, 316, and 318 respectively. Source audios 304, 308, and 312 are processed using audio graph 314 to generate effect audio 320. Source audios 304, 308, and 312 are processed using audio graph 316 to generate effect audio 322. Source audios 304, 308, and 312 are processed using audio graph 318 to generate effect audio 324. It should be noted that the number of audio graphs shown in Figure 3 is only exemplary; in actual application scenarios, user 140 may select two effects, or four or more effects. The embodiments disclosed herein do not limit the number of effects selected by user 140, nor do they limit the number of multiple audio graphs corresponding to multiple effects.
[0053] In some examples, an audio graph can contain multiple processing nodes, which are connected to form an audio processing architecture. Each processing node performs a specific processing operation on the audio flowing through it. These processing nodes can include, but are not limited to, input nodes, output nodes, detection nodes, sound effect processing nodes, and so on. Input nodes receive the source audio input, and output nodes output the generated sound effect audio. Detection nodes can be configured to detect audio parameters, and sound effect processing nodes can be configured to modify the audio characteristics, such as filtering, sound effect processing, etc.
[0054] In some examples, an audio graph may include multiple input nodes, which can be configured to input different types of source audio into the audio graph. Server 130 can provide each source audio from this set of source audios to a corresponding input node in the audio graph. As an example, FIG4 shows a schematic diagram of an example structure 400 of an audio graph according to some embodiments of the present disclosure. As shown in FIG4, nodes 412A, 414A, 416A in audio graph 314 and nodes 412B, 414B, 416B in audio graph 316 can be input nodes. Nodes 412A and 412B can be configured to receive source audio from an audio database. Nodes 414A and 414B can be configured to receive source audio from effects. Nodes 416A and 416B can be configured to receive source audio captured by a microphone.
[0055] Nodes 418A, 418B, 424A, and 424B can be detection nodes, configured to detect at least one corresponding audio parameter, such as volume, rhythm, or speech. Nodes 422A and 422B can be audio effect processing nodes, used to perform audio effect processing on the corresponding source audio to change one or more audio characteristics, such as timbre or pitch. Nodes 420A and 420B can be output nodes, configured to output special effects audio for playback. Nodes 426A and 426B can also be output nodes, configured to output special effects audio for video synthesis. Nodes 428A and 428B can be termination nodes of the audio graph, used to output the generated special effects audio to the outside of the audio graph.
[0056] Server 130 can distribute source audio 312 to nodes 412A and 412B, source audio 308 to nodes 414A and 414B, and source audio 304 to nodes 416A and 416B. Audio graphs 314 and 316 utilize their respective multiple processing nodes to process source audio 304, 308, and 312, generating special effects audio 320 and special effects audio 322. It should be noted that the internal structure of the above audio graphs is merely exemplary. When the special effects and indicated audio processing strategies of the audio graphs differ, the processing nodes included in the audio graphs may also differ. The embodiments of this disclosure do not limit the internal structure of the audio graphs.
[0057] In some embodiments, if a selection of multiple effects is received on the first interface, the server 130 can distribute multiple selection instructions corresponding to the multiple effects to the audio rendering engine. The audio rendering engine then distributes a set of source audio to multiple audio graphs corresponding to the multiple selection instructions. As an example, FIG5 shows a schematic diagram of an example architecture 500 for audio processing according to some embodiments of the present disclosure. As shown in FIG5, in response to a selection of multiple effects, the effects view rendering engine 506 can distribute multiple selection instructions to the audio rendering engine 508. The audio rendering engine 508 can receive a set of source audio from the data link layer (not shown). The audio rendering engine 508 can then distribute this set of source audio to audio graphs 314 and 316 based on the multiple selection instructions. The audio graphs 314 and 316 respectively process this set of source audio to generate effects audio 320 and effects audio 322.
[0058] In some embodiments, for a second and a third special effect among multiple special effects, server 130 can determine at least one audio parameter of a set of source audio based on the second special effect. Then, server 130 can process the set of source audio based on the third special effect and at least one audio parameter to generate special effect audio corresponding to the third special effect. This enables interactive control and coordination between special effects, significantly enriching their functional characteristics. Here, the second and third special effects can be any two of the multiple special effects.
[0059] As an example, as shown in Figure 5, in response to the selection of a second and a third special effect, the special effect view rendering engine 506 can render a special effect view 502 corresponding to the second special effect and a special effect view 504 corresponding to the third special effect. Special effect views 502 and 504 can be displayed through the user interface 150. The audio rendering engine 508 can, in response to the selection of the second special effect, invoke audio graph 314 to detect at least one audio parameter (e.g., rhythm) of a set of source audio. The audio rendering engine 508 can feed back the detection result to the special effect view rendering engine 506, which can then feed back the detection result to the special effect view 502. The special effect view 502 can provide the detection result to the special effect view 504, which can then generate a special effect audio corresponding to the third special effect based on the detection result, such as creating a sound effect that matches the rhythm of the source audio.
[0060] In the return process 200, at box 220, server 130 can perform audio synthesis on multiple special effects audios and provide the target audio on the first interface. The target audio contains the target sound effect formed by mixing the sound effects corresponding to each of the multiple special effects. As an example, referring to Figure 3, at box 326, server 130 can perform audio synthesis on special effects audios 320, 322, and 324 to generate target audio 328. Assume that special effects audio 320 is superimposed with the sound of waves on the source audio 304 captured by microphone 302, special effects audio 322 is superimposed with the sound of wind on the source audio 304 captured by microphone 302, and special effects audio 324 is superimposed with the sound of seabirds on the source audio 304 captured by microphone 302. By performing audio synthesis on special effects audios 320, 322, and 324, the resulting target audio can include the sound of waves, the sound of seabirds, and the sound of wind. It should be understood that the sound effects in the above examples are merely illustrative. In real-world applications, any other appropriate effects can be selected to create other sound effects as needed. The embodiments disclosed herein do not limit this.
[0061] In some examples, as shown in Figure 5, server 130 can use audio rendering engine 508 to perform audio synthesis on multiple special effects audios. Of course, server 130 can also use any other suitable functional layer, such as an audio output layer, to perform audio synthesis on multiple special effects audios.
[0062] In some embodiments, server 130 can perform time synchronization processing on multiple special effects audio files. Then, by performing audio synthesis on the time-synchronized audio files, the target audio is provided on a first interface. This avoids problems such as overlapping or misaligned sounds caused by time asynchrony between multiple audio files, thus improving the quality of audio synthesis.
[0063] In some embodiments, after acquiring the target audio, the server 130 can play the target audio on a first interface. Specifically, the server 130 can transmit the target audio to the terminal device 110, and the terminal device 110 can play the target audio on the first interface.
[0064] Alternatively or additionally, server 130 may also encode the target audio to obtain an audio stream for push to terminal devices associated with the live stream. Server 130 can then push the audio stream to terminal device 110 and other terminal devices associated with the live stream to achieve the purpose of live audio streaming. Of course, server 130 can also push the audio stream to terminal device 110, which can then push the audio stream to other terminal devices associated with the live stream. For example, terminal device 110 could be a terminal device associated with the broadcaster, and terminal device 110 could push the audio stream to terminal devices associated with guest participants or terminal devices associated with viewers.
[0065] Alternatively or additionally, server 130 can generate a target video based on target audio and multiple video frames corresponding to the target audio. As an example, terminal device 110 can capture source audio using an associated microphone, and terminal device 110 can also capture multiple video frames using an associated camera. Terminal device 110 can, in response to the selection of effects, pass the source audio, multiple video frames, and selection instructions for multiple effects to server 130. Server 130 can process the source audio based on the multiple selection instructions to generate target audio. Server 130 can then synthesize the target audio and multiple video frames into a target video. Afterwards, server 130 can feed the target video back to terminal device 110, or server 130 can push the target video to terminal devices associated with live streaming to achieve the purpose of live video streaming.
[0066] In some embodiments, server 130 can perform time synchronization processing on target audio and multiple video frames along a predetermined timeline. Then, server 130 can generate a target video based on the time-synchronized target audio and multiple video frames. This avoids problems such as lip-syncing issues between display and audio, and improves the quality of video synthesis.
[0067] As an example, Figure 6 illustrates a schematic diagram of an example scenario 600 for target video generation according to some embodiments of the present disclosure. As shown in Figure 6, example scenario 600 illustrates a frame timeline, a system timeline, and an audio timeline. The server 130 can perform time synchronization on video frames 602 on the frame timeline and audio frames 606 on the audio timeline along the system timeline. The video frames 602 and audio frames 606 are synchronized to time node 604 on the system timeline to achieve time synchronization between video frames 602 and audio frames 606. Of course, the predetermined timeline here is not limited to the system timeline; it can be any one of the audio timeline or the frame timeline, or it can be other timelines. Embodiments of the present disclosure do not limit this.
[0068] It should be noted that although the above example describes the execution of process 200 from the perspective of server 130, it should not be construed as the improved solutions of the embodiments of this disclosure being limited to execution by server 130. In practical application scenarios, the improved solutions of the embodiments of this disclosure can be implemented locally on terminal device 110, or they can be implemented through interaction and cooperation between terminal device 110 and server 130. That is, the above process 200 can also be executed by terminal device 110, or it can be executed collaboratively by terminal device 110 and server 130.
[0069] In this way, in the embodiments of this disclosure, multiple special effects (also known as props) can be used to process audio. By mixing the sound effects corresponding to multiple special effects, a mixed special effect (i.e., the target sound effect) different from a single sound effect can be formed. Therefore, users can freely use multiple special effects to combine rich sound effects, which is beneficial to improving audio processing capabilities and enhancing the audio-based interactive experience.
[0070] Example devices and equipment
[0071] Embodiments of this disclosure also provide corresponding apparatus for implementing the methods or processes described above. Figure 7 shows a schematic structural block diagram of an example apparatus 700 for audio processing according to certain embodiments of this disclosure. Apparatus 700 may be implemented as or included in server 130. The various modules / components in apparatus 700 may be implemented by hardware, software, firmware, or any combination thereof.
[0072] As shown in Figure 7, the device 700 includes: a processing module 710 configured to, in response to receiving a selection of multiple effects on a first interface, process a set of source audios associated with the first interface based on the multiple effects, and generate multiple effect audios corresponding to the multiple effects, each effect audio containing a sound effect corresponding to the corresponding effect; and a synthesis module 720 configured to provide a target audio on the first interface by performing audio synthesis on the multiple effect audios, the target audio containing a target sound effect formed by mixing the sound effects corresponding to each of the multiple effects.
[0073] In some embodiments, the set of source audios includes multiple source audios, and the processing module 710 is further configured to: process the multiple source audios based on the first effect for a first effect among the multiple effects, and generate an effect audio corresponding to the first effect.
[0074] In some embodiments, the set of source audio includes at least one of the following: audio captured using an audio capture device, audio selected according to a selection instruction, or audio included in a corresponding effect.
[0075] In some embodiments, the processing module 710 is further configured to: in response to receiving a selection of the plurality of effects on the first interface, distribute the set of source audio to a plurality of audio graphs corresponding to the plurality of effects, each audio graph indicating an audio processing strategy for the corresponding effect; and process the respective set of source audio using the plurality of audio graphs to generate the plurality of effect audio corresponding to the plurality of effects.
[0076] In some embodiments, the processing module 710 is further configured to: in response to receiving a selection of the plurality of effects on the first interface, distribute a plurality of selection instructions corresponding to the plurality of effects to the audio rendering engine; and utilize the audio rendering engine to distribute the set of source audio to the plurality of audio graphs corresponding to the plurality of selection instructions.
[0077] In some embodiments, the processing module 710 is further configured to: for the second and third special effects among the plurality of special effects, determine at least one audio parameter of the set of source audio based on the second special effect; and process the set of source audio based on the third special effect and the at least one audio parameter to generate special effect audio corresponding to the third special effect.
[0078] In some embodiments, the synthesis module 720 is further configured to: perform time synchronization processing on the plurality of special effects audios; and provide the target audio on the first interface by performing audio synthesis on the plurality of special effects audios after time synchronization processing.
[0079] In some embodiments, the device 700 further includes at least one of the following: a playback module configured to play the target audio on the first interface; an encoding module configured to perform encoding on the target audio to obtain an audio stream for push to a terminal device associated with the live stream; or a generation module configured to generate a target video based on the target audio and a plurality of video frames corresponding to the target audio.
[0080] In some embodiments, the generation module is further configured to: perform time synchronization processing on the target audio and the plurality of video frames along a predetermined time axis; and generate the target video based on the time-synchronized target audio and the plurality of video frames.
[0081] The units and / or modules included in device 700 can be implemented in various ways, including software, hardware, firmware, or any combination thereof. In some embodiments, one or more units and / or modules can be implemented using software and / or firmware, such as machine-executable instructions stored on a storage medium. In addition to or as an alternative to machine-executable instructions, some or all of the units and / or modules in device 500 can be implemented at least partially by one or more hardware logic components. By way of example and not limitation, exemplary types of hardware logic components that can be used include field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), complex programmable logic devices (CPLDs), and so on.
[0082] Figure 8 shows a block diagram of an electronic device 800 in which one or more embodiments of the present disclosure may be implemented. It should be understood that the electronic device 800 shown in Figure 8 is merely exemplary and should not constitute any limitation on the functionality and scope of the embodiments described herein. The electronic device 800 shown in Figure 8 may include or be implemented as the server 130 in Figure 1 or the device 700 in Figure 7.
[0083] As shown in Figure 8, the electronic device 800 is in the form of a general-purpose electronic device. Components of the electronic device 800 may include, but are not limited to, one or more processors or processing units 810, memory 820, storage device 830, one or more communication units 840, one or more input devices 850, and one or more output devices 860. The processor or processing unit 810 may be a physical or virtual processor and is capable of performing various processes according to executable instructions stored in memory 820. In a multiprocessor system, multiple processors execute computer-executable instructions in parallel to improve the parallel processing capability of the electronic device 800.
[0084] Electronic device 800 typically includes multiple computer storage media. Such media can be any accessible media that is accessible to electronic device 800, including but not limited to volatile and non-volatile media, removable and non-removable media. Memory 820 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. Storage device 830 can be removable or non-removable media and can include machine-readable media, such as flash drives, disks, or any other media that can be used to store information and / or data and can be accessed within electronic device 800.
[0085] Electronic device 800 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in FIG8, disk drives for reading or writing from removable, non-volatile disks (e.g., "floppy disks") and optical disk drives for reading or writing from removable, non-volatile optical disks may be provided. In these cases, each drive may be connected to a bus (not shown) via one or more data media interfaces. Memory 820 may include computer program product 825 having one or more executable instruction modules configured to perform various methods or actions of various embodiments of the present disclosure.
[0086] The communication unit 840 enables communication with other electronic devices via a communication medium. Additionally, the functionality of the components of the electronic device 800 can be implemented using a single computing cluster or multiple computing machines capable of communicating via communication connections. Therefore, the electronic device 800 can operate in a networked environment using logical connections to one or more other servers, network personal computers (PCs), or another network node.
[0087] Input device 850 can be one or more input devices, such as a mouse, keyboard, trackball, etc. Output device 860 can be one or more output devices, such as a monitor, speaker, printer, etc. Electronic device 800 can also communicate with one or more external devices (not shown) via communication unit 840 as needed. These external devices include storage devices, display devices, etc., and can communicate with one or more devices that enable user interaction with electronic device 800, or with any device that enables electronic device 800 to communicate with one or more other electronic devices (e.g., network card, modem, etc.). Such communication can be performed via input / output (I / O) interface (not shown).
[0088] According to an exemplary implementation of this disclosure, a computer-readable storage medium is provided that stores computer-executable instructions thereon, wherein the computer-executable instructions are executed by a processor to implement the methods described above. According to an exemplary implementation of this disclosure, a computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, which are executed by a processor to implement the methods described above.
[0089] Various aspects of this disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatuses, devices, and computer program products implemented according to this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable and executable instructions.
[0090] These computer-executable instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processor of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-executable instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner; thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0091] Computer-executable instructions can be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions that execute on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.
[0092] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, executable instruction, or portion of instructions, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than those marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.
[0093] Various implementations of this disclosure have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is chosen to best explain the principles, practical applications, or improvements to technology in the market, or to enable others skilled in the art to understand the various implementations disclosed herein.
Claims
1. An audio processing method, comprising: In response to receiving a selection of multiple effects on the first interface, the system processes a set of source audio associated with the first interface based on the multiple effects, generating multiple effect audios corresponding to the multiple effects, each effect audio containing a sound effect corresponding to the specific effect; and By performing audio synthesis on the multiple special effects audios, a target audio is provided on the first interface. The target audio contains a target sound effect formed by mixing the sound effects corresponding to the multiple special effects.
2. The method according to claim 1, wherein the set of source audio includes multiple source audios, and generating the multiple effect audios includes: Regarding the first of the multiple special effects, Based on the first special effect processing, the multiple source audios are processed to generate special effect audio corresponding to the first special effect.
3. The method according to claim 1 or 2, wherein the set of source audio includes at least one of the following: Audio collected using audio acquisition equipment. The audio selected according to the selection command, or The audio included in the corresponding special effects.
4. The method according to claim 1, wherein generating the plurality of special effects audio includes: In response to receiving a selection of the plurality of effects on the first interface, the set of source audio is distributed to a plurality of audio graphs corresponding to the plurality of effects, each audio graph indicating the audio processing strategy of the corresponding effect; as well as The multiple audio graphs are used to process their respective sets of source audio to generate the multiple special effects audio corresponding to the multiple special effects.
5. The method of claim 4, wherein distributing the set of source audio to the plurality of audio graphs comprises: In response to receiving a selection of the plurality of effects on the first interface, the plurality of selection instructions corresponding to the plurality of effects are distributed to the audio rendering engine; as well as The audio rendering engine is used to distribute the set of source audio to the multiple audio graphs corresponding to the multiple selection instructions.
6. The method according to claim 1, wherein generating the plurality of special effects audio includes: Regarding the second and third special effects among the aforementioned special effects Based on the second special effect, at least one audio parameter of the set of source audio is determined; as well as Based on the third special effect and the at least one audio parameter, the set of source audio is processed to generate special effect audio corresponding to the third special effect.
7. The method according to any one of claims 1-6, wherein performing audio synthesis on the plurality of special effects audios comprises: The multiple special effects audios are processed for time synchronization. as well as The target audio is provided on the first interface by performing audio synthesis on multiple time-synchronized audio effects.
8. The method according to any one of claims 1-7, further comprising at least one of the following: Play the target audio on the first interface. Encoding is performed on the target audio to obtain an audio stream for push to terminal devices associated with the live stream, or A target video is generated based on the target audio and multiple video frames corresponding to the target audio.
9. The method according to any one of claims 1-8, wherein generating the target video comprises: Perform time synchronization processing on the target audio and the plurality of video frames along a predetermined time axis; as well as The target video is generated based on the target audio after time synchronization processing and the multiple video frames.
10. An apparatus for audio processing, comprising: The processing module is configured to, in response to receiving a selection of multiple effects on a first interface, process a set of source audios associated with the first interface based on the multiple effects, and generate multiple effect audios corresponding to the multiple effects, each effect audio containing a sound effect corresponding to the specific effect; and The synthesis module is configured to perform audio synthesis on the plurality of special effects audios and provide a target audio on the first interface, wherein the target audio includes a target sound effect formed by mixing the sound effects corresponding to the plurality of special effects.
11. An electronic device, comprising: At least one processor; as well as At least one memory coupled to the at least one processor and storing instructions for execution by the at least one processor, the instructions causing the electronic device to perform the method according to any one of claims 1 to 9 when executed by the at least one processor.
12. A computer-readable storage medium having stored thereon computer-executable instructions that can be executed by a processor to implement the method according to any one of claims 1 to 9.
13. A computer program product comprising computer-executable instructions, wherein the computer-executable instructions, when executed by a processor, implement the method according to any one of claims 1 to 9.