Audio processing method and device, equipment and storage medium

By obtaining the audio content input by the user and generating the second media content based on the target style, the problem that traditional audio processing methods cannot change the tone is solved, and the voice change effect and sound and picture synchronization are improved on the basis of retaining the tone.

CN120375841APending Publication Date: 2025-07-25BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202410096064.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-23
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

Traditional audio processing methods can only change the speaker's tone, and cannot achieve voice change effect on the basis of retaining the tone.

Method used

By obtaining the first media content input by the user, the second media content is generated based on the user's selection of the target style, so that it has the same tone as the first audio content, and has audio attributes corresponding to the target style, and synchronous processing of audio and video content is performed in combination with the target model.

Benefits of technology

It realizes that while retaining the tone, it improves the sound change effect, enhances the sound and picture synchronization, and provides low-cost sound change solutions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120375841A_ABST
    Figure CN120375841A_ABST
Patent Text Reader

Abstract

The embodiment of the invention relates to an audio processing method and device, equipment and a storage medium. The method provided by the invention comprises the following steps: acquiring first media content input by a user, wherein the first media content comprises first audio content; and providing a second media content based on the user's selection of the target style, the second media content including a second audio content generated based on the first audio content, the second audio content having the same tone as the first audio content, and the second audio content having at least one audio attribute corresponding to the target style. Through the mode, the embodiment of the invention can improve the voice changing effect on the basis of keeping the tone quality.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Example embodiments of the present disclosure generally relate to the field of computers, and particularly to methods, apparatuses, devices, and computer-readable storage media for audio processing. Background Art

[0002] With the development of computer technology, the Internet has become an important platform for people's information interaction. In the process of people's information interaction through the Internet, various types of audio have become an important medium for people's social expression and information exchange. Therefore, there is an expectation for a voice conversion technology that can change the speaking style while preserving the timbre by processing audio. Summary of the Invention

[0003] In a first aspect of the present disclosure, a method for audio processing is provided. The method includes: obtaining first media content input by a user, the first media content including first audio content; and providing second media content based on a user's selection of a target style, the second media content including second audio content generated based on the first audio content, the second audio content having the same timbre as the first audio content, and the second audio content having at least one audio attribute corresponding to the target style.

[0004] In a second aspect of the present disclosure, an apparatus for audio processing is provided. The apparatus includes: an obtaining module configured to obtain first media content input by a user, the first media content including first audio content; and a providing module configured to provide second media content based on a user's selection of a target style, the second media content including second audio content generated based on the first audio content, the second audio content having the same timbre as the first audio content, and the second audio content having at least one audio attribute corresponding to the target style.

[0005] In a third aspect of the present disclosure, an electronic device is provided. The device includes at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit. The instructions, when executed by the at least one processing unit, cause the device to execute the method according to any one of the first aspect to the fourth aspect.

[0006] In a fourth aspect of the present disclosure, a computer-readable storage medium is provided. A computer program is stored on the computer-readable storage medium and can be executed by a processor to implement the method according to any one of the first aspect to the fourth aspect.

[0007] It should be understood that the content described in this content part is not intended to limit the key features or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0008] In conjunction with the accompanying drawings and with reference to the following detailed description, the above and other features, advantages, and aspects of the various embodiments of the present disclosure will become more apparent. In the drawings, the same or similar reference numerals denote the same or similar elements, where:

[0009] Figure 1 FIG. shows a schematic diagram of an example environment in which embodiments of the present disclosure can be implemented;

[0010] Figure 2 FIG. shows a flowchart of an example process for audio processing according to some embodiments of the present disclosure;

[0011] Figures 3A to 3C FIG. shows a schematic diagram of an example interface according to some embodiments of the present disclosure;

[0012] Figure 4 FIG. shows a schematic structural block diagram of an example apparatus for audio processing according to some embodiments of the present disclosure; and

[0013] Figure 5 FIG. shows a block diagram of an electronic device capable of implementing multiple embodiments of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0014] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.

[0015] It should be noted that the titles of any sections / subsections provided herein are not restrictive. Various embodiments are described throughout this document, and any type of embodiment can be included under any section / subsection. Additionally, the embodiments described in any section / subsection can be combined with any other embodiments described in the same section / subsection and / or different sections / subsections in any manner.

[0016] In the description of the embodiments of the present disclosure, the term "including" and its similar terms should be understood as open inclusion, that is, "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". There may also be other explicit and implicit definitions hereinafter. The terms "first", "second", etc. may refer to different or the same objects. There may also be other explicit and implicit definitions hereinafter.

[0017] In the embodiments of the present disclosure, it may involve the user's data, data acquisition and / or use, etc. These aspects all comply with the corresponding laws, regulations and related provisions. In the embodiments of the present disclosure, all data collection, acquisition, processing, processing, forwarding, use, etc. are carried out on the premise that the user is aware of and confirms. Accordingly, when implementing the embodiments of the present disclosure, the type, scope of use, usage scenarios, etc. of the data or information that may be involved should be informed to the user and the user's authorization should be obtained through appropriate means in accordance with the relevant laws and regulations. The specific notification and / or authorization methods may vary according to the actual situation and application scenarios, and the scope of the present disclosure is not limited in this regard.

[0018] For the solutions in this specification and the embodiments, if they involve personal information processing, they will all be processed on the premise of having a legal basis (such as obtaining the consent of the personal information subject, or being necessary for performing a contract, etc.), and will only be processed within the specified or agreed scope. If the user refuses to process personal information other than the necessary information required for the basic functions, it will not affect the user's use of the basic functions.

[0019] In the process of people's information interaction through the Internet, people expect to use high-quality audio processing methods to conveniently achieve the desired voice-changing effect. The traditional audio processing method is to use a voice changer to change the voice. However, a voice changer can only change the timbre of the speaker.

[0020] In view of this, the embodiments of the present disclosure propose an audio processing solution. According to this solution, the first media content input by the user can be obtained, and the first media content includes the first audio content. Further, based on the user's selection of the target style, the second media content can be provided, and the second media content includes the second audio content generated based on the first audio content. The second audio content has the same timbre as the first audio content, and the second audio content has at least one audio attribute corresponding to the target style. In this way, the embodiments of the present disclosure can improve the voice-changing effect while retaining the timbre.

[0021] The following further describes various example implementations of this solution in detail with reference to the accompanying drawings.

[0022] Example environment

[0023] Figure 1 FIG. 1 is a schematic diagram of an exemplary environment 100 in which embodiments of the present disclosure can be implemented. As Figure 1 shown, the exemplary environment 100 may include a terminal device 110.

[0024] In this exemplary environment 100, the terminal device 110 may run a platform that supports audio processing. For example, voice conversion processing of audio. The user 140 may interact with the platform via the terminal device 110 and / or its attached devices.

[0025] In Figure 1 the environment 100, if the platform is active, the terminal device 110 may present an interface 150 for supporting interface interaction through the platform.

[0026] In some embodiments, the terminal device 110 communicates with the server 130 to implement the supply of services of the platform. The terminal device 110 may be any type of mobile terminal, fixed terminal or portable terminal, including mobile phones, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, media computers, multimedia tablets, handheld computers, portable game terminals, VR / AR devices, Personal Communication System (PCS) devices, personal navigation devices, Personal Digital Assistant (PDA), audio / video players, digital cameras / cameras, positioning devices, television receivers, radio broadcast receivers, e-book devices, game devices or any combination of the foregoing, including accessories and peripherals of these devices or any combination thereof. In some embodiments, the terminal device 110 can also support any type of interface for users (such as "wearable" circuits, etc.).

[0027] The server 130 may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks, and big data and artificial intelligence platforms. The server 130 may include, for example, a computing system / server, such as a mainframe, an edge computing node, a computing device in a cloud environment, and so on. The server 130 may provide background services for the application 120 that supports virtual scenarios in the terminal device 110.

[0028] A communication connection can be established between server 130 and terminal device 110. The communication connection can be established by wired or wireless means. The communication connection can include, but is not limited to, a Bluetooth connection, a mobile network connection, a Universal Serial Bus (USB) connection, a Wireless Fidelity (WiFi) connection, etc. Embodiments of the present disclosure are not limited in this regard. In an embodiment of the present disclosure, server 130 and terminal device 110 can implement signaling interaction through the communication connection therebetween.

[0029] It should be understood that the structures and functions of the various elements in environment 100 are described only for exemplary purposes and do not imply any limitation on the scope of the present disclosure.

[0030] Some exemplary embodiments of the present disclosure will be further described below with reference to the accompanying drawings.

[0031] Example process

[0032] Figure 2 A flowchart of an example process 200 for audio processing according to some embodiments of the present disclosure is shown. Process 200 can be implemented at terminal device 110. The following will refer to Figure 1 to describe process 200.

[0033] As Figure 2 shown, at block 210, terminal device 110 obtains first media content input by user 140. The first media content includes first audio content.

[0034] In some embodiments, the first media content input by user 140 obtained by terminal device 110 can be first media content recorded by user 140. For example, a video shot by the user, or a voice recorded as the first media content. The first media content can include an audio content, such an audio content can include, for example, the user's speech content, or singing content, etc.

[0035] In some embodiments, the first media content input by user 140 obtained by terminal device 110 can be first media content uploaded by user 140. For example, a previous video stored on terminal device 110, or a previously recorded voice stored as the first media content.

[0036] The following will refer to Figures 3A to 3C to describe process 200. Figures 3A to 3C Schematic diagrams of example interfaces 301 to 303 according to some embodiments of the present disclosure are shown. Interfaces 301 to 303 can be provided, for example, by Figure 1 the terminal device 110 shown.

[0037] As shown Figure 3A in FIG. 1, the terminal device 110 obtains a video recorded by the user 140 on the shooting page, and the video includes first audio. For example, the speech information spoken by the user 140 (the shooter). Subsequently, the user 140 clicks on the sound control 311 in the interface 301.

[0038] Based on the click of the user 140, the terminal device 110 will present a selection panel 320. As shown Figure 3B in FIG. 2, the terminal device 110 can display the selection panel 320 in the interface 302. As an example, the selection panel 320 can display different styles of audio effects, such as style one 321, style two 322, style three, and so on.

[0039] In some embodiments, the selection panel 320 displayed by the terminal device 110 can provide a set of candidate audio effects. In some examples, when the terminal device 110 obtains the speech content input by the user 140, the selection panel 320 can provide different styles that can be converted for the speech content.

[0040] In some examples, when the terminal device 110 obtains a passage spoken by the user 140 in Mandarin, the selection panel 320 can provide audio effects in the "English" style, the "dialect" style, and so on. It can be understood that if the user 140 selects the audio effect in the "English" style presented on the selection panel 320, the terminal device 110 can convert the Mandarin spoken by the user 140 into a passage in the English version with the tone of the user 140. If the user 140 selects the audio effect in the "dialect" style presented on the selection panel 320, the terminal device 110 can convert the Mandarin spoken by the user 140 in dialect into a passage in the dialect version with the tone of the user 140. This is only exemplary and the present disclosure does not limit this.

[0041] In some embodiments, the terminal device 110 receives the user's selection of a target audio effect from a set of candidate audio effects. The target audio effect corresponds to a target style. For example, the terminal device 110 can receive the "English" style, or the "certain dialect" style, etc. selected by the user 140 from the set of candidate audio effects.

[0042] Continue to refer to Figure 2, at block 220, the terminal device 110 provides the second media content based on the user 140's selection of the target style. In some embodiments, the second media content includes second audio content. The second audio content includes second audio content generated based on the first audio content. The second audio content has the same timbre as the first audio content. The second audio content has at least one audio attribute corresponding to the target style. In some examples, the terminal device 110 converts the first audio content into the second audio content based on the target style selected by the user 140 while retaining the timbre corresponding to the first audio. The second audio content has at least one audio attribute corresponding to the target style.

[0043] In some embodiments, the second audio content has at least one audio attribute corresponding to the target style. The target audio attributes include at least pitch, rhythm, and so on. For example, if the pitch of the first audio content includes key A, the second audio content at least retains key A.

[0044] For Figure 3C as an example, the user 140 selects Style Three 331 as the target style on the selection panel 320. The terminal device 110 receives the user 140's selection and will convert the first audio content in the first media content into the second audio content in the second media content.

[0045] In some examples, if the terminal device 110 obtains the speech content input by the user 140, it can change the speech content in different styles. For example, convert the "a passage" (the first audio content) spoken by the user 140 in Mandarin into "a passage" (the second audio content) spoken in the "dialect" style with the same timbre as the user 140. The terminal device 110 calls the server 130 to convert the first audio content and provides the second audio content based on the user's selection of the target style.

[0046] In the embodiments of the present disclosure, in this way, the user can hear at low cost what it is like to speak in other styles with their own timbre.

[0047] The generation of the second media content is described below. The terminal device 110 generates the second audio content according to the first audio content included in the first media content. In some embodiments, the terminal device 110 adjusts the playback speed of the video content of the first video content included in the first media content according to the second audio content. In some embodiments, the terminal device 110 determines the audio part corresponding to the target content in the second audio content and the video part corresponding to the target content in the visual content. Thus, the playback speed of the video part is adjusted so that the video part is synchronized with the audio part. In some examples, the target content can be "a passage", "a sentence", "a character", or "a word", and so on.

[0048] In some examples, the terminal device 110 converts the "paragraph" spoken in Mandarin input by the user 140 into a "paragraph" spoken in the "dialect" style with the same voice as the user 140. Then, the terminal device 110 determines the audio part corresponding to the "paragraph" in the second audio content, and determines the video part corresponding to the "paragraph" in the visual content of the first video content. Subsequently, the terminal device 110 adjusts the playback speed of the video part according to the determined audio part and video part, so that the video part is synchronized with the audio part.

[0049] In some embodiments, the terminal device 110 generates a second video content based on the second audio content and the adjusted visual content as the second media content. In the embodiments of the present disclosure, by adjusting the playback speed of the visual content of the first video content to adapt to the second audio content after voice conversion, the inconsistency between audio and video can be reduced while retaining the voice.

[0050] The generation of the second audio content is described below. The terminal device 110 extracts the first audio content from the first media content. Subsequently, the terminal device 110 inputs the first audio content into the target model to obtain the second audio content. In some embodiments, the target model may be a model trained by the server according to the sample data corresponding to the target style.

[0051] In some embodiments, the target model includes a speech recognition module, a style conversion module, and a speech generation module. The language recognition module is configured to: the terminal device 110 determines the text content corresponding to the first audio content. In some examples, the terminal device 110 sends the first audio content input by the user 140 to the server 130. The server 130 converts the first audio content into text content by calling the speech recognition module in the target model. The server 130 sends the text content corresponding to the first audio content to the terminal device 110. The terminal device 110 determines the text content corresponding to the first audio content.

[0052] In some embodiments, the style conversion module is configured to: convert the text content into a first feature corresponding to the target style. In some embodiments, the first feature at least indicates the stress conversion style corresponding to the target style. In some examples, the terminal device 110 sends the text content corresponding to the first audio content and the target style selected by the user 140 to the server 130. The server 130 uses the style conversion module in the target model to convert the text content into a first feature corresponding to the target style. For example, map the time-varying content features of the words in the source accent or speech style to the content features in the target accent or speech style.

[0053] In some embodiments, the speech generation module is configured to generate intermediate audio content based on a first feature and a second feature. The second feature is used to characterize the timbre of the first audio content. In some embodiments, the speech generation module further includes a diffusion model. In some examples, the server 130 generates intermediate audio content according to the first feature it generates and the second feature used to characterize the timbre of the first audio content. The server 130 sends the generated intermediate audio to the terminal device 110.

[0054] For example, the target model bottleneck-to-bottleneck (i.e., BN2BN) model maps the time-varying content features of utterances in the source accent or speech style to the content features in the target accent or speech style. Then, the zero-shot voice conversion condition is performed through a denoising diffusion probability model (diffusion model).

[0055] For product-oriented models, intermediate bottleneck features extracted from a pre-trained automatic speech recognition (ASR) model are used to process time-varying content features (e.g., BN L10). The diffusion model adopts the accent conversion content features of the BN2BN model and the utterance-level speaker embedding as an adjustment signal, enabling fast voice conversion, i.e., preserving the timbre of any arbitrary source speaker not seen during training.

[0056] Embodiments of the present disclosure can achieve zero-shot identity-protected accent and speech style conversion by using a target model to process the first audio content and generate the second audio content.

[0057] In summary, embodiments of the present disclosure obtain the first audio content included in the first media content input by the user and provide the second media content including the second audio content having the same timbre as the first audio content according to the user's selection of the target style. Correspondingly, the second audio content has at least one audio attribute corresponding to the target style. Thus, the voice conversion effect can be improved while preserving the timbre. Further, by adjusting the playback speed of the visual content of the first video content to adapt to the voice-converted second audio content, the effect of lip-sync can be achieved while preserving the timbre.

[0058] Example device and equipment

[0059] Embodiments of the present disclosure also provide corresponding devices for implementing the above methods or processes. Figure 4 A schematic structural block diagram of an example device 400 for audio processing according to certain embodiments of the present disclosure is shown. The device 400 can be implemented as or included in the terminal device 110. Each module / component in the device 400 can be implemented by hardware, software, firmware, or any combination thereof.

[0060] AsFigure 4 As shown, device 400 includes an acquisition module 410 configured to acquire first media content input by a user, where the first media content includes first audio content.

[0061] Device 400 further includes a provision module 420 configured to provide second media content based on a user's selection of a target style, where the second media content includes second audio content generated based on the first audio content, the second audio content having the same timbre as the first audio content, and the second audio content having at least one audio attribute corresponding to the target style.

[0062] In some embodiments, the at least one audio attribute includes at least one of the following: pitch, rhythm.

[0063] In some embodiments, provision module 420 further includes a selection module configured to display a selection panel, where the selection panel provides at least a set of candidate audio effects; and receive a user's selection of a target audio effect from the set of audio effects, where the target audio effect corresponds to the target style.

[0064] In some embodiments, provision module 420 further includes a generation module configured to generate second audio content based on the first audio content; adjust the playback speed of the visual content of the first video content based on the second audio content; and generate second video content based on the second audio content and the adjusted visual content as the second media content.

[0065] In some embodiments, provision module 420 further includes an adjustment module configured to determine an audio portion of the second audio content corresponding to a target content and a video portion of the visual content corresponding to the target content; and adjust the playback speed of the video portion such that the video portion is synchronized with the audio portion.

[0066] In some embodiments, the first media content includes first video content, and acquisition module 410 is further configured to acquire first media content recorded by the user; and acquire first media content uploaded by the user.

[0067] In some embodiments, the generation module is further configured to extract the first audio content from the first media content; and process the first audio content using a target model to generate the second audio content, where the target model is trained based on sample data corresponding to the target style.

[0068] In some embodiments, the target model includes: a speech recognition module configured to determine text content corresponding to first audio content; a style conversion module configured to convert the text content into first features corresponding to a target style; and a speech generation module configured to generate intermediate audio content based on the first features and second features, where the second features are used to characterize the timbre of the first audio content.

[0069] In some embodiments, the speech generation module includes a diffusion model.

[0070] In some embodiments, the first features at least indicate an accent conversion style corresponding to the target style.

[0071] Figure 5 The block diagram of an electronic device 500 in which one or more embodiments of the present disclosure may be implemented is shown. It should be understood that Figure 5 The illustrated electronic device 500 is merely exemplary and should not constitute any limitation to the functions and scope of the embodiments described herein. Figure 5 The illustrated electronic device 500 can be used to implement Figure 1 the terminal device 110.

[0072] As Figure 5 shown, the electronic device 500 is in the form of a general-purpose electronic device. The components of the electronic device 500 may include, but are not limited to, one or more processors or processing units 510, a memory 520, a storage device 530, one or more communication units 540, one or more input devices 550, and one or more output devices 560. The processing unit 510 can be an actual or virtual processor and is capable of performing various processes according to the programs stored in the memory 520. In a multi-processor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing ability of the electronic device 500.

[0073] The electronic device 500 generally includes multiple computer storage media. Such media can be any accessible media that can be obtained by the electronic device 500, including but not limited to volatile and non-volatile media, removable and non-removable media. The memory 520 can be a volatile memory (such as registers, caches, random access memory (RAM)), a non-volatile memory (such as read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. The storage device 530 can be a removable or non-removable medium and can include machine-readable media, such as a flash drive, a magnetic disk, or any other medium that can be used to store information and / or data and can be accessed within the electronic device 500.

[0074] The electronic device 500 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in Figure 5 , a disk drive for reading from and writing to a removable, non-volatile disk (such as a "floppy disk") and an optical disk drive for reading from and writing to a removable, non-volatile optical disk may be provided. In these cases, each drive may be connected to a bus (not shown) by one or more data media interfaces. The memory 520 may include a computer program product 525 having one or more program modules configured to execute the various methods or actions of the various embodiments of the present disclosure.

[0075] The communication unit 540 enables communication with other electronic devices via a communication medium. Additionally, the functions of the components of the electronic device 500 may be implemented in a single computing cluster or multiple computer machines capable of communicating via a communication connection. Thus, the electronic device 500 may operate in a networked environment using a logical connection to one or more other servers, network personal computers (PCs), or another network node.

[0076] The input device 550 may be one or more input devices such as a mouse, keyboard, trackball, etc. The output device 560 may be one or more output devices such as a display, speaker, printer, etc. The electronic device 500 may also communicate with one or more external devices (not shown) as needed via the communication unit 540, such as storage devices, display devices, etc., communicate with one or more devices that enable a user to interact with the electronic device 500, or communicate with any device that enables the electronic device 500 to communicate with one or more other electronic devices (e.g., a network card, a modem, etc.). Such communication may be performed via an input / output (I / O) interface (not shown).

[0077] According to an exemplary implementation of the present disclosure, a computer-readable storage medium is provided, on which computer-executable instructions are stored, where the computer-executable instructions are executed by a processor to implement the methods described above. According to an exemplary implementation of the present disclosure, a computer program product is also provided, the computer program product being tangibly stored on a non-transitory computer-readable medium and including computer-executable instructions, and the computer-executable instructions being executed by a processor to implement the methods described above.

[0078] Aspects of the present disclosure are described herein with reference to the flowcharts and / or block diagrams of methods, apparatuses, devices, and computer program products implemented in accordance with the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and the combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.

[0079] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that the instructions, when executed by the processing unit of the computer or other programmable data processing apparatus, result in an apparatus that implements the functions / acts specified in one or more boxes of the flowchart and / or block diagram. These computer-readable program instructions may also be stored in a computer-readable storage medium that causes a computer, a programmable data processing apparatus, and / or other devices to function in a particular manner, such that the computer-readable medium storing the instructions comprises a manufacture including instructions that implement various aspects of the functions / acts specified in one or more boxes of the flowchart and / or block diagram.

[0080] The computer-readable program instructions may be loaded onto a computer, other programmable data processing apparatus, or other device, such that a series of operational steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, whereby the instructions executed on the computer, other programmable data processing apparatus, or other device implement the functions / acts specified in one or more boxes of the flowchart and / or block diagram.

[0081] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various implementations of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of code, or a portion of an instruction, which comprises one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two consecutive blocks may in fact be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending upon the functionality involved. It should also be noted that each block of the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented by special-purpose hardware-based systems that perform the specified functions or acts, or combinations of special-purpose hardware and computer instructions.

[0082] The various implementations of the present disclosure have been described above. The above description is exemplary, not exhaustive, and is not limited to the disclosed implementations. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described implementations. The choice of terms used herein is intended to best explain the principles of the implementations, the practical application, or improvements made to the technology in the marketplace, or to enable other ordinary skill in the art to understand the various implementations disclosed herein.

Claims

1. An audio processing method, comprising: Obtaining first media content input by a user, the first media content including first audio content; And Based on the user's selection of a target style, providing second media content, the second media content including second audio content generated based on the first audio content, the second audio content having the same timbre as the first audio content, and the second audio content having at least one audio attribute corresponding to the target style.

2. The method according to claim 1, wherein the at least one audio attribute includes at least one of the following: pitch, rhythm.

3. The method according to claim 1, further comprising: Displaying a selection panel, wherein the selection panel provides a set of candidate audio effects; And Receiving the user's selection of a target audio effect from the set of audio effects, wherein the target audio effect corresponds to the target style.

4. The method according to claim 1, wherein the first media content includes first video content, and the second media content is generated based on the following process: Generate the second audio content based on the first audio content; And Based on the second audio content, adjusting the playback speed of the visual content of the first video content; And Based on the second audio content and the adjusted visual content, generating second video content as the second media content.

5. The method according to claim 4, wherein adjusting the playback speed of the visual part of the first video content based on the second audio content includes: Determining an audio part corresponding to a target content in the second audio content and a video part corresponding to the target content in the visual content; And Adjusting the playback speed of the video part so that the video part is synchronized with the audio part.

6. The method according to claim 1, wherein obtaining first media content input by a user includes: Obtaining the first media content recorded by the user; And Obtaining the first media content uploaded by the user.

7. The method according to claim 1, wherein the second audio content is generated through the following process: Extracting the first audio content from the first media content; and Processing the first audio content using a target model to generate the second audio content, wherein the target model is trained based on sample data corresponding to the target style.

8. The method according to claim 7, wherein the target model includes: A speech recognition module configured to determine text content corresponding to the first audio content; A style conversion module configured to convert the text content into a first feature corresponding to the target style; And A speech generation module configured to generate intermediate audio content based on the first feature and a second feature, the second feature being used to characterize the timbre of the first audio content.

9. The method according to claim 8, wherein the speech generation module includes a diffusion model.

10. The method according to claim 8, wherein the first feature at least indicates a stress conversion style corresponding to the target style.

11. An apparatus for audio processing, comprising: An acquisition module, configured to acquire first media content input by a user, where the first media content includes first audio content; and A provision module, configured to provide second media content based on a user's selection of a target style, where the second media content includes second audio content generated based on the first audio content, the second audio content has the same timbre as the first audio content, and the second audio content has at least one audio attribute corresponding to the target style.

12. An electronic device, comprising: At least one processing unit; and At least one memory, the at least one memory being coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions, when executed by the at least one processing unit, causing the electronic device to perform the method according to any one of claims 1 to 10.

13. A computer-readable storage medium, having stored thereon a computer program, the computer program being executable by a processor to implement the method according to any one of claims 1 to 10.

Citation Information

Cited By

  • Audio processing method and apparatus, and device and storage medium

    WO2025156808A1