Audio processing method and apparatus, and device and storage medium
By obtaining the audio content input by the user and using the target model to generate audio content with the target style, the problem of single tone change in traditional audio processing methods is solved, and the style conversion and audio and video synchronization of audio content is realized.
Patent Information
- Application Number
- PCT/CN2024/134607
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-23
- Filing Date
- 2024-11-26
- Publication Date
- 2025-07-31
AI Technical Summary
Traditional audio processing methods can only change the speaker's tone, and it is difficult to achieve voice change effect on the basis of retaining the tone.
By acquiring the first media content input by the user, the second media content is generated based on the user's selection of the target style. The second audio content has the same tone as the first audio content and has audio attributes corresponding to the target style. The first audio content is processed using the target model to generate the second audio content, and combined with adjusting the playback speed of the video content to achieve audio and video synchronization.
On the basis of retaining the tone, the voice change effect is improved, and the style conversion of audio content and audio and video synchronization are achieved.
Smart Images

Figure CN2024134607_31072025_PF_FP_ABST
Abstract
Description
Audio processing method, device, equipment and storage medium
[0001] This application claims priority to the Chinese invention patent application entitled “Method, device, equipment and storage medium for audio processing” and application number 202410096064.X filed on January 23, 2024. The entire contents of that application are incorporated by reference into this application. Technical Field
[0002] Example embodiments of the present disclosure generally relate to the field of computers, and more particularly, to methods, devices, apparatuses, and computer-readable storage media for audio processing. Background Art
[0003] With the development of computer technology, the internet has become a vital platform for information exchange. In this process, various audio formats have become crucial media for social expression and information exchange. Therefore, voice-changing technologies are being developed that can modify speaking styles while preserving timbre through audio processing. Summary of the Invention
[0004] In a first aspect of the present disclosure, a method for audio processing is provided. The method comprises: obtaining first media content input by a user, the first media content comprising first audio content; and providing second media content based on a target style selected by the user, the second media content comprising second audio content generated based on the first audio content, the second audio content having the same timbre as the first audio content and at least one audio attribute corresponding to the target style.
[0005] In a second aspect of the present disclosure, a device for audio processing is provided. The device includes: an acquisition module configured to acquire first media content input by a user, the first media content including first audio content; and a provision module configured to provide second media content based on a user selection of a target style, the second media content including second audio content generated based on the first audio content, the second audio content having the same timbre as the first audio content and at least one audio attribute corresponding to the target style.
[0006] In a third aspect of the present disclosure, an electronic device is provided. The device includes at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit. When executed by the at least one processing unit, the instructions cause the device to perform the method of any one of the first to fourth aspects.
[0007] In a fourth aspect of the present disclosure, a computer-readable storage medium is provided, wherein a computer program is stored on the computer-readable storage medium, and the computer program can be executed by a processor to implement the method of any one of the first to fourth aspects.
[0008] It should be understood that the content described in this summary section is not intended to limit the key features or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] The above and other features, advantages and aspects of the embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. In the accompanying drawings, the same or similar reference numerals represent the same or similar elements, wherein:
[0010] FIG1 shows a schematic diagram of an example environment in which embodiments according to the present disclosure may be implemented;
[0011] FIG2 illustrates a flow chart of an example process for audio processing according to some embodiments of the present disclosure;
[0012] 3A to 3C are schematic diagrams illustrating example interfaces according to some embodiments of the present disclosure;
[0013] FIG4 shows a schematic structural block diagram of an example apparatus for audio processing according to some embodiments of the present disclosure; and
[0014] FIG5 shows a block diagram of an electronic device capable of implementing various embodiments of the present disclosure. DETAILED DESCRIPTION
[0015] The following describes embodiments of the present disclosure in more detail with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments described herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.
[0016] It should be noted that the titles of any section / subsection provided herein are not limiting. Various embodiments are described throughout this document, and any type of embodiment may be included under any section / subsection. Furthermore, the embodiments described in any section / subsection may be combined in any manner with any other embodiments described in the same section / subsection and / or in different sections / subsections.
[0017] In the description of the embodiments of the present disclosure, the term "including" and similar terms should be understood as open inclusion, that is, "including but not limited to". The term "based on" should be understood as "based at least in part on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". Other explicit and implicit definitions may be included below. The terms "first", "second", etc. may refer to different or the same objects. Other explicit and implicit definitions may be included below.
[0018] The embodiments of the present disclosure may involve user data, data acquisition and / or use, etc. These aspects shall comply with the corresponding laws, regulations and relevant provisions. In the embodiments of the present disclosure, all data collection, acquisition, processing, processing, forwarding, use, etc. are carried out on the premise that the user is aware of and confirms them. Accordingly, when implementing the various embodiments of the present disclosure, the types, scope of use, and usage scenarios of the data or information that may be involved should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with the relevant laws and regulations. The specific notification and / or authorization method may vary according to the actual situation and application scenario, and the scope of the present disclosure is not limited in this respect.
[0019] If this specification and the solutions in the examples involve the processing of personal information, such processing will be done only with a legitimate basis (such as with the consent of the subject of personal information or as necessary for the performance of a contract) and only within the prescribed or agreed scope. A user's refusal to process personal information other than that required for basic functions will not affect the user's use of basic functions.
[0020] As people interact with each other online, they expect to use high-quality audio processing methods to easily achieve the desired voice-changing effects. Traditional audio processing methods rely on voice changers. However, voice changers can only alter the speaker's timbre.
[0021] In view of this, an embodiment of the present disclosure proposes an audio processing solution. According to the solution, the first media content input by the user can be obtained, and the first media content includes the first audio content. Furthermore, based on the user's selection of the target style, the second media content can be provided, and the second media content includes the second audio content generated based on the first audio content, and the second audio content has the same timbre as the first audio content, and the second audio content has at least one audio attribute corresponding to the target style. In this way, the embodiment of the present disclosure can improve the voice changing effect while retaining the timbre.
[0022] Various example implementations of this solution are described in detail below in conjunction with the accompanying drawings.
[0023] Sample Environment
[0024] FIG1 shows a schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented. As shown in FIG1 , the example environment 100 may include a terminal device 110 .
[0025] In this example environment 100, the terminal device 110 may run a platform that supports audio processing, such as voice modification, and the user 140 may interact with the platform via the terminal device 110 and / or its attached devices.
[0026] In the environment 100 of FIG. 1 , if the platform is in an active state, the terminal device 110 may present an interface 150 for supporting interface interaction through the platform.
[0027] In some embodiments, the terminal device 110 communicates with the server 130 to enable the provision of services to the platform. The terminal device 110 can be any type of mobile terminal, fixed terminal or portable terminal, including a mobile phone, a desktop computer, a laptop computer, a notebook computer, a netbook computer, a tablet computer, a media computer, a multimedia tablet, a handheld computer, a portable game terminal, a VR / AR device, a personal communication system (Personal Communication System, PCS) device, a personal navigation device, a personal digital assistant (Personal Digital Assistant, PDA), an audio / video player, a digital camera / camcorder, a positioning device, a television receiver, a radio broadcast receiver, an e-book device, a gaming device or any combination thereof, including accessories and peripherals of these devices or any combination thereof. In some embodiments, the terminal device 110 can also support any type of interface for the user (such as a "wearable" circuit, etc.).
[0028] Server 130 can be a standalone physical server, a server cluster or distributed system consisting of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content distribution networks, and big data and artificial intelligence platforms. Server 130 can include, for example, computing systems / servers such as mainframes, edge computing nodes, computing devices in a cloud environment, and the like. Server 130 provides backend services for application 120 that supports virtual scenarios in terminal device 110.
[0029] A communication connection may be established between the server 130 and the terminal device 110. The communication connection may be established via a wired or wireless method. The communication connection may include, but is not limited to, a Bluetooth connection, a mobile network connection, a Universal Serial Bus (USB) connection, a Wireless Fidelity (WiFi) connection, etc., and the embodiments of the present disclosure are not limited in this respect. In the embodiments of the present disclosure, the server 130 and the terminal device 110 may implement signaling interaction via the communication connection between the two.
[0030] It should be understood that the structure and function of the various elements in the environment 100 are described for illustrative purposes only and do not imply any limitation on the scope of the present disclosure.
[0031] Some example embodiments of the present disclosure will be described below with continued reference to the accompanying drawings.
[0032] Example Process
[0033] FIG2 shows a flow chart of an example process 200 for audio processing according to some embodiments of the present disclosure. The process 200 may be implemented at the terminal device 110. The process 200 is described below with reference to FIG1.
[0034] As shown in Figure 2, in block 210, the terminal device 110 obtains first media content input by the user 140. The first media content includes first audio content.
[0035] In some embodiments, the first media content input by the user 140 and obtained by the terminal device 110 may be first media content recorded by the user 140. For example, the first media content may be a video shot by the user or a voice recording made by the user. The first media content may include an audio content, such as a user's speech or singing.
[0036] In some embodiments, the first media content input by the user 140 and acquired by the terminal device 110 may be the first media content uploaded by the user 140. For example, a previously stored video or a previously recorded voice on the terminal device 110 may be used as the first media content.
[0037] The process 200 will be described below with reference to Figures 3A to 3C . Figures 3A to 3C show schematic diagrams of example interfaces 301 to 303 according to some embodiments of the present disclosure. Interfaces 301 to 303 may be provided by the terminal device 110 shown in Figure 1 , for example.
[0038] As shown in FIG3A , the terminal device 110 obtains a video recorded by the user 140 on the shooting page, and the video includes a first audio, for example, the speech information of the user 140 (photographer). Subsequently, the user 140 clicks the sound control 311 in the interface 301 .
[0039] Based on the click of the user 140, the terminal device 110 presents a selection panel 320. As shown in FIG3B , the terminal device 110 may display the selection panel 320 in the interface 302. As an example, the selection panel 320 may display audio effects of different styles, such as style 1 321, style 2 322, style 3, and so on.
[0040] In some embodiments, the selection panel 320 displayed by the terminal device 110 can be used to provide a set of candidate audio effects. In some examples, the terminal device 110 obtains the spoken content input by the user 140, and the selection panel 320 can provide different styles that can be converted to the spoken content.
[0041] In some examples, the terminal device 110 obtains a passage spoken by the user 140 in Mandarin, and the selection panel 320 may provide an audio effect in an "English" style, an audio effect in a "dialect" style, and so on. It is understood that if the user 140 selects the "English" style audio effect presented on the selection panel 320, the terminal device 110 may convert the Mandarin spoken by the user 140 into an English version of the passage using the user 140's timbre. If the user 140 selects the "dialect" style audio effect presented on the selection panel 320, the terminal device 110 may convert the Mandarin spoken by the user 140 in a dialect into a dialect version of the passage using the user 140's timbre. This is merely exemplary and the present disclosure is not limited thereto.
[0042] In some embodiments, the terminal device 110 receives a user's selection of a target audio effect from a set of candidate audio effects. The target audio effect corresponds to a target style. For example, the terminal device 110 may receive the user 140's selection of an "English" style or a "certain dialect" style, etc., from the set of candidate audio effects.
[0043] Continuing with reference to FIG2 , in block 220 , the terminal device 110 provides second media content based on the user 140's selection of a target style. In some embodiments, the second media content includes second audio content. The second audio content includes second audio content generated based on the first audio content. The second audio content has the same timbre as the first audio content. The second audio content has at least one audio attribute corresponding to the target style. In some examples, the terminal device 110 converts the first audio content into second audio content based on the target style selected by the user 140 while retaining the timbre corresponding to the first audio content. The second audio content has at least one audio attribute corresponding to the target style.
[0044] In some embodiments, the second audio content has at least one audio attribute corresponding to the target style. The target audio attribute includes at least pitch, tempo, etc. For example, if the pitch of the first audio content includes A, the second audio content at least retains A.
[0045] 3C as an example, the user 140 selects style three 331 as the target style on the selection panel 320. Upon receiving the selection of the user 140, the terminal device 110 converts the first audio content in the first media content into the second audio content in the second media content.
[0046] In some examples, the terminal device 110 receives the spoken content input by the user 140 and can modify the spoken content to a different style. For example, a "sentence" (first audio content) input by the user 140, spoken in Mandarin, can be converted into a "sentence" (second audio content) with the same timbre as the user 140 but spoken in a "dialect" style. Based on the user's selection of the target style, the terminal device 110 calls the server 130 to convert the first audio content and provide the second audio content.
[0047] In the disclosed embodiment, in this way, the user can hear the effect of speaking in other styles with his own voice at a low cost.
[0048] The following describes the generation of the second media content. The terminal device 110 generates the second audio content based on the first audio content included in the first media content. In some embodiments, the terminal device 110 adjusts the playback speed of the video content of the first video content included in the first media content based on the second audio content. In some embodiments, the terminal device 110 determines the audio portion corresponding to the target content in the second audio content, and the video portion corresponding to the target content in the visual content. Thereby, the playback speed of the video portion is adjusted so that the video portion is synchronized with the audio portion. In some examples, the target content can be "a paragraph", "a sentence", "a character", or "a word", etc.
[0049] In some examples, terminal device 110 converts a "speech" input by user 140, spoken in Mandarin, into a "speech" with the same timbre as user 140, but spoken in a dialect style. Terminal device 110 then determines the audio portion of the second audio content that corresponds to the "speech," and the video portion of the visual content of the first video content that corresponds to the "speech." Terminal device 110 then adjusts the playback speed of the video portion based on the determined audio and video portions, thereby synchronizing the video portion with the audio portion.
[0050] In some embodiments, terminal device 110 generates second video content based on the second audio content and the adjusted visual content as the second media content. In the disclosed embodiments, by adjusting the playback speed of the visual content of the first video content to match the voice-changed second audio content, the inconsistency between the audio and the image can be minimized while preserving the timbre.
[0051] The following describes the generation of the second audio content. The terminal device 110 extracts the first audio content from the first media content. Subsequently, the terminal device 110 inputs the first audio content into a target model to obtain the second audio content. In some embodiments, the target model can be a model trained by the server based on sample data corresponding to the target style.
[0052] In some embodiments, the target model includes a speech recognition module, a style conversion module, and a speech generation module. The language recognition module is configured such that the terminal device 110 determines the text content corresponding to the first audio content. In some examples, the terminal device 110 sends the first audio content input by the user 140 to the server 130. The server 130 converts the first audio content into text content by calling the speech recognition module in the target model. The server 130 sends the text content corresponding to the first audio content to the terminal device 110. The terminal device 110 determines the text content corresponding to the first audio content.
[0053] In some embodiments, the style conversion module is configured to convert the text content into a first feature corresponding to a target style. In some embodiments, the first feature indicates at least an accent conversion style corresponding to the target style. In some examples, terminal device 110 sends the text content corresponding to the first audio content and the target style selected by user 140 to server 130. Server 130 utilizes the style conversion module in the target model to convert the text content into the first feature corresponding to the target style. For example, time-varying content features of speech in a source accent or voice style are mapped to content features in a target accent or voice style.
[0054] In some embodiments, the speech generation module is configured to generate intermediate audio content based on a first feature and a second feature. The second feature is used to characterize the timbre of the first audio content. In some embodiments, the speech generation module also includes a diffusion model. In some examples, server 130 generates intermediate audio content based on the generated first feature and the second feature used to characterize the timbre of the first audio content. Server 130 transmits the generated intermediate audio content to terminal device 110.
[0055] For example, the target model is a bottleneck-to-bottleneck (i.e., BN2BN) model that maps the time-varying content features of the utterance in the source accent or voice style to the content features in the target accent or voice style. Then, the condition of zero-shot voice conversion is performed by a denoising diffusion probability model (diffusion model).
[0056] For production-oriented models, we use the intermediate bottleneck features extracted from the pre-trained speech recognition technology ASR model to process time-varying content features (e.g., BN L10). The diffusion model uses the accent-converted content features of the BN2BN model and the utterance-level speaker embedding as a conditioning signal to enable fast voice conversion, that is, preserving the timbre of any arbitrary source speaker not seen during training.
[0057] The disclosed embodiment processes the first audio content using a target model to generate the second audio content, thereby achieving zero-shot identity-preserving accent and voice style conversion.
[0058] In summary, the embodiment of the present disclosure obtains the first audio content included in the first media content input by the user, and provides the second media content including the second audio content having the same timbre as the first audio content based on the user's selection of the target style. Correspondingly, the second audio content has at least one audio attribute corresponding to the target style, thereby improving the voice-changing effect while preserving the timbre. Furthermore, by adjusting the playback speed of the visual content of the first video content to adapt to the second audio content after voice change, the effect of audio-visual synchronization can be achieved while preserving the timbre.
[0059] Example devices and equipment
[0060] Embodiments of the present disclosure also provide corresponding apparatuses for implementing the above-described methods or processes. FIG4 shows a schematic block diagram of an example apparatus 400 for audio processing according to certain embodiments of the present disclosure. Apparatus 400 may be implemented as or included in terminal device 110. Each module / component in apparatus 400 may be implemented by hardware, software, firmware, or any combination thereof.
[0061] As shown in FIG4 , the apparatus 400 includes an acquisition module 410 configured to acquire first media content input by a user, where the first media content includes first audio content.
[0062] The device 400 also includes a providing module 420, which is configured to provide second media content based on the user's selection of the target style, where the second media content includes second audio content generated based on the first audio content, the second audio content has the same timbre as the first audio content, and the second audio content has at least one audio attribute corresponding to the target style.
[0063] In some embodiments, the at least one audio attribute includes at least one of the following: pitch, rhythm.
[0064] In some embodiments, the providing module 420 also includes a selection module, which is configured to display a selection panel, wherein the selection panel provides at least a set of candidate audio effects; and receive a user's selection of a target audio effect from the set of audio effects, wherein the target audio effect corresponds to a target style.
[0065] In some embodiments, the providing module 420 also includes a generating module, which is configured to generate second audio content based on the first audio content; adjust the playback speed of the visual content of the first video content based on the second audio content; and generate second video content based on the second audio content and the adjusted visual content as the second media content.
[0066] In some embodiments, the providing module 420 also includes an adjustment module, which is configured to determine an audio portion in the second audio content corresponding to the target content and a video portion in the visual content corresponding to the target content; and adjust the playback speed of the video portion so that the video portion is synchronized with the audio portion.
[0067] In some embodiments, the first media content includes first video content, and the acquisition module 410 is further configured to acquire first media content recorded by a user; and acquire first media content uploaded by a user.
[0068] In some embodiments, the generation module is further configured to extract first audio content from the first media content; and process the first audio content using a target model to generate second audio content, wherein the target model is trained based on sample data corresponding to a target style.
[0069] In some embodiments, the target model includes: a speech recognition module configured to determine text content corresponding to the first audio content; a style conversion module configured to convert the text content into a first feature corresponding to the target style; and a speech generation module configured to generate intermediate audio content based on the first feature and a second feature, the second feature being used to characterize the timbre of the first audio content.
[0070] In some embodiments, the speech generation module includes a diffusion model.
[0071] In some embodiments, the first feature indicates at least an accent conversion style corresponding to a target style.
[0072] FIG5 shows a block diagram of an electronic device 500 in which one or more embodiments of the present disclosure may be implemented. It should be understood that the electronic device 500 shown in FIG5 is merely exemplary and should not be construed as limiting the functionality and scope of the embodiments described herein. The electronic device 500 shown in FIG5 can be used to implement the terminal device 110 of FIG1 .
[0073] As shown in FIG5 , electronic device 500 is a general-purpose electronic device. Components of electronic device 500 may include, but are not limited to, one or more processors or processing units 510, memory 520, storage device 530, one or more communication units 540, one or more input devices 550, and one or more output devices 560. Processing unit 510 may be a real or virtual processor and is capable of performing various processes according to programs stored in memory 520. In a multi-processor system, multiple processing units execute computer-executable instructions in parallel to enhance the parallel processing capabilities of electronic device 500.
[0074] The electronic device 500 typically includes a plurality of computer storage media. Such media can be any accessible media that can be obtained by the electronic device 500, including but not limited to volatile and non-volatile media, removable and non-removable media. The memory 520 can be a volatile memory (e.g., registers, cache, random access memory (RAM)), a non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. The storage device 530 can be a removable or non-removable medium and can include a machine-readable medium, such as a flash drive, a disk, or any other medium that can be used to store information and / or data and can be accessed within the electronic device 500.
[0075] The electronic device 500 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in FIG. 5 , a disk drive for reading from or writing to a removable, non-volatile disk (e.g., a “floppy disk”) and an optical drive for reading from or writing to a removable, non-volatile optical disk may be provided. In these cases, each drive may be connected to a bus (not shown) by one or more data media interfaces. The memory 520 may include a computer program product 525 having one or more program modules configured to perform various methods or actions of various embodiments of the present disclosure.
[0076] The communication unit 540 enables communication with other electronic devices via a communication medium. Additionally, the functions of the components of the electronic device 500 can be implemented in a single computing cluster or multiple computing machines that can communicate via a communication connection. Thus, the electronic device 500 can operate in a networked environment using a logical connection with one or more other servers, a network personal computer (PC), or another network node.
[0077] Input device 550 may be one or more input devices, such as a mouse, keyboard, or trackball. Output device 560 may be one or more output devices, such as a display, a speaker, or a printer. Electronic device 500 may also communicate with one or more external devices (not shown) via communication unit 540 as needed, such as a storage device, a display device, or the like, with one or more devices that allow a user to interact with electronic device 500, or with any device that allows electronic device 500 to communicate with one or more other electronic devices (e.g., a network card, a modem, etc.). Such communication may be performed via an input / output (I / O) interface (not shown).
[0078] According to an exemplary implementation of the present disclosure, a computer-readable storage medium is provided, on which computer-executable instructions are stored, wherein the computer-executable instructions are executed by a processor to implement the method described above. According to an exemplary implementation of the present disclosure, a computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, and the computer-executable instructions are executed by a processor to implement the method described above.
[0079] Various aspects of the present disclosure are described herein with reference to flowcharts and / or block diagrams of methods, apparatuses, devices, and computer program products implemented according to the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.
[0080] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine, such that when these instructions are executed by the processing unit of the computer or other programmable data processing device, a device is generated that implements the functions / actions specified in one or more blocks in the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium, where these instructions cause the computer, programmable data processing device, and / or other device to operate in a specific manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing various aspects of the functions / actions specified in one or more blocks in the flowchart and / or block diagram.
[0081] Computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device so that a series of operational steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to implement the functions / actions specified in one or more boxes in the flowchart and / or block diagram.
[0082] The flow charts and block diagrams in the accompanying drawings show the possible architecture, functions and operations of the systems, methods and computer program products according to multiple implementations of the present disclosure. In this regard, each box in the flow chart or block diagram can represent a part for a module, program segment or instruction, and a part for a module, program segment or instruction comprises one or more executable instructions for realizing the logical function of the specification. In some alternative implementations, the functions marked in the box can also occur in a sequence different from that marked in the accompanying drawings. For example, two continuous boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be realized by a special hardware-based system that performs the function or action of the specification, or can be realized by a combination of special hardware and computer instructions.
[0083] While various implementations of the present disclosure have been described above, the foregoing description is intended to be illustrative, not exhaustive, and not limited to the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is selected to best explain the principles of the implementations, their practical applications, or improvements to existing technologies, or to enable others skilled in the art to understand the various implementations disclosed herein.
Claims
1. An audio processing method, comprising: Obtaining first media content input by a user, the first media content including first audio content; And Based on the user's selection of a target style, providing second media content, the second media content including second audio content generated based on the first audio content, the second audio content having the same timbre as the first audio content, and the second audio content having at least one audio attribute corresponding to the target style.
2. The method according to claim 1, wherein the at least one audio attribute includes at least one of the following: pitch, rhythm.
3. The method according to claim 1 or 2, further comprising: Displaying a selection panel, wherein the selection panel provides a set of candidate audio effects; And Receiving the user's selection of a target audio effect from the set of audio effects, wherein the target audio effect corresponds to the target style.
4. The method according to claim 1 or 2, wherein the first media content includes first video content, and the second media content is generated based on the following process: Generate the second audio content based on the first audio content; And Based on the second audio content, adjusting the playback speed of the visual content of the first video content; And Based on the second audio content and the adjusted visual content, generating second video content as the second media content.
5. The method according to claim 4, wherein adjusting the playback speed of the visual part of the first video content based on the second audio content includes: Determining an audio part corresponding to a target content in the second audio content and a video part corresponding to the target content in the visual content; And Adjusting the playback speed of the video part so that the video part is synchronized with the audio part.
6. The method according to claim 1 or 4, wherein obtaining first media content input by a user includes: Obtaining the first media content recorded by the user; And Obtaining the first media content uploaded by the user.
7. The method according to claim 1 or 5, wherein the second audio content is generated through the following process: Extracting the first audio content from the first media content; and Processing the first audio content using a target model to generate the second audio content, wherein the target model is trained based on sample data corresponding to the target style.
8. The method according to claim 7, wherein the target model includes: A speech recognition module configured to determine text content corresponding to the first audio content; A style conversion module configured to convert the text content into a first feature corresponding to the target style; And A speech generation module configured to generate intermediate audio content based on the first feature and a second feature, the second feature being used to characterize the timbre of the first audio content.
9. The method according to claim 8, wherein the speech generation module includes a diffusion model.
10. The method according to claim 8, wherein the first feature at least indicates a stress conversion style corresponding to the target style.
11. An apparatus for audio processing, comprising: An acquisition module, configured to acquire first media content input by a user, where the first media content includes first audio content; And A provision module, configured to provide second media content based on a user's selection of a target style, where the second media content includes second audio content generated based on the first audio content, the second audio content has the same timbre as the first audio content, and the second audio content has at least one audio attribute corresponding to the target style.
12. An electronic device, comprising: At least one processing unit; And At least one memory, the at least one memory being coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions, when executed by the at least one processing unit, causing the electronic device to perform the method according to any one of claims 1 to 10.
13. A computer-readable storage medium, having stored thereon a computer program, the computer program being executable by a processor to implement the method according to any one of claims 1 to 10.
Citation Information
Patent Citations
Audio processing method and device, equipment and storage medium
CN120375841A
Method and device for generating audio, equipment and medium
CN111899719A
Speech synthesis method and device, electronic equipment and storage medium
CN112365877A
Live broadcast interaction method and device, electronic equipment and readable storage medium
CN112562705A
Voice conversion method and related equipment
CN114299908A