Audio processing method and apparatus, and device and storage medium
By extracting and adjusting the tone information of the audio content, the problem of time-consuming and mismatch of audio modification in traditional methods is solved, and efficient media content editing and audio authenticity are achieved.
Patent Information
- Application Number
- PCT/CN2024/135302
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-02-19
- Filing Date
- 2024-11-28
- Publication Date
- 2025-08-28
AI Technical Summary
Traditional methods require re-recording when modifying audio content, which is time-consuming and may cause the audio to mismatch with other media content.
By extracting background audio and text audio content from the first media content, audio segments corresponding to the replacement text are generated using the tone information, and the background audio content is adjusted to generate new media content.
It realizes efficient editing of media content, and the generated audio content is more realistic, improving the quality of edited media content.
Smart Images

Figure CN2024135302_28082025_PF_FP_ABST
Abstract
Description
Audio processing method, device, equipment and storage medium
[0001] This application claims priority to the Chinese invention patent application entitled “Audio processing method, device, equipment and storage medium” filed on February 19, 2024, with application number: 202410185739.8. The entire contents of this application are incorporated by reference into this application. Technical Field
[0002] Example embodiments of the present disclosure generally relate to the field of computers, and more particularly, to methods, devices, apparatuses, and computer-readable storage media for audio processing. Background Art
[0003] With the development of computer technology, the Internet has become an important platform for people to exchange information. In the process of people interacting with information through the Internet, various types of audio have become an important medium for people to express themselves socially and exchange information. Summary of the Invention
[0004] In a first aspect of the present disclosure, a method for audio processing is provided. The method includes: extracting background audio content and text audio content corresponding to text content from first media content; based on a request to replace first text in the text content with second text, generating a second audio segment corresponding to the second text using timbre information associated with a first audio segment of the text audio content, wherein the first audio segment corresponds to the first text; adjusting a third audio segment corresponding to the first text in the background audio content based on the second audio segment; and generating second media content based on the first media content, the second audio segment, and the adjusted third audio segment.
[0005] In a second aspect of the present disclosure, a device for audio processing is provided. The device includes: an audio extraction module configured to extract background audio content and text audio content corresponding to text content from first media content; an audio generation module configured to generate a second audio segment corresponding to the second text based on a request to replace a first text in the text content with a second text, using timbre information associated with a first audio segment of the text audio content, wherein the first audio segment corresponds to the first text; an audio adjustment module configured to adjust a third audio segment corresponding to the first text in the background audio content based on the second audio segment; and a media generation module configured to generate second media content based on the first media content, the second audio segment, and the adjusted third audio segment.
[0006] In a third aspect of the present disclosure, an electronic device is provided. The device includes at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit. When executed by the at least one processing unit, the instructions cause the device to perform the method of the first aspect.
[0007] In a fourth aspect of the present disclosure, a computer-readable storage medium is provided, wherein a computer program is stored on the computer-readable storage medium, and the computer program can be executed by a processor to implement the method of the first aspect.
[0008] It should be understood that the content described in this summary section is not intended to limit the key features or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] The above and other features, advantages and aspects of the embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. In the accompanying drawings, the same or similar reference numerals represent the same or similar elements, wherein:
[0010] FIG1 shows a schematic diagram of an example environment in which embodiments according to the present disclosure may be implemented;
[0011] FIG2 illustrates a flow chart of an example process for audio processing according to some embodiments of the present disclosure;
[0012] FIG3 shows a schematic diagram of an example interface according to some embodiments of the present disclosure;
[0013] FIG4 shows a schematic structural block diagram of an example apparatus for audio processing according to some embodiments of the present disclosure; and
[0014] FIG5 shows a block diagram of an electronic device capable of implementing various embodiments of the present disclosure. DETAILED DESCRIPTION
[0015] The following describes embodiments of the present disclosure in more detail with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments described herein. Instead, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.
[0016] It should be noted that the titles of any section / subsection provided herein are not limiting. Various embodiments are described throughout this document, and any type of embodiment may be included under any section / subsection. Furthermore, the embodiments described in any section / subsection may be combined in any manner with any other embodiments described in the same section / subsection and / or in different sections / subsections.
[0017] In the description of the embodiments of the present disclosure, the term "including" and similar terms should be understood as open inclusion, that is, "including but not limited to". The term "based on" should be understood as "based at least in part on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". Other explicit and implicit definitions may be included below. The terms "first", "second", etc. may refer to different or the same objects. Other explicit and implicit definitions may be included below.
[0018] The embodiments of the present disclosure may involve user data, data acquisition and / or use, etc. These aspects shall comply with the corresponding laws, regulations and relevant provisions. In the embodiments of the present disclosure, all data collection, acquisition, processing, processing, forwarding, use, etc. are carried out on the premise that the user is aware of and confirms them. Accordingly, when implementing the various embodiments of the present disclosure, the types, scope of use, and usage scenarios of the data or information that may be involved should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with the relevant laws and regulations. The specific notification and / or authorization method may vary according to the actual situation and application scenario, and the scope of the present disclosure is not limited in this respect.
[0019] If this specification and the solutions in the examples involve the processing of personal information, such processing will be done only with a legitimate basis (such as with the consent of the subject of personal information or as necessary for the performance of a contract) and only within the prescribed or agreed scope. A user's refusal to process personal information other than that required for basic functions will not affect the user's use of basic functions.
[0020] When people create media content, the audio content may not meet their expectations. For example, the user may mispronounce individual words during recording, or may wish to modify a specific sentence in the media content. Traditionally, this requires the original user to re-record the corresponding audio content, which is time-consuming and may cause the re-recorded audio content to not match other elements in the original media content (e.g., background sound or images).
[0021] Embodiments of the present disclosure provide an audio processing solution. According to this solution, background audio content and text audio content corresponding to text content can be extracted from first media content. Furthermore, based on a request to replace a first text in the text content with a second text, timbre information associated with a first audio segment of the text audio content can be used to generate a second audio segment corresponding to the second text, where the first audio segment corresponds to the first text.
[0022] Furthermore, a third audio segment corresponding to the first text in the background audio content may be adjusted based on the second audio segment. Accordingly, the second media content may be generated based on the first media content, the second audio segment, and the adjusted third audio segment.
[0023] In this way, the embodiments of the present disclosure can support editing of media content by modifying the text corresponding to the media content, and can make the modified audio content more realistic, thereby improving the quality of the edited media content.
[0024] Various example implementations of this solution are described in detail below in conjunction with the accompanying drawings.
[0025] Sample Environment
[0026] FIG1 shows a schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented. As shown in FIG1 , the example environment 100 may include a terminal device 110 .
[0027] In this example environment 100, the terminal device 110 may run a platform that supports audio processing, such as voice modification, and the user 140 may interact with the platform via the terminal device 110 and / or its attached devices.
[0028] In the environment 100 of FIG. 1 , if the platform is in an active state, the terminal device 110 may present an interface 150 for supporting interface interaction through the platform.
[0029] In some embodiments, the terminal device 110 communicates with the server 130 to enable the provision of services to the platform. The terminal device 110 can be any type of mobile terminal, fixed terminal or portable terminal, including a mobile phone, a desktop computer, a laptop computer, a notebook computer, a netbook computer, a tablet computer, a media computer, a multimedia tablet, a handheld computer, a portable game terminal, a VR / AR device, a personal communication system (Personal Communication System, PCS) device, a personal navigation device, a personal digital assistant (Personal Digital Assistant, PDA), an audio / video player, a digital camera / camcorder, a positioning device, a television receiver, a radio broadcast receiver, an e-book device, a gaming device or any combination thereof, including accessories and peripherals of these devices or any combination thereof. In some embodiments, the terminal device 110 can also support any type of interface for the user (such as a "wearable" circuit, etc.).
[0030] Server 130 can be a standalone physical server, a server cluster or distributed system consisting of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content distribution networks, and big data and artificial intelligence platforms. Server 130 can include, for example, computing systems / servers such as mainframes, edge computing nodes, computing devices in a cloud environment, and the like. Server 130 provides backend services for application 120 that supports virtual scenarios in terminal device 110.
[0031] A communication connection may be established between the server 130 and the terminal device 110. The communication connection may be established via a wired or wireless method. The communication connection may include, but is not limited to, a Bluetooth connection, a mobile network connection, a Universal Serial Bus (USB) connection, a Wireless Fidelity (WiFi) connection, etc., and the embodiments of the present disclosure are not limited in this respect. In the embodiments of the present disclosure, the server 130 and the terminal device 110 may implement signaling interaction via the communication connection between the two.
[0032] It should be understood that the structure and function of the various elements in the environment 100 are described for illustrative purposes only and do not imply any limitation on the scope of the present disclosure.
[0033] Some example embodiments of the present disclosure will be described below with continued reference to the accompanying drawings.
[0034] Example Process
[0035] FIG2 shows a flow chart of an example process 200 for audio processing according to some embodiments of the present disclosure. Process 200 may be implemented at server 130 or other suitable electronic device. Process 200 is described below with reference to FIG1 .
[0036] As shown in FIG. 2 , at block 210 , the server 130 extracts background audio content and text audio content corresponding to the text content from the first media content.
[0037] In some embodiments, the server 130 may obtain audio content of the media content to be processed. Such media content may include, for example, a video file or an audio file with audio information.
[0038] Furthermore, the server 130 can perform sound source separation on the audio content to extract the text audio content corresponding to the human voice portion. Accordingly, the portion of the audio content from which the text audio content is extracted can be used as background audio content.
[0039] In block 220 , the server 130 generates a second audio segment corresponding to the second text using timbre information associated with the first audio segment of the text audio content based on the request to replace the first text in the text content with the second text, wherein the first audio segment corresponds to the first text.
[0040] The example process of block 220 will be described below with reference to Figure 3. Figure 3 shows an example interface 300 according to some embodiments of the present disclosure. The interface 300 may be provided by the terminal device 110 shown in Figure 1, for example.
[0041] 3 , the terminal device 110 may present an editing interface 300 for editing media content 305 . In the interface 300 , the terminal device 110 may present text content 310 determined based on audio information of the media content 305 .
[0042] Furthermore, the terminal device 110 may receive a user selection of the first text 315 in the text content 310 and may receive a request to replace the first text 315 with the second text 320. In some embodiments, the second text 320 may be, for example, text input by the user or text selected by the user.
[0043] Accordingly, the terminal device 110 may send a request to the server 130 to replace the first text 315 with the second text 320. After receiving the request, the server 130 may use the timbre information associated with the audio segment corresponding to the first text 315 (also referred to as the first audio segment) in the text audio segment to re-upload the audio segment corresponding to the second text 320 (also referred to as the second audio segment). This process may also be referred to as a timbre cloning process.
[0044] In some embodiments, to improve the accuracy of voice cloning, the server 130 may determine at least one sentence associated with the first text 315. For example, the server 130 may identify a sentence in which the first text 315 is located, or may identify multiple sentences in a paragraph in which the first text is located.
[0045] Furthermore, the server 130 may determine timbre information associated with the first audio segment 315 based on the timbre of the at least one sentence.
[0046] For example, the server 130 may perform voice cloning using the timbre information of the sentence containing the first text 315 to generate a second audio segment corresponding to the second text 320. It should be understood that the voice cloning process may be performed using any appropriate method, examples of which may include but are not limited to machine learning-based voice cloning technology.
[0047] In some embodiments, in order to ensure the consistency between the regenerated audio segment and the original audio content, the server 130 can also determine the speaking rate information corresponding to the first audio segment, and can generate a second audio segment based on the speaking rate information and timbre information, so that the generated second audio segment can match the speaking rate information.
[0048] For example, the first text 315 may include a short word, while the second text 320 may include multiple words. In this case, if the generated audio segment is guaranteed to be consistent with the original audio segment time, it may cause a large change in the speaking speed, thereby affecting the quality of the audio content.
[0049] By ensuring that the speech rate of the generated audio content is consistent with the speech rate of the original audio content, the embodiments of the present disclosure can further improve the quality of voice cloning.
[0050] 2 , in block 230 , the server 130 adjusts a third audio segment corresponding to the first text in the background audio content based on the second audio segment.
[0051] In some embodiments, server 130 may determine speed adjustment information based on the first duration of the first audio segment and the second duration of the second audio segment. For example, while ensuring speech rate consistency, server 130 may determine that the duration of the second audio segment may be twice the original duration. Accordingly, server 130 may determine that the speed of the corresponding background audio portion needs to be slowed down by two times.
[0052] Furthermore, the server 130 may adjust the speed of the third audio segment based on the speed adjustment information. For example, the server 130 may slow down the speed of the third audio segment by two times so that the adjusted duration of the third audio segment matches the duration of the second audio segment.
[0053] At block 240 , the server 130 generates second media content based on the first media content, the second audio segment, and the adjusted third audio segment.
[0054] In some embodiments, the first media content may include, for example, picture content. Since the speed adjustment of the audio content mentioned above may cause the duration of the media content to change, the server 130 may further adjust the picture segment corresponding to the first audio segment in the picture content based on the second audio segment so that the duration of the picture segment matches the second audio segment.
[0055] For example, the server 130 may adjust the time length of the picture segment by inserting picture frames, extracting picture frames, or other appropriate methods, so that the time length of the picture segment can match the second audio segment.
[0056] Furthermore, the server 130 may generate the second media content by fusing the adjusted picture segment, the first audio content, and the adjusted third audio segment. For example, the server 130 may replace the picture segment corresponding to the first text in the first media content with the adjusted picture segment, and may replace the audio segment corresponding to the first text with a fusion of the second audio segment and the adjusted third audio segment.
[0057] In some embodiments, during the process of generating the media content, the server 130 may further fuse the second audio segment with the adjusted third audio segment to obtain intermediate audio content. Furthermore, the server 130 may further adjust the reverberation effect or volume of the intermediate audio content and generate the second media content based on the adjusted intermediate audio content.
[0058] In some examples, adjusting the volume of the intermediate audio content includes performing global volume equalization on the intermediate audio content. The terminal device 110 generates the second media content according to the adjusted intermediate audio content.
[0059] In some embodiments, the terminal device 110 may first determine a reverberation parameter based on the first media content, and then adjust the reverberation effect of the intermediate audio content based on the reverberation parameter.
[0060] For example, the reverberation matching model and the support vector classifier (SVC) model are the models that require engineering for the entire link. Their inputs and outputs are independent and can run in parallel. The reverberation matching module is a lightweight convolutional neural network (CNN) and long short-term memory (LSTM) module, and generally completes before the support vector classifier (SVC) model. The reverberation matching model is solely responsible for calculating reverberation parameters from the original user voice (outputting three scalars). The actual addition of audio reverberation is completed by the central processing unit (CPU) as the final step before the link output. The biggest difference between the overall link and the VC link is probably the logic for reverberation matching and volume equalization, as well as the additional robust model extractor (RMVPE f0extractor) for estimating high-pitched sounds in polyphonic music.
[0061] Based on the process described above, the embodiments of the present disclosure can support editing of media content by modifying the text corresponding to the media content, and can make the modified audio content more realistic, thereby improving the quality of the edited media content.
[0062] Example devices and equipment
[0063] Embodiments of the present disclosure also provide corresponding apparatuses for implementing the above-described methods or processes. FIG4 shows a schematic block diagram of an example apparatus 400 for audio processing according to certain embodiments of the present disclosure. Apparatus 400 may be implemented as or included in server 130. Each module / component in apparatus 400 may be implemented by hardware, software, firmware, or any combination thereof.
[0064] As shown in Figure 4, the device 400 includes an audio extraction module 410, which is configured to extract background audio content and text audio content corresponding to the text content from the first media content; an audio generation module 420, which is configured to generate a second audio segment corresponding to the second text based on a request to replace the first text in the text content with the second text, using the timbre information associated with the first audio segment of the text audio content, the first audio segment corresponding to the first text; an audio adjustment module 430, which is configured to adjust the third audio segment corresponding to the first text in the background audio content based on the second audio segment; and a media generation module 440, which is configured to generate the second media content based on the first media content, the second audio segment and the adjusted third audio segment.
[0065] In some embodiments, the audio generation module 420 is further configured to: determine speech rate information corresponding to the first audio segment; and generate a second audio segment based on the timbre information and the speech rate information, so that the second audio segment matches the speech rate information.
[0066] In some embodiments, the apparatus 400 further includes a timbre determination module configured to: determine at least one sentence associated with the first text; and determine timbre information associated with the first audio segment based on the timbre of the at least one sentence.
[0067] In some embodiments, the audio adjustment module 430 is further configured to: determine speed adjustment information based on the first time length of the first audio segment and the second time length of the second audio segment; and adjust the speed of the third audio segment based on the speed adjustment information.
[0068] In some embodiments, the first media content includes picture content, and the media generation module 440 is further configured to: adjust the picture segment in the picture content corresponding to the first audio segment based on the second audio segment; and generate the second media content by fusing the adjusted picture segment, the first audio content and the adjusted third audio segment.
[0069] In some embodiments, the media generation module 440 is further configured to: fuse the second audio segment and the adjusted third audio segment to obtain intermediate audio content; adjust the reverberation effect or volume of the intermediate audio content; and generate second media content based on the adjusted intermediate audio content.
[0070] In some embodiments, the media generation module 440 is further configured to: determine a reverberation parameter based on the first media content; and adjust a reverberation effect of the intermediate audio content based on the reverberation parameter.
[0071] In some embodiments, the media generation module 440 is further configured to perform global volume equalization on the intermediate audio content.
[0072] FIG5 shows a block diagram of an electronic device 500 in which one or more embodiments of the present disclosure may be implemented. It should be understood that the electronic device 500 shown in FIG5 is merely exemplary and should not be construed as limiting the functionality and scope of the embodiments described herein. The electronic device 500 shown in FIG5 can be used to implement the server 130 of FIG1 .
[0073] As shown in FIG5 , electronic device 500 is a general-purpose electronic device. Components of electronic device 500 may include, but are not limited to, one or more processors or processing units 510, memory 520, storage device 530, one or more communication units 540, one or more input devices 550, and one or more output devices 560. Processing unit 510 may be a real or virtual processor and is capable of performing various processes according to programs stored in memory 520. In a multi-processor system, multiple processing units execute computer-executable instructions in parallel to enhance the parallel processing capabilities of electronic device 500.
[0074] The electronic device 500 typically includes a plurality of computer storage media. Such media can be any accessible media that can be obtained by the electronic device 500, including but not limited to volatile and non-volatile media, removable and non-removable media. The memory 520 can be a volatile memory (e.g., registers, cache, random access memory (RAM)), a non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. The storage device 530 can be a removable or non-removable medium and can include a machine-readable medium, such as a flash drive, a disk, or any other medium that can be used to store information and / or data and can be accessed within the electronic device 500.
[0075] The electronic device 500 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in FIG. 5 , a disk drive for reading from or writing to a removable, non-volatile disk (e.g., a “floppy disk”) and an optical drive for reading from or writing to a removable, non-volatile optical disk may be provided. In these cases, each drive may be connected to a bus (not shown) by one or more data media interfaces. The memory 520 may include a computer program product 525 having one or more program modules configured to perform various methods or actions of various embodiments of the present disclosure.
[0076] The communication unit 540 enables communication with other electronic devices via a communication medium. Additionally, the functions of the components of the electronic device 500 can be implemented in a single computing cluster or multiple computing machines that can communicate via a communication connection. Thus, the electronic device 500 can operate in a networked environment using a logical connection with one or more other servers, a network personal computer (PC), or another network node.
[0077] Input device 550 may be one or more input devices, such as a mouse, keyboard, or trackball. Output device 560 may be one or more output devices, such as a display, a speaker, or a printer. Electronic device 500 may also communicate with one or more external devices (not shown) via communication unit 540 as needed, such as a storage device, a display device, or the like, with one or more devices that allow a user to interact with electronic device 500, or with any device that allows electronic device 500 to communicate with one or more other electronic devices (e.g., a network card, a modem, etc.). Such communication may be performed via an input / output (I / O) interface (not shown).
[0078] According to an exemplary implementation of the present disclosure, a computer-readable storage medium is provided, on which computer-executable instructions are stored, wherein the computer-executable instructions are executed by a processor to implement the method described above. According to an exemplary implementation of the present disclosure, a computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, and the computer-executable instructions are executed by a processor to implement the method described above.
[0079] Various aspects of the present disclosure are described herein with reference to flowcharts and / or block diagrams of methods, apparatuses, devices, and computer program products implemented according to the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.
[0080] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine, such that when these instructions are executed by the processing unit of the computer or other programmable data processing device, a device is generated that implements the functions / actions specified in one or more blocks in the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium, where these instructions cause the computer, programmable data processing device, and / or other device to operate in a specific manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing various aspects of the functions / actions specified in one or more blocks in the flowchart and / or block diagram.
[0081] Computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device so that a series of operational steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to implement the functions / actions specified in one or more boxes in the flowchart and / or block diagram.
[0082] The flow charts and block diagrams in the accompanying drawings show the possible architecture, functions and operations of the systems, methods and computer program products according to multiple implementations of the present disclosure. In this regard, each box in the flow chart or block diagram can represent a part for a module, program segment or instruction, and a part for a module, program segment or instruction comprises one or more executable instructions for realizing the logical function of the specification. In some alternative implementations, the functions marked in the box can also occur in a sequence different from that marked in the accompanying drawings. For example, two continuous boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be realized by a special hardware-based system that performs the function or action of the specification, or can be realized by a combination of special hardware and computer instructions.
[0083] While various implementations of the present disclosure have been described above, the foregoing description is intended to be illustrative, not exhaustive, and not limited to the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is selected to best explain the principles of the implementations, their practical applications, or improvements to existing technologies, or to enable others skilled in the art to understand the various implementations disclosed herein.
Claims
1. An audio processing method, comprising: extracting background audio content and text audio content corresponding to the text content from the first media content; generating, based on a request to replace a first text in the text content with a second text, a second audio segment corresponding to the second text using timbre information associated with a first audio segment of the text audio content, the first audio segment corresponding to the first text; Based on the second audio segment, adjusting a third audio segment in the background audio content corresponding to the first text; as well as Second media content is generated based on the first media content, the second audio segment, and the adjusted third audio segment.
2. The method according to claim 1, wherein generating a second audio segment corresponding to the second text using timbre information associated with the first audio segment of the text audio content comprises: Determining speech rate information corresponding to the first audio segment; as well as Based on the timbre information and the speech rate information, the second audio segment is generated so that the second audio segment matches the speech rate information.
3. The method according to claim 1 or 2, further comprising: determining at least one sentence associated with the first text; as well as The timbre information associated with the first audio segment is determined based on the timbre of the at least one sentence.
4. The method according to claim 1 or 2, wherein adjusting a third audio segment corresponding to the first text in the background audio content based on the second audio segment comprises: determining speed adjustment information based on a first time length of the first audio segment and a second time length of the second audio segment; as well as Based on the speed adjustment information, the speed of the third audio segment is adjusted.
5. The method according to claim 1 or 2, wherein the first media content comprises picture content, and generating the second media content based on the first media content, the second audio segment and the adjusted third audio segment comprises: Based on the second audio segment, adjusting the picture segment corresponding to the first audio segment in the picture content; as well as The second media content is generated by fusing the adjusted picture segment, the first audio content, and the adjusted third audio segment.
6. The method according to claim 1 or 2, wherein generating second media content based on the first media content, the second audio segment and the adjusted third audio segment comprises: fusing the second audio segment and the adjusted third audio segment to obtain intermediate audio content; Adjusting the reverberation effect or volume of the intermediate audio content; as well as The second media content is generated based on the adjusted intermediate audio content.
7. The method of claim 6, wherein adjusting the reverberation effect of the intermediate audio content comprises: determining a reverberation parameter based on the first media content; as well as The reverberation effect of the intermediate audio content is adjusted based on the reverberation parameter.
8. The method according to claim 6, wherein adjusting the volume of the intermediate audio content comprises: Global volume equalization is performed on the intermediate audio content.
9. An apparatus for audio processing, comprising: an audio extraction module configured to extract background audio content and text audio content corresponding to the text content from the first media content; an audio generation module configured to generate a second audio segment corresponding to the second text using timbre information associated with a first audio segment of the text audio content based on a request to replace a first text in the text content with a second text, the first audio segment corresponding to the first text; an audio adjustment module, configured to adjust a third audio segment corresponding to the first text in the background audio content based on the second audio segment; as well as The media generation module is configured to generate second media content based on the first media content, the second audio segment and the adjusted third audio segment.
10. An electronic device comprising: at least one processing unit; as well as At least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions causing the electronic device to perform the method according to any one of claims 1 to 8 when executed by the at least one processing unit.
11. A computer-readable storage medium having a computer program stored thereon, wherein the computer program can be executed by a processor to implement the method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Audio processing method and device, equipment and storage medium
CN120510855A
Voice processing method and device, electronic equipment and storage medium
CN115497451A
Outbound broadcast method and device, storage medium and electronic equipment
CN115910030A
Text-based voice editing method and system, electronic equipment and storage medium
CN115966196A
Audio optimization method and device, electronic equipment and storage medium
CN116013303A