Audio processing method and device, equipment and medium

Through the audio generation model, the characteristics and text content of the target object are extracted from the audio of multiple objects, the virtual sound is generated, and the target audio is synthesized, which solves the problem of noise interference in the prior art, realizes efficient audio processing in complex environments, and improves speech clarity and user experience.

CN120299468APending Publication Date: 2025-07-11LENOVO (BEIJING) LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510724498.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-30
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

In audio and video conferencing and real-time communication devices, noise interference has a negative impact on user experience. The existing audio noise reduction technology has limited effect when dealing with non-stationary noise and complex environments, with high computational complexity and poor real-time performance.

Method used

Through the audio generation model, the audio features and text content of the target object are extracted from the audio of multiple objects, the virtual sound corresponding to the target object is generated, the target audio is synthesized, the noise of non-target objects is removed, and the audio quality is improved.

Benefits of technology

Effectively remove noise in complex environments, improve speech clarity, improve user experience, retain the real sound of the target object, and reduce voice distortion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120299468A_ABST
    Figure CN120299468A_ABST
Patent Text Reader

Abstract

The invention provides an audio processing method and device, equipment and a medium, and is applied to the technical field of audio processing. The audio processing method comprises the steps that a first audio is obtained, and the first audio comprises first sounds generated by a plurality of objects in the same scene; generating a second audio based on the first audio and an audio generation model, the second audio including a second sound corresponding to a target object in the plurality of objects, the audio feature of the second sound corresponding to the audio feature of the first sound generated by the target object; and processing the first audio according to the second audio to obtain a target audio, the target audio comprising the first sound generated by the target object.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the technical field of audio processing, and in particular, to an audio processing method, apparatus, device, and medium. Background Art

[0002] In application scenarios such as audio-video conferences and real-time communication devices, noise interference will have a negative impact on the user experience. Audio noise reduction is a technology aimed at reducing or eliminating background noise in audio to improve audio quality, and it plays a very important role in improving speech clarity in noisy environments. Summary of the Invention

[0003] In view of this, the present disclosure provides an audio processing method, apparatus, device, and medium.

[0004] According to a first aspect of the present disclosure, there is provided an audio processing method, including: obtaining a first audio, the first audio including first sounds generated by multiple objects in the same scene; generating a second audio based on the first audio and an audio generation model, the second audio including a second sound corresponding to a target object among the multiple objects, and the audio feature of the second sound corresponding to the audio feature of the first sound generated by the target object; processing the first audio according to the second audio to obtain a target audio, the target audio including the first sound generated by the target object.

[0005] According to an embodiment of the present disclosure, generating the second audio based on the first audio and the audio generation model includes: extracting the audio feature and text content of the target object from the first audio; generating the second audio corresponding to the audio feature of the target object through the audio generation model, and the content of the second audio includes the text content.

[0006] According to an embodiment of the present disclosure, before extracting the audio feature and text content of the target object from the first audio, the method further includes: obtaining position information, the position information being used to indicate the position where each object among the multiple objects is located; determining at least one object as the target object according to the position where each object is located.

[0007] According to an embodiment of the present disclosure, generating the second audio based on the first audio and the audio generation model includes: extracting the text content corresponding to the target object from the first audio; generating the second audio corresponding to the target audio feature through the audio generation model, and the content of the second audio includes the text content, and the target audio feature corresponds to the audio feature of the first sound generated by the target object.

[0008] According to an embodiment of the present disclosure, before extracting the text content corresponding to the target object from the first audio, the method further includes: obtaining the audio feature of each object among the multiple objects; determining the similarity between the audio feature of each object and the target audio feature; determining at least one object as the target object according to the similarity.

[0009] According to an embodiment of the present disclosure, processing the first audio according to the second audio to obtain the target audio includes: determining a third audio according to the first audio and the second audio, where the third audio is the difference part between the first audio and the second audio; removing the same part as the third audio from the first audio to obtain the target audio.

[0010] According to an embodiment of the present disclosure, processing the first audio according to the second audio to obtain the target audio includes: extracting the same part as the second audio from the first audio to obtain the target audio.

[0011] A second aspect of the present disclosure provides an audio processing apparatus, including: an obtaining module, configured to obtain a first audio, where the first audio includes first sounds generated by multiple objects in the same scene; a generating module, configured to generate a second audio based on the first audio and an audio generation model, where the second audio includes a second sound corresponding to a target object among the multiple objects, and the audio feature of the second sound corresponds to the audio feature of the first sound generated by the target object; a processing module, configured to process the first audio according to the second audio to obtain a target audio, where the target audio includes the first sound generated by the target object.

[0012] A third aspect of the present disclosure provides an electronic device, including: one or more processors; a memory, configured to store one or more programs, and the one or more processors execute the one or more computer programs to implement: obtaining a first audio, where the first audio includes first sounds generated by multiple objects in the same scene; generating a second audio based on the first audio and an audio generation model, where the second audio includes a second sound corresponding to a target object among the multiple objects, and the audio feature of the second sound corresponds to the audio feature of the first sound generated by the target object; processing the first audio according to the second audio to obtain a target audio, where the target audio includes the first sound generated by the target object.

[0013] A fourth aspect of the present disclosure further provides a computer-readable storage medium, on which executable instructions are stored, and when the computer program or instructions are executed by a processor, the following is implemented: obtaining a first audio, where the first audio includes first sounds generated by multiple objects in the same scene; generating a second audio based on the first audio and an audio generation model, where the second audio includes a second sound corresponding to a target object among the multiple objects, and the audio feature of the second sound corresponds to the audio feature of the first sound generated by the target object; processing the first audio according to the second audio to obtain a target audio, where the target audio includes the first sound generated by the target object.

[0014] The fifth aspect of the present disclosure further provides a computer program product, including a computer program which, when executed by a processor, implements the above audio processing method.

[0015] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Through the following description of the embodiments of the present disclosure with reference to the drawings, the above and other objects, features, and advantages of the present disclosure will become clearer. In the drawings:

[0017] Figure 1A Schematically shows a flowchart of an audio processing method in the related art;

[0018] Figure 1B Schematically shows a schematic diagram of an audio processing method in the related art;

[0019] Figure 2 Schematically shows an application scenario diagram of the audio processing method according to an embodiment of the present disclosure;

[0020] Figure 3 Schematically shows a flowchart of the audio processing method according to an embodiment of the present disclosure;

[0021] Figure 4A Schematically shows a schematic diagram of the audio processing method according to an embodiment of the present disclosure;

[0022] Figure 4B Schematically shows a schematic diagram of the audio processing method according to another embodiment of the present disclosure;

[0023] Figure 5 Schematically shows a block diagram of the structure of the audio processing device according to an embodiment of the present disclosure; and

[0024] Figure 6 Schematically shows a block diagram of an electronic device suitable for implementing the audio processing method according to an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0025] Hereinafter, embodiments of the present disclosure will be described with reference to the drawings. However, it should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present disclosure. In the following detailed description, for the sake of explanation, many specific details are set forth to provide a comprehensive understanding of the embodiments of the present disclosure. However, obviously, one or more embodiments can also be implemented without these specific details. In addition, in the following description, descriptions of well-known structures and technologies are omitted to avoid unnecessarily confusing the concepts of the present disclosure.

[0026] The terms used herein are merely for describing specific embodiments and are not intended to limit the present disclosure. The terms "including", "comprising", etc. as used herein indicate the presence of features, steps, operations, and / or components, but do not preclude the presence or addition of one or more other features, steps, operations, or components.

[0027] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification, and should not be interpreted in an idealized or overly rigid manner.

[0028] In the case of using expressions such as "at least one of A, B, and C, etc.", generally, it should be interpreted according to the meaning commonly understood by those skilled in the art (for example, "a system having at least one of A, B, and C" should include, but is not limited to, a system having only A, only B, only C, having A and B, having A and C, having B and C, and / or having A, B, and C, etc.).

[0029] Embodiments of the present disclosure provide an audio processing method, apparatus, device, and medium. Before introducing the technical solutions provided by the embodiments of the present disclosure, the related technologies involved in the present disclosure will be described first.

[0030] In application scenarios such as audio - video conferences and real - time communication devices, noise interference will have a negative impact on the user experience. Audio noise reduction is a technology aimed at reducing or eliminating background noise in audio to improve audio quality, and it plays a very important role in improving speech clarity in noisy environments.

[0031] In one example, during a real - time call in a complex environment (such as a noisy office, subway car, etc.), there is usually a problem of background noise interference, resulting in a poor user call experience.

[0032] In order to suppress the background noise in the call voice and improve the clarity of the call. Referring to Figure 1A and Figure 1B , after obtaining the audio, noise estimation is performed through a noise reduction algorithm, and then noise cancellation is performed to obtain the target audio that meets the user's requirements. The noise reduction algorithm can be implemented by the following methods:

[0033] 1. Traditional signal processing techniques: such as spectral subtraction, Wiener filtering, etc. These methods mainly separate the speech signal and the noise signal through a mathematical model, and are suitable for suppressing stationary noise, but have limited effects when dealing with non - stationary noise (such as burst noise, human voice interference, etc.).

[0034] 2. Noise reduction algorithms based on statistical models: These algorithms estimate the noise characteristics through statistical methods and remove noise from speech signals. However, their noise reduction effects are often limited by the accuracy of the statistical model and it is difficult to adapt to complex and changing noise environments.

[0035] 3. Deep learning noise reduction methods: By training a neural network model to learn the feature differences between speech signals and noise signals, noise suppression is achieved. However, existing deep learning noise reduction methods may have problems such as high computational complexity, poor real-time performance, and insufficient model generalization ability.

[0036] Embodiments of the present disclosure provide an audio processing method, including: obtaining a first audio, the first audio including a first sound generated by multiple objects in the same scene; generating a second audio based on the first audio and an audio generation model, the second audio including a second sound corresponding to a target object among the multiple objects, and the audio features of the second sound corresponding to the audio features of the first sound generated by the target object; processing the first audio according to the second audio to obtain a target audio, the target audio including the first sound generated by the target object. By synthesizing a virtual sound of the second audio corresponding to the target object that needs to be retained in the real sound of the first audio, and retaining the real sound about the target object according to the synthesized virtual sound to obtain a target audio that only includes the real sound of the target object, audio data in different complex environments can be processed, improving the effect of audio processing.

[0037] Figure 2 FIG. schematically shows an application scenario diagram of the audio processing method according to an embodiment of the present disclosure.

[0038] As Figure 2 shown, the application scenario 100 according to this embodiment may include a first terminal device 101, a second terminal device 102, a third terminal device 103, a network 104, and a server 105. The network 104 is used to provide a medium for communication links between the first terminal device 101, the second terminal device 102, the third terminal device 103, and the server 105. The network 104 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.

[0039] Users can use the first terminal device 101, the second terminal device 102, and the third terminal device 103 to interact with the server 105 through the network 104 to receive or send messages, etc. Various communication client applications, such as audio processing applications, instant messaging tools, etc. (only for example), may be installed on the first terminal device 101, the second terminal device 102, and the third terminal device 103.

[0040] The first terminal device 101, the second terminal device 102, and the third terminal device 103 may be various electronic devices with a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, laptop computers, desktop computers, and the like.

[0041] The server 105 may be a server that provides various services, such as a background management server (only for example) that supports the websites browsed by the user using the first terminal device 101, the second terminal device 102, and the third terminal device 103. The background management server may analyze and process data such as user requests received, and feedback the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal device.

[0042] It should be noted that the audio processing method provided by the embodiments of the present disclosure can generally be executed by the server 105. Correspondingly, the audio processing device provided by the embodiments of the present disclosure can generally be set in the server 105. The audio processing method provided by the embodiments of the present disclosure can also be executed by a server or a server cluster different from the server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or the server 105. Correspondingly, the audio processing device provided by the embodiments of the present disclosure can also be set in a server or a server cluster different from the server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or the server 105.

[0043] In other embodiments, the audio processing method can also be directly executed by the processors of the first terminal device 101, the second terminal device 102, and the third terminal device 103.

[0044] It should be understood that Figure 2 the numbers of the terminal devices, the network, and the servers in

[0045] The audio processing method of the embodiments of the present disclosure can be applied to an electronic device with an audio processing device. The electronic device can be a mobile phone, a smart screen, a tablet computer, a wearable electronic device, a vehicle-mounted electronic device, an augmented reality (AR) device, a virtual reality (VR) device, a notebook computer, an ultra-mobile personal computer (UMPC), a netbook, a personal digital assistant (PDA), a projector, an augmented reality (AR) / virtual reality (VR) device, a media player, a television, etc. The embodiments of the present disclosure do not impose special restrictions on the specific form of the device. The operating system of the electronic device can include but is not limited to Flyme operating system, Android operating system, IOS operating system, Symbian operating system, Black Berry operating system, Windows Phone operating system, and so on.

[0046] Next, based on Figure 2 the described scenario, through Figures 3 to 4B the audio processing method of the embodiments of the present disclosure will be described in detail.

[0047] Figure 3 FIG. schematically shows a flowchart of the audio processing method according to the embodiments of the present disclosure.

[0048] As Figure 3 shown, the audio processing method of this embodiment includes operation S210 to operation S230.

[0049] In operation S210, a first audio is obtained. The first audio includes first sounds generated by multiple objects in the same scene.

[0050] In operation S220, a second audio is generated based on the first audio and an audio generation model. The second audio includes a second sound corresponding to a target object among the multiple objects, and the audio feature of the second sound corresponds to the audio feature of the first sound generated by the target object.

[0051] In operation S230, the first audio is processed according to the second audio to obtain a target audio. The target audio includes the first sound generated by the target object.

[0052] Exemplarily, the object can be an object, an animal, a person, etc. that can produce sound. The first sound can be an expression of the real sound of the object in the real environment.

[0053] The first audio may be a mixed audio formed by the first sounds generated by multiple objects in the same scenario. For example, in a meeting scenario, the first audio may be an audio including the voices of multiple speakers and the sounds generated by other objects in the environment; in audio recording, the first audio may be an audio including the sound of user A plucking a guitar, the sound of user A singing, and the sounds generated by other objects in the environment.

[0054] In a real-time call scenario, the first audio may also be an audio including the voice of user A and environmental noise.

[0055] The first audio can be obtained in real time or non-real time. For example, in a real-time call scenario between user A and user B, the first audio may be the voice of user B speaking and the sounds generated by other objects in the environment where user B is located, which are collected in real time by user B's audio collection device (e.g., microphone). The first audio can be an audio downloaded from an audio database. In the embodiments of the present disclosure, the acquisition method of the first audio is not specifically limited.

[0056] The target object can be at least one object among the multiple objects of the first audio. That is, the object to which the sound to be retained in the first audio belongs.

[0057] The audio feature can be a physical or perceptual feature used to describe sound. For example, the audio feature can be the voiceprint feature of the object's sound, can also be the timbre of the object's sound, can also be the intonation of the object's sound, and can also be the speech rate of the object's sound. In the embodiments of the present disclosure, the type of the audio feature is not specifically limited.

[0058] The second sound can be a virtual expression of the first sound of the target object. For example, the second sound can be a sound synthesized by an audio generation model according to the audio feature of the target object and having the same sound content as the first sound of the target object.

[0059] The target audio can be an audio formed according to the first sound of the target object. The target audio only contains the expression of the real sound of the target object. The target audio can be obtained by extracting the content in the first audio that has the same audio feature as the second audio; it can also be obtained by eliminating the content in the first audio that has different audio features from the second audio.

[0060] It can be understood that by synthesizing, through an audio generation model, the virtual sound of the second audio corresponding to the target object that needs to be retained in the real sound of the first audio, and retaining the real sound of the target object according to the synthesized virtual sound, a target audio only containing the real sound of the target object is obtained. This embodiment can process audio data in different complex environments and improve the audio processing effect.

[0061] In some embodiments, the audio generation model can be obtained through the following operations:

[0062] Obtain a plurality of sample data, where the sample data includes the audio features, text content of the sample object, and the original audio of the sample object corresponding to the text content. Among them, the audio features and text content of the sample object can be extracted according to the original audio of each sample object.

[0063] Input the plurality of sample data into the audio generation model to obtain a predicted synthesized audio corresponding to the text content corresponding to each sample.

[0064] Adjust the parameters of the audio generation model according to the difference between the predicted synthesized audio and the original audio until the difference between the predicted synthesized audio and the original audio is less than a threshold or the model converges, to obtain a trained audio synthesis model.

[0065] As described above, in operation S220, a second audio is generated based on the first audio and the audio generation model. In one implementable manner, this operation can further include operations S221a to S222a.

[0066] In operation S221a, extract the audio features and text content of the target object from the first audio.

[0067] In operation S222a, generate a second audio corresponding to the audio features of the target object through the audio generation model, and the content of the second audio includes the text content.

[0068] Exemplarily, the text content can be the speech content corresponding to each target object in the first audio. For example, the content spoken by user A during a call.

[0069] In one example, taking the real-time call between user A and user B as an example, obtain the first audio of user A during the call (the voice of user A speaking and the sounds generated by other objects in the environment where user A is located), and extract the audio features and the content of what user A said (text content) from the first audio. Input the audio features and the content of what user A said (text content) into the audio generation model to obtain a synthesized audio with the audio features of user A and the same content of what is said.

[0070] As described above, in operation S221a, extract the audio features and text content of the target object from the first audio. In one implementable manner, before performing this operation, operations S2211a to S2212a can be further included.

[0071] In operation S2211a, obtain position information, where the position information is used to indicate the position where each object among a plurality of objects is located.

[0072] In operation S2212a, at least one object is determined as a target object according to the position where each object is located.

[0073] Exemplarily, the position information may be the relative position between multiple objects. The position information may be determined based on the image information of multiple objects during the acquisition of the first audio by an image acquisition device. The position information may also be the relative position between multiple objects determined by the magnitudes of the sounds generated by multiple objects in the first audio.

[0074] The target object to which the sound to be retained belongs may be determined according to the position information of multiple objects that generate multiple first sounds in the real environment. That is, different positions indicate the importance of the sounds generated by objects in the current audio acquisition environment. For example, in a video recording of user A singing, user A is located in the middle of the video image, and user B is located at the edge of the video image, then user A is the target object.

[0075] At different times, the target object may be the same or different. For example, during a call between user A and user B, for the first audio about user A, the target object may always be user A. In a meeting where multiple people are speaking, at the first moment, the target object in the first audio is object 1, and at the second moment, the target object in the first audio is object 2.

[0076] In one example, in a scenario of a multi-person meeting, image information (video / photo) about the multi-person meeting may be captured by a camera. At the first moment, in the image information, object 1 is located at the speaker's position, and object 2 is located at the audience's position. At this time, in the first audio where the voices of object 1 and object 2 appear simultaneously, according to the image information at the first moment, the target object can be determined as object 1, and object 2 is a non-target object. At the second moment, in the image information, object 2 is located at the speaker's position, and object 1 is located at the audience's position. At this time, in the first audio where the voices of object 1 and object 2 appear simultaneously, according to the image information at the second moment, the target object can be determined as object 2, and object 1 is a non-target object.

[0077] In some other embodiments, description information about the target object is obtained, and at least one object among multiple objects is determined as the target object according to the description information, where the description information characterizes the features of the target object.

[0078] Exemplarily, the description information may be input by voice or by text. The embodiments of the present disclosure do not make specific limitations on this.

[0079] In one example, in the audio recording of user A singing, in the environment where user A is located, there are the barks of a puppy and the voices of user B and user C talking. At this time, descriptive information about the target object is input through voice: Keep the singing voice. Then the target object can be determined to be user A.

[0080] In other embodiments, application information regarding the first audio collection is obtained, and at least one object among multiple objects is determined to be the target object according to that information, where the application information characterizes the scene in which the first audio is generated.

[0081] In one example, during the call between user A and user B, the first audio includes the voice of user A speaking, and there are also the voices of multiple users discussing in the environment where user A is located. If the first audio of user A is the audio generated during the operation of the call application, then the target object can be determined to be user A.

[0082] In other embodiments, the location information, descriptive information, and application information can be combined simultaneously to determine the target object.

[0083] As described above, in operation S220, a second audio is generated based on the first audio and the audio generation model. In another implementable manner, this operation can further include operation S221b to operation S222b.

[0084] In operation S221b, the text content corresponding to the target object is extracted from the first audio.

[0085] In operation S222b, a second audio corresponding to the target audio feature is generated through the audio generation model. The content of the second audio includes the text content, and the target audio feature corresponds to the audio feature of the first sound generated by the target object.

[0086] When the target audio feature is known, the audio generation model can be pre-trained according to the target audio feature. In subsequent processing, only the speech content (text content) of the target object needs to be extracted from the first audio. By inputting the speech content into the audio generation model, a synthetic audio (second audio) with the target audio feature and the same speech content as the target object can be synthesized.

[0087] Among them, knowing the audio feature of the target can be that the target object is determined by the user in advance before processing the first audio, and the audio feature of the target object is determined.

[0088] In one example, the target audio features can be pre-extracted, such as the audio features of user B, and the model to be trained can be trained with the audio features of user B to obtain an audio generation model. In this way, when user A and user B are on a call, if there are sounds emitted by different objects in the environment where user B is located, then the call sound needs to be denoised. During the denoising process, the text content corresponding to the sound with the same audio features as those of user B is extracted from the sounds emitted by different objects, and then a second audio is generated based on the extracted text content and the audio generation model, and the first audio transmitted by user B to user A is denoised with the second audio. Thus, during the call, user B can transmit a target audio containing only the voice of user B to user A.

[0089] In one example, before user A and user B make a call, user B has previously determined the operation for processing the call audio, and the audio features of user A have been extracted during the previous process of processing the call audio between user A and user B. The model to be trained is trained with the audio features of user A to obtain an audio generation model. Therefore, during the call between user A and user B, the target object is user A and the audio features of user A are known. For the voices of multiple users discussing in the environment where user A is located during the call between user A and user B, the first audio transmitted by user A to user B includes the voice of user A speaking and noise. Then, the content of user A's speech can be extracted based on the target object: user A and the audio features of user A. Then, a second audio is generated based on the extracted content of user A's speech and the audio generation model, and the first audio transmitted by user A to user B is denoised with the second audio. Thus, during the call, user B can receive a target audio containing only the voice of user A sent by user A.

[0090] As described above, in operation S221b, the text content corresponding to the target object is extracted from the first audio. In one implementable manner, before performing this operation, operations S2211b to S2213b may be further included.

[0091] In operation S2211b, the audio features of each object among multiple objects are obtained.

[0092] In operation S2212b, the similarity between the audio features of each object and the target audio features is determined.

[0093] In operation S2213b, at least one object is determined as the target object according to the similarity.

[0094] In one example, given the target audio features of a known target object, before extracting the text content of the target object in the first audio, the first audio features of multiple objects in the first audio can be obtained, the similarity between the target audio features and the first audio features of each object can be calculated, and when the similarity between the target audio features and the first audio features of a certain object is greater than a preset threshold (for example, 95%), this object can be regarded as the target object in the first audio. And the text content corresponding to this object can be extracted as the text content of the target object.

[0095] It should be noted that the embodiments of the present disclosure do not specifically limit the size of the preset threshold, and the size of the preset threshold can be adjusted according to the actual situation.

[0096] The real voice contains the complete physiological acoustic features (such as formants) of the target object, while the synthesized voice depends on the accuracy of the audio synthesis model and may have subtle spectral distortions. The audio synthesized by the audio generation model may have problems such as over-smoothing of the spectrum and unnatural prosody. For example, for the content spoken by the object after hesitation, the hesitant part may be missing in the synthesized audio. Therefore, retaining the most real voice of the target object is more conducive to ensuring the sound quality after audio processing, thereby improving the user experience.

[0097] Figure 4A A schematic diagram of an audio processing method according to an embodiment of the present disclosure is schematically shown.

[0098] As described above, in operation S230, the first audio is processed according to the second audio to obtain the target audio. In one implementable manner, this operation may further include operations S231a to S232a.

[0099] In operation S231a, according to the first audio and the second audio, a third audio is determined, and the third audio is the different part between the first audio and the second audio.

[0100] In operation S232a, the same part in the first audio as in the third audio is removed to obtain the target audio.

[0101] Exemplarily, the third audio may be other sounds in the first audio except for the first sound of the target object. That is, the third audio is the real sound generated by other objects except the target object among multiple objects.

[0102] The first audio includes the third audio and the target audio. As Figure 4AAs shown, the first audio is compared with the synthesized second audio of the target object to obtain the noise part (the third audio) in the first audio excluding the voice of the target object. By using signal processing techniques (such as spectrum analysis, filter design, etc.), the third audio in the first audio can be eliminated to obtain the audio of the true voice of the target object (the target audio).

[0103] Figure 4B FIG. schematically shows a schematic diagram of an audio processing method according to another embodiment of the present disclosure.

[0104] As described above, in operation S230, the first audio is processed according to the second audio to obtain the target audio. In another implementable manner, this operation may further include operation S232b.

[0105] In operation S232b, the same part in the first audio as that in the second audio is extracted to obtain the target audio.

[0106] Exemplarily, the target audio is the audio formed by the true voice corresponding to the target object. Then, there are the same text content and the same audio features between the second audio and the target audio.

[0107] In one example, as Figure 4B shown, the synthesized second audio of the target object is used as a reference, and the same part (text information + audio features) of the first audio and the second audio is extracted to obtain the audio of the true voice of the target object (the target audio).

[0108] It can be understood that by using the synthesized audio as a reference for the audio elimination part, it is convenient to better identify the part to be eliminated in the audio, eliminate the sound specifically, and achieve the accuracy of separation between the eliminated audio and the retained audio. At the same time, retaining the audio corresponding to the true voice that is the same as the synthesized audio in the retained audio can avoid the problem of speech distortion caused by audio processing.

[0109] Based on the above audio processing method, the present disclosure further provides an audio processing device. The following will be combined with Figure 5 to describe this device in detail.

[0110] Figure 5 FIG. schematically shows a structural block diagram of an audio processing device according to an embodiment of the present disclosure.

[0111] As Figure 5 shown, the audio processing device 300 in this embodiment includes an obtaining module 310, a generating module 320, and a processing module 330.

[0112] The obtaining module 310 is configured to obtain a first audio, where the first audio includes first sounds generated by multiple objects in the same scene. In one embodiment, the obtaining module 310 may be configured to perform the operation S210 described above, which will not be elaborated herein.

[0113] The generating module 320 is configured to generate a second audio based on the first audio and an audio generation model, where the second audio includes a second sound corresponding to a target object among the multiple objects, and the audio feature of the second sound corresponds to the audio feature of the first sound generated by the target object. In one embodiment, the generating module 320 may be configured to perform the operation S220 described above, which will not be elaborated herein.

[0114] The processing module 330 is configured to process the first audio according to the second audio to obtain a target audio, where the target audio includes the first sound generated by the target object. In one embodiment, the processing module 330 may be configured to perform the operation S230 described above, which will not be elaborated herein.

[0115] According to an embodiment of the present disclosure, any multiple of the obtaining module 310, the generating module 320, and the processing module 330 may be combined and implemented in one module, or any one of them may be split into multiple modules. Or, at least part of the functions of one or more of these modules may be combined with at least part of the functions of other modules and implemented in one module. According to an embodiment of the present disclosure, at least one of the obtaining module 310, the generating module 320, and the processing module 330 may be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on chip, a system on substrate, a system on package, an application specific integrated circuit (ASIC), or may be implemented by any other reasonable way of integrating or packaging circuits, etc., in hardware or firmware, or implemented in any one of the three implementation manners of software, hardware, and firmware, or in an appropriate combination of any several of them. Or, at least one of the obtaining module 310, the generating module 320, and the processing module 330 may be at least partially implemented as a computer program module, and when the computer program module runs, it may perform corresponding functions.

[0116] Figure 6 A block diagram of an electronic device suitable for implementing the audio processing method according to an embodiment of the present disclosure is schematically shown.

[0117] As Figure 6As shown, an electronic device 400 according to an embodiment of the present disclosure includes a processor 401, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 402 or a program loaded from a storage section 408 into a random access memory (RAM) 403. The processor 401 may include, for example, a general microprocessor (such as a CPU), an instruction set processor, and / or a related chipset, and / or a dedicated microprocessor (such as an application specific integrated circuit (ASIC)), etc. The processor 401 may also include on-board memory for caching purposes. The processor 401 may include a single processing unit or multiple processing units for performing different actions of a method flow according to an embodiment of the present disclosure.

[0118] In the RAM 403, various programs and data required for the operation of the electronic device 400 are stored. The processor 401, the ROM 402, and the RAM 403 are connected to each other via a bus 404. The processor 401 performs various operations of a method flow according to an embodiment of the present disclosure by executing a program in the ROM 402 and / or the RAM 403. It should be noted that the program may also be stored in one or more memories other than the ROM 402 and the RAM 403. The processor 401 may also perform various operations of a method flow according to an embodiment of the present disclosure by executing a program stored in one or more memories.

[0119] According to an embodiment of the present disclosure, the electronic device 400 may further include an input / output (I / O) interface 405, and the input / output (I / O) interface 405 is also connected to the bus 404. The electronic device 400 may further include one or more of the following components connected to the I / O interface 405: an input section 406 including a keyboard, a mouse, etc.; an output section 407 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 408 including a hard disk, etc.; and a communication section 409 including a network interface card such as a LAN card, a modem, etc. The communication section 409 performs communication processing via a network such as the Internet. A drive 410 is also connected to the I / O interface 405 as needed. A removable medium 411, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 410 as needed so that a computer program read from it can be installed into the storage section 408 as needed.

[0120] The present disclosure also provides a computer-readable storage medium, which may be included in the device / device / system described in the above embodiments; or may exist separately without being assembled into the device / device / system. The above computer-readable storage medium carries one or more programs, and when the above one or more programs are executed, a method according to an embodiment of the present disclosure is implemented.

[0121] According to an embodiment of the present disclosure, the computer-readable storage medium may be a non-volatile computer-readable storage medium, which may include, for example, but not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In the present disclosure, the computer-readable storage medium may be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device. For example, according to an embodiment of the present disclosure, the computer-readable storage medium may include the ROM 402 and / or RAM 403 described above and / or one or more memories other than the ROM 402 and RAM 403.

[0122] An embodiment of the present disclosure further includes a computer program product, which includes a computer program that contains program code for executing the method shown in the flowchart. When the computer program product runs in a computer system, the program code is used to enable the computer system to implement the audio processing method provided by the embodiment of the present disclosure.

[0123] When the computer program is executed by the processor 401, it executes the above functions defined in the system / apparatus of the embodiment of the present disclosure. According to an embodiment of the present disclosure, the above-described systems, apparatuses, modules, units, etc. can be implemented by computer program modules.

[0124] In one embodiment, the computer program may rely on tangible storage media such as optical storage devices and magnetic storage devices. In another embodiment, the computer program may also be transmitted and distributed in the form of a signal on a network medium and be downloaded and installed through the communication part 409, and / or be installed from the removable medium 411. The program code included in the computer program can be transmitted by any suitable network medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.

[0125] In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 409, and / or be installed from the removable medium 411. When the computer program is executed by the processor 401, it executes the above functions defined in the system of the embodiment of the present disclosure. According to an embodiment of the present disclosure, the above-described systems, devices, apparatuses, modules, units, etc. can be implemented by computer program modules.

[0126] In accordance with embodiments of the present disclosure, program code for executing the computer programs provided by the embodiments of the present disclosure can be written in any combination of one or more programming languages. Specifically, these computing programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. The programming languages include, but are not limited to, such as Java, C++, Python, the "C" language, or similar programming languages. The program code can be executed entirely on the user's computing device, partially on the user's device, partially on a remote computing device, or entirely on a remote computing device or server. In cases involving a remote computing device, the remote computing device can be connected to the user's computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., by connecting through the Internet using an Internet service provider).

[0127] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks can occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and the combination of blocks in the block diagram or flowchart, can be implemented using a dedicated hardware-based system for performing the specified functions or operations, or can be implemented using a combination of dedicated hardware and computer instructions.

[0128] Those skilled in the art can understand that the features recited in the various embodiments and / or claims of the present disclosure can be combined or combined in various ways, even if such combinations or combinations are not explicitly recited in the present disclosure. In particular, without departing from the spirit and teachings of the present disclosure, the features recited in the various embodiments and / or claims of the present disclosure can be combined and combined in various ways. All such combinations and / or combinations fall within the scope of the present disclosure.

[0129] The embodiments of the present disclosure have been described above. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of the present disclosure. Although the embodiments have been described separately above, this does not mean that the measures in each embodiment cannot be used advantageously in combination. The scope of the present disclosure is defined by the appended claims and their equivalents. Without departing from the scope of the present disclosure, those skilled in the art can make various substitutions and modifications, and these substitutions and modifications should fall within the scope of the present disclosure.

Claims

1. An audio processing method, comprising: Obtaining a first audio, where the first audio includes first sounds generated by multiple objects in the same scene; Generating a second audio based on the first audio and an audio generation model, where the second audio includes a second sound corresponding to a target object among the multiple objects, and the audio features of the second sound correspond to the audio features of the first sound generated by the target object; Processing the first audio according to the second audio to obtain a target audio, where the target audio includes the first sound generated by the target object.

2. The method according to claim 1, wherein generating a second audio based on the first audio and an audio generation model includes: Extracting the audio features and text content of the target object from the first audio; Generating the second audio corresponding to the audio features of the target object through the audio generation model, where the content of the second audio includes the text content.

3. The method according to claim 2, before extracting the audio features and text content of the target object from the first audio, the method further includes: Obtaining position information, where the position information is used to indicate the position of each object among the multiple objects; Determining at least one object as the target object according to the position where each object is located.

4. The method according to claim 1, wherein generating a second audio based on the first audio and an audio generation model includes: Extracting the text content corresponding to the target object from the first audio; Generating the second audio corresponding to the target audio features through the audio generation model, where the content of the second audio includes the text content, and the target audio features correspond to the audio features of the first sound generated by the target object.

5. The method according to claim 4, before extracting the text content corresponding to the target object from the first audio, the method further includes: Obtaining the audio features of each object among the multiple objects; Determining the similarity between the audio features of each object and the target audio features; Determining at least one object as the target object according to the similarity.

6. The method according to any one of claims 1-5, wherein processing the first audio according to the second audio to obtain a target audio includes: Determining a third audio according to the first audio and the second audio, where the third audio is the different part between the first audio and the second audio; Removing the same part in the first audio as that in the third audio to obtain the target audio.

7. The method according to any one of claims 1-5, wherein processing the first audio according to the second audio to obtain a target audio includes: Extracting the same part in the first audio as that in the second audio to obtain the target audio.

8. An audio processing apparatus, comprising: An obtaining module, configured to obtain a first audio, where the first audio includes first sounds generated by multiple objects in the same scene; A generating module, configured to generate a second audio based on the first audio and an audio generation model, where the second audio includes a second sound corresponding to a target object among the multiple objects, and an audio feature of the second sound corresponds to an audio feature of a first sound generated by the target object; A processing module, configured to process the first audio according to the second audio to obtain a target audio, where the target audio includes the first sound generated by the target object.

9. An electronic device, comprising: One or more processors; A memory, configured to store one or more computer programs, Characterized in that the one or more processors execute the one or more computer programs to implement: obtaining a first audio, where the first audio includes first sounds generated by multiple objects in the same scene; Generating a second audio based on the first audio and an audio generation model, where the second audio includes a second sound corresponding to a target object among the multiple objects, and an audio feature of the second sound corresponds to an audio feature of a first sound generated by the target object; Processing the first audio according to the second audio to obtain a target audio, where the target audio includes the first sound generated by the target object.

10. A computer-readable storage medium having a computer program or instructions stored thereon, characterized in that, When the computer program or instruction is executed by a processor, it implements: obtaining a first audio, where the first audio includes first sounds generated by multiple objects in the same scene; generating a second audio based on the first audio and an audio generation model, where the second audio includes a second sound corresponding to a target object among the multiple objects, and an audio feature of the second sound corresponds to an audio feature of a first sound generated by the target object; Processing the first audio according to the second audio to obtain a target audio, where the target audio includes the first sound generated by the target object.