Audio processing method and apparatus, and storage medium
By adjusting multi-channel audio signals with single-channel audio signals to align sound levels, the method addresses the misalignment of sound sources in earpiece mode, enhancing spatial audio communication quality and user experience.
Patent Information
- Application Number
- PCT/CN2023/143453
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-12-29
- Publication Date
- 2025-07-03
AI Technical Summary
Existing audio processing technologies in terminals fail to support spatial audio communication in earpiece mode, leading to misalignment of sound sources and poor user experience due to inability to maintain relative positional relationships between sound sources in the acoustic scene.
The method involves processing audio signals in earpiece mode by combining multi-channel audio signals in stereo format with single-channel audio signals to adjust the position of the target sound source, using phase-based adjustments to align sound levels across channels, ensuring accurate spatial audio representation.
This approach enhances the spatial audio experience by accurately positioning the target sound source, improving user communication quality and experience in earpiece mode.
Smart Images

Figure CN2023143453_03072025_PF_FP_ABST
Abstract
Description
Audio processing method, device and storage medium Technical Field
[0001] The present disclosure relates to the field of audio technology, and in particular to an audio processing method, device, and storage medium. Background Art
[0002] Currently, the basic functions of terminals include mono calls, which can collect audio signals of users with high signal-to-noise ratio in handset mode. The collected audio signals contain almost only the target audio components, and try to eliminate ambient sounds and other audio components unrelated to the target audio.
[0003] With the rise of spatial audio services, devices are increasingly supporting spatial audio capture. The goal of spatial audio capture technology is to reproduce the complete sound field in the environment where the device is located, while preserving the relative positions of all sound sources within the field.
[0004] Summary of the Invention
[0005] In handset mode, the terminal supports single-channel audio communication and cannot support spatial audio communication.
[0006] The embodiments of the present disclosure provide an audio processing method, an audio processing device, and a storage medium.
[0007] According to a first aspect of an embodiment of the present disclosure, an audio processing method is proposed, the method comprising: in response to a terminal being in a handset mode, obtaining a first audio signal and a second audio signal; wherein the first audio signal is a multi-channel audio signal collected based on a stereo format, and the first audio signal includes audio signals of multiple sound sources, the multiple sound sources including a target sound source, and the second audio signal is a single-channel audio signal collected based on a non-stereo format, and the second audio signal is an audio signal of the target sound source; and using the second audio signal, adjusting a portion of the audio signal corresponding to the target sound source in the first audio signal.
[0008] According to a second aspect of an embodiment of the present disclosure, an audio processing device is proposed, which includes: an acquisition module for acquiring a first audio signal and a second audio signal in response to a terminal being in a handset mode; wherein the first audio signal is a multi-channel audio signal collected based on a stereo format, and the first audio signal includes audio signals of multiple sound sources, the multiple sound sources including a target sound source, and the second audio signal is a single-channel audio signal collected based on a non-stereo format, and the second audio signal is an audio signal of the target sound source; and a processing module for using the second audio signal to adjust a portion of the audio signal corresponding to the target sound source in the first audio signal.
[0009] According to a third aspect of an embodiment of the present disclosure, an electronic device is provided, comprising: a memory for storing instructions; and
[0010] The processor is configured to call the instructions stored in the memory to execute the first aspect and any one of the audio processing methods in the first aspect.
[0011] According to a fourth aspect of an embodiment of the present disclosure, a storage medium is proposed, comprising: instructions stored in the storage medium, and when the instructions are executed by a processor, the first aspect and any one of the audio processing methods in the first aspect are executed.
[0012] The present disclosure obtains a multi-channel audio signal collected in a stereo format and a single-channel audio signal collected in a non-stereo format when the terminal is in handset mode, wherein the multi-channel audio signal includes multiple sound sources including a target sound source, and the single-channel audio signal is the audio signal of the target sound source. The single-channel audio signal can be used to adjust the portion of the multi-channel audio signal that is the target sound source, so that the portion of the multi-channel audio signal that is the target sound source is more reasonable and is positioned at a desired location, thereby realizing communication based on spatial audio. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure, the following drawings required for describing the embodiments are introduced. The following drawings are merely some embodiments of the present disclosure and do not impose specific limitations on the protection scope of the present disclosure.
[0014] FIG1 is a schematic diagram showing a communication system architecture according to an embodiment of the present disclosure.
[0015] FIG2 is a schematic diagram showing an interaction of a communication method according to an embodiment of the present disclosure.
[0016] FIG3 a is a flow chart of an audio processing method according to an embodiment of the present disclosure.
[0017] FIG3 b is a flow chart of an audio processing method according to an embodiment of the present disclosure.
[0018] FIG4 is a flow chart of an audio processing method according to an embodiment of the present disclosure.
[0019] FIG5 is a schematic diagram showing an interaction of a communication method according to an embodiment of the present disclosure.
[0020] FIG6 a is a schematic structural diagram of a terminal according to an embodiment of the present disclosure.
[0021] FIG6 b is a schematic structural diagram of a terminal according to an embodiment of the present disclosure.
[0022] Fig. 7a is a schematic structural diagram of a communication device according to an exemplary embodiment.
[0023] FIG7 b is a schematic diagram showing a chip structure according to an exemplary embodiment.
[0024] Fig. 8 is a block diagram showing an audio processing apparatus according to an exemplary embodiment. DETAILED DESCRIPTION
[0025] The embodiments of the present disclosure provide an audio processing method, an audio processing device, and a storage medium.
[0026] In a first aspect, an embodiment of the present disclosure proposes an audio processing method, the method comprising: in response to a terminal being in a handset mode, obtaining a first audio signal and a second audio signal; wherein, the first audio signal is a multi-channel audio signal collected based on a stereo format, and the first audio signal includes audio signals of multiple sound sources, the multiple sound sources including a target sound source, the second audio signal is a single-channel audio signal collected based on a non-stereo format, and the second audio signal is an audio signal of the target sound source; using the second audio signal, adjusting a portion of the audio signal corresponding to the target sound source in the first audio signal.
[0027] In the above embodiment, when the terminal is in handset mode, a multi-channel audio signal collected in a stereo format and a single-channel audio signal collected in a non-stereo format are obtained, wherein the multi-channel audio signal includes multiple sound sources including a target sound source, and the single-channel audio signal is the audio signal of the target sound source. The single-channel audio signal can be used to adjust the portion of the multi-channel audio signal that is the target sound source, so that the portion of the multi-channel audio signal that is the target sound source is more reasonable and is positioned at the desired location, thereby realizing a call based on spatial audio.
[0028] In some optional embodiments of the first aspect, a difference between a frequency of a portion of the audio signal corresponding to the target sound source in the first audio signal and a frequency of the second audio signal is less than or equal to a threshold.
[0029] In the above embodiment, the frequency difference between the portion of the audio signal corresponding to the target sound source in the first audio signal and the second audio signal is less than or equal to the threshold. That is, when a portion of the audio signal in the first audio signal has a frequency difference with the second audio signal that is less than or equal to the threshold, it can be determined as the portion of the audio signal corresponding to the target sound source, thereby efficiently determining the portion of the audio signal corresponding to the target sound source from the first audio signal.
[0030] In some optional embodiments of the first aspect, the target sound source is at a first position of the terminal, and using the second audio signal to adjust the portion of the audio signal corresponding to the target sound source in the first audio signal includes at least one of the following: using the second audio signal to weaken the portion of the audio signal corresponding to the target sound source in the first audio signal of the channel corresponding to the first position; using the second audio signal to enhance the portion of the audio signal corresponding to the target sound source in the first audio signal of the channel corresponding to other positions other than the first position.
[0031] In the above embodiment, when the target sound source is located in a certain direction of the terminal, the portion of the target sound source contained in the audio signal of the channel corresponding to that direction can be weakened. Alternatively, the portion of the target sound source contained in the audio signal of channels other than the channel corresponding to that direction can be strengthened. This ensures that the amplitude of the target sound source is as close as possible on the side where the target sound source is located and on the side where the target sound source is not located. In other words, the target sound source can be positioned in the center of the space, improving the user experience during calls.
[0032] In some optional embodiments of the first aspect, the weakening of the portion of the audio signal corresponding to the target sound source in the first audio signal of the channel corresponding to the first orientation includes at least one of the following: using a second audio signal with the same phase as the portion of the audio signal to subtract the portion of the audio signal from the portion of the audio signal to obtain a weakened portion of the audio signal; using a second audio signal with an opposite phase to the portion of the audio signal to add the portion of the audio signal to obtain a weakened portion of the audio signal.
[0033] In the above embodiment, the partial audio signal corresponding to the target sound source is weakened by subtracting it from a second audio signal with the same phase as the partial audio signal, or adding it to a second audio signal with the opposite phase to the partial audio signal, so as to efficiently obtain the weakened target sound source.
[0034] In some optional embodiments of the first aspect, enhancing the partial audio signal corresponding to the target sound source in the first audio signal of the channel corresponding to other directions other than the first direction includes at least one of the following: using a second audio signal with the same phase as the partial audio signal to add to the partial audio signal to obtain an enhanced partial audio signal; using a second audio signal with an opposite phase to the partial audio signal to subtract from the partial audio signal to obtain the enhanced partial audio signal.
[0035] In the above embodiment, the part of the audio signal corresponding to the target sound source is enhanced by adding a second audio signal with the same phase as the part of the audio signal, or subtracting a second audio signal with the opposite phase to the part of the audio signal, so as to efficiently obtain the enhanced target sound source.
[0036] According to a second aspect, an audio processing device is provided, comprising: an acquisition module for acquiring a first audio signal and a second audio signal in response to a terminal being in a handset mode; wherein the first audio signal is a multi-channel audio signal collected based on a stereo format, and the first audio signal includes audio signals of multiple sound sources, the multiple sound sources including a target sound source; the second audio signal is a single-channel audio signal collected based on a non-stereo format, and the second audio signal is an audio signal of the target sound source; and a processing module for adjusting, using the second audio signal, a portion of the audio signal corresponding to the target sound source in the first audio signal.
[0037] In some optional embodiments of the second aspect, a difference between a frequency of a portion of an audio signal corresponding to a target sound source in the first audio signal and a frequency of the second audio signal is less than or equal to a threshold.
[0038] In some optional embodiments of the second aspect, the target sound source is at the first position of the terminal, and the processing module uses the second audio signal to adjust the portion of the audio signal corresponding to the target sound source in the first audio signal using at least one of the following methods: using the second audio signal to weaken the portion of the audio signal corresponding to the target sound source in the first audio signal of the channel corresponding to the first position; using the second audio signal to enhance the portion of the audio signal corresponding to the target sound source in the first audio signal of the channel corresponding to other positions other than the first position.
[0039] In some optional embodiments of the second aspect, the processing module uses at least one of the following methods to weaken the partial audio signal corresponding to the target sound source in the first audio signal of the channel corresponding to the first orientation: using a second audio signal with the same phase as the partial audio signal to subtract it from the partial audio signal to obtain a weakened partial audio signal; using a second audio signal with an opposite phase to the partial audio signal to add it to the partial audio signal to obtain a weakened partial audio signal.
[0040] In some optional embodiments of the second aspect, the processing module enhances the partial audio signal corresponding to the target sound source in the first audio signal of the channel corresponding to other directions other than the first direction by at least one of the following methods, including: using a second audio signal with the same phase as the partial audio signal to add to the partial audio signal to obtain an enhanced partial audio signal; using a second audio signal with an opposite phase to the partial audio signal to subtract from the partial audio signal to obtain an enhanced partial audio signal.
[0041] In a third aspect, an electronic device is provided, comprising: a memory for storing instructions; and a processor for calling the instructions stored in the memory to execute the first aspect and any one of the audio processing methods in the first aspect.
[0042] In a fourth aspect, a storage medium is provided, comprising: instructions stored in the storage medium, and when the instructions are executed by a processor, the audio processing method according to the first aspect and any one of the audio processing methods in the first aspect is executed.
[0043] In a fifth aspect, an embodiment of the present disclosure proposes a program product. When the program product is executed by a communication device, the communication device executes the method described in the optional implementation manner of the first aspect or the second aspect.
[0044] In a sixth aspect, an embodiment of the present disclosure proposes a computer program, which, when executed on a computer, enables the computer to execute the method described in the optional implementation of the first aspect or the second aspect.
[0045] In a seventh aspect, an embodiment of the present disclosure provides a chip or a chip system, which includes a processing circuit configured to execute the method described in the optional implementation of the first or second aspect.
[0046] It is understood that the terminals, communication systems, storage media, program products, computer programs, chips, or chip systems involved in each embodiment of the present disclosure are all used to perform the methods proposed in the embodiments of the present disclosure. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects of the corresponding methods and will not be repeated here.
[0047] The embodiments of the present disclosure are not exhaustive and are merely illustrative of some embodiments, and are not intended to be a specific limitation on the scope of protection of the present disclosure. In the absence of contradiction, each step in a certain embodiment can be implemented as an independent embodiment, and the steps can be arbitrarily combined. For example, a solution after removing some steps in a certain embodiment can also be implemented as an independent embodiment, and the order of the steps in a certain embodiment can be arbitrarily exchanged. In addition, the optional implementation methods in a certain embodiment can be arbitrarily combined; in addition, the embodiments can be arbitrarily combined. For example, some or all steps of different embodiments can be arbitrarily combined, and a certain embodiment can be arbitrarily combined with the optional implementation methods of other embodiments.
[0048] In each embodiment of the present disclosure, unless otherwise specified or provided for, the terms and / or descriptions between the embodiments are consistent and may be referenced by each other. The technical environments in different embodiments may be combined to form new embodiments based on their inherent logical relationships.
[0049] The terms used in the embodiments of the present disclosure are only for the purpose of describing specific embodiments and are not intended to limit the present disclosure.
[0050] In the embodiments of the present disclosure, unless otherwise specified, elements expressed in the singular, such as "a", "an", "the", "above", "said", "the", "the", etc., may mean "one and only one", or "one or more", "at least one", etc. For example, when using articles such as "a", "an", "the" in English in translation, the noun following the article may be understood as a singular expression or a plural expression.
[0051] In the embodiments of the present disclosure, “plurality” refers to two or more.
[0052] In some embodiments, the terms "at least one," "one or more," "a plurality of," "multiple," etc. may be used interchangeably.
[0053] In some embodiments, descriptions such as "at least one of A and B," "A and / or B," "A in one case, B in another case," or "in response to one case A, in response to another case B" may include the following technical solutions depending on the situation: in some embodiments, A (A is executed independently of B); in some embodiments, B (B is executed independently of A); in some embodiments, execution is selected from A and B (A and B are selectively executed); and in some embodiments, A and B (both A and B are executed). The above is also applicable when there are more branches such as A, B, and C.
[0054] In some embodiments, "A or B" and other descriptions may include the following technical solutions depending on the situation: in some embodiments, A (A is executed independently of B); in some embodiments, B (B is executed independently of A); in some embodiments, execution is selected from A and B (A and B are selectively executed). The above is also applicable when there are more branches such as A, B, C, etc.
[0055] The prefixes such as "first" and "second" in the embodiments of the present disclosure are only used to distinguish different description objects and do not constitute any restriction on the position, order, priority, quantity or content of the description objects. For the statement of the description object, please refer to the description in the context of the claims or embodiments, and no unnecessary restriction should be constituted due to the use of prefixes. For example, if the description object is a "field", the ordinal number before the "field" in the "first field" and the "second field" does not limit the position or order between the "fields". "First" and "second" do not limit whether the "fields" they modify are in the same message, nor do they limit the order of the "first field" and the "second field". For another example, if the description object is a "level", the ordinal number before the "level" in the "first level" and the "second level" does not limit the priority between the "levels". For another example, the number of description objects is not limited by the ordinal number and can be one or more. Taking "first device" as an example, the number of "devices" can be one or more. In addition, the objects modified by different prefixes can be the same or different. For example, if the description object is "device", then the "first device" and the "second device" can be the same device or different devices, and their types can be the same or different; for example, if the description object is "information", then the "first information" and "the performance of each AI model" can be the same information or different information, and their contents can be the same or different.
[0056] In some embodiments, “including A,” “comprising A,” “used to indicate A,” and “carrying A” can be interpreted as directly carrying A or indirectly indicating A.
[0057] In some embodiments, terms such as "in response to...", "in response to determining...", "in the case of...", "at the time of...", "when...", "if...", "if...", etc. can be used interchangeably.
[0058] In some embodiments, terms such as "greater than", "greater than or equal to", "not less than", "more than", "more than or equal to", "not less than", "higher than", "higher than or equal to", "not less than", and "above" can be replaced with each other, and terms such as "less than", "less than or equal to", "not greater than", "less than", "less than or equal to", "not more than", "lower than", "lower than or equal to", "not higher than", and "below" can be replaced with each other.
[0059] In some embodiments, devices and equipment can be interpreted as physical or virtual, and their names are not limited to the names recorded in the embodiments. In some cases, they can also be understood as "equipment", "device", "circuit", "network element", "node", "function", "unit", "section", "system", "network", "chip", "chip system", "entity", "subject", etc.
[0060] In some embodiments, "network" can be interpreted as devices included in the network, such as access network equipment, core network equipment, etc.
[0061] In some embodiments, "access network device (AN device)" may also be referred to as "radio access network device (RAN device)", "base station (BS)", "radio base station", "fixed station", and in some embodiments may also be understood as "node", "access point", "transmission point (TP)", "reception point (RP)", "transmission and / or reception point (TRP)" "panel", "antenna panel", "antenna array", "cell", "macro cell", "small cell", "femto cell", "pico cell", "sector", "cell group", "serving cell", "carrier", "component carrier", "bandwidth part (BWP)", etc.
[0062] In some embodiments, "terminal" or "terminal device" may be referred to as "user equipment (UE)", "user terminal" "mobile station (MS)", "mobile terminal (MT)", subscriber station, mobile unit, subscriber unit, wireless unit, remote unit, mobile device, wireless device, wireless communication device, remote device, mobile subscriber station, access terminal, mobile terminal, wireless terminal, remote terminal, handset, user agent, mobile client, client, etc.
[0063] In some embodiments, obtaining data, information, etc. may comply with the laws and regulations of the country where the data is obtained.
[0064] In some embodiments, data, information, etc. may be obtained with the user's consent.
[0065] In addition, each element, each row, or each column in the table of the embodiment of the present disclosure can be implemented as an independent embodiment, and the combination of any elements, any rows, and any columns can also be implemented as an independent embodiment.
[0066] FIG1 is a schematic diagram showing a communication system architecture according to an embodiment of the present disclosure.
[0067] As shown in FIG1 , a communication system 100 includes a terminal 101 and a terminal 102 .
[0068] In some embodiments, terminal 101 or terminal 102 includes, for example, a mobile phone, a wearable device, an Internet of Things device, a car with communication function, a smart car, a tablet computer, a computer with wireless transceiver function, a virtual reality (VR) terminal device, an augmented reality (AR) terminal device, a wireless terminal device in industrial control, a wireless terminal device in self-driving, a wireless terminal device in remote medical surgery, a wireless terminal device in a smart grid, a wireless terminal device in transportation safety, a wireless terminal device in a smart city, and at least one of a wireless terminal device in a smart home, but is not limited thereto.
[0069] In some embodiments, the technical solution of the present disclosure can be applied to the Open RAN architecture. In this case, the interfaces between or within the access network devices involved in the embodiments of the present disclosure can be transformed into internal interfaces of the Open RAN, and the processes and information interactions between these internal interfaces can be implemented through software or programs.
[0070] It can be understood that the communication system described in the embodiment of the present disclosure is for the purpose of more clearly illustrating the technical solution of the embodiment of the present disclosure, and does not constitute a limitation on the technical solution proposed in the embodiment of the present disclosure. Ordinary technicians in this field can know that with the evolution of the system architecture and the emergence of new business scenarios, the technical solution proposed in the embodiment of the present disclosure is also applicable to similar technical problems.
[0071] The following embodiments of the present disclosure may be applied to the communication system 100 shown in FIG1 , or a portion thereof, but are not limited thereto. The entities shown in FIG1 are illustrative only. The communication system may include all or part of the entities shown in FIG1 , or may include other entities outside of FIG1 . The number and form of the entities are arbitrary, and the entities may be physical or virtual. The connection relationships between the entities are illustrative only. The entities may be connected or disconnected, and the connection may be in any manner, including direct or indirect, wired or wireless.
[0072] The embodiments of the present disclosure can be applied to Long Term Evolution (LTE), LTE-Advanced (LTE-A), LTE-Beyond (LTE-B), SUPER 3G, IMT-Advanced, 4th generation mobile communication system (4G), 5th generation mobile communication system (5G), 5G new radio (NR), future radio access (FRA), new radio access technology (RAT), new radio (NR), new radio access (NX), future generation radio access (FX), Global System for Mobile communications (GSM (registered trademark)), CDMA2000, Ultra Mobile Broadband (UMB), IEEE 802.11 (Wi-Fi (registered trademark)), IEEE 802.16 (WiMAX (registered trademark)), IEEE 802.20, Ultra-WideBand (UWB), Bluetooth (registered trademark), Public Land Mobile Network (PLMN) networks, Device-to-Device (D2D) systems, Machine-to-Machine (M2M) systems, Internet of Things (IoT) systems, Vehicle-to-Everything (V2X), systems utilizing other audio processing methods, and next-generation systems based on and extending these systems. Furthermore, multiple systems may be combined (for example, a combination of LTE or LTE-A with 5G).
[0073] Currently, to meet the needs of traditional monophonic communication and the emerging demand for spatial audio, terminals are typically equipped with multiple microphones (mics), for example, at least two mics. Spatial audio can be understood as collecting sound from all directions through multiple channels, allowing users to perceive sound from all directions and obtain an audio experience consistent with the real scene. The goal of spatial audio collection technology is to reproduce the complete sound field in the terminal's environment, maintaining the relative positional relationship between all sound sources in the sound field.
[0074] With the rise of spatial audio services, terminals have gradually begun to support spatial audio capture. Current technologies support capturing spatial audio during video recording, or capturing spatial audio and communicating in hands-free mode. However, in handset mode, communication based on spatial audio is not possible. Or the captured spatial audio effect is poor. For example, assuming that the terminal has left and right channels, when the user is on the left side of the terminal, the user's voice captured by the left channel is louder than the user's voice captured by the right channel, resulting in a misalignment of the user's voice in the entire sound field, and a poor communication experience for the user. That is, because current spatial audio capture technology cannot change the relative relationship of the sound sources in the sound field, if the position where the user speaks is regarded as the front, the left and right sides of the sound field captured by the terminal are the front and back directions of the user, which is different from the audio service experienced by the user on the capture end, resulting in a poor audio experience for the receiving end user using the audio signal sent by the transmitter.
[0075] Therefore, the present disclosure provides an audio processing method, which obtains a multi-channel audio signal collected in a stereo format and a single-channel audio signal collected in a non-stereo format when the terminal is in handset mode, wherein the multi-channel audio signal includes multiple sound sources including a target sound source, and the single-channel audio signal is the audio signal of the target sound source. The single-channel audio signal can be used to adjust the portion of the multi-channel audio signal with the target sound source, so that the portion of the multi-channel audio signal with the target sound source is more reasonable and is positioned at a desired location, thereby realizing communication based on spatial audio.
[0076] FIG2 is a schematic diagram illustrating an interaction of a communication method according to an embodiment of the present disclosure. As shown in FIG2 , the present disclosure embodiment relates to a communication method for use in a communication system 100, the method comprising:
[0077] Step S2101: Terminal 101 obtains a first audio signal and a second audio signal.
[0078] In some embodiments, in response to the terminal being in a handset mode, the first audio signal and the second audio signal are acquired.
[0079] In some embodiments, the first audio signal is a multi-channel audio signal captured in a stereo format. That is, the terminal may capture the multi-channel audio signal in a stereo format to obtain the first audio signal. The first audio signal can be understood as a spatial audio signal. The first audio signal includes multiple sound sources, including a target sound source. For example, if the target sound source is a user's voice, the first audio signal may be the overall ambient sound including the user's voice.
[0080] In some embodiments, the stereo format includes, for example, a 5.1.4 format, a 7.1.4 format, a stereo format, an ambisonic format, etc. The present disclosure does not provide examples one by one, but is not limited to these. For example, a microphone array composed of multiple microphones on the terminal can be used to collect the first audio signal.
[0081] In some embodiments, the second audio signal is a single-channel audio signal collected in a non-stereo format. That is, the terminal can collect the single-channel audio signal in a non-stereo format to obtain the second audio signal. The second audio signal is a signal of a target sound source. For example, the target sound source can be the user's voice. Of course, in different scenarios, the target sound source can also be any other sound, such as the sound of an animal or the sound of a machine, etc., and this disclosure does not limit this.
[0082] In some embodiments, a single-channel audio signal can be captured using the terminal's own microphone. Speech enhancement techniques, such as beamforming, noise reduction, and blind source separation (BSS), can also be used to process the signal to produce a second audio signal with a high signal-to-noise ratio. It will be appreciated that speech enhancement techniques can remove ambient sound, noise, and reverberation from the original single-channel audio signal captured by the terminal, retaining only the components of the target sound source.
[0083] In some embodiments, a single-channel audio signal can be captured using an external microphone connected to the terminal. Furthermore, speech enhancement techniques such as beamforming, noise reduction, and BSS can be used to process the signal to produce a second audio signal with a high signal-to-noise ratio. Examples of external microphones include headsets and handheld microphones, though this disclosure does not provide a comprehensive list of these.
[0084] In step S2102 , the terminal 101 uses the second audio signal to adjust a portion of the audio signal corresponding to the target sound source in the first audio signal.
[0085] In some embodiments, the difference in frequency between the portion of the audio signal corresponding to the target sound source in the first audio signal and the second audio line signal is less than or equal to a threshold. That is, the portion of the audio signal in the first audio signal whose frequency difference with the second audio signal is less than or equal to the threshold can be determined as the portion of the audio signal corresponding to the target sound source. The terminal can use the second audio signal to adjust the portion of the audio signal in the first audio signal corresponding to the target sound source.
[0086] In some embodiments, the first audio signal is a multi-channel audio signal. When the target sound source is at a first position of the terminal, adjusting a portion of the first audio signal corresponding to the target sound source using the second audio signal includes at least one of the following: using the second audio signal to weaken a portion of the first audio signal corresponding to the target sound source in a channel corresponding to the first position; and using the second audio signal to enhance a portion of the first audio signal corresponding to the target sound source in channels other than the first position.
[0087] Optionally, the second audio signal can be used to attenuate a portion of the audio signal corresponding to the target sound source in the first audio signal of the channel corresponding to the first position. Attenuating a portion of the audio signal may mean attenuating the amplitude of the portion of the audio signal. For example, assuming a terminal has left and right channels and collects first audio signals from both channels, when the user is on the left side of the terminal, the amplitude of the portion of the audio signal corresponding to the target sound source in the first audio signal of the left channel can be attenuated so that the amplitude of the portion of the audio signal corresponding to the target sound source in the first audio signal of the left channel is substantially the same as the amplitude of the portion of the audio signal corresponding to the target sound source in the first audio signal of the right channel. For another example, when the user is on the right side of the terminal, the amplitude of the portion of the audio signal corresponding to the target sound source in the first audio signal of the right channel can be attenuated so that the amplitude of the portion of the audio signal corresponding to the target sound source in the first audio signal of the left channel is substantially the same as the amplitude of the portion of the audio signal corresponding to the target sound source in the first audio signal of the right channel. For another example, assuming a terminal has multiple channels and the user is on the left front side of the terminal, the amplitude of the portion of the audio signal corresponding to the target sound source in the first audio signal of the left front channel can be attenuated so that the amplitude of the portion of the audio signal corresponding to the target sound source in the first audio signals of multiple channels is substantially the same.
[0088] Optionally, the second audio signal can be used to enhance the portion of the audio signal corresponding to the target sound source in the first audio signal of other channels other than the channel corresponding to the first orientation. The enhanced portion of the audio signal can be the amplitude of the enhanced portion of the audio signal. For example, assuming that the terminal has left and right channels, the first audio signals of the left and right channels are collected. When the user is on the left side of the terminal, the amplitude of the portion of the audio signal corresponding to the target sound source in the first audio signal of the right channel can be enhanced, so that the amplitude of the portion of the audio signal of the target sound source in the first audio signal of the left channel is almost the same as the amplitude of the portion of the audio signal of the target sound source in the first audio signal of the right channel. For another example, when the user is on the right side of the terminal, the amplitude of the portion of the audio signal corresponding to the target sound source in the first audio signal of the left channel can be enhanced, so that the amplitude of the portion of the audio signal of the target sound source in the first audio signal of the left channel is almost the same as the amplitude of the portion of the audio signal of the target sound source in the first audio signal of the right channel. For another example, assuming a terminal has multiple audio channels, if the user is standing in front of the left side of the terminal, the amplitude of the portion of the audio signal corresponding to the target sound source in the first audio signal of the right channel and the right front channel can be enhanced so that the amplitude of the portion of the audio signal corresponding to the target sound source in the first audio signals of multiple channels is almost the same. Of course, this disclosure only gives a few specific examples and cannot list them all, but it is not limited to these.
[0089] Optionally, the second audio data can be used to simultaneously weaken the portion of the audio signal corresponding to the target sound source in the first audio signal of the channel corresponding to the first orientation, and the second audio signal can be used to enhance the portion of the audio signal corresponding to the target sound source in the first audio signal of other channels other than the channel corresponding to the first orientation.
[0090] In some embodiments, weakening a portion of the audio signal corresponding to the target sound source in the first audio signal of the channel corresponding to the first orientation includes at least one of the following: using a second audio signal with the same phase as the partial audio signal to subtract the partial audio signal from the partial audio signal to obtain a weakened partial audio signal; using a second audio signal with an opposite phase to the partial audio signal to add the partial audio signal to obtain a weakened partial audio signal.
[0091] Alternatively, a second audio signal having the same phase as the partial audio signal may be subtracted from the partial audio signal to obtain a weakened partial audio signal. Due to the same phase, the amplitude of the partial audio signal after subtraction is smaller than that of the original partial audio signal. Thus, a weakened partial audio signal is obtained.
[0092] Alternatively, a second audio signal having an opposite phase to the partial audio signal may be added to the partial audio signal to obtain an added partial audio signal. Due to the opposite phase, the added partial audio signal has a smaller amplitude than the original partial audio signal. In other words, a weakened partial audio signal is obtained.
[0093] In some embodiments, enhancing the partial audio signal corresponding to the target sound source in the first audio signal of the channel corresponding to the orientation other than the first orientation includes at least one of the following: using a second audio signal with the same phase as the partial audio signal to add to the partial audio signal to obtain the enhanced partial audio signal; using a second audio signal with the opposite phase to the partial audio signal to subtract from the partial audio signal to obtain the enhanced partial audio signal.
[0094] Alternatively, a second audio signal having the same phase as the partial audio signal may be added to the partial audio signal to obtain an enhanced partial audio signal. Due to the same phase, the amplitude of the added partial audio signal is larger than that of the original partial audio signal. Thus, an enhanced partial audio signal is obtained.
[0095] Alternatively, a second audio signal having an opposite phase to the partial audio signal may be subtracted from the partial audio signal to obtain a subtracted partial audio signal. Due to the opposite phase, the amplitude of the subtracted partial audio signal is larger than that of the original partial audio signal. In other words, an enhanced partial audio signal is obtained.
[0096] Step S2103 : Terminal 101 sends the second audio signal and the adjusted first audio signal to terminal 102 .
[0097] In some embodiments, terminal 102 receives the second audio signal and the adjusted first audio signal sent by terminal 101 .
[0098] In some embodiments, terminal 101 can be understood as a transmitting terminal, i.e., the end that sends audio signals. Terminal 102 can be understood as a receiving terminal, i.e., the end that receives audio signals. Terminal 101 can send the adjusted first and second audio signals to terminal 102 to enable communication based on spatial audio in handset mode. This can maximize the positioning of the target sound source at the desired location, for example, positioning the user's voice in the middle of the ambient sound, to enhance the user experience.
[0099] In some embodiments, terminal 101 may directly send the second audio signal and the adjusted first audio signal to terminal 102 , or may mix the second audio signal and the adjusted first audio signal to obtain a mixed audio signal, and send the mixed audio signal to terminal 102 .
[0100] The audio processing method involved in the embodiment of the present disclosure may include at least one of steps S2101 to S2103. For example, step S2101 and step S2102 may be implemented as independent embodiments.
[0101] In some embodiments, step S2103 is optional, and one or more of these steps may be omitted or replaced in different embodiments.
[0102] In some embodiments, reference may be made to other optional implementations described before or after the description corresponding to FIG. 2 .
[0103] FIG3a is a flow chart of an audio processing method according to an embodiment of the present disclosure. As shown in FIG3a, the embodiment of the present disclosure relates to an audio processing method, which is executed by terminal 101 and includes:
[0104] Step S3101: Acquire a first audio signal and a second audio signal.
[0105] The optional implementation of step S3101 can refer to the optional implementation of step S2101 in Figure 2 and other related parts in the embodiment involved in Figure 2, which will not be repeated here.
[0106] Step S3102: Using the second audio signal, adjust a portion of the audio signal corresponding to the target sound source in the first audio signal.
[0107] The optional implementation of step S3102 can refer to the optional implementation of step S2102 in Figure 2 and other related parts in the embodiment involved in Figure 2, which will not be repeated here.
[0108] Step S3103 : Send the second audio signal and the adjusted first audio signal to the terminal 102 .
[0109] The optional implementation of step S3103 can refer to the optional implementation of step S2103 in Figure 2 and other related parts in the embodiment involved in Figure 2, which will not be repeated here.
[0110] In some embodiments, terminal 101 sends the second audio signal and the adjusted first audio signal to terminal 102, but is not limited thereto. The second audio signal and the adjusted first audio signal may also be sent to other entities.
[0111] FIG3b is a flow chart of an audio processing method according to an embodiment of the present disclosure. As shown in FIG3b , the embodiment of the present disclosure relates to an audio processing method, which is executed by terminal 101 and includes:
[0112] Step S3201: Acquire a first audio signal and a second audio signal.
[0113] The optional implementation of step S3201 can refer to the optional implementation of step S2101 in Figure 2 and other related parts in the embodiment involved in Figure 2, which will not be repeated here.
[0114] Step S3202: Using the second audio signal, adjust a portion of the audio signal corresponding to the target sound source in the first audio signal.
[0115] The optional implementation of step S3202 can refer to the optional implementation of step S2102 in Figure 2 and other related parts in the embodiment involved in Figure 2, which will not be repeated here.
[0116] In some embodiments, a difference between the frequencies of the portion of the audio signal corresponding to the target sound source in the first audio signal and the second audio signal is less than or equal to a threshold.
[0117] In some embodiments, when a target sound source is at a first position of a terminal, adjusting a portion of the first audio signal corresponding to the target sound source using the second audio signal includes at least one of the following: using the second audio signal to weaken the portion of the first audio signal corresponding to the target sound source in the first audio signal of the channel corresponding to the first position; or using the second audio signal to enhance the portion of the first audio signal corresponding to the target sound source in the first audio signal of the channel corresponding to positions other than the first position.
[0118] In some embodiments, attenuating a portion of the audio signal corresponding to the target sound source in the first audio signal of the channel corresponding to the first orientation includes at least one of the following: subtracting a second audio signal having the same phase as the partial audio signal from the partial audio signal to obtain the attenuated partial audio signal; or adding a second audio signal having an opposite phase to the partial audio signal to the partial audio signal to obtain the attenuated partial audio signal.
[0119] In some embodiments, enhancing a portion of the audio signal corresponding to the target sound source in the first audio signal of the channel corresponding to directions other than the first direction includes at least one of the following: adding a second audio signal having the same phase as the partial audio signal to the partial audio signal to obtain the enhanced partial audio signal; or subtracting a second audio signal having an opposite phase to the partial audio signal from the partial audio signal to obtain the enhanced partial audio signal.
[0120] FIG4 is a flow chart of an audio processing method according to an embodiment of the present disclosure. As shown in FIG4 , the embodiment of the present disclosure relates to an audio processing method, which is executed by terminal 102 and includes:
[0121] Step S4101: Acquire the second audio signal and the adjusted first audio signal sent by the terminal 101.
[0122] The optional implementation of step S4101 can refer to the optional implementation of step S2103 in Figure 2 and other related parts in the embodiment involved in Figure 2, which will not be repeated here.
[0123] FIG5 is a flow chart of a communication method according to an embodiment of the present disclosure. As shown in FIG5 , the embodiment of the present disclosure relates to a communication method, and the method includes:
[0124] Step S5101 : Terminal 101 sends a second audio signal and an adjusted first audio signal to terminal 102 .
[0125] The optional implementation of step S5101 can be found in S2103 of FIG. 2 and other related parts of the embodiment involved in FIG. 2 , which will not be described in detail here.
[0126] In some embodiments, the above method may include the method of the above embodiments related to the communication system 100, terminal 101, terminal 102, etc., which will not be repeated here.
[0127] Step S5102: Terminal 102 receives the second audio signal and the adjusted first audio signal.
[0128] The optional implementation of step S5102 can be found in S2103 of FIG. 2 and other related parts of the embodiment involved in FIG. 2 , which will not be described in detail here.
[0129] In some embodiments, the above method may include the method of the above embodiments related to the communication system 100, terminal 101, terminal 102, etc., which will not be repeated here.
[0130] The present disclosure also provides an audio processing method as follows:
[0131] In some embodiments, the UE at the transmitting end is set in the handset mode, and the UE collects the first audio signal as a multi-channel spatial audio signal, and at the same time, collects the second audio signal as the audio signal of the target sound source. Since they are all collected at the same position, the signals of each channel in the first audio signal contain audio components of the target sound source, and the components of the target sound source in the first audio signal are highly correlated with the target sound source in the second audio signal. Therefore, the amplitude of the target sound source in each channel of the first audio signal can be controlled by increasing the amplitude of the addition of signals with the same phase, reducing the amplitude of the addition of signals with opposite phases, or reducing the amplitude of the subtraction of signals with the same phase.
[0132] In some embodiments, the UE at the transmitting end may be terminal 101.
[0133] In some embodiments, the component of the target sound source in the first audio signal is the portion of the audio signal corresponding to the target sound source in the first audio signal, where high correlation can be achieved by having high frequency similarity. For example, the difference between the frequency of the portion of the audio signal corresponding to the target sound source in the first audio signal and the frequency of the second audio signal can be less than or equal to a threshold.
[0134] In some embodiments, the target sound source signal is sent to different channels of the spatial audio signal, and the intensity of the target sound source component in each channel of the spatial audio is enhanced or weakened so that the intensity ratio of the target signal in each channel meets the requirements, thereby controlling the positioning of the target signal at a desired location.
[0135] In some embodiments, in the handset mode, the UE is on the left side of the human head, collects audio in a stereo format, and collects a mono audio signal. The left and right directions of the user are consistent with the direction of the sound field. Since the user is on the right side of the UE, the collected user audio signal will be heard on the right side. Collect a mono audio signal. For example, in the audio signal collected by traditional mono communication, the mono audio signal is the main audio component in the target audio signal. Since the mono audio signal is highly correlated with some audio signal components in the dual channel, the mono audio signal can be used to enhance the relevant audio components in the left channel of the dual channel, or to weaken the relevant audio components in the right channel of the dual channel, so that the signal strength of the main audio components in the left and right channels is approximately the same, thereby positioning the main audio components in the middle of the sound field.
[0136] In some embodiments, in handset mode, the UE is placed on the left side of the head to collect audio in stereo format, and an external microphone, such as a headset or handheld microphone, is used to collect a mono audio signal. The processing method is consistent with the above embodiment 1.
[0137] In some embodiments, the user is on the left front side of the UE, and the user's voice needs to be positioned in the upper front. Spatial audio is captured in a 5.1.4 format, and multiple microphones on the UE are used to form a microphone array. Speech enhancement technologies such as beamforming and BSS are used to capture a mono audio signal. Since the user is on the left front side of the UE, the captured user audio signal will be perceived as being on the left front side. The captured mono signal is used to enhance the user audio signal in the right channel and the right front upper channel, positioning the user audio signal in the upper front.
[0138] Figure 6a is a schematic diagram of the structure of the terminal proposed in an embodiment of the present disclosure. As shown in Figure 6a, the terminal 6100 may include: a processing module 6101. In some embodiments, the processing module 6101 is used to obtain a first audio signal and a second audio signal in response to the terminal being in a handset mode; wherein the first audio signal is a multi-channel audio signal collected based on a stereo format, and the first audio signal includes audio signals of multiple sound sources, the multiple sound sources include a target sound source, and the second audio signal is a single-channel audio signal collected based on a non-stereo format, and the second audio signal is an audio signal of the target sound source; using the second audio signal, adjust the portion of the audio signal corresponding to the target sound source in the first audio signal.
[0139] In some embodiments, a difference between the frequencies of the portion of the audio signal corresponding to the target sound source in the first audio signal and the second audio signal is less than or equal to a threshold.
[0140] In some embodiments, when the target sound source is at a first position of the terminal, the processing module uses the second audio signal to adjust a portion of the first audio signal corresponding to the target sound source using at least one of the following methods: using the second audio signal to weaken a portion of the first audio signal corresponding to the target sound source in a channel corresponding to the first position; or using the second audio signal to enhance a portion of the first audio signal corresponding to the target sound source in a channel corresponding to positions other than the first position.
[0141] In some embodiments, the processing module attenuates a portion of the audio signal corresponding to the target sound source in the first audio signal of the channel corresponding to the first orientation using at least one of the following methods: subtracting a second audio signal having the same phase as the partial audio signal from the partial audio signal to obtain the attenuated partial audio signal; or adding a second audio signal having an opposite phase to the partial audio signal to the partial audio signal to obtain the attenuated partial audio signal.
[0142] In some embodiments, the processing module enhances a portion of the audio signal corresponding to the target sound source in the first audio signal of the channel corresponding to directions other than the first direction using at least one of the following methods, including: adding a second audio signal having the same phase as the partial audio signal to the partial audio signal to obtain the enhanced partial audio signal; and subtracting a second audio signal having an opposite phase to the partial audio signal from the partial audio signal to obtain the enhanced partial audio signal.
[0143] In some embodiments, the terminal 6100 may further include a transceiver module 6102 , which may be configured to send the second audio signal and the adjusted first audio signal to the terminal 102 , for example.
[0144] FIG6b is a schematic diagram of the structure of a terminal according to an embodiment of the present disclosure. As shown in FIG6b , the terminal 6200 may include a transceiver module 6201. The transceiver module 6201 is configured to receive the second audio signal and the adjusted first audio signal sent by the terminal 101.
[0145] Figure 7a is a schematic diagram of the structure of a communication device 7100 proposed in an embodiment of the present disclosure. Communication device 7100 can be terminal 101 or terminal 102, or a chip, chip system, or processor that supports the terminal in implementing any of the above methods. Alternatively, the terminal can be user equipment, etc. Communication device 7100 can be used to implement the methods described in the above method embodiments. For details, please refer to the description of the above method embodiments.
[0146] As shown in Figure 7a, communication device 7100 includes one or more processors 7101. Processor 7101 can be a general-purpose processor or a dedicated processor, for example, a baseband processor or a central processing unit. The baseband processor can be used to process communication protocols and communication data, while the central processing unit can be used to control the communication device, execute programs, and process program data. Communication device 7100 is used to perform any of the above methods. Optionally, the communication device can be a base station, a baseband chip, a terminal device, a terminal device chip, a DU or CU, etc.
[0147] In some embodiments, the communication device 7100 further includes one or more memories 7102 for storing instructions. Optionally, all or part of the memories 7102 may be located outside the communication device 7100.
[0148] In some embodiments, the communication device 7100 further includes one or more transceivers 7103. When the communication device 7100 includes one or more transceivers 7103, the transceiver 7103 performs the communication step S2101 such as sending and / or receiving in the above method, and the processor 7101 performs other steps.
[0149] In some embodiments, a transceiver may include a receiver and / or a transmitter. The receiver and transmitter may be separate or integrated. Optionally, the terms transceiver, transceiver unit, transceiver, and transceiver circuit may be used interchangeably; the terms transmitter, transmitting unit, transmitter, and transmitting circuit may be used interchangeably; and the terms receiver, receiving unit, receiver, and receiving circuit may be used interchangeably.
[0150] In some embodiments, the communication device 7100 may include one or more interface circuits 7104. Optionally, the interface circuit 7104 is connected to the memory 7102. The interface circuit 7104 may be configured to receive signals from the memory 7102 or other devices, and may be configured to send signals to the memory 7102 or other devices. For example, the interface circuit 7104 may read instructions stored in the memory 7102 and send the instructions to the processor 7101.
[0151] The communication device 7100 described in the above embodiment may be a network device or a terminal, but the scope of the communication device 7100 described in the present disclosure is not limited thereto, and the structure of the communication device 7100 may not be limited by FIG. 7a. The communication device may be an independent device or may be part of a larger device. For example, the communication device may be: 1) an independent integrated circuit IC, or a chip, or a chip system or subsystem; (2) a collection of one or more ICs, optionally, the above IC collection may also include a storage component for storing data or programs; (3) an ASIC, such as a modem; (4) a module that can be embedded in other devices; (5) a receiver, a terminal device, an intelligent terminal device, a cellular phone, a wireless device, a handheld device, a mobile unit, an in-vehicle device, a network device, a cloud device, an artificial intelligence device, etc.; (6) others, etc.
[0152] FIG7 b is a schematic diagram of the structure of a chip 7200 according to an embodiment of the present disclosure. If the communication device 7100 can be a chip or a chip system, please refer to the schematic diagram of the structure of the chip 7200 shown in FIG7 b , but the present disclosure is not limited thereto.
[0153] The chip 7200 includes one or more processors 7201 , and the chip 7200 is configured to execute any of the above methods.
[0154] In some embodiments, chip 7200 further includes one or more interface circuits 7202. Optionally, interface circuit 7202 is connected to memory 7203. Interface circuit 7202 can be used to receive signals from memory 7203 or other devices, or to send signals to memory 7203 or other devices. For example, interface circuit 7202 can read instructions stored in memory 7203 and send the instructions to processor 7201.
[0155] In some embodiments, the interface circuit 7202 executes the communication step S2101 of sending and / or receiving in the above method, and the processor 7201 executes other steps.
[0156] In some embodiments, terms such as interface circuit, interface, transceiver pin, and transceiver may be used interchangeably.
[0157] In some embodiments, the chip 7200 further includes one or more memories 7203 for storing instructions. Alternatively, all or part of the memories 7203 may be located outside the chip 7200.
[0158] The present disclosure also proposes a storage medium having instructions stored thereon. When the instructions are executed on the communication device 7100, the communication device 7100 executes any of the above methods. Optionally, the storage medium is an electronic storage medium. Optionally, the storage medium is a computer-readable storage medium, but is not limited thereto and may also be a storage medium readable by other devices. Optionally, the storage medium may be a non-transitory storage medium, but is not limited thereto and may also be a temporary storage medium.
[0159] The present disclosure also provides a program product, which, when executed by the communication device 7100, enables the communication device 7100 to perform any of the above methods. Optionally, the program product is a computer program product.
[0160] The present disclosure also proposes a computer program, which, when executed on a computer, causes the computer to perform any one of the above methods.
[0161] Fig. 8 is a block diagram showing an audio processing apparatus 800 according to an exemplary embodiment.
[0162] As shown in FIG. 8 , apparatus 800 may include one or more of the following components: a processing component 802 , a memory 804 , a power component 806 , a multimedia component 808 , an audio component 810 , an input / output (I / O) interface 812 , a sensor component 814 , and a communication component 816 .
[0163] The processing component 802 generally controls the overall operation of the device 800, such as operations associated with display, phone calls, data communications, camera operation, and recording operations. The processing component 802 may include one or more processors 820 to execute instructions to perform all or part of the steps of the above-described method. In addition, the processing component 802 may include one or more modules to facilitate interaction between the processing component 802 and other components. For example, the processing component 802 may include a multimedia module to facilitate interaction between the multimedia component 808 and the processing component 802.
[0164] The memory 804 is configured to store various types of data to support the operations of the device 800. Examples of such data include instructions for any application or method operating on the device 800, contact data, phone book data, messages, pictures, videos, etc. The memory 804 can be implemented by any type of volatile or non-volatile storage device, or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk, or optical disk.
[0165] The power component 806 provides power to the various components of the device 800. The power component 806 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to the device 800.
[0166] The multimedia component 808 includes a screen that provides an output interface between the device 800 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, slides, and gestures on the touch panel. The touch sensor can not only sense the boundaries of a touch or slide action, but also detect the duration and pressure associated with the touch or slide operation. In some embodiments, the multimedia component 808 includes a front camera and / or a rear camera. When the device 800 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera can receive external multimedia data. Each front camera and rear camera can be a fixed optical lens system or have focal length and optical zoom capabilities.
[0167] The audio component 810 is configured to output and / or input audio signals. For example, the audio component 810 includes a microphone (MIC), which is configured to receive external audio signals when the device 800 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signal can be further stored in the memory 804 or transmitted via the communication component 816. In some embodiments, the audio component 810 also includes a speaker for outputting audio signals.
[0168] I / O interface 812 provides an interface between processing component 802 and peripheral interface modules, such as a keyboard, click wheel, buttons, etc. These buttons may include but are not limited to: a home button, volume buttons, a start button, and a lock button.
[0169] The sensor assembly 814 includes one or more sensors for providing various aspects of the status assessment of the device 800. For example, the sensor assembly 814 can detect the open / closed state of the device 800, the relative positioning of components, such as the display and keypad of the device 800. The sensor assembly 814 can also detect changes in the position of the device 800 or a component of the device 800, the presence or absence of user contact with the device 800, the orientation or acceleration / deceleration of the device 800, and temperature changes of the device 800. The sensor assembly 814 may include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor assembly 814 may also include an optical sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor assembly 814 may also include an accelerometer, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.
[0170] The communication component 816 is configured to facilitate wired or wireless communication between the device 800 and other devices. The device 800 can access a wireless network based on a communication standard, such as WiFi, 2G or 3G, or a combination thereof. In an exemplary embodiment, the communication component 816 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 816 also includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology and other technologies.
[0171] In an exemplary embodiment, the apparatus 800 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the above-described method.
[0172] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 804 including instructions, which can be executed by the processor 820 of the apparatus 800 to perform the above method. For example, the non-transitory computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.
Claims
1. An audio processing method, characterized in that The method includes: In response to the terminal being in the handset mode, obtaining a first audio signal and a second audio signal; Wherein, the first audio signal is a multi-channel audio signal collected based on a stereo format, and the first audio signal includes audio signals of multiple sound sources, among which there is a target sound source, the second audio signal is a single-channel audio signal collected based on a non-stereo format, and the second audio signal is the audio signal of the target sound source; Using the second audio signal to adjust a partial audio signal corresponding to the target sound source in the first audio signal.
2. The method according to claim 1, characterized in that The difference between the frequency of the partial audio signal corresponding to the target sound source in the first audio signal and the frequency of the second audio signal is less than or equal to a threshold.
3. The method according to claim 1, characterized in that, The target sound source is at a first azimuth of the terminal, and using the second audio signal to adjust the partial audio signal corresponding to the target sound source in the first audio signal includes at least one of the following: Using the second audio signal to weaken the partial audio signal corresponding to the target sound source in the first audio signal of the channel corresponding to the first azimuth; Using the second audio signal to enhance the partial audio signal corresponding to the target sound source in the first audio signal of the channels corresponding to other azimuths except the first azimuth.
4. The method according to claim 3, wherein Weakening the partial audio signal corresponding to the target sound source in the first audio signal of the channel corresponding to the first azimuth includes at least one of the following: Using the second audio signal with the same phase as the partial audio signal to subtract from the partial audio signal to obtain a weakened partial audio signal; Using the second audio signal with the opposite phase to the partial audio signal to add to the partial audio signal to obtain a weakened partial audio signal.
5. The method according to claim 3, wherein Enhancing the partial audio signal corresponding to the target sound source in the first audio signal of the channels corresponding to other azimuths except the first azimuth includes at least one of the following: Using the second audio signal with the same phase as the partial audio signal to add to the partial audio signal to obtain an enhanced partial audio signal; Using the second audio signal with the opposite phase to the partial audio signal to subtract from the partial audio signal to obtain an enhanced partial audio signal.
6. An audio processing device, characterized in that, The apparatus includes: An obtaining module, configured to obtain a first audio signal and a second audio signal in response to the terminal being in the handset mode; Wherein, the first audio signal is a multi-channel audio signal collected based on a stereo format, and the first audio signal includes audio signals of multiple sound sources, among which there is a target sound source, the second audio signal is a single-channel audio signal collected based on a non-stereo format, and the second audio signal is the audio signal of the target sound source; A processing module, configured to use the second audio signal to adjust a partial audio signal corresponding to the target sound source in the first audio signal.
7. The device according to claim 6, characterized in that, The difference between the frequency of the partial audio signal corresponding to the target sound source in the first audio signal and the frequency of the second audio signal is less than or equal to a threshold.
8. The device according to claim 6, characterized in that, The target sound source is at a first azimuth of the terminal, and the processing module uses the second audio signal to adjust the partial audio signal corresponding to the target sound source in the first audio signal in at least one of the following ways: Using the second audio signal, attenuate the partial audio signal corresponding to the target sound source in the first audio signal of the channel corresponding to the first azimuth; Using the second audio signal, enhance the partial audio signal corresponding to the target sound source in the first audio signals of the channels corresponding to the azimuths other than the first azimuth.
9. The device according to claim 8, wherein The processing module attenuates the partial audio signal corresponding to the target sound source in the first audio signal of the channel corresponding to the first azimuth by using at least one of the following methods: Using a second audio signal having the same phase as the partial audio signal, subtract the partial audio signal to obtain the attenuated partial audio signal; Using a second audio signal having a phase opposite to that of the partial audio signal, add the partial audio signal to obtain the attenuated partial audio signal.
10. The device according to claim 8, characterized in that, The processing module enhances the partial audio signal corresponding to the target sound source in the first audio signals of the channels corresponding to the azimuths other than the first azimuth by using at least one of the following methods, including: Using a second audio signal having the same phase as the partial audio signal, add the partial audio signal to obtain the enhanced partial audio signal; Using a second audio signal having a phase opposite to that of the partial audio signal, subtract the partial audio signal to obtain the enhanced partial audio signal.
11. An electronic device, characterized in that, Comprising: A memory for storing instructions; And A processor for calling the instructions stored in the memory to execute the method according to any one of claims 1-5.
12. A storage medium, characterized in that, Instructions are stored in the storage medium, and when the instructions are executed by a processor, the method according to any one of claims 1-5 is executed.
Citation Information
Patent Citations
Distributed audio capture and mixing
CN108369811A
Audio processing method, audio processing device and electronic equipment
CN113767432A
Audio generation method and related device
CN116939473A
Method and apparatus for enhancing sound sources
US20170287499A1