Audio signal processing method, apparatus and electronic device

By determining the orientation information of the audio signal and its changes in the audio acquisition device, determining whether the sound source is the same, and voice enhancement is performed based on the orientation information determined by two adjacent times, the voice enhancement error problem caused by burst noise interference is solved, and the audio processing effect is improved.

CN114627888BActive Publication Date: 2025-06-24LENOVO (BEIJING) LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210324148.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-28
Publication Date
2025-06-24
Estimated Expiration
2042-03-28

AI Technical Summary

Technical Problem

During the audio acquisition process, sudden interference noise may lead to voice enhancement errors, affecting the audio processing effect.

Method used

By determining the orientation information and changes of the audio signal currently collected by the audio acquisition device, it is determined whether the sound source is the same, and the audio signal is voice-enhanced based on the orientation information determined by the adjacent two times.

Benefits of technology

It effectively reduces misjudgment caused by noise interference, ensures that the effective audio signal is not weakened, and improves the audio processing effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114627888B_ABST
    Figure CN114627888B_ABST
Patent Text Reader

Abstract

The present application discloses an audio signal processing method, apparatus and electronic device. The method includes: determining first azimuth information of a first sound source of an audio signal currently collected by an audio acquisition device relative to the audio acquisition device; determining difference information between the first azimuth information and second azimuth information, where the first azimuth information and the second azimuth information are two adjacent azimuth information determined; determining whether the first sound source is the same as the second sound source according to the difference information; in the case where the first sound source is different from the second sound source, respectively performing voice enhancement on the audio signal collected by the audio acquisition device based on the first azimuth information and the second azimuth information to obtain a first voice enhancement signal corresponding to the first azimuth information and a second voice enhancement signal corresponding to the second azimuth information. The solution of the present application can improve the audio processing effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of audio processing, and more specifically, to an audio signal processing method, apparatus, and electronic device. Background Art

[0002] In the process of audio processing, the direction of arrival (DOA) estimation algorithm can be used to determine the direction information of a sound source relative to an audio acquisition device. Based on this direction information, the audio signal in the corresponding direction can be enhanced in speech.

[0003] In audio processing scenarios such as audio conferences or voice control, it is necessary to accurately locate the direction of a user's voice. However, during audio acquisition, if there is sudden interference noise in the environment where the audio acquisition device is located, then the direction where the noise source is located may be determined as the direction where speech enhancement is required, resulting in incorrect speech enhancement and further affecting the audio processing effect. Summary of the Invention

[0004] The present application provides an audio signal processing method, apparatus, and electronic device.

[0005] Among them, an audio signal processing method includes:

[0006] Determine the first azimuth information of the first sound source of the audio signal currently collected by the audio acquisition device relative to the audio acquisition device;

[0007] Determine the difference information between the first azimuth information and the second azimuth information, where the second azimuth information is the azimuth information of the second sound source of the audio signal previously collected by the audio acquisition device relative to the audio acquisition device, and the first azimuth information and the second azimuth information are two adjacent azimuth information determined;

[0008] Determine whether the first sound source is the same as the second sound source according to the difference information;

[0009] When the first sound source is different from the second sound source, perform speech enhancement on the audio signal collected by the audio acquisition device based on the first azimuth information and the second azimuth information respectively, to obtain a first speech enhancement signal corresponding to the first azimuth information and a second speech enhancement signal corresponding to the second azimuth information.

[0010] In a possible implementation, the apparatus further includes:

[0011] When the first sound source is the same as the second sound source, perform voice enhancement on the audio signal collected by the audio acquisition device based on the second orientation information to obtain a third voice enhancement signal corresponding to the second orientation information.

[0012] In another possible implementation, determining whether the first sound source is the same as the second sound source according to the difference information includes:

[0013] Detecting whether the amount of orientation change represented by the difference information exceeds a set threshold;

[0014] Wherein, the amount of orientation change exceeding the set threshold indicates that the first sound source is different from the second sound source;

[0015] The amount of orientation change not exceeding the set threshold indicates that the first sound source is the same as the second sound source.

[0016] In another possible implementation, respectively performing voice enhancement on the audio signal collected by the audio acquisition device based on the first orientation information and the second orientation information includes:

[0017] Within a set duration, respectively perform voice enhancement on the audio signal collected by the audio acquisition device based on the first orientation information and the second orientation information;

[0018] The method further includes:

[0019] Within the set duration, reduce the second voice enhancement signal by a set amplitude to obtain a second voice enhancement signal after reduction;

[0020] Determine the first voice enhancement signal and the second voice enhancement signal after reduction as the voice signal after voice enhancement of the audio signal collected by the audio acquisition device.

[0021] In another possible implementation, respectively performing voice enhancement on the audio signal collected by the audio acquisition device based on the first orientation information and the second orientation information includes:

[0022] Within a set duration, respectively perform voice enhancement on the audio signal collected by the audio acquisition device based on the first orientation information and the second orientation information;

[0023] The method further includes:

[0024] After reaching the set duration, combine the first voice enhancement signal and the second voice enhancement signal, and determine at least one target orientation information that needs voice enhancement from the first orientation information and the second orientation information;

[0025] Perform voice enhancement on the audio signals collected by the audio collection device respectively based on the at least one target orientation information.

[0026] In another possible implementation manner, the combining the first voice enhancement signal and the second voice enhancement signal, and determining at least one target orientation information that needs voice enhancement from the first orientation information and the second orientation information includes:

[0027] Combining the first voice enhancement signal and the second voice enhancement signal, and determining at least one target orientation information at the current moment that the audio collection device can collect audio signals from the first orientation information and the second orientation information, where the target orientation information is the orientation information that needs voice enhancement.

[0028] In another possible implementation manner, combining the first voice enhancement signal and the second voice enhancement signal, and determining at least one target orientation information at the current moment that the audio collection device can collect audio signals from the first orientation information and the second orientation information includes:

[0029] If the first voice enhancement signals determined in the most recent set number of times are not all empty, and the second voice enhancement signals determined in the most recent set number of times are not all empty, determine the first orientation information and the second orientation information as the target orientation information that needs voice enhancement;

[0030] If the first voice enhancement signals determined in the most recent set number of times are all empty, and the second voice enhancement signals determined in the most recent set number of times are not all empty, determine the second orientation information as the target orientation information that needs voice enhancement;

[0031] If the first voice enhancement signals determined in the most recent set number of times are not all empty, and the second voice enhancement signals determined in the most recent set number of times are all empty, determine the first orientation information as the target orientation information that needs voice enhancement.

[0032] In another possible implementation manner, before determining the first orientation information, it further includes: establishing a voice conference channel between the electronic device and at least one other electronic device;

[0033] The determining the first voice enhancement signal and the second voice enhancement signal after amplitude reduction as the voice signal after voice enhancement of the audio signals collected by the audio collection device includes:

[0034] Transmit the first voice enhancement signal and the second voice enhancement signal after amplitude reduction to the at least one other electronic device through the voice conference channel;

[0035] After performing voice enhancement on the audio signal collected by the audio collection device respectively based on the at least one target orientation information, the following is further included:

[0036] Based on the voice conference channel, transmit the fourth voice enhanced signal obtained by performing voice enhancement on the audio signal based on the target orientation information to the at least one other electronic device.

[0037] Wherein, an audio signal processing device includes:

[0038] An orientation determination unit, configured to determine the first orientation information of the first sound source of the audio signal currently collected by the audio collection device relative to the audio collection device;

[0039] A difference determination unit, configured to determine the difference information between the first orientation information and the second orientation information, where the second orientation information is the orientation information of the second sound source of the audio signal previously collected by the audio collection device relative to the audio collection device, and the first orientation information and the second orientation information are two adjacent orientation information determined;

[0040] A sound source discrimination unit, configured to determine whether the first sound source is the same as the second sound source according to the difference information;

[0041] A first voice enhancement unit, configured to, when the first sound source is different from the second sound source, perform voice enhancement on the audio signal collected by the audio collection device respectively based on the first orientation information and the second orientation information, to obtain a first voice enhanced signal corresponding to the first orientation information and a second voice enhanced signal corresponding to the second orientation information.

[0042] Wherein, an electronic device includes at least a memory and a processor;

[0043] Wherein, the processor is configured to execute the audio signal processing method described in any one of the above;

[0044] The memory is configured to store a program required for the processor to perform operations.

[0045] As can be seen from the above solution, after determining the first azimuth information of the current sound source, the present application will determine the difference information between the first azimuth information and the second azimuth information of the sound source determined last time. If the difference information indicates that the audio signals detected twice come from different sound sources, voice enhancement will be performed on the audio signals collected by the audio acquisition device based on the first azimuth information and the second azimuth information respectively to obtain the voice enhancement signals in these two azimuth information respectively. Therefore, even if sudden noise interference occurs, the effective audio signals emitted by the original sound source will not be weakened, thereby reducing misjudgment caused by noise interference and preventing the situation where the interference noise is enhanced while the effective audio is weakened, and further improving the audio processing effect. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0047] Figure 1 It is a schematic flowchart of an audio signal processing method provided by an embodiment of the present application;

[0048] Figure 2 It is another schematic flowchart of an audio signal processing method provided by an embodiment of the present application;

[0049] Figure 3 It is another schematic flowchart of an audio signal processing method provided by an embodiment of the present application;

[0050] Figure 4 It is a schematic flowchart of an audio signal processing method provided by an embodiment of the present application in an application scenario;

[0051] Figure 5 It is a schematic structural diagram of a composition of an audio signal processing device provided by an embodiment of the present application;

[0052] Figure 6 It is a schematic structural diagram of a composition of an electronic device provided by an embodiment of the present application.

[0053] Terms such as "first", "second", "third", "fourth", etc. (if any) in the specification, claims and the above drawings are used to distinguish similar parts and do not necessarily describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application described here can be implemented in an order other than those illustrated here. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0054] The solution of the present application can be applied to any scenario involving audio signal processing, so as to more reasonably enhance the speech in the audio signal and improve the audio signal processing effect.

[0055] The solution of the present application can be applied to any scenario that requires speech enhancement processing of speech signals. For example, the solution of the present application can be applied to enhance the speech in the audio signal in a speech conference scenario, and can also be used for speech signal enhancement before audio signal recognition in a smart speaker, or audio enhancement processing during audio recording based on an electronic device, etc., without limitation.

[0056] Next, the technical solutions in the embodiments of the present application will be clearly and completely described with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.

[0057] As Figure 1 shown, it shows a schematic flowchart of another embodiment of an audio signal processing method provided by an embodiment of the present application. The method of this embodiment is applied to an electronic device, which can be a mobile terminal such as a mobile phone or a tablet computer, or a desktop computer, a smart speaker or a voice conference terminal, etc., without limitation.

[0058] The method of this embodiment may include:

[0059] S101, determining first orientation information of a first sound source of an audio signal currently collected by an audio collection device relative to the audio collection device.

[0060] It can be understood that during the process of the audio collection device continuously collecting audio signals, the audio processing device of the electronic device determines the orientation information of the collected audio signals at a set time interval or irregularly.

[0061] For the convenience of distinguishing from the audio signals collected in the subsequent history, the sound source of the audio signal currently collected in the present application is called the first sound source. Correspondingly, for the convenience of distinguishing from the orientation information of the sound sources determined in the history, the orientation information of the determined first sound source relative to the audio collection device is called the first orientation information.

[0062] It can be understood that the first orientation information may at least include: the direction information of the first sound source relative to the audio acquisition device. For example, the direction information of the first sound source to which the audio signal belongs relative to the audio acquisition device can be determined through a direction of arrival (DOA) estimation algorithm. Of course, the direction information can also be determined by other means.

[0063] Of course, the first orientation information may also include other orientation information such as the distance between the first sound source and the audio acquisition device, and there is no limitation on this.

[0064] It can be understood that after obtaining the audio signal, there can be multiple possible specific implementation manners for determining the orientation information of the audio signal relative to the audio acquisition device that acquires the audio signal, and this application does not limit this.

[0065] S102. Determine the difference information between the first orientation information and the second orientation information.

[0066] Wherein, the second orientation information is the orientation information of the second sound source of the audio signal historically acquired by the audio acquisition device relative to the audio acquisition device, and the first orientation information and the second orientation information are two adjacent orientation information determined. That is to say, the second orientation information is the orientation information determined most recently by the electronic device before determining the first orientation information.

[0067] Wherein, the difference information refers to the orientation information difference between the first orientation information and the second orientation information.

[0068] S103. Determine whether the first sound source is the same as the second sound source according to the difference information.

[0069] It can be understood that since the first orientation information and the second orientation information are two adjacent orientation information determined, therefore, based on the difference information between the first orientation information and the second orientation information, it can be determined whether the sound sources corresponding to the two adjacent orientation information determined are the same, that is, determine whether the first sound source corresponding to the first orientation information is the same as the second sound source corresponding to the second orientation information determined last time.

[0070] In a possible implementation manner, since the time interval between the orientation information of the sound sources detected by the electronic device twice adjacent is short, and the difference information can also characterize the orientation change information of the sound sources determined twice adjacent, therefore, this application can combine the orientation change amount characterized by the difference information to determine whether the sound sources determined twice adjacent are the same.

[0071] Specifically, it can be detected whether the amount of change in the orientation represented by the difference information exceeds a set threshold. If the amount of change in the orientation exceeds the set threshold, it indicates that the first sound source is different from the second sound source; conversely, if the amount of change in the orientation does not exceed the set threshold, it indicates that the first sound source is the same as the second sound source.

[0072] Among them, the set threshold can be set as needed, and there is no restriction on this.

[0073] It can be understood that the same sound source cannot have a large change in a short time. Therefore, if the change in the orientation information determined twice in succession is large, it means that the sound sources corresponding to the orientation information of the audio signals determined twice are not the same.

[0074] For example, taking the orientation information of the sound source of the audio signal relative to the audio acquisition device as the direction information as an example for illustration.

[0075] After determining the first direction information of the first sound source of the currently acquired audio signal relative to the audio acquisition device, the second direction information of the second sound source of the historical audio signal determined last time relative to the audio acquisition device can be obtained. On this basis, it can be detected whether the amount of change in the direction between the first direction information and the second direction information exceeds the set amount of change in the direction. Correspondingly, if the amount of change in the direction exceeds the set amount of change in the direction, it is determined that the first sound source is different from the second sound source.

[0076] For example, if the change angle between the first direction information and the second direction information exceeds the set angle threshold, it can be determined that the first sound source and the second sound source are not the same.

[0077] It can be understood that during the process of detecting the direction of arrival of audio signals twice in succession, the direction information of the same sound source relative to the audio acquisition device will not change much. For example, it is impossible for a user to move from one direction relative to the audio acquisition device to another direction with a large change in the direction angle in a short time.

[0078] S104, in the case where the first sound source is different from the second sound source, perform speech enhancement on the audio signal collected by the audio acquisition device respectively based on the first orientation information and the second orientation information, and obtain a first speech enhancement signal corresponding to the first orientation information and a second speech enhancement signal corresponding to the second orientation information.

[0079] Among them, speech enhancement refers to extracting as pure an original speech as possible from noisy speech. Therefore, the purpose of speech enhancement is to enhance the quality of the audio signal and reduce background noise interference, etc.

[0080] It can be understood that if the first sound source is different from the second sound source, it indicates that the first sound source is a newly emerged sound source. If this new sound source is a valid sound source other than the second sound source, then it indicates that there may be a possibility of two sound sources occurring simultaneously. Based on this, if the audio signal is enhanced only based on the first azimuth information of the first sound source, then within the time period from the detection of the first azimuth information to the next azimuth detection, even if the second sound source still emits sound, the sound emitted by the second sound source will not be enhanced, resulting in the weakening or elimination of some valid information, and the audio signal emitted by multiple sound sources simultaneously cannot be retained.

[0081] In addition, when the first sound source is different from the second sound source, since the first sound source is a newly emerged sound source, if the first sound source is a noise source, enhancing the speech of the audio signal only based on the first azimuth information of the first sound source will inevitably enhance the audio from the noise source in the audio signal, while weakening or even eliminating the audio emitted by the non-noise second sound source, so that the user cannot obtain the required audio signal from the enhanced speech signal.

[0082] Based on this, in order to reduce the situation where only the audio signal of the noise sound source is extracted due to sudden noise, or in the case where multiple sound sources emit speech signals simultaneously and the multi-channel audio signals overlap, to improve the speech enhancement effect of the audio signal, in this application, when a new sound source (i.e., the first sound source) is detected, the speech is not directly enhanced only using the azimuth information of the new sound source, but the azimuth information determined in the two most recent adjacent times is respectively used for speech enhancement simultaneously to retain and enhance the speech signals emitted by these two sound sources at the same time.

[0083] Among them, after the azimuth information of the sound source of the audio signal is determined, there can be various specific implementation manners for speech enhancement based on the azimuth information, and this application does not limit the specific implementation manner of speech enhancement.

[0084] It can be understood that if it is determined that the first sound source is the same as the second sound source based on the difference information between the first azimuth information and the second azimuth information, then it can be explained that only the azimuth information of the sound source has changed. Correspondingly, the audio signal collected by the audio acquisition device can be enhanced for speech based on the second azimuth information to obtain a third speech enhanced signal corresponding to the second azimuth information.

[0085] It should be noted that, for the convenience of distinction in the embodiments of the present application, when the first sound source and the second sound source are different, the voice enhancement signal obtained by enhancing the audio signal based on the first azimuth information is called the first voice enhancement signal, and the voice enhancement signal obtained by enhancing the audio signal based on the second azimuth information is called the second voice enhancement signal. At the same time, when the first sound source is the same as the second sound source, the voice enhancement signal obtained based on the second azimuth information is called the third voice enhancement signal.

[0086] As can be seen from the above, after the first azimuth information of the current sound source is determined in the present application, the difference information between the first azimuth information and the second azimuth information of the sound source determined last time will be determined. If the difference information indicates that the audio signals detected twice come from different sound sources, the audio signal collected by the audio acquisition device will be enhanced separately based on the first azimuth information and the second azimuth information to obtain the voice enhancement signals in these two azimuth information respectively. Therefore, even if sudden noise interference occurs, the effective audio signal emitted by the original sound source will not be weakened, thereby reducing the misjudgment caused by noise interference and avoiding the situation where the interference noise is enhanced while the effective audio is weakened, and further improving the audio processing effect.

[0087] It can be understood that considering that the first sound source may be a noise source, on this basis, in order to ensure the quality of the audio signal emitted by the original second sound source to the greatest extent, after the second voice enhancement signal is obtained in the present application, the second voice enhancement signal can also be reduced by a set amplitude. On this basis, the first voice enhancement signal and the second voice enhancement signal after amplitude reduction can be determined as the voice signal after voice enhancement of the audio signal collected by the audio acquisition device.

[0088] For example, in scenarios such as audio recording, the first voice enhancement signal and the second voice signal after amplitude reduction can be determined as the voice signal after voice enhancement for storage, etc.

[0089] It can be understood that considering that a sudden noise source may last for a period of time, in order to avoid wrongly enhancing the noise signal emitted by the noise source when the azimuth information of the audio signal needs to be located next time, the present application can enhance the audio signal collected by the audio acquisition device based on the first azimuth information and the second azimuth information respectively within a certain duration.

[0090] For example, a set duration can be set. The set duration can be set according to needs.

[0091] On this basis, when it is determined that the first sound source and the second sound source are different, within the set duration, the audio signals collected by the audio acquisition device can be enhanced separately based on the first azimuth information and the second azimuth information.

[0092] Correspondingly, within the set duration, the second voice enhancement signal can be reduced by a set amplitude to obtain a second voice enhancement signal after reduction. On this basis, within the set duration, the first voice enhancement signal and the second voice enhancement signal after reduction will be determined as the voice signal after voice enhancement of the audio signals collected by the audio acquisition device.

[0093] Furthermore, after reaching the set duration, the present application can also combine the first voice enhancement signal and the second voice enhancement signal to determine whether both the first azimuth information and the second azimuth information need voice enhancement. The following uses a flowchart to illustrate this situation.

[0094] As Figure 2 shown, it shows a schematic flowchart of another embodiment of an audio signal processing method of the present application. The method of this embodiment may include:

[0095] S201, determining the first azimuth information of the first sound source of the audio signal currently collected by the audio acquisition device relative to the audio acquisition device.

[0096] S202, determining the difference information between the first azimuth information and the second azimuth information.

[0097] Among them, the second azimuth information is the azimuth information of the second sound source of the audio signal historically collected by the audio acquisition device relative to the audio acquisition device, and the first azimuth information and the second azimuth information are two adjacent azimuth information determined.

[0098] The above steps can refer to the relevant introduction in the previous embodiment and will not be elaborated here.

[0099] S203, when it is determined that the first sound source and the second sound source are different according to the difference information, within the set duration, the audio signals collected by the audio acquisition device are enhanced separately based on the first azimuth information and the second azimuth information to obtain a first voice enhancement signal corresponding to the first azimuth information and a second voice enhancement signal corresponding to the second azimuth information.

[0100] S204, within the set duration, reducing the second voice enhancement signal by a set amplitude to obtain a second voice enhancement signal after reduction.

[0101] S205, determining the first voice enhancement signal and the second voice enhancement signal after reduction as the voice signal after voice enhancement of the audio signals collected by the audio acquisition device.

[0102] It can be understood that within the set duration, the audio acquisition device continuously acquires audio signals. In this application, within the set duration, the audio signals acquired by the audio acquisition device are continuously subjected to voice enhancement based on the first azimuth information and the second azimuth information, which can reduce the situation of misidentifying the noise source as the only valid audio signal for enhancement due to the continuous emission of sound by the noise source.

[0103] In an optional manner, the set duration can be greater than the azimuth detection period for detecting the azimuth information of the audio signal.

[0104] It can be understood that after the azimuth information is determined in each azimuth detection period, the electronic device performs voice enhancement on the audio signal based on the determined azimuth information. Then, after performing voice enhancement based on the first azimuth information and the second azimuth information respectively, when reaching the next azimuth detection period, if the second sound source is noise and this sound source still continuously emits audio signals, then the second azimuth information corresponding to the audio signals emitted by the sound source may be determined as the azimuth information that needs voice enhancement, thus wrongly enhancing the noise signal.

[0105] Based on this, if the set duration is set to be greater than one azimuth detection period, then the situation of only enhancing the noise signals emitted by the noise source within the next azimuth detection period after detecting the first azimuth information can be reduced.

[0106] S206, after reaching the set duration, combine the first voice enhancement signal and the second voice enhancement signal, and determine at least one target azimuth information that needs voice enhancement from the first azimuth information and the second azimuth information.

[0107] It can be understood that the first voice enhancement signal is obtained by performing voice enhancement on the audio signal based on the first azimuth information. Therefore, the first voice enhancement signal mainly represents the characteristics of the audio signals emitted by the first sound source. Based on the characteristics of the audio signals emitted by the first sound source represented by the first voice enhancement signal, it can be judged whether the first sound source is still emitting sound, and the volume of the audio signals emitted by the first sound source, etc., such as the quality of the audio signals.

[0108] If the first sound source does not continue to emit an audio signal after reaching the set duration, then there is naturally no need to continue performing voice enhancement based on the first azimuth information. Or, if it is determined based on the first voice enhancement signal that the volume of the audio signal emitted by the first sound source is small or the audio quality is poor, etc., it indicates that the first sound source is not the sound source that needs to be concerned about or may be a suddenly appearing noise source, and the first azimuth information also needs to be used as the target azimuth information for which voice enhancement is required. On the contrary, if the first voice enhancement signal indicates that the first sound source is still emitting an audio signal, or that the audio quality of the emitted audio signal is relatively high, etc., then the first azimuth information can be determined as the target azimuth information.

[0109] Similarly, the second voice enhancement signal can indicate the characteristics of the audio signal emitted by the second sound source. After reaching the set duration, if it is determined based on the second voice enhancement signal that the second sound source does not output an audio signal, or that the volume of the audio signal emitted by the second sound source is too small or the audio quality is poor, then there is no need to continue using the second sound source as the sound source of the effective audio signal, and naturally there is no need to use the second azimuth information as the target azimuth information for which voice enhancement is required. On the contrary, the second azimuth information can be determined as the target azimuth information for which voice enhancement is required.

[0110] S207, perform voice enhancement on the audio signal collected by the audio collection device respectively based on the at least one target azimuth information.

[0111] For example, if the target azimuth information is the first azimuth information, perform voice enhancement on the audio signal collected by the audio collection device based on the first azimuth information; if the target azimuth information is the second azimuth information, perform voice enhancement on the audio signal collected by the audio collection device based on the second azimuth information.

[0112] Of course, if the target azimuth information includes the first azimuth information and the second azimuth information, voice enhancement can be performed on the audio signal collected by the audio collection device respectively based on the first azimuth information and the second azimuth information.

[0113] It can be understood that in step S207, after performing voice enhancement on the audio signal collected by the audio collection device based on the target azimuth information, the present application can, according to the requirements of the actual scenario, store the voice-enhanced audio signal or transmit it to the opposite-end electronic device of the audio conference, etc., without limitation in this regard.

[0114] It can be understood that the set duration described above can be set as needed. In an alternative way, the set duration is set to be greater than the duration of one azimuth detection period. When the set duration is set to be greater than the duration of one azimuth detection period, then whether the first voice enhancement signal and the second voice enhancement signal continuously emit audio signals, or the quality of the continuously emitted audio signals, can be used to more reasonably and accurately determine that the sounds emitted by the first sound source and the second sound source belong to valid audio signals, and then the azimuth information for audio enhancement can be adjusted.

[0115] It can be understood that after step S207, if the azimuth information detection moment arrives based on the azimuth detection period, then the electronic device can re-detect the azimuth information corresponding to the sound source of the audio signal collected by the audio collection device. At the same time, the audio signal can be enhanced by voice in combination with the detected azimuth information. On this basis, the electronic device can also store the detected azimuth information as historical azimuth information, so as to continue to detect whether there is a newly emerged sound source in combination with the historical azimuth information later and perform the operations of this embodiment, which will not be elaborated here.

[0116] From the above content, it can be seen that in this embodiment, after determining the first azimuth information of the first sound source of the collected audio signal relative to the audio collection device, if the first sound source and the second sound source corresponding to the second azimuth information determined most recently are not the same, the present application will perform voice enhancement on the audio signal collected by the audio collection device based on the first azimuth information and the second azimuth information respectively within the set duration, so as to reduce the situation where noise signals are enhanced while effective audio signals are weakened.

[0117] In addition, after the set duration is reached, the present application will also combine the characteristics of the first voice enhancement signal and the second voice enhancement signal to determine one or two target azimuth information that actually needs to be enhanced among the first azimuth information and the second azimuth information, so as to more reasonably determine the azimuth information for voice enhancement.

[0118] It can be understood that after the set duration is reached, for any one of the first sound source and the second sound source, whether the sound source still emits an audio signal is a reliable basis for judging whether the sound source is a valid sound source. Based on this, the present application takes the example of selecting target azimuth information based on whether the sound source emits an audio signal after the set duration is reached for illustration.

[0119] As Figure 3 shown, it shows another schematic flowchart of a method for processing an audio signal according to the present application. The method of this embodiment may include:

[0120] S301. Determine the first azimuth information of the first sound source of the audio signal currently collected by the audio acquisition device relative to the audio acquisition device.

[0121] S302. Determine the difference information between the first azimuth information and the second azimuth information.

[0122] The second azimuth information is the azimuth information of the second sound source of the audio signal historically collected by the audio acquisition device relative to the audio acquisition device, and the first azimuth information and the second azimuth information are two adjacent azimuth information determined.

[0123] S303. Detect whether the azimuth change amount represented by the difference information exceeds a set threshold. If so, execute step S304; if not, execute step S308.

[0124] The set threshold can be set as needed.

[0125] Among them, the azimuth change amount exceeding the set threshold indicates that the first sound source and the second sound source are different; the azimuth change amount not exceeding the set threshold indicates that the first sound source and the second sound source are the same.

[0126] It can be understood that this embodiment takes an implementation manner of judging whether the first sound source and the second sound source are the same as an example for illustration, and the other implementation manners mentioned above are also applicable to this embodiment, which will not be elaborated here.

[0127] S304. Within a set duration, perform voice enhancement on the audio signal collected by the audio acquisition device respectively based on the first azimuth information and the second azimuth information to obtain a first voice enhancement signal corresponding to the first azimuth information and a second voice enhancement signal corresponding to the second azimuth information.

[0128] S305. Reduce the second voice enhancement signal by a set amplitude, and determine the first voice enhancement signal and the second voice enhancement signal after the reduction as the voice signal after voice enhancement of the audio signal collected by the audio acquisition device.

[0129] It can be understood that within the set duration, for the audio signal sampled by the audio acquisition device each time, after step S304 is executed, step S305 needs to be executed.

[0130] S306. After reaching the set duration, combine the first voice enhancement signal and the second voice enhancement signal, and determine at least one target azimuth information of the audio acquisition device that can collect the audio signal at the current moment from the first azimuth information and the second azimuth information.

[0131] Among them, the target azimuth information is the azimuth information that needs voice enhancement.

[0132] S307, perform voice enhancement on the audio signals collected by the audio collection device respectively based on the at least one target orientation information.

[0133] It can be understood that after reaching the set duration, if the first voice enhancement signal is not empty, it indicates that the first sound source still emits audio signals. On this basis, the probability that the first sound source is a burst noise source is relatively small, and the first orientation information corresponding to the first sound source needs to be retained as the target orientation information; otherwise, the first orientation information may not be determined as the target orientation information.

[0134] Similarly, if the second voice enhancement signal is not empty, it indicates that the second sound source is also still emitting sounds continuously. Then, the audio signals emitted by the second sound source need to be extracted from the collected audio signals. Therefore, the second orientation information corresponding to the second sound source also needs to be determined as the target orientation information. Conversely, if the second voice enhancement signal is empty, it indicates that the second sound source here has stopped emitting audio signals, and voice enhancement based on the second orientation information needs to be performed again.

[0135] In an alternative manner, in order to more accurately determine whether the first sound source and the second sound source still continue to emit audio, the first voice enhancement signal and the second voice enhancement signal determined in the most recent set number of times can also be combined for analysis. Specifically, it can be divided into the following three possible situations:

[0136] In a possible situation, if the first voice enhancement signals determined in the most recent set number of times are not all empty, and the second voice enhancement signals determined in the most recent set number of times are not all empty, the first orientation information and the second orientation information are determined as the target orientation information that needs voice enhancement. Voice enhancement can be performed on the audio signals collected by the audio collection device respectively based on the first orientation information and the second orientation information to perform reasonable voice enhancement on the overlapping audio emitted by the two sound sources.

[0137] In another possible situation, if the first voice enhancement signals determined in the most recent set number of times are all empty, and the second voice enhancement signals determined in the most recent set number of times are not all empty, the second orientation information is determined as the target orientation information that needs voice enhancement. In this case, it indicates that there is no audio signal output from the first sound source. Therefore, only voice enhancement is performed on the collected audio signals based on the second orientation information.

[0138] In another possible situation, if the first voice enhancement signals determined in the most recent set number of times are not all empty, and the second voice enhancement signals determined in the most recent set number of times are all empty, the first orientation information is determined as the target orientation information that needs voice enhancement. Correspondingly, only voice enhancement is performed on the collected audio signals based on the first orientation information.

[0139] Among them, the most recent setting times can be set as needed. For example, the most recent setting times can be the most recent time, or the most recent two times, etc.

[0140] It can be understood that if the target azimuth information includes the first azimuth information and the second azimuth information, it means that the first sound source and the second sound source emit sounds simultaneously. That is, the audio signal collected by the audio acquisition device is a superimposed signal of the audio signals emitted by the two sound sources. On this basis, the audio signals emitted by these two sound sources should both belong to the audio signals that need to be extracted. Therefore, after reaching the set duration, after performing speech enhancement on the audio signal collected by the audio acquisition device based on the first azimuth information and the second azimuth information, the present application will no longer reduce the amplitude of the speech enhancement signal on any azimuth information, but will determine the speech enhancement signals obtained on these two azimuth information as the final speech enhancement signals after speech enhancement.

[0141] S308. When the first sound source is the same as the second sound source, perform speech enhancement on the audio signal collected by the audio acquisition device based on the second azimuth information to obtain a third speech enhancement signal corresponding to the second azimuth information.

[0142] It can be understood that when the first sound source is the same as the second sound source, within one azimuth detection period, speech enhancement can be performed on the audio signal collected by the audio acquisition device based on the second azimuth information. When reaching the next azimuth detection period, the above operations of the embodiments of the present application can be repeatedly executed based on the detected azimuth information, which will not be elaborated here.

[0143] It can be understood that the method of the embodiments of the present application can be applied to various application scenarios. For the sake of easy understanding, the following takes the scenario of a voice conference as an example for illustration.

[0144] It can be understood that in the voice conference scenario, it is necessary to first establish a voice conference channel between the electronic device and at least one other electronic device. On this basis, after performing speech enhancement on the audio signal based on the solution of the present application, the speech enhancement signal after speech enhancement in any of the above-mentioned situations can be transmitted to the at least one other electronic device in the voice conference based on the voice conference channel.

[0145] For example, based on the voice conference channel, send the first speech enhancement signal and the second speech enhancement signal after amplitude reduction to at least one other electronic device; or, based on the voice conference channel, transmit the third speech enhancement signal after performing speech enhancement on the audio signal based on the target azimuth information to at least one other electronic device.

[0146] For ease of understanding, the following will be described in conjunction with this application scenario, taking the azimuth information of the sound source of the determined audio signal relative to the audio collection device as the direction information as an example.

[0147] As Figure 4 shown, it shows another schematic flowchart of the audio signal processing method provided by the embodiments of the present application. The method of this embodiment is applied to an electronic device, and this embodiment may include:

[0148] S401, establish a voice conference channel between the electronic device and at least one other electronic device.

[0149] S402, determine the first direction information of the first sound source of the audio signal currently collected by the audio collection device of the electronic device relative to the audio collection device.

[0150] S403, determine the angle difference between the first direction information and the second direction information corresponding to the audio signal determined most recently in history.

[0151] The second direction information is the direction information of the second sound source of the audio signal historically collected by the audio collection device relative to the audio collection device, and the first direction information and the second direction information are the directions determined in two adjacent times.

[0152] S404, if the angle difference is greater than the angle threshold, within a set duration, perform voice enhancement on the audio signal collected by the audio collection device respectively based on the first direction information and the second direction information to obtain a first voice enhancement signal corresponding to the first direction information and a second voice enhancement signal corresponding to the second direction information.

[0153] The angle threshold can be set as needed.

[0154] S405, reduce the second voice enhancement signal by a set amplitude, and send the first voice enhancement signal and the second voice enhancement signal after the reduction through the voice conference channel to the at least one other electronic device.

[0155] S405, at the end of the set duration, combine the first voice enhancement signal and the second voice enhancement signal, and determine at least one target direction information at which the audio collection device can collect an audio signal at the current moment from the first direction information and the second direction information.

[0156] This step can refer to the relevant introduction in the previous Figure 3 embodiments and will not be elaborated here.

[0157] S406, perform voice enhancement on the audio signal collected by the audio collection device respectively based on the at least one target direction information, and transmit the voice enhancement signal after voice enhancement to at least one other electronic device.

[0158] It can be understood that, for the convenience of distinction, the voice enhancement signal obtained by performing voice enhancement on the audio signal based on the target direction information may be referred to as the fourth voice enhancement signal. Of course, if there are multiple pieces of target direction information, the voice enhancement signal obtained based on each piece of target direction information is a fourth voice enhancement signal.

[0159] It should be noted that in this embodiment, the azimuth information is taken as an example of the direction information. However, it can be understood that replacing the direction information with the azimuth information is also applicable to this embodiment and will not be elaborated here.

[0160] S407, if the angle difference is not greater than the angle threshold, perform voice enhancement on the audio signal collected by the audio collection device based on the second direction information, and transmit the obtained third voice enhancement signal to the at least one other electronic device through the voice conference channel.

[0161] Corresponding to an audio signal processing method provided by an embodiment of the present application, the present application also provides an audio signal processing device.

[0162] As Figure 5 shown, it shows a schematic structural diagram of a composition of an audio signal processing device of the present application. The device is applied to an electronic device, and the device may include:

[0163] An azimuth determination unit 501, configured to determine the first azimuth information of the first sound source of the audio signal currently collected by the audio collection device relative to the audio collection device;

[0164] A difference determination unit 502, configured to determine the difference information between the first azimuth information and the second azimuth information, where the second azimuth information is the azimuth information of the second sound source of the audio signal historically collected by the audio collection device relative to the audio collection device, and the first azimuth information and the second azimuth information are two adjacent pieces of determined azimuth information;

[0165] A sound source discrimination unit 503, configured to determine whether the first sound source is the same as the second sound source according to the difference information;

[0166] A first voice enhancement unit 504, configured to, when the first sound source is different from the second sound source, perform voice enhancement on the audio signal collected by the audio collection device based on the first azimuth information and the second azimuth information respectively, to obtain a first voice enhancement signal corresponding to the first azimuth information and a second voice enhancement signal corresponding to the second azimuth information.

[0167] In a possible implementation manner, the device further includes:

[0168] A second voice enhancement unit, configured to perform voice enhancement on the audio signal collected by the audio acquisition device based on the second azimuth information when the first sound source is the same as the second sound source, so as to obtain a third voice enhancement signal corresponding to the second azimuth information.

[0169] In another possible implementation manner, the sound source discrimination unit includes:

[0170] A change detection unit, configured to detect whether the azimuth change amount characterized by the difference information exceeds a set threshold; wherein, that the azimuth change amount exceeds the set threshold indicates that the first sound source is different from the second sound source; that the azimuth change amount does not exceed the set threshold indicates that the first sound source is the same as the second sound source.

[0171] In another possible implementation manner, the first voice enhancement unit includes:

[0172] A first voice enhancement subunit, configured to perform voice enhancement on the audio signal collected by the audio acquisition device based on the first azimuth information and the second azimuth information respectively within a set duration;

[0173] The device further includes:

[0174] An amplitude reduction unit, configured to reduce the second voice enhancement signal by a set amplitude within the set duration to obtain a second voice enhancement signal after amplitude reduction;

[0175] An enhancement signal confirmation unit, configured to determine the first voice enhancement signal and the second voice enhancement signal after amplitude reduction as the voice signal after voice enhancement of the audio signal collected by the audio acquisition device.

[0176] In another possible implementation manner, the first voice enhancement unit includes:

[0177] A first voice enhancement subunit, configured to perform voice enhancement on the audio signal collected by the audio acquisition device based on the first azimuth information and the second azimuth information respectively within a set duration;

[0178] The device further includes:

[0179] An azimuth re-determination unit, configured to, after reaching the set duration, combine the first voice enhancement signal and the second voice enhancement signal to determine at least one target azimuth information that needs voice enhancement from the first azimuth information and the second azimuth information;

[0180] A third voice enhancement unit, configured to perform voice enhancement on the audio signal collected by the audio acquisition device based on the at least one target azimuth information respectively.

[0181] In a possible implementation, the azimuth re-determination unit is specifically configured to, after reaching the set duration, combine the first voice enhancement signal and the second voice enhancement signal, and determine, from the first azimuth information and the second azimuth information, at least one target azimuth information at which the audio acquisition device can acquire an audio signal at the current moment, where the target azimuth information is the azimuth information that needs voice enhancement.

[0182] In another possible implementation, the azimuth re-determination unit includes:

[0183] The first azimuth re-determination unit is configured to, if the first voice enhancement signals determined in the most recent set number of times are not all empty, and the second voice enhancement signals determined in the most recent set number of times are not all empty, determine the first azimuth information and the second azimuth information as the target azimuth information that needs voice enhancement;

[0184] The second azimuth re-determination unit is configured to, if the first voice enhancement signals determined in the most recent set number of times are all empty, and the second voice enhancement signals determined in the most recent set number of times are not all empty, determine the second azimuth information as the target azimuth information that needs voice enhancement;

[0185] The third azimuth re-determination unit is configured to, if the first voice enhancement signals determined in the most recent set number of times are not all empty, and the second voice enhancement signals determined in the most recent set number of times are all empty, determine the first azimuth information as the target azimuth information that needs voice enhancement.

[0186] In another possible implementation, the device further includes: a conference establishment unit configured to establish a voice conference channel between the electronic device and at least one other electronic device before the azimuth determination unit determines the first azimuth information;

[0187] The enhancement signal confirmation unit includes:

[0188] The enhancement signal confirmation sub-unit is configured to transmit the first voice enhancement signal and the second voice enhancement signal after amplitude reduction to the at least one other electronic device through the voice conference channel;

[0189] The device further includes:

[0190] The signal transmission unit is configured to, after the third voice enhancement unit performs voice enhancement on the audio signal acquired by the audio acquisition device respectively based on the at least one target azimuth information, transmit the fourth voice enhancement signal obtained by performing voice enhancement on the audio signal based on the target azimuth information to the at least one other electronic device through the voice conference channel.

[0191] In another aspect, the present application further provides an electronic device, such as Figure 6As shown, it shows a schematic diagram of a composition structure of the electronic device. The electronic device can be any type of electronic device, and the electronic device at least includes a processor 601 and a memory 602;

[0192] Among them, the processor 601 is used to execute the audio signal processing method in any one of the above embodiments.

[0193] The memory 602 is used to store the programs required for the processor to perform operations.

[0194] It can be understood that the electronic device may further include a display unit 603 and an input unit 604.

[0195] Of course, the electronic device may also have Figure 6 more or fewer components, and no restrictions are imposed on this.

[0196] On the other hand, the present application also provides a computer-readable storage medium. At least one instruction, at least one segment of program, code set or instruction set is stored in the computer-readable storage medium. The at least one instruction, the at least one segment of program, the code set or instruction set is loaded and executed by the processor to implement the audio signal processing method described in any one of the above embodiments.

[0197] The present application also proposes a computer program. The computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. When the computer program runs on the electronic device, it is used to execute the audio signal processing method in any one of the above embodiments.

[0198] It should be noted that the various embodiments in this specification are all described in a progressive manner. Each embodiment focuses on the differences from other embodiments, and the same or similar parts between the various embodiments can be referred to each other. At the same time, the features recorded in the various embodiments in this specification can be replaced or combined with each other, enabling those skilled in the art to implement or use the present application. For device-type embodiments, since they are basically similar to method embodiments, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiments.

[0199] Finally, it should also be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements not only includes those elements but also includes other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.

[0200] The foregoing description of the disclosed embodiments enables those skilled in the art to implement or use the present application. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. An audio signal processing method, comprising: Determining first azimuth information of a first sound source of an audio signal currently collected by an audio acquisition device relative to the audio acquisition device; The first azimuth information at least includes: direction information of the first sound source relative to the audio acquisition device; Determining difference information between the first azimuth information and second azimuth information, wherein the second azimuth information is azimuth information of a second sound source of an audio signal historically collected by the audio acquisition device relative to the audio acquisition device, and the first azimuth information and the second azimuth information are two adjacent azimuth information determined; Determining whether the first sound source and the second sound source are the same according to the difference information; When the first sound source and the second sound source are different, respectively performing voice enhancement on the audio signal collected by the audio acquisition device based on the first azimuth information and the second azimuth information to obtain a first voice enhancement signal corresponding to the first azimuth information and a second voice enhancement signal corresponding to the second azimuth information.

2. The method according to claim 1, further comprising: When the first sound source and the second sound source are the same, performing voice enhancement on the audio signal collected by the audio acquisition device based on the second azimuth information to obtain a third voice enhancement signal corresponding to the second azimuth information.

3. The method according to claim 1, wherein determining whether the first sound source and the second sound source are the same according to the difference information comprises: Detecting whether an azimuth change amount represented by the difference information exceeds a set threshold; Wherein, the azimuth change amount exceeding the set threshold indicates that the first sound source and the second sound source are different; The azimuth change amount not exceeding the set threshold indicates that the first sound source and the second sound source are the same.

4. The method according to claim 1, wherein respectively performing voice enhancement on the audio signal collected by the audio acquisition device based on the first azimuth information and the second azimuth information comprises: Within a set time period, respectively performing voice enhancement on the audio signal collected by the audio acquisition device based on the first azimuth information and the second azimuth information; The method further comprises: Within the set time period, reducing the second voice enhancement signal by a set amplitude to obtain a second voice enhancement signal after reduction; Determining the first voice enhancement signal and the second voice enhancement signal after reduction as the voice signal after voice enhancement of the audio signal collected by the audio acquisition device.

5. The method according to claim 1 or 4, wherein respectively performing voice enhancement on the audio signal collected by the audio acquisition device based on the first azimuth information and the second azimuth information comprises: Within a set time period, respectively performing voice enhancement on the audio signal collected by the audio acquisition device based on the first azimuth information and the second azimuth information; The method further comprises: After reaching the set time period, combining the first voice enhancement signal and the second voice enhancement signal, and determining at least one target azimuth information that needs voice enhancement from the first azimuth information and the second azimuth information; Perform voice enhancement on the audio signal collected by the audio collection device respectively based on the at least one target orientation information.

6. The method according to claim 5, wherein the combining the first voice enhancement signal and the second voice enhancement signal to determine at least one target orientation information that requires voice enhancement from the first orientation information and the second orientation information includes: Combining the first voice enhancement signal and the second voice enhancement signal to determine at least one target orientation information at which the audio collection device can collect an audio signal at the current moment from the first orientation information and the second orientation information, where the target orientation information is the orientation information that requires voice enhancement.

7. The method according to claim 5, wherein combining the first voice enhancement signal and the second voice enhancement signal to determine at least one target orientation information at which the audio collection device can collect an audio signal at the current moment from the first orientation information and the second orientation information includes: If the first voice enhancement signals determined in the most recent set number of times are not all empty, and the second voice enhancement signals determined in the most recent set number of times are not all empty, determine the first orientation information and the second orientation information as the target orientation information that requires voice enhancement; If the first voice enhancement signals determined in the most recent set number of times are all empty, and the second voice enhancement signals determined in the most recent set number of times are not all empty, determine the second orientation information as the target orientation information that requires voice enhancement; If the first voice enhancement signals determined in the most recent set number of times are not all empty, and the second voice enhancement signals determined in the most recent set number of times are all empty, determine the first orientation information as the target orientation information that requires voice enhancement.

8. The method according to claim 5, further comprising, before determining the first orientation information: Establish a voice conference channel between the electronic device and at least one other electronic device; The determining the first voice enhancement signal and the second voice enhancement signal after amplitude reduction as the voice signal after voice enhancement of the audio signal collected by the audio collection device includes: Transmit the first voice enhancement signal and the second voice enhancement signal after amplitude reduction to the at least one other electronic device through the voice conference channel; After the performing voice enhancement on the audio signal collected by the audio collection device respectively based on the at least one target orientation information, further includes: Based on the voice conference channel, transmit the fourth voice enhancement signal after performing voice enhancement on the audio signal based on the target orientation information to the at least one other electronic device.

9. An audio signal processing device, including: An orientation determination unit, configured to determine the first orientation information of the first sound source of the audio signal currently collected by the audio collection device relative to the audio collection device; The first orientation information at least includes: the direction information of the first sound source relative to the audio collection device; A difference determination unit, configured to determine the difference information between the first orientation information and the second orientation information, where the second orientation information is the orientation information of the second sound source of the audio signal historically collected by the audio collection device relative to the audio collection device, and the first orientation information and the second orientation information are two adjacent orientation information determined; A sound source discrimination unit, configured to determine whether the first sound source is the same as the second sound source according to the difference information; A first voice enhancement unit, configured to, when the first sound source is different from the second sound source, perform voice enhancement on the audio signal collected by the audio acquisition device respectively based on the first orientation information and the second orientation information, so as to obtain a first voice enhancement signal corresponding to the first orientation information and a second voice enhancement signal corresponding to the second orientation information.

10. An electronic device, at least including a memory and a processor; Among them, The processor is configured to execute the audio signal processing method according to any one of claims 1 to 8 above; The memory is configured to store a program required for the processor to perform operations.

Citation Information

Patent Citations

  • Voice signal processing method, device and system and computer readable storage medium

    CN113012700A

  • Voice processing device, audio and video output apparatus, communication system, and sound processing method

    US20180040332A1