Stereo processing method and apparatus

By acquiring the inter-channel difference values ​​of the stereo audio signal and adjusting the left and right channel audio signals, the problem of inconsistent stereo output caused by irregular microphone arrays was solved, thus improving the performance consistency and effect of stereo.

WO2026156511A1PCT designated stage Publication Date: 2026-07-30BEIJING XIAOMI MOBILE SOFTWARE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
BEIJING XIAOMI MOBILE SOFTWARE CO LTD
Filing Date
2025-01-21
Publication Date
2026-07-30

AI Technical Summary

Technical Problem

Due to the irregularity and inconsistent performance of multiple microphone arrays, the stereo audio signal output performance is inconsistent in the left and right directions, resulting in inconsistent sound field and reduced performance.

Method used

By acquiring stereo audio signals, the channel difference between the left and right audio signals is determined, and adjustments to the left and right audio signals are made based on the difference to ensure consistent performance of the stereo audio signals in the left and right directions.

Benefits of technology

It achieves consistent performance of stereo audio signals in the left and right directions, avoids performance degradation, and improves the effect of stereo output.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025073792_30072026_PF_FP_ABST
    Figure CN2025073792_30072026_PF_FP_ABST
Patent Text Reader

Abstract

Provided in the present disclosure are a stereo processing method and apparatus. The method comprises: acquiring a stereo audio signal, wherein the stereo audio signal comprises a left-channel audio signal and a right-channel audio signal; determining an inter-channel difference value between the left-channel audio signal and the right-channel audio signal; and on the basis of the inter-channel difference value, determining whether to adjust at least one of the left-channel audio signal and the right-channel audio signal. In this way, whether to adjust at least one of the left-channel audio signal and the right-channel audio signal can be determined on the basis of the inter-channel difference value between the left-channel audio signal and the right-channel audio signal in the stereo audio signal, so as to ensure that the output performance of the stereo audio signal in a left direction is consistent with the output performance of the stereo audio signal in a right direction, thereby avoiding degradation in the performance of the stereo audio signal.
Need to check novelty before this filing date? Find Prior Art

Description

Stereo processing methods and devices Technical Field

[0001] This disclosure relates to the field of communication technology, and in particular to a stereo processing method and apparatus. Background Technology

[0002] With the rapid development of technology, stereo playback technology has been widely used in multimedia devices. From modern mobile phones, televisions, headphones, and tablets to automobiles, almost all terminals support stereo playback. The growing demand for immersive stereo formats has driven continuous progress in audio technology. To provide a higher quality audio experience, many smartphones are equipped with multiple microphones, such as two to four microphones, forming a large microphone array. This provides a rich data source for audio algorithms, laying the hardware foundation for the application of stereo algorithms on mobile devices. Summary of the Invention

[0003] This disclosure provides a stereo processing method and apparatus that can determine whether to adjust at least one of the left and right channel audio signals based on the channel difference value between the left and right channel audio signals of the stereo audio signal, so as to ensure that the output performance of the stereo audio signal is consistent in the left and right directions and avoid the performance degradation of the stereo audio signal.

[0004] This disclosure presents a stereo processing method and apparatus.

[0005] According to a first aspect of the present disclosure, a stereo processing method is proposed, comprising: acquiring a stereo audio signal, wherein the stereo audio signal includes a left channel audio signal and a right channel audio signal; determining an inter-channel difference value between the left channel audio signal and the right channel audio signal; and determining, based on the inter-channel difference value, whether to adjust at least one of the left channel audio signal and the right channel audio signal.

[0006] In the above embodiments, it is possible to determine whether to adjust at least one of the left and right channel audio signals based on the channel difference value between the left and right channel audio signals of the stereo audio signal, so as to ensure that the output performance of the stereo audio signal is consistent in the left and right directions and avoid the performance degradation of the stereo audio signal.

[0007] According to a second aspect of the present disclosure, a stereo processing apparatus is provided, comprising: a signal acquisition module for acquiring a stereo audio signal, wherein the stereo audio signal includes a left channel audio signal and a right channel audio signal; a difference determination module for determining an inter-channel difference value between the left channel audio signal and the right channel audio signal; and an adjustment determination module for determining, based on the inter-channel difference value, whether to adjust at least one of the left channel audio signal and the right channel audio signal.

[0008] According to a third aspect of the present disclosure, a communication device is provided, comprising: one or more processors, wherein the communication device is configured to perform the method described in the first aspect.

[0009] According to a fourth aspect of the present disclosure, a communication device is provided, comprising: one or more processors; and a memory coupled to the processors, the memory storing instructions which, when executed by the processors, cause the communication device to perform the method described in the first aspect.

[0010] According to a fifth aspect of the present disclosure, a computer storage medium is provided, wherein the computer storage medium stores computer-executable instructions; the computer-executable instructions, when executed by a processor, are capable of implementing the method described in the first aspect.

[0011] According to a sixth aspect of the present disclosure, a computer program product is provided, wherein the computer program product stores a computer program; after being executed by a processor, the computer program is able to implement the method described in the first aspect.

[0012] Additional aspects and advantages of this disclosure will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of this disclosure. Attached Figure Description

[0013] To more clearly illustrate the technical solutions in the embodiments of this disclosure, the accompanying drawings required for the description of the embodiments are introduced below. The following drawings are only some embodiments of this disclosure and do not impose specific limitations on the protection scope of this disclosure.

[0014] Figure 1 is an architecture diagram of a communication system provided in an embodiment of this disclosure;

[0015] Figure 2 is a flowchart of a stereo processing method provided in an embodiment of this disclosure;

[0016] Figure 3 is a flowchart of another stereo processing method provided in an embodiment of this disclosure;

[0017] Figure 4 is a flowchart of another stereo processing method provided in an embodiment of this disclosure;

[0018] Figure 5 is a structural diagram of a stereo processing device provided in an embodiment of this disclosure;

[0019] Figure 6A is a structural diagram of a communication device provided in an embodiment of this disclosure;

[0020] Figure 6B is a structural diagram of a chip provided in an embodiment of this disclosure. Detailed Implementation

[0021] This disclosure presents a stereo processing method and apparatus.

[0022] In a first aspect, embodiments of this disclosure propose a stereo processing method, comprising: acquiring a stereo audio signal, wherein the stereo audio signal includes a left channel audio signal and a right channel audio signal; determining an inter-channel difference value between the left channel audio signal and the right channel audio signal; and determining, based on the inter-channel difference value, whether to adjust at least one of the left channel audio signal and the right channel audio signal.

[0023] In the above embodiments, it is possible to determine whether to adjust at least one of the left and right channel audio signals based on the channel difference value between the left and right channel audio signals of the stereo audio signal, so as to ensure that the output performance of the stereo audio signal is consistent in the left and right directions and avoid the performance degradation of the stereo audio signal.

[0024] In conjunction with some embodiments of the first aspect, in some embodiments, acquiring a stereo audio signal includes: acquiring the original audio signal of a sound source through multiple microphones; and processing the original audio signal to generate a stereo audio signal.

[0025] In the above embodiments, stereo audio signals can be obtained by collecting the original audio signals of the sound source through multiple microphones and processing them.

[0026] In conjunction with some embodiments of the first aspect, in some embodiments, determining whether to adjust at least one of the left channel audio signal and the right channel audio signal based on the inter-channel difference value includes:

[0027] Based on the differences between channels and the mapping relationship, determine the first offset value and the second offset value;

[0028] Based on the first offset value and the second offset value, determine whether to adjust at least one of the left channel audio signal and the right channel audio signal;

[0029] The mapping relationship includes the correspondence between the first offset value of the left channel audio signal and the inter-channel difference value, and the correspondence between the second offset value of the right channel audio signal and the inter-channel difference value.

[0030] In the above embodiments, the mapping relationship between the inter-channel difference value and the first offset value for adjusting the left channel audio signal and the second offset value for adjusting the right channel audio signal can be predetermined. Based on the inter-channel difference value and the mapping relationship, the first offset value and the second offset value are determined. Then, based on the first offset value and the second offset value, it is determined whether to adjust at least one of the left channel audio signal and the right channel audio signal. This can ensure that the stereo audio signal output performance is consistent in the left and right directions and avoid the performance degradation of the stereo audio signal.

[0031] In conjunction with some embodiments of the first aspect, in some embodiments, determining whether to adjust at least one of the left channel audio signal and the right channel audio signal based on the inter-channel difference value includes:

[0032] Based on the target inter-channel difference value and the inter-channel difference value, determine the first offset value corresponding to the left channel audio signal and the second offset value corresponding to the right channel audio signal;

[0033] Based on the first offset value and the second offset value, determine whether to adjust at least one of the left channel audio signal and the right channel audio signal.

[0034] In conjunction with some embodiments of the first aspect, in some embodiments, the above method further includes: determining the target location of the sound source based on the original audio signal;

[0035] Based on the target location of the sound source, determine the target channel difference value between the target left channel audio signal and the target right channel audio signal of the desired target stereo audio signal.

[0036] In the above embodiments, the target position of the sound source can be determined based on the original audio signal, and the target channel difference value between the target left channel audio signal and the target right channel audio signal of the desired target stereo audio signal can be determined based on the target position of the sound source. Furthermore, based on the target channel difference value and the channel difference value, the first offset value corresponding to the left channel audio signal and the second offset value corresponding to the right channel audio signal can be determined. Then, based on the first offset value and the second offset value, it is determined whether to adjust at least one of the left channel audio signal and the right channel audio signal, which can ensure that the output performance of the stereo audio signal is consistent in the left and right directions and avoid the performance degradation of the stereo audio signal.

[0037] In conjunction with some embodiments of the first aspect, in some embodiments, determining the target location of the sound source based on the original audio signal includes: processing the original audio signal based on a sound source localization algorithm to determine the target location of the sound source.

[0038] In conjunction with some embodiments of the first aspect, in some embodiments, the target channel difference value is the difference in the ratio of the target left channel audio signal to the target right channel audio signal, and the channel difference value is the difference in the ratio of the left channel audio signal to the right channel audio signal. Determining a first offset value corresponding to the left channel audio signal and a second offset value corresponding to the right channel audio signal based on the target channel difference value and the channel difference value includes:

[0039] Based on the target inter-channel difference value and the inter-channel difference value, the first offset value is determined as the ratio of the target inter-channel difference value to the inter-channel difference value, and the second offset value is 1; or

[0040] Based on the target inter-channel difference value and the inter-channel difference value, the first offset value is determined to be 1, and the second offset value is the ratio of the inter-channel difference value to the target inter-channel difference value.

[0041] In conjunction with some embodiments of the first aspect, in some embodiments, the target channel difference value is the difference in the ratio of the target right channel audio signal to the target left channel audio signal, and the channel difference value is the difference in the ratio of the right channel audio signal to the left channel audio signal. Determining a first offset value corresponding to the left channel audio signal and a second offset value corresponding to the right channel audio signal based on the target channel difference value and the channel difference value includes:

[0042] Based on the target inter-channel difference value and the inter-channel difference value, the first offset value is determined as the ratio of the inter-channel difference value to the target inter-channel difference value, and the second offset value is 1; or

[0043] Based on the target inter-channel difference value and the inter-channel difference value, the first offset value is determined to be 1, and the second offset value is the ratio of the target inter-channel difference value to the inter-channel difference value.

[0044] In conjunction with some embodiments of the first aspect, in some embodiments, determining whether to adjust at least one of the left channel audio signal and the right channel audio signal based on a first offset value and a second offset value includes at least one of the following:

[0045] If the first offset value is not 1, the left channel audio signal will be adjusted according to the first offset value.

[0046] The first offset value is 1, which determines that the left channel audio signal will not be adjusted;

[0047] If the second offset value is not 1, then the right channel audio signal will be adjusted according to the second offset value.

[0048] The second offset value is 1, which determines that the right channel audio signal will not be adjusted;

[0049] If the first offset value is not 0, the left channel audio signal will be adjusted according to the first offset value.

[0050] The first offset value is 0, which determines that the left channel audio signal will not be adjusted;

[0051] If the second offset value is not 0, the right channel audio signal will be adjusted according to the second offset value.

[0052] The second offset value is 0, which means that the right channel audio signal will not be adjusted.

[0053] In the above embodiments, it is possible to determine whether to adjust at least one of the left channel audio signal and the right channel audio signal based on the first offset value and the second offset value, which can ensure that the output performance of the stereo audio signal is consistent in the left and right directions and avoid the performance degradation of the stereo audio signal.

[0054] In conjunction with some embodiments of the first aspect, in some embodiments, the inter-channel difference value includes one of the following:

[0055] Difference in strength between channels;

[0056] Time difference between channels;

[0057] The parameter values ​​are determined based on the intensity difference and time difference between channels.

[0058] In conjunction with some embodiments of the first aspect, in some embodiments, the difference values ​​between target channels include one of the following:

[0059] Intensity difference between target channels;

[0060] Time difference between target channels;

[0061] The target parameter values ​​are determined based on the intensity difference and time difference between the target channels.

[0062] Secondly, embodiments of this disclosure provide a stereo processing apparatus, comprising: a signal acquisition module for acquiring stereo audio signals, wherein the stereo audio signals include a left channel audio signal and a right channel audio signal; a difference determination module for determining the inter-channel difference value between the left channel audio signal and the right channel audio signal; and an adjustment determination module for determining, based on the inter-channel difference value, whether to adjust at least one of the left channel audio signal and the right channel audio signal.

[0063] Thirdly, a communication device is proposed, comprising: one or more processors, wherein the communication device is used to execute the method described in the first aspect.

[0064] Fourthly, embodiments of this disclosure provide a communication device, which includes: one or more processors; and a memory coupled to the processors, the memory storing instructions that, when executed by the processors, cause the communication device to perform the method described in the first aspect.

[0065] Fifthly, embodiments of this disclosure provide a storage medium storing instructions that, when executed on a communication device, cause the communication device to perform the method described in the first aspect.

[0066] In a sixth aspect, embodiments of this disclosure provide a program product that, when executed by a communication device, causes the communication device to perform the method described in the first aspect.

[0067] In a seventh aspect, embodiments of this disclosure provide a computer program that, when run on a computer, causes the computer to perform the method described in the first aspect.

[0068] Eighthly, embodiments of this disclosure provide a chip or chip system. The chip or chip system includes processing circuitry configured to perform the method described in the first aspect.

[0069] It is understood that the aforementioned communication equipment, communication system, storage medium, program product, etc., are all used to execute the methods proposed in the embodiments of this disclosure. Therefore, the beneficial effects they can achieve can be referred to the beneficial effects in the corresponding methods, and will not be repeated here.

[0070] This disclosure provides a stereo processing method and apparatus. In some embodiments, the terms stereo processing method, information processing method, signal processing method, etc., can be used interchangeably.

[0071] This disclosure is not exhaustive, but merely illustrative of some embodiments, and is not intended to limit the scope of protection of this disclosure. Unless otherwise specified, each step in a particular embodiment can be implemented as an independent embodiment, and the steps can be arbitrarily combined. For example, a solution after removing some steps in a particular embodiment can also be implemented as an independent embodiment, and the order of the steps in a particular embodiment can be arbitrarily interchanged. Furthermore, the optional implementation methods in a particular embodiment can be arbitrarily combined; moreover, the embodiments can be arbitrarily combined, for example, some or all steps of different embodiments can be arbitrarily combined, and a particular embodiment can be arbitrarily combined with the optional implementation methods of other embodiments. In all embodiments of this disclosure, unless otherwise specified or logically conflicting, the terminology and / or descriptions between the embodiments are consistent and can be mutually referenced. Technical features in different embodiments can be combined to form new embodiments based on their inherent logical relationships.

[0072] The terminology used in the embodiments of this disclosure is for the purpose of describing particular embodiments only and is not intended to limit the scope of this disclosure.

[0073] In this embodiment of the disclosure, unless otherwise stated, elements expressed in the singular form, such as "a," "an," "the," "the," "the," "the," "the," "the," "this," etc., can mean "one and only one," or "one or more," "at least one," etc. For example, when using articles such as "a," "an," "the," etc. in translation, the noun following the article can be understood as either a singular expression or a plural expression.

[0074] In the embodiments of this disclosure, "multiple" refers to two or more.

[0075] In some embodiments, the terms “at least one of A or B, at least one of A and B”, “one or more”, “a plurality of”, “multiple”, etc., may be used interchangeably.

[0076] In some embodiments, the notation "at least one of A and B", "A and / or B", "A in one case, B in another", "in response to one case A, in response to another case B", etc., may include the following technical solutions depending on the situation: in some embodiments, A (execute A regardless of whether there is a branch B); in some embodiments, B (execute B regardless of whether there is a branch A); in some embodiments, execution is selected from A and B (A and B are selectively executed); in some embodiments, both A and B are executed. The same applies when there are more branches such as A, B, C, etc.

[0077] In some embodiments, the notation "A or B" may include the following technical solutions, depending on the situation: in some embodiments, A (execute A regardless of whether a branch B exists); in some embodiments, B (execute B regardless of whether a branch A exists); in some embodiments, execution is selected from A and B (A and B are selectively executed). The same applies when there are more branches such as A, B, and C.

[0078] The prefixes "first," "second," etc., used in the embodiments of this disclosure are merely for distinguishing different descriptive objects and do not impose restrictions on the position, order, priority, quantity, or content of the descriptive objects. The description of the descriptive objects is found in the claims or the context of the embodiments, and the use of prefixes should not constitute unnecessary restrictions. For example, if the descriptive object is a "field," the ordinal numbers preceding "field" in "first field" and "second field" do not restrict the position or order of the "fields." "First" and "second" do not restrict whether the "fields" they modify are in the same message, nor do they restrict the order of "first field" and "second field." Similarly, if the descriptive object is a "level," the ordinal numbers preceding "level" in "first level" and "second level" do not restrict the priority between "levels." Furthermore, the number of descriptive objects is not limited by ordinal numbers and can be one or more. For example, in "first device," the number of "devices" can be one or more. Furthermore, the objects modified by different prefixes can be the same or different. For example, if the object being described is "device", then "first device" and "second device" can be the same device or different devices, and their types can be the same or different. Similarly, if the object being described is "information", then "first information" and "second information" can be the same information or different information, and their content can be the same or different.

[0079] In some embodiments, “including A,” “containing A,” “for indicating A,” and “carrying A” can be interpreted as directly carrying A or indirectly indicating A.

[0080] In some embodiments, terms such as "time / frequency" and "time-frequency domain" refer to the time domain and / or frequency domain.

[0081] In some embodiments, terms such as “in response to…”, “in response to determining…”, “in the case of…”, “when…”, “when…”, “if…”, etc. can be used interchangeably. These descriptions all refer to the device making a corresponding action under certain objective circumstances. They do not necessarily limit the time, nor do they require the device to make a judgment action when implementing it, nor do they mean that there must be other limitations.

[0082] In some embodiments, the terms “greater than,” “greater than or equal to,” “not less than,” “more than,” “more than or equal to,” “not less than,” “higher than,” “higher than or equal to,” “not lower than,” and “above” can be used interchangeably, as can the terms “less than,” “less than or equal to,” “not greater than,” “less than,” “less than or equal to,” “not more than,” “lower than,” “lower than or equal to,” “not higher than,” and “below”.

[0083] In some embodiments, devices, etc., may be interpreted as physical or virtual, and their names are not limited to those described in the embodiments. Terms such as “device,” “equipment,” “circuit,” “network element,” “network function,” “network device,” “function,” “node,” “unit,” “section,” “system,” “network,” “chip,” “chip system,” “entity,” and “subject” are interchangeable.

[0084] In some embodiments, "network" can be interpreted as devices included in a network (e.g., access network devices, core network devices, etc.).

[0085] In some embodiments, the terms "access network device (AN device)," "radio access network device (RAN device)," "base station (BS)," "radio base station," "fixed station," "node," "access point," "transmission point (TP)," "reception point (RP)," "transmission / reception point (TRP)," "panel," "antenna panel," "antenna array," "cell," "macro cell," "small cell," "femto cell," "pico cell," "sector," "cell group," "serving cell," "carrier," "component carrier," and "bandwidth part (BWP)" can be used interchangeably.

[0086] In some embodiments, the terms "terminal", "terminal device", "user equipment (UE)", "user terminal", "mobile station (MS)", "mobile terminal (MT)", "subscriber station", "mobile unit", "subscriber unit", "wireless unit", "remote unit", "mobile device", "wireless device", "wireless communication device", "remote device", "mobile subscriber station", "access terminal", "mobile terminal", "wireless terminal", "remote terminal", "handset", "user agent", "mobile client", and "client" can be used interchangeably.

[0087] In some embodiments, access network devices, core network devices, or network devices can be replaced by terminals. For example, embodiments of this disclosure can also be applied to structures where communication between access network devices, core network devices, or network devices and terminals is replaced by communication between multiple terminals (e.g., device-to-device (D2D), vehicle-to-everything (V2X), etc.). In this case, the structure can also be configured such that the terminal has all or part of the functions of the access network device. Furthermore, terms such as "uplink" and "downlink" can be replaced with terms corresponding to communication between terminals (e.g., "sidelink"). For example, uplink channel, downlink channel, etc., can be replaced with sidelink channel, and uplink link, downlink, etc., can be replaced with sidelink link.

[0088] In some embodiments, the terminal may be replaced by an access network device, a core network device, or a network device. In this case, the access network device, core network device, or network device may also be configured to have all or some of the functions of the terminal.

[0089] In some embodiments, the acquisition of data, information, etc., may comply with the laws and regulations of the country where the location is situated.

[0090] In some embodiments, data, information, etc., may be obtained with the user's consent.

[0091] Furthermore, each element, each row, or each column in the table of this disclosure can be implemented as an independent embodiment, and any combination of any element, any row, or any column can also be implemented as an independent embodiment.

[0092] Figure 1 is a schematic diagram of the architecture of a communication system according to an embodiment of the present disclosure.

[0093] Figure 1 is an architecture diagram of a communication system provided in an embodiment of this disclosure.

[0094] As shown in Figure 1, the communication system 100 includes a terminal 101 and a network device 102.

[0095] In some embodiments, terminal 101 includes, but is not limited to, at least one of the following: mobile phone, wearable device, Internet of Things device, car with communication function, smart car, tablet computer, computer with wireless transceiver function, virtual reality (VR) terminal, augmented reality (AR) terminal, wireless terminal in industrial control, wireless terminal in self-driving, wireless terminal in remote medical surgery, wireless terminal in smart grid, wireless terminal in transportation safety, wireless terminal in smart city, and wireless terminal in smart home.

[0096] In some embodiments, network device 102 may include at least one of access network device and core network device.

[0097] In some embodiments, the access network device is, for example, a node or device that connects a terminal to a wireless network. The access network device may include, but is not limited to, at least one of the following in a 5G communication system: evolved Node B (eNB), next-generation eNB (ng-eNB), next-generation Node B (gNB), node B (NB), home node B (HNB), home evolved node B (HeNB), radio backhaul device, radio network controller (RNC), base station controller (BSC), base transceiver station (BTS), base band unit (BBU), mobile switching center, base station in a 6G communication system, open RAN, cloud RAN, base station in other communication systems, and access node in a Wi-Fi system.

[0098] In some embodiments, the access network device may be a satellite.

[0099] In some embodiments, the core network equipment may be a single device, multiple devices, or a group of devices, including all or part of a first network element, a second network element, a third network element, a fourth network element, etc. Network elements may be virtual or physical. The core network may include, for example, at least one of an Evolved Packet Core (EPC), a 5G Core Network (5GCN), and a Next Generation Core (NGC).

[0100] In some embodiments, the first network element is, for example, an access and mobility management function (AMF) network element.

[0101] In some embodiments, the second network element is, for example, a Location Management Function (LMF) network element.

[0102] In some embodiments, the third network element is, for example, a sensing function (SF) network element.

[0103] In some embodiments, the fourth network element is, for example, a network repository function (NRF) network element.

[0104] In some embodiments, the first network element is used to implement terminal access management and mobility management. It is responsible for terminal state maintenance, terminal reachability management, mobility management (MM), forwarding of non-access stratum (NAS) messages, and forwarding of session management (SM) N2 messages.

[0105] In some embodiments, the second network element is used to coordinate and schedule the resources required for the location of the terminal.

[0106] In some embodiments, the third network element is used to perform wireless sensing using access network equipment or terminals to realize sensing services.

[0107] In some embodiments, the fourth network element is used for dynamic registration of network function service capabilities and network function discovery.

[0108] In some embodiments, at least one of the first network element, the second network element, and the third network element can be independent of the core network equipment.

[0109] In some embodiments, at least one of the first network element, the second network element, and the third network element may be part of the core network equipment.

[0110] It is understood that the communication system described in this disclosure is for the purpose of more clearly illustrating the technical solutions of this disclosure, and does not constitute a limitation on the technical solutions proposed in this disclosure. As those skilled in the art will know, with the evolution of system architecture and the emergence of new business scenarios, the technical solutions proposed in this disclosure are also applicable to similar technical problems.

[0111] The following embodiments of this disclosure can be applied to the communication system 100 shown in FIG1, or to some of the main bodies, but are not limited thereto. The main bodies shown in FIG1 are illustrative. The communication system may include all or some of the main bodies in FIG1, or may include other main bodies outside of FIG1. ​​The number and form of each main body are arbitrary. Each main body may be physical or virtual. The connection relationship between the main bodies is illustrative. The main bodies may not be connected or may be connected. The connection can be in any way, it can be a direct connection or an indirect connection, it can be a wired connection or a wireless connection.

[0112] The embodiments disclosed herein can be applied to Long Term Evolution (LTE), LTE-Advanced (LTE-A), LTE-Beyond (LTE-B), Super 3G, IMT-Advanced, 4th Generation Mobile Communication System (4G), 5th Generation Mobile Communication System (5G), 5G New Radio (NR), Future Radio Access (FRA), New-Radio Access Technology (RAT), New Radio (NR), New Radio Access (NX), Future Generation Radio Access (FX), Global System for Mobile Communications (GSM), CDMA2000, Ultra Mobile Broadband (UMB), IEEE 802.11 (Wi-Fi), IEEE 802.16 (WiMAX), IEEE 802.20, and Ultra-Wideband. The technologies used include UWB (Ultra-Wideband), Bluetooth (a registered trademark), public land mobile network (PLMN) networks, device-to-device (D2D) systems, machine-to-machine (M2M) systems, Internet of Things (IoT) systems, vehicle-to-everything (V2X) systems, systems utilizing other communication methods, and next-generation systems built upon them. Furthermore, multiple systems can be combined (e.g., a combination of LTE or LTE-A with 5G).

[0113] In some embodiments, 1. Stereo audio is a two-channel audio format (left channel, right channel) that constructs the desired sound field localization through the relative amplitude and phase (delay) relationship between the left and right channel audio.

[0114] In some embodiments, stereo acquisition uses multiple microphones (mics) that are spatially symmetrical and have consistent performance.

[0115] In related technologies, due to issues such as cost and hardware space limitations, terminals generally suffer from irregular microphone arrays, poor microphone hardware performance, and inconsistent performance between different microphones and the same microphone in different directions.

[0116] The problem of asymmetrical microphone performance can lead to inconsistent output performance of stereo audio signals in the left and right directions, resulting in inconsistent sound fields in the final output stereo audio signal.

[0117] Poor microphone performance will lead to a decrease in the performance of the final stereo audio signal, which will manifest as a smaller final sound field width than the target.

[0118] Therefore, it is urgent to solve the problem that stereo audio signals collected by multiple microphones with irregular distribution, different performance and different hardware configurations are inconsistent in the left and right directions and have insufficient performance when output.

[0119] Based on this, this disclosure provides a stereo processing method and apparatus. The method includes: acquiring a stereo audio signal, wherein the stereo audio signal includes a left channel audio signal and a right channel audio signal; determining the inter-channel difference value between the left channel audio signal and the right channel audio signal; and determining, based on the inter-channel difference value, whether to adjust at least one of the left channel audio signal and the right channel audio signal. This enables the determination of whether to adjust at least one of the left channel audio signal and the right channel audio signal based on the inter-channel difference value of the stereo audio signal, thereby ensuring consistent output performance of the stereo audio signal in the left-right direction and avoiding performance degradation of the stereo audio signal.

[0120] Figure 2 is a schematic diagram illustrating a stereo processing method according to an embodiment of the present disclosure. As shown in Figure 2, the present disclosure relates to a stereo processing method, which includes:

[0121] S201, acquire stereo audio signal.

[0122] In this embodiment of the disclosure, the stereo processing method is executed by a stereo processing device, wherein the stereo processing device can be a communication device, and the communication device can be a terminal, such as a mobile phone, computer, tablet, car, etc.

[0123] In this embodiment of the disclosure, the communication device can execute the stereo processing method by installing a specific application. For example, if the communication device is a terminal, a specific application can be installed on the terminal to execute the stereo processing method.

[0124] In some embodiments, S201, acquiring a stereo audio signal includes: acquiring the original audio signal of a sound source through multiple microphones; and processing the original audio signal to generate a stereo audio signal.

[0125] In some embodiments, the stereo processing device has multiple microphones that can acquire the original audio signal of the sound source and process the original audio signal through a stereo algorithm to generate a stereo audio signal.

[0126] In some embodiments, the stereo processing device directly acquires stereo audio signals.

[0127] For example, a stereo acquisition device, which is different from a stereo processing device, has multiple microphones to acquire the original audio signal of the sound source, processes the original audio signal through a stereo algorithm to generate a stereo audio signal, and provides the generated stereo audio signal to the stereo processing device, so that the stereo processing device can obtain the stereo audio signal.

[0128] In some embodiments, the stereo processing device and the stereo acquisition device may be integrated into one device, or the stereo processing device and the stereo acquisition device may be separate devices.

[0129] The stereo audio signal includes the left channel audio signal and the right channel audio signal.

[0130] S202, determine the channel difference value between the left channel audio signal and the right channel audio signal.

[0131] In this embodiment of the disclosure, a stereo audio signal is obtained, which includes a left channel audio signal and a right channel audio signal. The channel difference value between the left channel audio signal and the right channel audio signal can be determined based on the left channel audio signal and the right channel audio signal.

[0132] In some embodiments, the inter-channel difference is the ratio difference between the left channel audio signal and the right channel audio signal. Examples include differences in energy ratio, amplitude ratio, delay ratio, and ratios of parameter values ​​determined based on intensity (amplitude) and delay values.

[0133] In some embodiments, the inter-channel difference is the magnitude difference between the left channel audio signal and the right channel audio signal. Exemplarily, this includes differences in energy magnitude, amplitude magnitude, delay value magnitude, and differences in parameter values ​​determined based on intensity (amplitude) and delay value.

[0134] In some embodiments, the inter-channel difference value includes one of the following:

[0135] Difference in strength between channels;

[0136] Time difference between channels;

[0137] The parameter values ​​are determined based on the intensity difference and time difference between channels.

[0138] In this embodiment of the disclosure, the inter-channel difference value between the left channel audio signal and the right channel audio signal can be determined based on the left channel audio signal and the right channel audio signal.

[0139] In one possible implementation, the inter-channel intensity difference is the ratio of the intensity of the left channel audio signal to the intensity of the right channel audio signal.

[0140] In another possible implementation, the inter-channel intensity difference is the ratio of the intensity of the right channel audio signal to the intensity of the left channel audio signal.

[0141] In one possible implementation, the inter-channel intensity difference is the absolute value of the difference between the intensity of the left channel audio signal and the intensity of the right channel audio signal.

[0142] In another possible implementation, the inter-channel intensity difference is the difference between the intensity of the left channel audio signal and the intensity of the right channel audio signal.

[0143] In another possible implementation, the inter-channel intensity difference is the difference between the intensity of the right channel audio signal and the intensity of the left channel audio signal.

[0144] In one possible implementation, the inter-channel time difference is the delay value between the left channel audio signal and the right channel audio signal, or the absolute value of the delay value.

[0145] In another possible implementation, the inter-channel time difference is the delay value between the right channel audio signal and the left channel audio signal, or the absolute value of the delay value.

[0146] In another possible implementation, the inter-channel time difference is the difference between the delay value of the left channel audio signal and the delay value of the right channel audio signal.

[0147] In another possible implementation, the inter-channel time difference is the difference between the delay value of the right channel audio signal and the delay value of the left channel audio signal.

[0148] In one possible implementation, the inter-channel difference value is a parameter value determined based on the inter-channel intensity difference and the inter-channel time difference. The parameter value determined based on the inter-channel intensity difference and the inter-channel time difference is determined according to the ratio of the intensity of the left channel audio signal to the intensity of the right channel audio signal, and the delay value between the left channel audio signal and the right channel audio signal.

[0149] In another possible implementation, the inter-channel difference value is a parameter value determined based on the inter-channel intensity difference and the inter-channel time difference. The parameter value determined based on the inter-channel intensity difference and the inter-channel time difference is determined according to the ratio of the intensity of the right channel audio signal to the intensity of the left channel audio signal, and the delay value between the right channel audio signal and the left channel audio signal.

[0150] In one possible implementation, the inter-channel difference value is a parameter value determined based on the inter-channel intensity difference and the inter-channel time difference. The parameter value determined based on the inter-channel intensity difference and the inter-channel time difference is determined by the ratio of the intensity of the left channel audio signal to the intensity of the right channel audio signal, and the absolute value of the delay between the left channel audio signal and the right channel audio signal.

[0151] In another possible implementation, the inter-channel difference value is a parameter value determined based on the inter-channel intensity difference and the inter-channel time difference. The parameter value determined based on the inter-channel intensity difference and the inter-channel time difference is determined by the ratio of the intensity of the right channel audio signal to the intensity of the left channel audio signal, and the absolute value of the delay between the right channel audio signal and the left channel audio signal.

[0152] S203, based on the inter-channel difference value, determine whether to adjust at least one of the left channel audio signal and the right channel audio signal.

[0153] In this embodiment of the disclosure, when the inter-channel difference value of the left channel audio signal and the right channel audio signal is determined, it can be determined whether to adjust at least one of the left channel audio signal and the right channel audio signal based on the inter-channel difference value.

[0154] In some embodiments, S203, determining whether to adjust at least one of the left channel audio signal and the right channel audio signal based on the inter-channel difference value includes: determining a first offset value and a second offset value based on the inter-channel difference value and a mapping relationship; determining whether to adjust at least one of the left channel audio signal and the right channel audio signal based on the first offset value and the second offset value; wherein the mapping relationship includes the correspondence between the first offset value of the left channel audio signal and the inter-channel difference value, and the correspondence between the second offset value of the right channel audio signal and the inter-channel difference value.

[0155] In this embodiment of the disclosure, a mapping relationship is pre-stored, that is, the correspondence between the inter-channel difference value and the first offset value of the left channel audio signal, and the correspondence between the inter-channel difference value and the second offset value of the right channel audio signal are stored. The first offset value and the second offset value can be determined according to the inter-channel difference value and the mapping relationship. Then, based on the first offset value and the second offset value, it is determined whether to adjust at least one of the left channel audio signal and the right channel audio signal.

[0156] For example, as shown in Figure 3, the stereo processing device is equipped with multiple microphones to collect the original audio signal from the sound source. After processing by a stereo algorithm, a stereo audio signal is obtained, which includes a left channel audio signal and a right channel audio signal. The inter-channel difference value V1 is determined based on the left channel audio signal and the right channel audio signal.

[0157] Among them, by pre-storing the mapping relationship between the difference values ​​between different channels and the first offset value for adjusting the left channel audio signal and the second offset value for adjusting the right channel audio signal, that is, when pre-storing the difference values ​​between different channels, the fixed coefficient vl(V1) (i.e. the first offset value) for adjusting the left channel audio signal and the fixed coefficient vr(V1) (i.e. the second offset value) for adjusting the right channel audio signal.

[0158] After determining the first offset value and the second offset value, it is determined whether to adjust at least one of the left channel audio signal and the right channel audio signal by multiplying the left channel audio signal by the first offset value and the right channel audio signal by the second offset value.

[0159] It should be noted that the example in Figure 3 is for illustration only and is not intended to limit the specific embodiments of this disclosure. After determining the first offset value and the second offset value, the delay corresponding to the first offset value can be increased or decreased for the left channel audio signal, and the delay corresponding to the second offset value can be increased or decreased for the right channel audio signal.

[0160] In some embodiments, S203, determining whether to adjust at least one of the left channel audio signal and the right channel audio signal based on the inter-channel difference value includes:

[0161] Determine the target location of the sound source based on the original audio signal;

[0162] Based on the target location of the sound source, determine the target channel difference value between the target left channel audio signal and the target right channel audio signal of the desired target stereo audio signal;

[0163] Based on the target inter-channel difference value and the inter-channel difference value, determine the first offset value corresponding to the left channel audio signal and the second offset value corresponding to the right channel audio signal;

[0164] Based on the first offset value and the second offset value, determine whether to adjust at least one of the left channel audio signal and the right channel audio signal.

[0165] In this embodiment of the present disclosure, when the stereo processing device acquires the original audio signal, it can determine the target location of the sound source, such as the direction of arrival (DOA), based on the original audio signal.

[0166] In some embodiments, determining the target location of a sound source based on the original audio signal includes: processing the original audio signal based on a sound source localization algorithm to determine the target location of the sound source.

[0167] In this embodiment of the disclosure, the target location of the sound source can be determined by running a sound source localization method on the raw audio signals acquired by multiple microphones.

[0168] In some embodiments, a sound source localization algorithm is an algorithm that determines the location of a sound source by analyzing sound signals. This algorithm has applications in many fields, including robot navigation, environmental monitoring, and security surveillance. The basic principle of a sound source localization algorithm is to utilize the characteristics of sound propagation in space, and to deduce the location of the sound source by measuring the differences (relationships) in reception time or intensity of sound at different locations (mics at different locations).

[0169] In some embodiments, the sound source localization algorithm processes the following steps: One is a cross-correlation analysis method: The sound source localization algorithm first performs cross-correlation processing on each pair of microphones in the microphone array. By analyzing the correlation between the received signals, the relative position of the sound source is determined. This step utilizes the basic principle of signal processing, namely, estimating the angle between the sound source and the receiver by comparing the phase difference of the signals. Another method is multi-angle scanning: To comprehensively assess the location of the sound source, the sound source localization algorithm performs a detailed scan of all possible angles. By evaluating the signal response at each angle, the sound source localization algorithm can identify the peaks of the spatial spectrum. The location and intensity of these peaks provide valuable information about the direction of signal arrival.

[0170] In some embodiments, the inter-channel difference value includes one of the following:

[0171] Difference in strength between channels;

[0172] Time difference between channels;

[0173] The parameter values ​​are determined based on the intensity difference and time difference between channels.

[0174] In some embodiments, the target channel difference value includes one of the following:

[0175] Intensity difference between target channels;

[0176] Time difference between target channels;

[0177] The target parameter values ​​are determined based on the intensity difference and time difference between the target channels.

[0178] In some embodiments, the inter-channel difference value is the inter-channel intensity difference, and the target inter-channel difference value is the target inter-channel intensity difference.

[0179] In some embodiments, the inter-channel difference value is the inter-channel time difference, and the target inter-channel difference value is the target inter-channel time difference.

[0180] In some embodiments, the inter-channel difference value is a parameter value determined based on the inter-channel intensity difference and the inter-channel time difference, and the target inter-channel difference value is a target parameter value determined based on the target inter-channel intensity difference and the target inter-channel time difference.

[0181] In this embodiment of the disclosure, the difference between target channels is the intensity difference between target channels, or the energy difference (amplitude difference) between target channels. The difference between target channels of the target left channel audio signal and the target right channel audio signal can be determined based on the target left channel audio signal and the target right channel audio signal.

[0182] In one possible implementation, the intensity difference between target channels is the ratio of the intensity of the target left channel audio signal to the intensity of the target right channel audio signal.

[0183] In another possible implementation, the intensity difference between target channels is the ratio of the intensity of the target right channel audio signal to the intensity of the target left channel audio signal.

[0184] In one possible implementation, the intensity difference between target channels is the absolute value of the difference between the intensity of the target left channel audio signal and the intensity of the target right channel audio signal.

[0185] In another possible implementation, the intensity difference between target channels is the difference between the intensity of the target left channel audio signal and the intensity of the target right channel audio signal.

[0186] In another possible implementation, the intensity difference between target channels is the difference between the intensity of the target right channel audio signal and the intensity of the target left channel audio signal.

[0187] In one possible implementation, the time difference between target channels is the delay value between the target left channel audio signal and the target right channel audio signal, or the absolute value of the delay value.

[0188] In another possible implementation, the time difference between the target channels is the delay value between the target right channel audio signal and the target left channel audio signal, or the absolute value of the delay value.

[0189] In one possible implementation, the target channel difference value is a target parameter value determined based on the target channel intensity difference and the target channel time difference. The target parameter value determined based on the target channel intensity difference and the target channel time difference is determined based on the ratio of the intensity of the target left channel audio signal to the intensity of the target right channel audio signal, and the delay value between the target left channel audio signal and the target right channel audio signal.

[0190] In another possible implementation, the target channel difference value is a parameter value determined based on the target channel intensity difference and the target channel time difference. The target parameter value determined based on the target channel intensity difference and the target channel time difference is determined based on the ratio of the intensity of the target right channel audio signal to the intensity of the target left channel audio signal, and the delay value between the target right channel audio signal and the target left channel audio signal.

[0191] In one possible implementation, the target channel difference value is a target parameter value determined based on the target channel intensity difference and the target channel time difference. The target parameter value determined based on the target channel intensity difference and the target channel time difference is determined based on the ratio of the intensity of the target left channel audio signal to the intensity of the target right channel audio signal, and the absolute value of the delay between the target left channel audio signal and the target right channel audio signal.

[0192] In another possible implementation, the target channel difference value is a target parameter value determined based on the target channel intensity difference and the target channel time difference. The target parameter value determined based on the target channel intensity difference and the target channel time difference is determined by the ratio of the intensity of the target right channel audio signal to the intensity of the target left channel audio signal, and the absolute value of the delay between the target right channel audio signal and the target left channel audio signal.

[0193] For example, as shown in Figure 4, the stereo processing device is equipped with multiple microphones to collect the original audio signal from the sound source. After processing by a stereo algorithm, a stereo audio signal is obtained, which includes a left channel audio signal and a right channel audio signal. The inter-channel difference value v1 is determined based on the left channel audio signal and the right channel audio signal.

[0194] Specifically, based on the original audio signal, the target location of the sound source is determined, and further based on the target location of the sound source, the target channel difference value v0 between the target left channel audio signal and the target right channel audio signal of the desired target stereo audio signal is determined.

[0195] Then, based on the target inter-channel difference value v0 and the inter-channel difference value v1, a first offset value of v0 / v1 is determined for adjusting the left channel audio signal, and a second offset value of 1 is determined for adjusting the right channel audio signal. Then, based on the first and second offset values, it is determined whether to adjust at least one of the left and right channel audio signals. Multiplying the left channel audio signal by the first offset value v0 / v1 and multiplying the right channel audio signal by the second offset value 1 is equivalent to determining that no adjustment is needed for the right channel audio signal.

[0196] It should be noted that the example in Figure 4 is for illustration only and is not intended to limit the specific embodiments of this disclosure. After determining the first offset value and the second offset value, the first offset value can be added to or subtracted from the left channel audio signal, and the second offset value can be added to or subtracted from the right channel audio signal.

[0197] In some embodiments, the target inter-channel difference value is the ratio of the intensity of the target left channel audio signal to the intensity of the target right channel audio signal, and the inter-channel difference value is the ratio of the intensity of the left channel audio signal to the intensity of the right channel audio signal. Determining a first offset value corresponding to the left channel audio signal and a second offset value corresponding to the right channel audio signal based on the target inter-channel difference value and the inter-channel difference value includes:

[0198] Based on the target inter-channel difference value and the inter-channel difference value, the first offset value is determined as the ratio of the target inter-channel difference value to the inter-channel difference value, and the second offset value is 1; or

[0199] Based on the target inter-channel difference value and the inter-channel difference value, the first offset value is determined to be 1, and the second offset value is the ratio of the inter-channel difference value to the target inter-channel difference value.

[0200] In some embodiments, the target inter-channel difference value is the ratio of the intensity of the target right channel audio signal to the intensity of the target left channel audio signal, and the inter-channel difference value is the ratio of the intensity of the right channel audio signal to the intensity of the left channel audio signal. Determining a first offset value corresponding to the left channel audio signal and a second offset value corresponding to the right channel audio signal based on the target inter-channel difference value and the inter-channel difference value includes:

[0201] Based on the target inter-channel difference value and the inter-channel difference value, the first offset value is determined as the ratio of the inter-channel difference value to the target inter-channel difference value, and the second offset value is 1; or

[0202] Based on the target inter-channel difference value and the inter-channel difference value, the first offset value is determined to be 1, and the second offset value is the ratio of the target inter-channel difference value to the inter-channel difference value.

[0203] For example, as shown in Figure 4, taking the inter-channel difference as the inter-channel intensity difference as an example, the inter-channel intensity difference is the ratio of the intensity value L1 of the left channel audio signal to the intensity value R1 of the right channel audio signal, L1 / R1. The target inter-channel intensity difference is the ratio of the intensity value L2 of the target left channel audio signal to the intensity value R2 of the target right channel audio signal, L2 / R2. Based on the target inter-channel intensity difference L2 / R2 and the inter-channel intensity difference L1 / R1, the first offset value is determined to be (L2*R1) / (R2*L1), which is the ratio of the target inter-channel intensity difference to the inter-channel intensity difference, and the second offset value is 1; or based on the target inter-channel intensity difference L2 / R2 and the inter-channel intensity difference L1 / R1, the first offset value is determined to be 1, and the second offset value is (R2*L1) / (L2*R1).

[0204] For example, as shown in Figure 4, taking the inter-channel difference value as the inter-channel intensity difference as an example, the inter-channel intensity difference is the ratio of the intensity value R1 of the right channel audio signal to the intensity value L1 of the left channel audio signal, R1 / L1. The target inter-channel intensity difference is the ratio of the intensity value R2 of the target right channel audio signal to the intensity value L2 of the target left channel audio signal, R2 / L2. Based on the target inter-channel intensity difference R2 / L2 and the inter-channel intensity difference R1 / L1, the target offset value is determined to be (L2*R1) / (R2*L1), and the second offset value is 1; or based on the target inter-channel intensity difference R2 / L2 and the inter-channel intensity difference R1 / L1, the target offset value is determined to be 1, and the second offset value is (L1*R2) / (R1*L2).

[0205] In some embodiments, determining whether to adjust at least one of the left channel audio signal and the right channel audio signal based on a first offset value and a second offset value includes at least one of the following:

[0206] If the first offset value is not 1, the left channel audio signal will be adjusted according to the first offset value.

[0207] The first offset value is 1, which determines that the left channel audio signal will not be adjusted;

[0208] If the second offset value is not 1, then the right channel audio signal will be adjusted according to the second offset value.

[0209] The second offset value is 1, which determines that the right channel audio signal will not be adjusted;

[0210] If the first offset value is not 0, the left channel audio signal will be adjusted according to the first offset value.

[0211] The first offset value is 0, which determines that the left channel audio signal will not be adjusted;

[0212] If the second offset value is not 0, the right channel audio signal will be adjusted according to the second offset value.

[0213] The second offset value is 0, which means that the right channel audio signal will not be adjusted.

[0214] In this embodiment, if the first offset value is not 1, the left channel audio signal is adjusted according to the first offset value. For example, the left channel audio signal is multiplied by the first offset value, or divided by the first offset value. If the first offset value is greater than 1, multiplying the left channel audio signal by the first offset value can be understood as amplifying the left channel audio signal; dividing the left channel audio signal by the first offset value can be understood as reducing the left channel audio signal. If the first offset value is less than 1 and greater than 0, dividing the left channel audio signal by the first offset value can be understood as amplifying the left channel audio signal; multiplying the left channel audio signal by the first offset value can be understood as reducing the left channel audio signal. This ensures consistent output performance of the stereo audio signal in the left and right directions, avoiding performance degradation of the stereo audio signal.

[0215] In this embodiment of the disclosure, the first offset value is 1, indicating that no adjustment is made to the left channel audio signal. For example,

[0216] In this embodiment, if the second offset value is not 1, it is determined that the right channel audio signal will be adjusted according to the second offset value. For example, the right channel audio signal is multiplied by the second offset value, or divided by the second offset value. Wherein, if the second offset value is greater than 1, multiplying the right channel audio signal by the second offset value can be understood as amplifying the right channel audio signal; dividing the right channel audio signal by the second offset value can be understood as reducing the right channel audio signal. Wherein, if the second offset value is less than 1 and greater than 0, dividing the right channel audio signal by the second offset value can be understood as amplifying the right channel audio signal; multiplying the right channel audio signal by the second offset value can be understood as reducing the right channel audio signal. This ensures consistent output performance of the stereo audio signal in the right-to-right direction, avoiding performance degradation of the stereo audio signal.

[0217] In this embodiment of the disclosure, the second offset value is 1, which determines that the right channel audio signal will not be adjusted.

[0218] In this embodiment of the disclosure, if the first offset value is not 0, it is determined that the left channel audio signal will be adjusted according to the first offset value. Exemplarily, the intensity of the left channel audio signal is increased by the first offset value, or the intensity of the left channel is decreased by the first offset value. Wherein, if the first offset value is greater than 0, increasing the intensity of the left channel audio signal by the first offset value can be understood as amplifying the left channel audio signal; decreasing the intensity of the left channel audio signal by the first offset value can be understood as reducing the intensity of the left channel audio signal. Wherein, if the first offset value is less than 0, decreasing the intensity of the left channel audio signal by the first offset value can be understood as amplifying the left channel audio signal; increasing the intensity of the left channel audio signal by the first offset value can be understood as reducing the intensity of the left channel audio signal. This ensures consistent output performance of the stereo audio signal in the left and right directions, avoiding performance degradation of the stereo audio signal.

[0219] In this embodiment of the disclosure, the first offset value is 0, which determines that the left channel audio signal will not be adjusted.

[0220] In this embodiment, if the second offset value is not 0, it is determined that the right channel audio signal will be adjusted according to the second offset value. For example, the second offset value is added to the right channel audio signal intensity, or subtracted from the right channel audio signal intensity. Wherein, if the second offset value is greater than 0, adding the second offset value to the right channel audio signal intensity can be understood as amplifying the right channel audio signal; subtracting the second offset value from the right channel audio signal intensity can be understood as reducing the right channel audio signal intensity. Wherein, if the second offset value is less than 0, subtracting the second offset value from the right channel audio signal intensity can be understood as amplifying the right channel audio signal; adding the second offset value to the right channel audio signal can be understood as reducing the right channel audio signal intensity. This ensures consistent output performance of the stereo audio signal in the right-to-right direction, avoiding performance degradation of the stereo audio signal.

[0221] In this embodiment of the disclosure, the second offset value is 0, which determines that the right channel audio signal will not be adjusted.

[0222] By implementing the embodiments of this disclosure, it is possible to determine whether to adjust at least one of the left and right channel audio signals based on the inter-channel difference value of the left and right channel audio signals of the stereo audio signal, so as to ensure that the output performance of the stereo audio signal is consistent in the left and right directions and avoid the performance degradation of the stereo audio signal.

[0223] Due to cost and hardware space limitations, terminal devices in related technologies generally suffer from irregular microphone arrays, poor microphone hardware performance, and inconsistent performance between different microphones and the same microphone in different directions.

[0224] The asymmetry in the microphone's performance can lead to inconsistent output performance in the left and right directions, resulting in inconsistent sound fields in the final stereo audio output.

[0225] Poor microphone performance leads to reduced stereo audio performance, manifesting as a smaller sound field width than the target sound field. How can we resolve the inconsistency and left-right asymmetry between the stereo audio sound field width and the target sound field?

[0226] The processing flow mainly consists of 3 steps:

[0227] 1. Determine the current sound field.

[0228] 2. Calculate the required offset.

[0229] 3. Add the corresponding electrical displacement (amplitude or delay) to the target channel.

[0230] Example 1:

[0231] The target mobile phone outputs stereo audio via 2 microphones.

[0232] Due to the structural and cost limitations of the mobile phone itself, there is a large error between the two microphones, and one microphone has a larger target frequency band than the other, resulting in insufficient sound field width and a narrower sound field on the left side.

[0233] In a real sound field, a sound source at 60° to the left is stereo reproduced at 10° (target position 60°), and a sound source at 60° to the right is reproduced at 20° (target position 60°).

[0234] 1. In a single-source environment, the left channel has greater energy when the source is on the left and the right channel has greater energy when the source is on the right. Therefore, the sound field can be simplified to be determined by the energy of the left and right channels.

[0235] 2. The offset can be a pre-stored parameter.

[0236] 3. Shift the amplitude of the corresponding audio channel.

[0237] Example 2:

[0238] Furthermore, by using DOA directional data to determine the accurate location of the sound source, the sound field is adjusted.

[0239] Assuming the obtained DOA direction data is 60°, the actual positioning position of the output stereo audio is 10°.

[0240] The difference in audio amplitude between the left and right channels at the target location is L / R = v0, and the difference in audio amplitude between the output left and right channels is L / R = v1.

[0241] Multiply the amplitude of the L channel by v0 / v1 to make the amplitude difference between the left and right channels satisfy the target position's amplitude difference v0.

[0242] This disclosure provides a method for processing audio signals:

[0243] Based on the difference between the actual output sound field and the target sound field, the output sound field is adjusted to match the target sound field.

[0244] The input signal can be a signal processed by a stereo algorithm to enhance the effect of the stereo algorithm.

[0245] It can also be used as a stereo algorithm to process the original audio signals from two microphones into stereo signals.

[0246] This disclosure also proposes an apparatus (also referred to as a communication device, etc.) for implementing any of the above methods. For example, an apparatus is proposed, which includes units or modules for implementing the steps performed by the terminal in any of the above methods.

[0247] It should be understood that the division of units or modules in the above device is only a logical functional division. In actual implementation, they can be fully or partially integrated into a single physical entity, or they can be physically separated. Furthermore, the units or modules in the device can be implemented by a processor calling software: for example, the device includes a processor connected to a memory containing instructions. The processor calls the instructions stored in the memory to implement any of the above methods or to implement the functions of the units or modules in the above device. The processor can be, for example, a general-purpose processor, such as a Central Processing Unit (CPU) or a microprocessor, and the memory can be internal or external to the device. Alternatively, the units or modules in the device can be implemented in the form of hardware circuits. The functionality of some or all of the units or modules can be achieved through the design of these hardware circuits, which can be understood as one or more processors. For example, in one implementation, the hardware circuit is an application-specific integrated circuit (ASIC). The functionality of some or all of the units or modules is achieved through the design of the logical relationships between the components within the circuit. In another implementation, the hardware circuit can be implemented using a programmable logic device (PLD). Taking a field-programmable gate array (FPGA) as an example, it can include a large number of logic gates. The connection relationships between the logic gates are configured through configuration files, thereby achieving the functionality of some or all of the units or modules. All units or modules of the above device can be implemented entirely through processor-called software, entirely through hardware circuits, or partially through processor-called software with the remaining parts implemented through hardware circuits.

[0248] In this embodiment, the processor is a circuit with signal processing capabilities. In one implementation, the processor can be a circuit with instruction read and execute capabilities, such as a Central Processing Unit (CPU), a microprocessor, a graphics processing unit (GPU) (which can be understood as a microprocessor), or a digital signal processor (DSP). In another implementation, the processor can implement certain functions through the logical relationships of hardware circuits. The logical relationships of the aforementioned hardware circuits are fixed or reconfigurable. For example, the processor is a hardware circuit implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD), such as an FPGA. In a reconfigurable hardware circuit, the process of the processor loading a configuration document and configuring the hardware circuit can be understood as the process of the processor loading instructions to implement the functions of some or all of the above units or modules. Furthermore, it can also be a hardware circuit designed for artificial intelligence, which can be understood as an ASIC, such as a Neural Network Processing Unit (NPU), a Tensor Processing Unit (TPU), or a Deep Learning Processing Unit (DPU).

[0249] Figure 5 is a schematic diagram of the structure of the stereo processing device proposed in an embodiment of this disclosure. As shown in Figure 5, the stereo processing device 10 may include at least one of the following: a signal acquisition module 11, a difference determination module 12, and an adjustment determination module 13.

[0250] In some embodiments, the signal acquisition module 11 is used to acquire a stereo audio signal, wherein the stereo audio signal includes a left channel audio signal and a right channel audio signal; the difference determination module 12 is used to determine the inter-channel difference value between the left channel audio signal and the right channel audio signal; and the adjustment determination module 13 is used to determine whether to adjust at least one of the left channel audio signal and the right channel audio signal based on the inter-channel difference value.

[0251] In some embodiments, the signal acquisition module 11 is specifically used to: acquire the original audio signal of the sound source through multiple microphones; process the original audio signal to generate a stereo audio signal.

[0252] In some embodiments, the adjustment determining module 13 is specifically used to: determine a first offset value and a second offset value based on the inter-channel difference value and the mapping relationship; and determine whether to adjust at least one of the left channel audio signal and the right channel audio signal based on the first offset value and the second offset value. The mapping relationship includes the correspondence between the first offset value of the left channel audio signal and the inter-channel difference value, and the correspondence between the second offset value of the right channel audio signal and the inter-channel difference value.

[0253] In some embodiments, the adjustment determining module 13 is specifically used to: determine a first offset value corresponding to the left channel audio signal and a second offset value corresponding to the right channel audio signal based on the target inter-channel difference value and the inter-channel difference value;

[0254] Based on the first offset value and the second offset value, determine whether to adjust at least one of the left channel audio signal and the right channel audio signal.

[0255] In some embodiments, the adjustment determining module 13 is further configured to: determine the target location of the sound source based on the original audio signal;

[0256] Based on the target location of the sound source, determine the target channel difference value between the target left channel audio signal and the target right channel audio signal of the desired target stereo audio signal.

[0257] In some embodiments, the adjustment determination module 13 is specifically used to: process the original audio signal based on the sound source localization algorithm to determine the target location of the sound source.

[0258] In some embodiments, the target inter-channel difference value is the difference in the ratio of the target left channel audio signal to the target right channel audio signal, and the inter-channel difference value is the difference in the ratio of the left channel audio signal to the right channel audio signal. Specifically, the adjustment determining module 13 is used to: determine a first offset value as the ratio of the target inter-channel difference value to the inter-channel difference value, and a second offset value of 1, based on the target inter-channel difference value and the inter-channel difference value; or determine a first offset value of 1 and a second offset value as the ratio of the inter-channel difference value to the target inter-channel difference value, based on the target inter-channel difference value and the inter-channel difference value.

[0259] In some embodiments, the target inter-channel difference value is the difference in the ratio of the target right channel audio signal to the target left channel audio signal, and the inter-channel difference value is the difference in the ratio of the right channel audio signal to the left channel audio signal. Specifically, the adjustment determining module 13 is used to: determine a first offset value as the ratio of the inter-channel difference value to the target inter-channel difference value and a second offset value of 1 based on the target inter-channel difference value and the inter-channel difference value; or determine a first offset value of 1 and a second offset value as the ratio of the target inter-channel difference value to the inter-channel difference value based on the target inter-channel difference value and the inter-channel difference value.

[0260] In some embodiments, the adjustment determining module 13 is specifically used to perform at least one of the following:

[0261] If the first offset value is not 1, the left channel audio signal will be adjusted according to the first offset value.

[0262] The first offset value is 1, which determines that the left channel audio signal will not be adjusted;

[0263] If the second offset value is not 1, then the right channel audio signal will be adjusted according to the second offset value.

[0264] The second offset value is 1, which determines that the right channel audio signal will not be adjusted;

[0265] If the first offset value is not 0, the left channel audio signal will be adjusted according to the first offset value.

[0266] The first offset value is 0, which determines that the left channel audio signal will not be adjusted;

[0267] If the second offset value is not 0, the right channel audio signal will be adjusted according to the second offset value.

[0268] The second offset value is 0, which means that the right channel audio signal will not be adjusted.

[0269] In some embodiments, the inter-channel difference value includes one of the following:

[0270] Difference in strength between channels;

[0271] Time difference between channels;

[0272] The parameter values ​​are determined based on the intensity difference and time difference between channels.

[0273] In some embodiments, the target channel difference value includes one of the following:

[0274] Intensity difference between target channels;

[0275] Time difference between target channels;

[0276] The target parameter values ​​are determined based on the intensity difference and time difference between the target channels.

[0277] Figure 6A is a schematic diagram of the structure of the communication device 5100 proposed in an embodiment of this disclosure. The communication device 5100 may be a stereo processing device (e.g., a terminal), or a chip, chip system, or processor that supports the stereo processing device in implementing any of the above methods. The communication device 5100 can be used to implement the methods described in the above method embodiments, and specific details can be found in the descriptions in the above method embodiments.

[0278] As shown in Figure 6A, the communication device 5100 is used to execute any of the above methods. In some embodiments, the communication device 5100 includes one or more processors 5101. The processor 5101 may be a general-purpose processor or a special-purpose processor, such as a baseband processor or a central processing unit. The baseband processor may be used to process communication protocols and communication data, and the central processing unit may be used to control communication devices (e.g., base stations, baseband chips, terminals, terminal chips, DUs or CUs, etc.), execute programs, and process program data. Optionally, the communication device 5100 is used to execute any of the above methods. Optionally, one or more processors 5101 are used to invoke instructions to cause the communication device 5100 to execute any of the above methods.

[0279] In some embodiments, the communication device 5100 further includes one or more transceivers 5102. When the communication device 5100 includes one or more transceivers 5102, the transceiver 5102 performs at least one of the communication steps such as sending and / or receiving in the above-described method (e.g., the sending and / or receiving steps in S201-S203, but not limited thereto), and the processor 5101 performs at least one of other steps (e.g., steps other than sending and / or receiving in S201-S203, but not limited thereto). In optional embodiments, the transceiver may include a receiver and / or a transmitter, which may be separate or integrated. Optionally, the terms transceiver, transceiver unit, transceiver, transceiver circuit, interface circuit, interface, etc., can be used interchangeably; the terms transmitter, sending unit, transmitter, sending circuit, etc., can be used interchangeably; the terms receiver, receiving unit, receiver, receiving circuit, etc., can be used interchangeably.

[0280] In some embodiments, the communication device 5100 further includes one or more memories 5103 for storing data and / or instructions. Optionally, one or more processors 5101 are used to invoke instructions stored in the memory 5103 to cause the communication device 5100 to perform any of the above methods. Optionally, all or part of the memory 5103 may also be located outside the communication device 5100. In an optional embodiment, the communication device 5100 may include one or more interface circuits 5104. Optionally, the interface circuit 5104 is connected to the memory 5103 and can be used to receive data and / or instructions from the memory 5103 or other devices, and can be used to send data and / or instructions to the memory 5103 or other devices. For example, the interface circuit 5104 can read data and / or instructions stored in the memory 5103 and send the data and / or instructions to the processor 5101.

[0281] The communication device 5100 described in the above embodiments may be a network device or a terminal, but the scope of the communication device 5100 described in this disclosure is not limited thereto, and the structure of the communication device 5100 may not be limited by FIG. 6A. The communication device may be a standalone device or a part of a larger device. For example, the communication device may be: (1) a standalone integrated circuit IC, or chip, or chip system or subsystem; (2) a collection of one or more ICs, optionally, the IC collection may also include storage components for storing data, programs and / or instructions; (3) an ASIC, such as a modem; (4) a module that can be embedded in other devices; (5) a receiver, terminal, smart terminal, cellular phone, wireless device, handheld device, mobile unit, vehicle device, network device, cloud device, artificial intelligence device, etc.; (6) others, etc.

[0282] Figure 6B is a schematic diagram of the structure of the chip 5200 proposed in an embodiment of this disclosure. For cases where the communication device 5100 can be a chip or a chip system, the schematic diagram of the chip 5200 shown in Figure 6B can be referenced, but is not limited thereto.

[0283] Chip 5200 includes one or more processors 5201. Chip 5200 is used to perform any of the methods described above.

[0284] In some embodiments, chip 5200 further includes one or more interface circuits 5202. Optionally, terms such as interface circuit, interface, and transceiver pin can be used interchangeably. In some embodiments, chip 5200 further includes one or more memories 5203 for storing data and / or instructions. Optionally, all or part of the memories 5203 may be located outside of chip 5200. Optionally, the interface circuit 5202 is connected to the memories 5203, and the interface circuit 5202 can be used to receive data and / or instructions from the memories 5203 or other devices, and the interface circuit 5202 can be used to send data and / or instructions to the memories 5203 or other devices. For example, the interface circuit 5202 can read data and / or instructions stored in the memories 5203 and send the data and / or instructions to the processor 5201.

[0285] In some embodiments, the interface circuit 5202 performs at least one of the communication steps such as sending and / or receiving in the above-described method (e.g., the sending and / or receiving steps in S201 to S203, but not limited thereto). The interface circuit 5202 performing the communication steps such as sending and / or receiving in the above-described method refers, for example, to the interface circuit 5202 performing data and / or instruction interaction between the processor 5201, the chip 5200, the memory 5203, or the transceiver device. In some embodiments, the processor 5201 performs at least one of other steps (e.g., steps other than sending and / or receiving in S201 to S203, but not limited thereto).

[0286] The modules and / or devices described in the various embodiments, such as virtual devices, physical devices, and chips, can be combined or separated arbitrarily as needed. Optionally, some or all steps can also be performed collaboratively by multiple modules and / or devices, which is not limited here.

[0287] This disclosure also proposes a storage medium storing instructions that, when executed on a communication device, cause the communication device to perform any of the above methods. Optionally, the storage medium is an electronic storage medium. Optionally, the storage medium is a computer-readable storage medium, but not limited thereto; it may also be a storage medium readable by other devices. Optionally, the storage medium may be a non-transitory storage medium, but not limited thereto; it may also be a temporary storage medium.

[0288] This disclosure also proposes a program product, including a program and / or instructions, which, when executed by a communication device, cause the communication device to perform any of the above methods. Optionally, the program product is a computer program product. Optionally, the program product is stored on the storage medium.

[0289] This disclosure also proposes a computer program that, when run on a computer, causes the computer to perform any of the above methods.

[0290] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this disclosure.

[0291] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0292] The above description is merely a specific embodiment of this disclosure, but the scope of protection of this disclosure is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this disclosure should be included within the scope of protection of this disclosure. Therefore, the scope of protection of this disclosure should be determined by the scope of the claims.

Claims

1. A method of stereo processing, characterized by, include: Acquire stereo audio signals, wherein the stereo audio signals include left channel audio signals and right channel audio signals; Determine the channel difference value between the left channel audio signal and the right channel audio signal; Based on the inter-channel difference value, determine whether to adjust at least one of the left channel audio signal and the right channel audio signal.

2. The method of claim 1, wherein, The acquisition of stereo audio signals includes: The raw audio signal from the sound source is acquired using multiple microphones; The original audio signal is processed to generate the stereo audio signal.

3. The method of claim 1 or 2, wherein, The step of determining whether to adjust at least one of the left channel audio signal and the right channel audio signal based on the inter-channel difference value includes: Based on the inter-channel difference values ​​and mapping relationship, determine the first offset value and the second offset value; Based on the first offset value and the second offset value, determine whether to adjust at least one of the left channel audio signal and the right channel audio signal; The mapping relationship includes the correspondence between the first offset value of the left channel audio signal and the inter-channel difference value, and the correspondence between the second offset value of the right channel audio signal and the inter-channel difference value.

4. The method of claim 2, wherein, The step of determining whether to adjust at least one of the left channel audio signal and the right channel audio signal based on the inter-channel difference value includes: Based on the target inter-channel difference value and the inter-channel difference value, determine the first offset value corresponding to the left channel audio signal and the second offset value corresponding to the right channel audio signal; Based on the first offset value and the second offset value, determine whether to adjust at least one of the left channel audio signal and the right channel audio signal.

5. The method of claim 4, wherein, The method further includes: Based on the original audio signal, determine the target location of the sound source; Based on the target location of the sound source, determine the target channel difference value between the target left channel audio signal and the target right channel audio signal of the desired target stereo audio signal; 6. The method of claim 5, wherein, Determining the target location of the sound source based on the original audio signal includes: The original audio signal is processed based on a sound source localization algorithm to determine the target location of the sound source.

7. The method of any one of claims 1 to 6, wherein, The inter-channel difference value includes one of the following: Difference in strength between channels; Time difference between channels; The parameter values ​​are determined based on the intensity difference and time difference between channels.

8. The method of any one of claims 4 to 6, wherein, The target channel difference value includes one of the following: Intensity difference between target channels; Time difference between target channels; The target parameter values ​​are determined based on the intensity difference and time difference between the target channels.

9. A stereo processing apparatus, characterized by comprising: The device includes: A signal acquisition module is used to acquire stereo audio signals, wherein the stereo audio signals include left channel audio signals and right channel audio signals; The difference determination module is used to determine the channel difference value between the left channel audio signal and the right channel audio signal; The adjustment determination module is used to determine whether to adjust at least one of the left channel audio signal and the right channel audio signal based on the inter-channel difference value.

10. A communication device comprising: One or more processors; A memory coupled to a processor, the memory storing instructions, which, when executed by the processor, cause the communication device to perform the method as described in any one of claims 1 to 8.

11. A storage medium, the storage medium storing instructions, wherein, When the instructions are executed on the communication device, the communication device performs the method as described in any one of claims 1 to 8.

12. A program product comprising at least one of a program, instructions, characterized in that When at least one of the programs or instructions is executed by the communication device, it implements the method of any one of claims 1 to 8.