Method and apparatus for calibrating multiple audio devices included in a system
The method and apparatus for calibrating multiple audio devices in an audio system address the challenge of optimizing audio performance by estimating device locations and listener positions to adjust volume levels, resulting in a superior audio experience.
Patent Information
- Application Number
- PCT/CN2023/139088
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-12-15
- Publication Date
- 2025-06-19
AI Technical Summary
Existing audio systems with multiple devices struggle to automatically calibrate and optimize audio performance based on the user's environment and device arrangements, leading to suboptimal audio experiences.
A method and apparatus for calibrating multiple audio devices, which involves estimating the location of each device, determining the listener's position, and adjusting the volume of each device based on their relative positions to ensure balanced audio output.
The solution enables automatic adjustment of volume in multi-channel speaker systems, balancing sound levels and providing a superior listening experience by considering the user's position and device arrangements.
Smart Images

Figure CN2023139088_19062025_PF_FP_ABST
Abstract
Description
METHOD AND APPARATUS FOR CALIBRATING MULTIPLE AUDIO DEVICES INCLUDED IN A SYSTEMTECHNICAL FIELD
[0001] The present disclosure relates to audio devices, and in particular, to a method and apparatus for calibrating multiple audio devices included in a system.BACKGROUND
[0002] Currently, an audio system can consist of multiple audio devices to achieve desired audio performances. Each of the audio devices can include one or more speakers to play audio for a user of the audio system. The audio devices can also include one or more microphones, for example, to listen to the user’s command, to sense the audio signals in the listening environment, and so on.
[0003] In particular, a one-platform configuration is provided in the audio system to allow the audio devices in the audio system to connect with each other and group as a multi-channel speaker system.
[0004] SUMMARY OF THE DISCLOSURE
[0005] The following presents a simplified summary relating to one or more aspects disclosed herein. Thus, the following summary should not be considered an extensive overview relating to all contemplated aspects, nor should the following summary be considered to identify key or critical elements relating to all contemplated aspects or to delineate the scope associated with any particular aspect. Accordingly, the following summary has the sole purpose of presenting certain concepts relating to one or more aspects relating to the mechanisms disclosed herein in a simplified form to precede the detailed description presented below.
[0006] In order to overcome the defects in the prior art, the disclosure provides a method, device, and apparatus for calibrating multiple audio devices included in a system.
[0007] According to a first aspect of the disclosure, a method for calibrating multiple audio devices included in a system is provided, comprising: estimating a location of each of the multiple audio devices; estimating a position of a listener for the system; and calibrating volume of each of the multiple audio devices based on the estimated location of each of the multiple audio devices and the estimated position of the listener.
[0008] According to embodiments of the disclosure, in case the system includes three or more audio devices, estimating the position of the listener for the system comprises: determining a centroid of a polygon formed by the multiple audio devices as the estimated position of the listener, based on the estimated location of each of the multiple audio devices.
[0009] According to embodiments of the disclosure, determining the centroid of the polygon comprises: dividing the polygon into multiple triangles; and determining the centroid of the polygon based on the centroids of the multiple triangles with different weights for different triangles.
[0010] According to embodiments of the disclosure, the weights are adjusted based on the orientations of the respective triangles relative to the listener.
[0011] According to embodiments of the disclosure, calibrating the volume of each of the multiple audio devices comprises: determining an actual distance between the listener and each of the multiple audio devices, based on the estimated location of each of the multiple audio devices and the estimated position of the listener; determining a target distance between the listener and each of the multiple audio devices; and calibrating the volume of each of the multiple audio devices, based on the actual distance and the target distance between the actual position of the listener and the respective audio devices.
[0012] According to embodiments of the disclosure, the estimated position of the listener is the determined centroid of the polygon formed by the multiple audio devices.
[0013] According to embodiments of the disclosure, the location of each of the multiple audio devices is estimated by: playing an audio signal from a speaker of a first audio device of the multiple audio devices; listening to the audio signal from the speaker of the first audio device with multiple microphones of a second audio device of the multiple audio devices; and calculating the location of the first audio device relative to the second audio device, based on the time at which a first microphone of the multiple microphones of the second audio device receives the audio signal from the speaker of the first audio device and the time at which a second microphone of the multiple microphones of the second audio device receives the audio signal from the speaker of the first audio device.
[0014] According to embodiments of the disclosure, the location of each of the multiple audio devices is estimated further by: playing an audio signal from a speaker of the second audio device; listening to the audio signal from the speaker of the second audio device with the multiple microphones of the second audio device; and calculating the location of the first audio device relative to the second audio device, based on the time at which the first microphone of the multiple microphones of the second audio device receives the audio signal from the speaker of the first audio device, the time at which the second microphone of the multiple microphones of the second audio device receives the audio signal from the speaker of the first audio device, the time at which the first microphone of the multiple microphones of the second audio device receives the audio signal from the speaker of the second audio device, and the time at which the second microphone of the multiple microphones of the second audio device receives the audio signal from the speaker of the second audio device.
[0015] According to embodiments of the disclosure, the multiple audio devices include a primary audio device and one or more secondary audio devices, and the method further comprises: transmitting a synchronized audio signal to the primary audio device and the secondary audio devices to ensure a time delay between the primary audio device and each of the secondary audio devices within a specified range.
[0016] According to embodiments of the disclosure, the multiple audio devices communicate via at least one of P2P, Wi-Fi, or Ethernet connections.
[0017] According to a second aspect of the disclosure, a device for calibrating multiple audio devices included in a system is provided, comprising: an audio device location estimation module, configured to estimate a location of each of the multiple audio devices; a listener position estimation module, configured to estimate a position of a listener for the system; and a volume calibration module, configured to calibrate volume of each of the multiple audio devices based on the estimated location of each of the multiple audio devices and the estimated position of the listener.
[0018] According to embodiments of the disclosure, in case the system includes three or more audio devices, the listener position estimation module is configured to: determine a centroid of a polygon formed by the multiple audio devices as the estimated position of the listener, based on the estimated location of each of the multiple audio devices.
[0019] According to embodiments of the disclosure, determining the centroid of the polygon comprises: dividing the polygon into multiple triangles; and determining the centroid of the polygon based on the centroids of the multiple triangles with different weights for different triangles.
[0020] According to embodiments of the disclosure, the weights are adjusted based on the orientations of the respective triangles relative to the listener.
[0021] According to embodiments of the disclosure, the volume calibration module is configured to: determine an actual distance between the listener and each of the multiple audio devices, based on the estimated location of each of the multiple audio devices and the estimated position of the listener; determine a target distance between the listener and each of the multiple audio devices; and calibrate the volume of each of the multiple audio devices, based on the actual distance and the target distance between the actual position of the listener and the respective audio devices.
[0022] According to embodiments of the disclosure, the estimated position of the listener is the determined centroid of the polygon formed by the multiple audio devices.
[0023] According to embodiments of the disclosure, the audio device location estimation module is configured to: play an audio signal from a speaker of a first audio device of the multiple audio devices; listen to the audio signal from the speaker of the first audio device with multiple microphones of a second audio device of the multiple audio devices; and calculate the location of the first audio device relative to the second audio device, based on the time at which a first microphone of the multiple microphones of the second audio device receives the audio signal from the speaker of the first audio device and the time at which a second microphone of the multiple microphones of the second audio device receives the audio signal from the speaker of the first audio device.
[0024] According to embodiments of the disclosure, the audio device location estimation module is configured to: play an audio signal from a speaker of the second audio device; listen to the audio signal from the speaker of the second audio device with the multiple microphones of the second audio device; and calculate the location of the first audio device relative to the second audio device, based on the time at which the first microphone of the multiple microphones of the second audio device receives the audio signal from the speaker of the first audio device, the time at which the second microphone of the multiple microphones of the second audio device receives the audio signal from the speaker of the first audio device, the time at which the first microphone of the multiple microphones of the second audio device receives the audio signal from the speaker of the second audio device, and the time at which the second microphone of the multiple microphones of the second audio device receives the audio signal from the speaker of the second audio device.
[0025] According to embodiments of the disclosure, the multiple audio devices include a primary audio device and one or more secondary audio devices, and the device is configured to: transmit a synchronized audio signal to the primary audio device and the secondary audio devices to ensure a time delay between the primary audio device and each of the secondary audio devices within a specified range.
[0026] According to embodiments of the disclosure, the multiple audio devices communicate via at least one of P2P, Wi-Fi, or Ethernet connections.BRIEF DESCRIPTION OF THE DRAWINGS
[0027] The accompanying drawings are presented to aid in the description of various aspects of the disclosure and are provided solely for illustration of the aspects and not limitation thereof.
[0028] FIG. 1 illustrates a schematic diagram for an audio system according to embodiments of the present disclosure;
[0029] FIG. 2 illustrates a flow chart of a method for calibrating multiple audio devices included in an audio system according to embodiments of the present disclosure;
[0030] FIG. 3 illustrates a schematic diagram for estimating the location of the audio device according to embodiments of the present disclosure;
[0031] FIG. 4 illustrates a schematic diagram for estimating a centroid of a polygon formed by the multiple audio devices according to embodiments of the present disclosure;
[0032] FIG. 5 illustrates a schematic diagram for determining the actual distance between the listener and each of the multiple audio devices according to embodiments of the present disclosure;
[0033] FIG. 6 illustrates a schematic diagram for determining the target distance between the listener and each of the multiple audio devices according to embodiments of the present disclosure;
[0034] FIG. 7 illustrates a schematic diagram for the synchronization in the audio system according to embodiments of the present disclosure;
[0035] FIG. 8 illustrates a schematic diagram for a device for calibrating multiple audio devices included in an audio system according to embodiments of the present disclosure; and
[0036] FIG. 9 illustrates a schematic diagram for an apparatus for calibrating multiple audio devices included in an audio system according to embodiments of the present disclosure.
[0037] DESCRIPTION OF THE EMBODIMENTS
[0038] Aspects of the disclosure are provided in the following description and related drawings directed to various examples provided for illustration purposes. Alternate aspects may be devised without departing from the scope of the disclosure. Additionally, well-known elements of the disclosure will not be described in detail or will be omitted so as not to obscure the relevant details of the disclosure.
[0039] A user of an audio system containing multiple audio devices can generally group and configure the multiple audio devices altogether to provide a desired audio performance. For example, the user may use a smartphone or any other device that can configure the audio devices as per the user’s commands. Also, the smartphone can also instruct the user to place each of the audio devices in a specific position, for example, according to information from an APP installed on the microphone.
[0040] However, the smartphone can only tell the user about the approximate position or direction of each of the audio devices, which may not take the user’s environment into consideration. As an example, different layouts of the room where the audio devices are positioned may have a significant impact on the desired audio performance. As another example, the different arrangements of the multiple audio devices in the audio system may also have a significant impact on the desired audio performance. Therefore, it would be desirable for the audio system to automatically calibrate the multiple audio devices in the audio system, so as to obtain an optimal performance.
[0041] FIG. 1 illustrates a schematic diagram for an audio system 100 according to embodiments of the present disclosure. As shown in FIG. 1, the system 100 can include four audio devices 110A, 110B, 110C, and 110D, arranged around the user of the audio system 100. It is to be understood that system 100 is not limited to including four audio devices. For example, the system can include two, three, five, or more audio devices. As shown in FIG. 1, the audio device 110D of the multiple audio devices can include a speaker 112D configured to play audio. According to embodiments of the present disclosure, the audio device 110D can further include microphones 114D and 116D configured, for example, to listen to the user’s command, to sense the audio signals in the listening environment, and so on. It is to be understood that the audio device 110D is not limited to including one speaker and two microphones, and can include two or more speakers and one, three, or more microphones. Although not illustrated in FIG. 1, it is to be understood that other audio devices 110A, 110B, or 110C can also include one or more speaker (s) and / or microphone (s) similarly.
[0042] FIG. 2 illustrates a flow chart of a method for calibrating multiple audio devices included in an audio system according to embodiments of the present disclosure. As shown in FIG. 2, in step 210, the location of each of the multiple audio devices can be estimated.
[0043] According to embodiments of the present application, the location of each of the multiple audio devices can be estimated by: playing an audio signal from a speaker of a first audio device of the multiple audio devices; listening to the audio signal from the speaker of the first audio device with multiple microphones of a second audio device of the multiple audio devices; and calculating the location of the first audio device relative to the second audio device, based on the time at which a first microphone of the multiple microphones of the second audio device receives the audio signal from the speaker of the first audio device and the time at which a second microphone of the multiple microphones of the second audio device receives the audio signal from the speaker of the first audio device.
[0044] FIG. 3 illustrates a schematic diagram for estimating the location of the audio device according to the embodiments of the present disclosure. In particular, FIG. 3 illustrates the method for estimating the distance and angle between the audio device 310A and the audio device 310B. As illustrated, the audio device 310A can be configured to listen to the audio signal from the audio device 310B and the audio device 310B can be configured to play an audio signal so as to estimate the relative location of the audio device 310B relative to the audio device 310A, i.e., the relative distance dBtoA and the relative angle θ between the two audio devices. The relative distance dBtoA and the relative angle θ between the two audio devices can be estimated based on the time of arrival T1BtoA from the speaker 312B of the audio device 310B to the microphone 314A of the audio device 310A, and the time of arrival T2BtoA from the speaker 312B of the audio device 310B to the microphone 316A of the audio device 310A. As a non-limiting example, a sweep signal can be played from the speaker 312B of the audio device 310B, and the respective impulse response can be recorded by the microphones 314A and 316A of the audio device 310A. The time at which the impulse response recorded by the microphone 314A reaches the peak can be recorded as the time of arrival at the microphone 314A, and the time at which the impulse response recorded by the microphone 314B reaches the peak can be recorded as the time of arrival at the microphone 314B. It is to be understood that the above-mentioned sweep signal is just exemplary, and other audio signals can also be utilized herein to estimate the relative distance and angle between any two of the multiple audio devices.
[0045] In particular, the latency can be regarded as the time difference of two impulse responses, and it can be calculated by: Tdiff_BtoA=T1BtoA-T2BtoA (1)
[0046] where T1BtoA represents the time of arrival of the audio signal from the speaker 312B of the audio device 310B to the microphone 314A of the audio device 310A,
[0047] T2BtoA represents the time of arrival of the audio signal from the speaker 312B of the audio device 310B to the microphone 316A of the audio device 310A.
[0048] where Tdiff_BtoA is the time difference of sound arrival of the microphones 314A and 316A of the audio device 310A, when the audio device 310B is playing.
[0049] T1BtoA and T2BtoA can be the time indices of the impulse-response peaks of the microphones 314A and 316A of the audio device 310A, respectively, when the audio device 310B is playing with angle θ. The microphones 314A and 316A of the audio device 310A can receive the signal with latency because of the distance dMic between the pair of microphones 314A and 316A. Thus, the relative angle θ and distance dBtoA between the two audio devices 310A and 310B can be calculated by: θ=sin-1 (Tdiff_BtoA*c / dMic) (2) dBtoA= (T1BtoA+T2BtoA) / 2*c (3)
[0050] where c represents the sound speed, and
[0051] dMic represents the distance between the microphones 314A and 316A of the audio device 310A.
[0052] The implementation as illustrated in FIG. 3 is directed to the distance and angle between two audio devices of the audio system. If the audio system consists of N audio devices, the above calculation can be performed between any two of the N audio devices to estimate the relative distance and angle between the two audio devices and thus the location of each of the audio devices can be determined, for example, in a coordinate with the original point set at random as illustrated in FIG. 4.
[0053] Further, the wireless system may have time drifted issues, meaning that T1BtoA and T2BtoA can vary all the time since the recording and playing of the audio signal may not be synchronized, which may result in inaccuracy of the above calculation. To solve the problem, self-recording of the audio devices can be utilized. According to embodiments of the present disclosure, the location of each of the multiple audio devices can be estimated further by: playing an audio signal from a speaker of the second audio device; listening to the audio signal from the speaker of the second audio device with the multiple microphones of the second audio device; and calculating the location of the first audio device relative to the second audio device, based on the time at which the first microphone of the multiple microphones of the second audio device receives the audio signal from the speaker of the first audio device, the time at which the second microphone of the multiple microphones of the second audio device receives the audio signal from the speaker of the first audio device, the time at which the first microphone of the multiple microphones of the second audio device receives the audio signal from the speaker of the second audio device, and the time at which the second microphone of the multiple microphones of the second audio device receives the audio signal from the speaker of the second audio device.
[0054] That is, taking the latency drifted issues into consideration, the above equation (3) can be rewritten as:
[0055] where T1BtoA represents the time of arrival of the audio signal from the speaker 312B of the audio device 310B to the microphone 314A of the audio device 310A,
[0056] T1AtoA represents the time of arrival of the audio signal from the speaker 312A of the audio device 310A to the microphone 314A of the audio device 310A,
[0057] T2BtoA represents the time of arrival of the audio signal from the speaker 312B of the audio device 310B to the microphone 316A of the audio device 310A, and
[0058] T2AtoA represents the time of arrival of the audio signal from the speaker 312A of the audio device 310A to the microphone 316A of the audio device 310A.
[0059] Referring back to FIG. 2, in step 220, a position of a listener can be estimated. According to embodiments of the present disclosure, since the multiple audio devices may be placed arbitrarily in user’s room, in case the system includes three or more audio devices, the audio devices can form an irregular polygon, and thus estimating the position of the listener can include determining the centroid of a polygon formed by the multiple audio devices as the estimated position of the listener, based on the estimated location of each of the multiple audio devices in step 210. According to embodiments of the present disclosure, determining the centroid of the polygon can include: dividing the polygon into multiple triangles; and determining the centroid of the polygon based on the centroids of the multiple triangles with different weights for different triangles.
[0060] FIG. 4 illustrates a schematic diagram for estimating a centroid of a polygon formed by the multiple audio devices according to embodiments of the present disclosure. As illustrated in FIG. 4, the coordinates (x, y) of the multiple audio devices 410A to 410D (i.e., (x1, y1) of the 1st audio device 410A, (x2, y2) of the 2nd audio device 410B, (x3, y3) of the 3rd audio device 410C, and (x4, y4) of the 4th audio device 410D) can be determined based on the method described above with respect to step 210 of method 200 and with the original point randomly set, and form an irregular quadrilateral. The irregular quadrilateral as illustrated in FIG. 4 can be divided into four triangles. The centroid of each triangle can be determined as the average of the vertices of the respective triangle. The centroid of the irregular quadrilateral can be determined based on the centroids of the four triangles with different weights for different triangles, as follows: Si= (xiyi+1-xi+1yi) / 2 (7)
[0061] where (xc, yc) is the coordinate of the centroid of the polygon;
[0062] is the coordinate of the centroid of the ith triangle of the multiple divided triangles; and Si is the area of the ith triangle of the multiple divided triangles and can be positive or negative to indicate the direction.
[0063] In case of the multi-channel content playback, the sound power of the front channel is usually higher, thus the weight for each of the triangles divided from the polygon can be adjusted based on the orientations of the respective triangle relative to the listener to weight each triangle considering the playback power differences. In this case, the above equation (7) can be rewritten as: Si=αi* (xiyi+1-xi+1yi) / 2 (8)
[0064] where αi is the weighting of each triangle, and normally the front ones should be smaller.
[0065] As an example, for the quadrilateral formed by the audio devices 410A through 410D as illustrated in FIG. 4, the user is assumed to face the direction of the y-axis. Thus, the divided triangle formed by the audio devices 410A, 410B and the centroid is assigned a relatively low weighting, since the longitudinal coordinates of the audio devices 410A, 410B are both positive; the divided triangle formed by the audio devices 410A, 410C and the centroid is assigned a medium weighting since the longitudinal coordinate of the audio device 410A is positive while the longitudinal coordinate of the audio device 410C is negative; the divided triangle formed by the audio devices 410B, 410D and the centroid are assigned a medium weighting since the longitudinal coordinate of the audio device 410B is positive while the longitudinal coordinate of the audio device 410D is negative; and the divided triangle formed by the audio device 410C, 410D and the centroid is assigned a relatively high weighting, since the longitudinal coordinates of the audio devices 410C, 410D are both negative.
[0066] As a further example, for a triangle formed by the audio devices 410A through 410C with the illustrated audio device 410D omitted, the user is assumed to face the direction of the y-axis. Thus, the divided triangle formed by the audio devices 410A, 410B and the centroid is assigned a relatively low weighting, since the longitudinal coordinates of the audio devices 410A, 410B are both positive; the divided triangle formed by the audio devices 410A, 410C and the centroid is assigned a medium weighting since the longitudinal coordinate of the audio device 410A is positive while the longitudinal coordinate of the audio device 410C is negative; and the divided triangle formed by the audio devices 410B, 410C and the centroid is assigned a medium weighting since the longitudinal coordinate of the audio device 410B is positive while the longitudinal coordinate of the audio device 410C is negative.
[0067] As another example, for a pentagon formed by the audio devices 410A, 410B, 410C, 410D, and 410E (not illustrated in FIG. 4 and positioned between 410A and 410B) , the user is assumed to face the direction of the y-axis. Thus, the divided triangle formed by the audio devices 410A, 410E and the centroid, and the divided triangle formed by the audio devices 410E, 410B and the centroid is assigned a relatively low weighting, since the longitudinal coordinates of the audio devices 410A, 410B, 410E are all positive; the divided triangle formed by the audio devices 410A, 410C and the centroid is assigned a medium weighting, since the longitudinal coordinate of the audio device 410A is positive while the longitudinal coordinate of the audio device 410C is negative; the divided triangle formed by the audio devices 410B, 410D and the centroid are assigned a medium weighting, since the longitudinal coordinate of the audio device 410B is positive while the longitudinal coordinate of the audio device 410D is negative; and the divided triangle formed by the audio device 410C, 410D and the centroid is assigned a relatively high weighting, since the longitudinal coordinates of the audio devices 410C, 410D are both negative.
[0068] The above examples are just exemplary and are intended to explain that the weightings of the divided triangles are arranged such that the assumed position of the listener is closer to the audio devices positioned on the rear side relative to the facing orientation of the listener.
[0069] Referring back to FIG. 2, in step 230, based on the estimated location of each of the multiple audio devices and the estimated position of the listener, the volume of each of the multiple audio devices can be calibrated. According to the embodiments of the present disclosure, calibrating the volume of each of the multiple audio devices can include: determining an actual distance between the listener and each of the multiple audio devices, based on the estimated location of each of the multiple audio devices and the estimated position of the listener; determining a target distance between the estimated position of the listener and each of the multiple audio devices; and calibrating the volume of each of the multiple audio devices, based on the actual distance and the target distance between the actual position of the listener and the respective audio devices. In particular, in case the system includes three or more audio devices to form a polygon, the estimated position of the listener can be the determined centroid of the polygon formed by the multiple audio devices, for example, (xc, yc) as illustrated in FIG. 4; and in case the system includes two audio devices, the estimated position of the listener can be the point that can form an equilateral triangle with the two audio devices, since the distance between the listening position and speaker is equal or close to the distance between two speakers in most cases.
[0070] FIG. 5 illustrates a schematic diagram for determining the actual distance between the listener and each of the multiple audio devices according to the embodiments of the present disclosure. The actual distance Di_actual between the listener and each of the multiple audio devices (for example, the actual distance D1_actual between the listener and the 1st audio device 510A, the actual distance D2_actual between the listener and the 2nd audio device 510B, the actual distance D3_actual between the listener and the 3rd audio device 510C, and the actual distance D4_actual between the listener and the 4th audio device 510D) can be obtained by:
[0071] where (xi, yi) is the location of the estimated location of the ith audio device as discussed above;
[0072] (xc, yc) is the estimated position of the listener as discussed above, for example, with reference to the equations (5) - (8) .
[0073] FIG. 6 illustrates a schematic diagram for determining the target distance between the listener and each of the multiple audio devices according to the embodiments of the present disclosure. During the product development, acoustics engineers will tune the default volume of multi-channel speaker system in a test lab, which is optimal to most of the users in a certain distance Di_target between the listening position and the ith audio device (for example, the target distance D1_target between the listener and the 1st audio device 510A, the target distance D2_target between the listener and the 2nd audio device 510A, the target distance D3_target between the listener and the 3rd audio device 510A, and the target distance D4_target between the listener and the 4th audio device 510A) . The target distance Di_target can be, for example, predetermined by the acoustics engineers.
[0074] After determining the target distance and the actual distance between the listener and each of the audio devices, the calibrated volume Vi_cal of the ith audio device can be obtained by:
[0075] where Vi_default is the default value of the ith audio device, which can be predetermined by the acoustics engineers; and
[0076] β is a tuning parameter within the range from 0.1 to 10.
[0077] It is to be understood that the steps as illustrated in FIG. 2 are just exemplary, and the method 200 for calibrating the multiple audio devices can include more or less steps. For example, in case the multiple audio devices include a primary audio device and one or more secondary audio devices, in addition to steps 210-230, the method 200 can further comprise the step of transmitting a synchronized audio signal to the primary audio device and the secondary audio devices to ensure a time delay between the primary audio device and each of the secondary audio devices within a specified range (although not illustrated in FIG. 2) . In particular, the multiple audio devices can communicate via at least one of P2P, Wi-Fi, or Ethernet connections.
[0078] In this case, each audio device will have two roles of source and sink. When the audio device is defined as primary, both source and sink will work. When the audio device is defined as secondary, it is the sink that works. The primary audio device is responsible for processing and distributing audio data, and therefore for calculating the reference audio data between audio devices. The audio latency between the primary and secondary audio devices in the multichannel should not be too high, for example, less than 150 microseconds. The distance that sound travels in the air is 0.051 m (340 m / s*0.00015 s) , which can be ignored when compared to the target distance between the audio devices, mostly greater than 1m. The purpose is to calculate the distance between two audio devices, and then the primary audio device will adjust the volume of each audio device based on the previously calculated Vi_cal to achieve volume balance.
[0079] FIG. 7 illustrates a schematic diagram for the synchronization in the audio system according to the embodiments of the present disclosure. As illustrated in FIG. 7, the audio device 710A can play the role of primary audio device, and the audio devices 710B to 710D can play the role of secondary audio device. The primary audio device can be coupled to the audio service module 720. In this case, the primary audio device can be configured to perform the method for calibrating multiple audio devices included in a system in coordination with the audio devices 710B to 710D, as discussed above with reference to FIG. 2, and the primary audio device 710A can receive the audio signal from the audio service module 720 and provide synchronized audio signal and parameters (for example, the parameters regarding the calibrated volume for the respective audio device) to the secondary audio devices 710B through 710D. Also, the method as discussed above with reference to FIG. 2 can be implemented in another device coupled to the primary audio device 710A, and the primary audio device can receive the synchronized audio signal from the other device and transmit the synchronized audio signal to the secondary audio devices 710B through 710D.
[0080] FIG. 8 illustrates a schematic diagram for a device 800 for calibrating multiple audio devices included in an audio system according to embodiments of the present disclosure. As illustrated in FIG. 8, the device 800 can include an audio device location estimation module 802, a listener position estimation module 804, and a volume calibration module 806. The audio device location estimation module 802 can be configured to estimate a location of each of the multiple audio devices. The listener position estimation module 804 can be configured to estimate a position of a listener for the system. The volume calibration module 806 can be configured to calibrate volume of each of the multiple audio devices based on the estimated location of each of the multiple audio devices and the estimated position of the listener. It is to be understood that the modules included in the device 800 as illustrated in FIG. 8 are just exemplary, and the device 800 can include more or fewer modules.
[0081] FIG. 9 illustrates a schematic diagram for an apparatus 900 for calibrating multiple audio devices included in an audio system according to embodiments of the present disclosure. As shown in FIG. 9, the apparatus 900 can comprise a processor 902 and a memory 904. In particular, the processor 902 can be coupled with the memory 904 and configured to perform the method described herein, for example, the method 200 as shown in FIG. 2. The processor 902 can be any device with processing power capable of implementing the functions of the embodiments of the present disclosure, such as a general-purpose processor, digital signal processor (DSP) , Application Specific Integrated Circuit (ASIC) , field programmable gate array signal (FPGA) or programmable logic device (PLD) , discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described herein The memory 904 can include computer system readable media in the form of volatile memory, such as random access memory (RAM) and / or cache memory, as well as other removable or non-removable, volatile or non-volatile memory, such as hard disk drives, floppy disks, CD-ROM, DVD-ROM, or any other optical storage media.
[0082] According to embodiments of the present disclosure, the memory 904 can store computer program instructions, and the processor 902 can execute the computer program instructions stored in the memory 904. When the computer program instructions are performed by the processor, causing the processor to execute the method for the system including multiple audio devices according to embodiments of the present disclosure. The method for the system including multiple audio devices can be the method described with reference to FIG. 2, which is omitted here for conciseness.
[0083] In conclusion, the method and apparatus according to embodiments of the present disclosure can automatically adjust volume of multi-channel speaker system based on the detected and estimated user’s listening position, which is aimed to balance the volume of each speaker and approach the target loudness recommended by acoustics engineer and help to provide superior listening experience for user in any speaker grouping cases.
[0084] An expression such as “according to” , “based on” , “dependent on” , and so on as used in the disclosure does not mean “according only to” , “based only on” , or “dependent only on” unless it is explicitly otherwise stated. In other words, such expression generally means “according at least to” , “based at least on” , or “dependent at least on” in the disclosure.
[0085] Any reference in the disclosure to an element using the designation “first” , “second” and so forth is not intended to comprehensively limit the number or order of such elements. These expressions can be used in the disclosure as a convenient method for distinguishing two or more units. Thus, a reference to a first unit and a second unit does not imply that only two units can be employed or that the first unit must precede the second unit in some form.
[0086] The term “determining” used in the disclosure can include various operations. For example, regarding “determining” , calculating, computing, processing, deriving, investigating, looking up (e.g., looking up in tables, databases, or other data structure) , ascertaining, and so forth are regarded as "determination" . In addition, regarding “determining” , receiving (for example, receiving information) , transmitting (for example, transmitting information) , input, output, accessing (for example, access to data in the memory) , and so forth, are also regarded as “determining” . In addition, regarding “determining” , resolving, selecting, choosing, establishing, comparing, and so forth can also be regarded as “determining” . That is, regarding "determining" , several actions can be regarded as “determining” .
[0087] The terms such as “connected” , “coupled” or any of their variants used in the disclosure refer to any connection or combination, direct or indirect, between two or more units, which can include the following situations: between two units that are “connected” or “coupled” with each other, there are one or more intermediate units. The coupling or connection between the units can be physical or logical, or can also be a combination of the two. As used in the disclosure, two units can be considered to be electrically connected through the use of one or more wires, cables, and / or printed, and as a number of non-limiting and non-exhaustive examples, and are “connected” or “coupled” with each other through the use of electromagnetic energy with wavelengths in a radio frequency region, the microwave region, and / or in the light (both visible and invisible) region, and so forth.
[0088] When used in the disclosure or the claims ‘including” , “comprising” , and variations thereof, these terms are as open-ended as the term “having” . Further, the term “or” used in the disclosure or in the claims is not an exclusive-or.
[0089] The present disclosure has been described in detail above, but it is obvious to those skilled in the art that the present disclosure is not limited to the embodiments described in the disclosure. The present disclosure can be implemented as a modified and changed form without departing from the spirit and scope of the present disclosure defined by the description of the claims. Therefore, the description in the disclosure is for illustration and does not have any limiting meaning to the present disclosure.
Claims
1.A method for calibrating multiple audio devices included in a system, comprising:estimating a location of each of the multiple audio devices;estimating a position of a listener for the system; andcalibrating volume of each of the multiple audio devices based on the estimated location of each of the multiple audio devices and the estimated position of the listener.2.The method of claim 1, wherein in case the system includes three or more audio devices, estimating the position of the listener for the system comprises:determining a centroid of a polygon formed by the multiple audio devices as the estimated position of the listener, based on the estimated location of each of the multiple audio devices.3.The method of claim 2, wherein determining the centroid of the polygon comprises:dividing the polygon into multiple triangles; anddetermining the centroid of the polygon based on the centroids of the multiple triangles with different weights for different triangles.4.The method of claim 3, wherein the weights are adjusted based on an orientation of each respective triangle in the multiple triangles relative to the listener.5.The method of claim 2, wherein calibrating the volume of each of the multiple audio devices comprises:determining an actual distance between the listener and each of the multiple audio devices, based on the estimated location of each of the multiple audio devices and the estimated position of the listener;determining a target distance between the listener and each of the multiple audio devices; andcalibrating the volume of each of the multiple audio devices, based on the actual distance and the target distance between the actual position of the listener and the respective audio devices.6.The method of claim 5, wherein the estimated position of the listener is the determined centroid of the polygon formed by the multiple audio devices.7.The method of claim 1, wherein the location of each of the multiple audio devices are estimated by:playing an audio signal from a speaker of a first audio device of the multiple audio devices;listening to the audio signal from the speaker of the first audio device with multiple microphones of a second audio device of the multiple audio devices; andcalculating the location of the first audio device relative to the second audio device, based on a time at which a first microphone of the multiple microphones of the second audio device receives the audio signal from the speaker of the first audio device and a time at which a second microphone of the multiple microphones of the second audio device receives the audio signal from the speaker of the first audio device.8.The method of claim 7, wherein the location of each of the multiple audio devices is estimated further by:playing an audio signal from a speaker of the second audio device;listening to the audio signal from the speaker of the second audio device with the multiple microphones of the second audio device; andcalculating the location of the first audio device relative to the second audio device, based on the time at which the first microphone of the multiple microphones of the second audio device receives the audio signal from the speaker of the first audio device, the time at which the second microphone of the multiple microphones of the second audio device receives the audio signal from the speaker of the first audio device, the time at which the first microphone of the multiple microphones of the second audio device receives the audio signal from the speaker of the second audio device, and the time at which the second microphone of the multiple microphones of the second audio device receives the audio signal from the speaker of the second audio device.9.The method of claim 1, wherein the multiple audio devices include a primary audio device and one or more secondary audio devices, and the method further comprises:transmitting a synchronized audio signal to the primary audio device and the secondary audio devices to ensure a time delay between the primary audio device and each of the secondary audio devices within a specified range.10.The method of claim 9, wherein the multiple audio devices communicate via at least one of P2P, Wi-Fi, or Ethernet connections.11.A device for calibrating multiple audio devices included in a system, comprising:an audio device location estimation module, configured to estimate a location of each of the multiple audio devices;a listener position estimation module, configured to estimate a position of a listener for the system; anda volume calibration module, configured to calibrate volume of each of the multiple audio devices based on the estimated location of each of the multiple audio devices and the estimated position of the listener.12.The device of claim 11, wherein in case the system includes three or more audio devices, the listener position estimation module is configured to:determine a centroid of a polygon formed by the multiple audio devices as the estimated position of the listener, based on the estimated location of each of the multiple audio devices.13.The device of claim 12, wherein determining the centroid of the polygon comprises:dividing the polygon into multiple triangles; anddetermining the centroid of the polygon based on the centroids of the multiple triangles with different weights for different triangles.14.The device of claim 13, wherein the weights are adjusted based on the orientations of the respective triangles relative to the listener.15.The device of claim 12, wherein the volume calibration module is configured to:determine an actual distance between the listener and each of the multiple audio devices, based on the estimated location of each of the multiple audio devices and the estimated position of the listener;determine a target distance between the listener and each of the multiple audio devices; andcalibrate the volume of each of the multiple audio devices, based on the actual distance and the target distance between the actual position of the listener and the respective audio devices.16.The device of claim 15, wherein the estimated position of the listener is the determined centroid of the polygon formed by the multiple audio devices.17.The device of claim 11, wherein the audio device location estimation module is configured to:play an audio signal from a speaker of a first audio device of the multiple audio devices;listen to the audio signal from the speaker of the first audio device with multiple microphones of a second audio device of the multiple audio devices; andcalculate the location of the first audio device relative to the second audio device, based on a time at which a first microphone of the multiple microphones of the second audio device receives the audio signal from the speaker of the first audio device and a time at which a second microphone of the multiple microphones of the second audio device receives the audio signal from the speaker of the first audio device.18.The device of claim 17, wherein the audio device location estimation module is further configured to:play an audio signal from a speaker of the second audio device;listen to the audio signal from the speaker of the second audio device with the multiple microphones of the second audio device; andcalculate the location of the first audio device relative to the second audio device, based on the time at which the first microphone of the multiple microphones of the second audio device receives the audio signal from the speaker of the first audio device, the time at which the second microphone of the multiple microphones of the second audio device receives the audio signal from the speaker of the first audio device, the time at which the first microphone of the multiple microphones of the second audio device receives the audio signal from the speaker of the second audio device, and the time at which the second microphone of the multiple microphones of the second audio device receives the audio signal from the speaker of the second audio device.19.The device of claim 11, wherein the multiple audio devices include a primary audio device and one or more secondary audio devices, and the device is configured to:transmit a synchronized audio signal to the primary audio device and the secondary audio devices to ensure a time delay between the primary audio device and each of the secondary audio devices within a specified range.20.The device of claim 19, wherein the multiple audio devices communicate via at least one of P2P, Wi-Fi, or Ethernet connections.
Citation Information
Patent Citations
Speaker assignment device, speaker assignment method and speaker assignment program
JP2015122612A
Sound field measuring apparatus and sound field measuring method
US20070019815A1
A device for and a method of processing audio data
WO2007135581A2