Method and system for automatic audio calibration
By using a camera in the audio system to capture video, extract environment and listener information, and generate compensation filters, the problem that existing audio systems are difficult to adapt to mobile users and complex environments during calibration, achieving stable sound quality and simplified calibration process.
Patent Information
- Application Number
- CN202280102009.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-02
- Publication Date
- 2025-06-27
AI Technical Summary
Existing audio systems are difficult to accurately adapt to mobile users and complex room environments during calibration, resulting in unstable sound quality.
By combining the camera with the audio system, video captures room environment and listener information, estimates the environmental impact in the sound field, and generates compensation filters to adjust the output of the audio system.
It realizes stable sound quality in different room environments and user movement, improves user listening experience, and simplifies the calibration process.
Smart Images

Figure CN120226386A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to audio processing, and more particularly, to a method and system for automatic audio calibration. Background Art
[0002] Generally, the sound field generated by a loudspeaker is determined not only by the loudspeaker itself but also greatly affected by the environment. There are inevitably many obstacles or reflectors in a room, such as walls, floors, tables, desks, etc. When sound waves reach the obstacles or reflectors, reflection, scattering, and diffraction will occur. The reflected waves often interfere with the original sound, which results in an increase or decrease in the frequency response of different frequency bands. This phenomenon is particularly obvious when the distance is small, indicating that the user is listening to the loudspeaker in the near field and there is a large reflector nearby. A typical situation of a home audio system is that the loudspeaker is placed on a desk and the listener is sitting in front of the loudspeaker. The listener can feel a drastic change in the timbre of the sound when leaning forward and backward.
[0003] Home audio products usually apply calibration to compensate for the environmental effects so that users can still hear similar sounds in their rooms despite the different room environments. Currently, calibration is mainly achieved by acoustic methods. An internal or external microphone is used to measure the sound field so that the output of the loudspeaker can be modified according to the measured results. However, the internal microphone can only measure the sound near the loudspeaker, so only a rough estimate of the sound field at the listener's position can be obtained without accurate information. In addition, the calibration performed by the internal microphone cannot adapt to a moving user. On the contrary, using an external microphone in the listening area is another effective calibration method. It can directly measure the sound field at the user's position but is often complained about for its inconvenience in use.
[0004] Therefore, it is necessary to develop other improved calibration methods to adjust the sound performance. Summary of the Invention
[0005] According to one aspect of the present disclosure, there is provided an automatic audio calibration method for an audio system in a room. The method can use a camera to capture a video of the room through the camera. The method can also retrieve environmental information and listener information from the video; estimate the environmental effects in the sound field at the listener's position based on the environmental information and listener information; and generate a compensation filter for the audio system to compensate for the estimated environmental effects.
[0006] According to another aspect of the present disclosure, a system for automatic audio correction of an audio system is provided. The system may include a camera and a processor. The camera may be configured to capture a video of a room via the camera. The processor may be coupled to the camera and may be configured to retrieve environmental information and listener information from the video. The processor may also be configured to estimate environmental effects in the sound field at the listener based on the environmental information and the listener information, and generate a compensation filter for the audio system to compensate for the estimated environmental effects.
[0007] According to yet another aspect of the present disclosure, a non-transitory computer-readable storage medium including computer-executable instructions is provided, which when executed by a computer cause the computer to perform the methods disclosed herein. BRIEF DESCRIPTION OF THE DRAWINGS
[0008] Figure 1 A schematic diagram of an audio system according to one or more embodiments of the present disclosure is shown;
[0009] Figure 2 A flowchart of an automatic audio calibration method for an audio system according to one or more embodiments of the present disclosure is shown;
[0010] Figure 3 A schematic diagram of information retrieval from a video according to one or more embodiments of the present disclosure is shown;
[0011] Figure 4 An example of a video captured by a TOF camera according to one or more embodiments of the present disclosure is shown;
[0012] Figure 5 A simple configuration of an audio calibration process is shown; and
[0013] Figure 6 An example of the amplitude responses of the direct sound, the total sound, and the compensation EQ filter is shown.
[0014] It should be contemplated that the elements disclosed in one embodiment may be advantageously used in other embodiments without specific recitation. The drawings referred to herein should not be construed as being drawn to scale unless specifically noted. Also, for clarity of presentation and explanation, the drawings are generally simplified and details or components are omitted. The drawings and the discussion are used to explain the principles discussed below, where like reference numerals represent like elements. DETAILED DESCRIPTION
[0015] Examples will be provided below for illustration. The descriptions of the various examples will be presented for illustrative purposes, but are not intended to be exhaustive or limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments.
[0016] In the present disclosure, an improved method and system for automatic audio calibration are provided. The method and system proposed in the present disclosure combine an audio system with at least one camera to provide a consistent and stable tone color for a listener regardless of the listener's movement and different room environments. Using the camera can provide a complete view of the room environment and enable continuous head tracking of a moving listener without any external devices. Specifically, the camera can be used to detect the room by recording a video of the room. From the video, the method and system can retrieve useful information for calibration, such as information about the room environment and information about the listener's position. Therefore, the method and system can estimate the impact of the environment on the sound field generated at the listener's location based on the useful information and adaptively adjust the audio system to compensate for the environmental impact, so that a stable tone color can be provided for the listener regardless of the room environment and the listener's movement. By combining the audio system with the camera and estimating and compensating for the environmental impact based on video detection, the proposed method can achieve automatic audio calibration without complex installation and operation, which can provide a better listening experience for users and greatly improve the user's product experience. The method will be explained in detail below with reference to Figures 1 to 6 explain the method.
[0017] Figure 1 A schematic diagram of an audio system according to one or more embodiments of the present disclosure is shown. Figure 1 The system 100 shown includes a camera 102, a memory 104, a processor 106, an audio source 108, and a speaker 110.
[0018] The camera 102 can be positioned at any location near the speaker 110. For example, the camera 102 can be positioned on the top or front of the speaker box including the speaker 110, or at any location near the speaker where the camera can detect and record information about the room. The camera 102 can be an optical camera (such as an RGB camera) with one or more viewing angles or a depth camera (such as a TOF (time of flight) camera). In some examples, the camera 102 can be a digital camera configured to acquire a video with a series of frames (e.g., images) at a programmable frame rate. In some examples, the frame rate can be selected based on the processing speed of the processor 106.
[0019] The memory 104 may include any non-transitory tangible computer-readable medium in which programming instructions are stored. As used herein, the term "tangible computer-readable medium" is expressly defined to include any type of computer-readable storage device. The example methods described herein may be implemented using encoded instructions (e.g., computer-readable instructions) stored on a non-transitory computer-readable medium such as flash memory, read-only memory (ROM), random access memory (RAM), cache, or any other storage medium in which information is stored for any duration (e.g., stored for an extended period of time, stored permanently, stored for a brief instance, stored for temporary buffering and / or stored for caching of information). The computer memory of a computer-readable storage medium as referenced herein may include volatile and non-volatile or removable and non-removable media for storing information in electronic format such as computer-readable program instructions or modules of computer-readable program instructions, data, etc., which may be standalone or may be part of a computing device. Examples of computer memory may include any other medium that can be used to store information in the required electronic format and that can be accessed by one or more processors or at least a portion of a computing device.
[0020] The processor 106 may be configured to execute machine-readable instructions stored in the memory 104. The processor 106 may be electronically and / or communicatively coupled to the camera 102, and may process and analyze video including images received from the camera 102. In some examples, the processor 106 may be configured to retrieve useful information from the video, estimate environmental effects based on the retrieved information, and generate a calibration / compensation filter with adaptive filter coefficients to compensate for the environmental effects. The processor 106 may perform the calibration method described above, as will be described in detail below with respect to Figures 2 to 6 Detailed description.
[0021] The processor 106 may be single-core or multi-core, and the programs executed by the processor 106 may be configured for parallel processing or distributed processing. The processor 106 may be any technically feasible hardware unit configured to implement processing functions and execute software applications, including but not limited to a central processing unit (CPU), a microcontroller unit (MCU), an application specific integrated circuit (ASIC), a digital signal processor (DSP) chip, a field programmable gate array (FPGA), a graphic board, etc.
[0022] In addition, Figure 1An audio pipeline from an audio source 108 to a speaker 110 is shown. It can be understood that some modules / functions (not shown) in the audio pipeline may include an EQ filter (equalizer filter), a limiter, a gain unit, a delay unit, an amplifier, etc., which can be implemented by software, hardware, or a combination of software and hardware. It can also be understood that the audio system may include more than one speaker and more than one corresponding camera for more than one speaker. Figure 1 It is only an example for clearly presenting and explaining the principles of the proposed methods and systems, which will be described in detail below.
[0023] Figure 2 A flowchart of an automatic audio calibration method for an audio system according to one or more embodiments of the present disclosure is shown. At S202, video in the room where the audio system is located can be obtained through a camera. At S204, useful information can be retrieved from the video. In some embodiments, the useful information may include environmental information and listener information in the room. In some examples, the environmental information includes location information about at least one reflector or obstacle in the room. In some embodiments, the listener information includes location information about the listener in the room, such as the head position or ear position of the listener. Then, at S206, based on the retrieved environmental information and listener information, the environmental impact in the sound field at the listener (e.g., at the listener's head or ear) can be estimated. The environmental impact is associated with at least one reflected sound caused by at least one reflector or obstacle. At S208, a compensation filter can be generated based on the estimated environmental impact. In some examples, the filter coefficients of the compensation filter can be generated and applied to the EQ filter in the audio system.
[0024] Figure 3 A schematic diagram of information retrieval from video by a processor according to one or more embodiments of the present disclosure is shown. At block 302, existing object recognition methods or algorithms can be used to roughly identify different objects in the video. Among the identified objects, the listener and large reflectors should be selected, and the large reflectors can be ignored. In some examples, some large reflectors can be determined as main reflectors by comparing the size of each reflector with a size threshold. The size threshold can be preset by engineers according to their practical experience. In some examples, reflectors with a size larger than the size threshold are selected as main reflectors. For example, the main reflectors may involve walls, floors, and furniture with large planes, such as tables and desks.
[0025] At block 304, environmental detection can be performed to obtain the position information of the identified main reflector. In some examples, the environmental detection can be performed only once, for example, when the audio system is first powered on. In some examples, the environmental detection can be performed at longer time intervals, such as once a month or several months, or once a year. This is because these large reflectors rarely move. Once the position information is obtained, other information such as room volume and shape can also be inferred.
[0026] At block 306, listener detection can be performed to obtain the position information of the listener. In some examples, existing head tracking methods or algorithms can be used to perform listener detection to obtain the position information of the listener's head. Compared with environmental detection, the detection of the listener should always run to track the movement of the listener. Knowing the real-time position of the listener's head or ears is necessary for the effectiveness of calibration.
[0027] The specific method or algorithm for information retrieval can vary according to the exact type of camera and video. When considering the position tracking of an individual, the common method is to use an optical camera (such as an RGB camera) in combination with face recognition. However, optical cameras are affected by environmental conditions (shadows, low light, sunlight, etc.) and cannot obtain accurate distance measurements. In addition, complex processing (such as face recognition algorithms) is also required. More importantly, there are also privacy issues with cameras.
[0028] In the present disclosure, the recommended example is to use a TOF camera. A TOF camera provides a 3-D image through a CMOS array together with an active modulated light source. It works by illuminating the scene with a modulated light source (a solid-state laser or an LED, usually near-infrared light invisible to the human eye) and observing the reflected light. The time delay of the light can reflect the distance information.
[0029] Figure 4 An example of a video captured by a TOF camera is shown. Although the video examples herein are shown as grayscale images, it can be understood that the videos captured by the TOF camera can be in color. Different colors are used to distinguish objects at different depths. For example, the listener can be identified with a red sketch, and other reflectors can be identified as yellow or green blocks. Even in grayscale images, the listener and reflectors can be identified with different gray levels. As an example, Figure 4 the coordinates shown in indicate the position and can be directly obtained by the TOF camera. Therefore, the processor can obtain the position information of the listener and the main reflector from the video received by the camera. Figure 4
[0030] The motion sensors used in this disclosure have the following advantages: being robust in various environments (especially dark environments), being easy to integrate with an audio system due to relatively simple object recognition and tracking and on-chip processing, and having no privacy issues. For example, a rather simple algorithm can be applied to detect listeners and large reflectors in the background. As Figure 4 shown, different objects can be relatively easily distinguished through depth information, while ordinary RGB cameras may require complex algorithms for face recognition. The TOF camera can continuously track the listener (e.g., the listener's head or ear) and provide a video to the processor that includes position information associated with the movement of the listener.
[0031] Once the processor retrieves the position information associated with the reflector and the listener from the video, the processor can analyze the position information to estimate the environmental impact on the sound field at the listener, and can derive the filter coefficients of a compensation filter (i.e., EQ filter coefficients suitable for the EQ filter in the audio system) to compensate for the estimated environmental impact. For example, the EQ filter in the audio system may include a high-pass filter, a low-pass filter, a band-pass filter, a peak filter, etc. In some embodiments, the environmental impact is associated with at least one reflected sound caused by at least one reflector. In some embodiments, the compensation filter (e.g., the compensation EQ filter) can be generated or designed through empirical methods as well as through physical modeling and calculations.
[0032] Examples of obtaining the compensation filter through modeling and calculation methods will be shown. For illustrative purposes, Figure 5 a simple setup of the calibration process is shown. Figure 5 The spatial position of the speaker is shown. As Figure 5 shown, the speaker box 502 with a speaker inside is positioned on a table with a large reflecting plane. The camera 504 (e.g., a TOF camera) is positioned near the speaker and on top of the speaker box. It can be understood that Figure 5 the examples are presented solely for illustrative purposes and are not intended to be exhaustive or limited to the examples disclosed herein. The camera 504 can be located at any position near the speaker where the camera can detect and record information about the room.
[0033] Figure 5 The sound propagation path 508 for propagating the direct sound from the speaker to the listener 506 is shown. The direct sound refers to the sound received by the listener, which emits from the speaker and reaches the listener directly without any reflection. There is another sound propagation path 510 for propagating the reflected sound. The reflected sound refers to the sound emitted by the speaker that reaches the listener after being reflected by the reflecting plane.
[0034] In some embodiments, the retrieved location information may include a distance L, which indicates the distance from the speaker to the listener's head or ear. In some embodiments, the retrieved location information may further include a distance H, which indicates the perpendicular distance from the listener's head or ear to the plane of the reflecting surface. In some embodiments, the retrieved location information may include a distance h, which indicates the perpendicular distance from the speaker to the reflecting surface of the plane where the speaker is located. In Figure 5 the example of Figure 5 , the listener is at a distance L from the speaker, the height of the listener's head or ear above the reflecting surface of the table is H, and the height of the driver of the speaker above the table plane is h. The positions of L, H, and h can all be obtained from the retrieved useful information. Alternatively, the height h can be obtained from the layout design of the speaker.
[0035] Based on the location information L, H, and h, the processor can estimate the environmental impact on the sound field at the listener. In some embodiments, the sound pressure is used as a parameter for estimating the environmental impact. For convenience, it is assumed that the damping factor in sound propagation and reflection is uniform. The total sound pressure P t is the superposition of the direct sound pressure P d and the reflected sound pressure P r , expressed as follows:
[0036]
[0037] where P0 is the sound pressure at the speaker, k is the wave number of the sound, and β is the damping factor. The approximation holds when L >> h. The travel distance of the reflected sound can be obtained by the Cosine Theorem. The angle θ is obtained based on the positions of the speaker and the listener, as follows:
[0038]
[0039] Therefore, the total sound pressure at the listener is expressed as
[0040]
[0041] It can be understood that the reflected sound pressure P r represents the interference caused by the reflected sound, which can be regarded as the environmental impact and can be estimated based on the above equation. Due to the interference caused by the reflected sound, the total sound heard by the listener is different from the direct sound, and the response of the total sound varies with frequency. Therefore, it is necessary to compensate for or eliminate the influence of the room environment (e.g., caused by certain main reflectors) on the sound field at the listener.
[0042] In some embodiments, a compensation filter can be generated or designed to compensate for the interference caused by the reflected sound. In some examples, based on the total sound pressure and the reflected sound pressure, a compensation filter can be generated or designed to compensate for the reflected sound interference. In some embodiments, the filter coefficients of the compensation filter can be generated based on the estimated environmental impact calculated by the above equations. The generated filter coefficients can be applied to the EQ filter in the audio system. The EQ filter to which the generated filter coefficients are applied can collectively correspond to the compensation filter. In some embodiments, the generation of the filter coefficients of the EQ filter in the audio system can include generating filter coefficients such that the response of the EQ filter to which the generated filter coefficients are applied (i.e., the response of the compensation filter) can compensate for or eliminate the difference between the frequency response of the total sound and the frequency response of the direct sound. In some examples, the generation of the filter coefficients of the EQ filter in the audio system can include generating filter coefficients such that the amplitude response of the compensation filter can compensate for or eliminate the difference between the amplitude response of the total sound and the amplitude response of the direct sound. In other words, some adaptive EQ filters can be selected to compensate for the environmental impact so that a relatively flat frequency response is achieved at the listener.
[0043] Figure 6 An example of the frequency response curves of the direct sound, the total sound, and the compensation EQ filter is shown. In Figure 6 it, the three curves 602, 604, and 606 represent the amplitudes of the direct sound response, the total sound response, and the compensation filter response respectively, which are simulated with the following parameters: h = 0.08 cm, H = 0.48 cm, L = 0.79 cm, β = 0.5. For a more intuitive illustration, Figure 6 the amplitude responses shown are normalized amplitude responses. For example, the normalized amplitude responses of the direct sound and the total sound are obtained by and respectively. It can be seen that due to the interference from the reflected sound waves, the total sound field at the listener varies with frequency. The interference from the reflected sound waves can cause an increase or decrease in the frequency response of different frequency bands, depending on the positions of the listener, the speaker, and the reflector. However, the generated compensation filter can compensate for the environmental impact. Additionally, since the position of the listener is continuously tracked, the generated compensation filter can compensate for the movement of the listener. As described above, the filter coefficients of the EQ filter in the audio system can be adjusted in real time according to the detected position of the listener's head or ears.
[0044] For clarity of presentation and explanation, a configuration of one camera and one loudspeaker is used as an example to illustrate how to retrieve information from a video and analyze the information, and how to estimate and compensate for the environmental impact on the sound field at the listener. However, it can be understood that the audio system may include multiple loudspeakers, and a corresponding camera may be present near each loudspeaker. For each configuration including one loudspeaker and one camera, the methods described in this disclosure can be employed. Additionally, it can be understood that there may be multiple reflectors in the room, and thus multiple reflected sounds are generated by these multiple reflectors. Figure 5 Only one reflected sound is shown, which is presented for illustrative purposes only and is not intended to be exhaustive or limited to the number of reflected sounds. In the case of multiple reflected sounds, the sound pressure P of the reflected sound r can represent the superposition of the sound pressures of each reflected sound, for example P r = P r1 + P r2 ..., + P rn . For each reflected sound, the methods described in this disclosure can be used to estimate the interference of the reflected sound on the sound field and generate appropriate filter coefficients to compensate for or cancel the interference caused by each reflected sound.
[0045] In this disclosure, a new method for acoustic calibration via video is provided. The environment and the listener in the room can be captured in the video. Thus, the position information of the environment and the position information of the listener can be retrieved, and the position of the listener can be continuously tracked. Then, based on the position information, the interference of the reflected sound waves from the listener can be predicted to generate an EQ filter to compensate for the environmental impact. By using the techniques described herein, an all-in-one form factor combining the loudspeaker and the camera can be obtained. The automatic audio calibration described herein can compensate for the effects of the room environment and the listener's position. Additionally, no additional hardware is required, and there are no privacy issues. Further, no complex algorithms are required, and thus computation time is saved and system robustness is improved. As a result, the listener can obtain a better listening experience.
[0046] Clause 1. In some embodiments, an automatic audio calibration method for an audio system in a room, comprising: capturing a video of the room via a camera; retrieving environmental information and listener information from the video; estimating an environmental impact in the sound field at the listener based on the environmental information and the listener information; and generating a compensation filter for the audio system to compensate for the estimated environmental impact.
[0047] Clause 2. The method according to Clause 1, wherein retrieving the environmental information and listener information from the video includes: identifying an object from the video; selecting at least one main reflector and one listener; and obtaining position information of the at least one main reflector and position information of the listener.
[0048] Clause 3. The method according to any one of Clauses 1 to 2, wherein estimating the environmental impact in the sound field at the listener includes estimating at least one reflected sound pressure of at least one reflected sound based on the environmental information and listener information, wherein the at least one reflected sound is caused by at least one reflection plane of the at least one main reflector.
[0049] Clause 4. The method according to any one of Clauses 1 to 3, wherein generating the compensation filter for the audio system includes generating filter coefficients based on the estimated environmental impact.
[0050] Clause 5. The method according to any one of Clauses 1 to 4, further comprising applying the generated filter coefficients to an EQ filter in the audio system.
[0051] Clause 6. The method according to any one of Clauses 1 to 5, wherein estimating at least one reflected sound pressure includes: obtaining a first distance indicating the distance from a speaker in the audio system to the head or ear of the listener; and for each reflector, obtaining a second distance indicating the perpendicular distance from the head or ear of the listener to the plane of the reflection plane of the reflector; obtaining a third distance indicating the perpendicular distance from the speaker to the reflection plane; and estimating the reflected sound pressure based on the first distance, the second distance, and the third distance.
[0052] Clause 7. The method according to any one of Clauses 1 to 6, wherein generating filter coefficients based on the estimated environmental impact includes generating filter coefficients such that the amplitude response of the compensation filter with the generated filter coefficients compensates for the difference between the amplitude response of the total sound and the amplitude response of the direct sound.
[0053] Clause 8. The method according to any one of Clauses 1 to 7, wherein the total sound includes the superposition of a direct sound and at least one reflected sound, and the direct sound represents a sound wave that is emitted from the speaker and reaches the listener directly without any reflection.
[0054] Clause 9. The method according to any one of Clauses 1 to 8, wherein selecting at least one main reflector includes selecting at least one reflector whose size is greater than a size threshold as the at least one main reflector.
[0055] Clause 10. The method according to any one of Clauses 1 to 9, wherein the environmental information includes position information of at least one main reflector, and wherein the listener information includes position information of the listener's head or ears.
[0056] Clause 11. In some embodiments, an automatic audio calibration system for an audio system in a room, comprising: a camera configured to capture video of the room through the camera; and a processor coupled to the camera and configured to: retrieve environmental information and listener information from the video; estimate environmental effects in the sound field at the listener based on the environmental information and the listener information; and generate a compensation filter for the audio system to compensate for the estimated environmental effects.
[0057] Clause 12. The system according to Clause 11, wherein the processor is further configured to: identify objects from the video; pick out at least one main reflector and one listener; and obtain position information of the at least one main reflector and position information of the listener.
[0058] Clause 13. The system according to any one of Clauses 11 to 12, wherein the processor is further configured to estimate at least one reflected sound pressure of at least one reflected sound based on the environmental information and the listener information, wherein the at least one reflected sound is caused by at least one reflecting plane of the at least one main reflector.
[0059] Clause 14. The system according to any one of Clauses 11 to 13, wherein the processor is further configured to generate filter coefficients based on the estimated environmental effects.
[0060] Clause 15. The system according to any one of Clauses 11 to 14, wherein the processor is further configured to apply the generated filter coefficients to an EQ filter in the audio system.
[0061] Clause 16. The system according to any one of Clauses 11 to 15, wherein the processor is further configured to: obtain a first distance indicating the distance from a speaker in the audio system to the listener's head or ears; and for each reflector, obtain a second distance indicating the perpendicular distance from the listener's head or ears to the plane of the reflecting plane of the reflector; obtain a third distance indicating the perpendicular distance from the speaker to the reflecting plane; and estimate the reflected sound pressure based on the first distance, the second distance, and the third distance.
[0062] Clause 17. The system according to any one of Clauses 11 to 16, wherein the processor is further configured to generate filter coefficients such that the magnitude response of a compensation filter with the generated filter coefficients compensates for the difference between the magnitude response of the total sound and the magnitude response of the direct sound.
[0063] Clause 18. The method according to any one of Clauses 11 to 17, wherein the total sound includes a superposition of a direct sound and at least one reflected sound, and the direct sound indicates a sound wave that is emitted from a speaker and reaches a listener directly without any reflection.
[0064] Clause 19. The system according to any one of Clauses 11 to 18, wherein the processor is further configured to select at least one reflector having a size greater than a size threshold as the at least one main reflector.
[0065] Clause 20. In some embodiments, a computer-readable storage medium includes computer-executable instructions that, when executed by a computer, cause the computer to perform the method according to any one of Clauses 1 to 10.
[0066] The descriptions of the various embodiments have been presented for purposes of illustration, but are not intended to be exhaustive or limited to the disclosed embodiments. The terms used herein have been chosen to best explain the principles of the embodiments, the practical application or technical improvement of technologies found in the market, or to enable those skilled in the art to understand the embodiments disclosed herein.
[0067] Previously, reference has been made to the embodiments presented in this disclosure. However, the scope of this disclosure is not limited to the specific described embodiments. Rather, any combination of the foregoing features and elements, whether or not they relate to different embodiments, is contemplated to implement and practice the contemplated embodiments. In addition, although the embodiments disclosed herein may achieve advantages over other possible solutions or over the prior art, the scope of this disclosure is not limited by whether a given embodiment achieves a specific advantage. Thus, the foregoing aspects, features, embodiments, and advantages are merely illustrative and are not to be considered elements or limitations of the appended claims, unless expressly recited in the claims.
[0068] Aspects of the present disclosure may take the form of: a fully hardware embodiment, a fully software embodiment (including firmware, resident software, microcode, etc.), or an embodiment combining software and hardware aspects, which are generally referred to herein as "circuitry", "module", "unit", or "system".
[0069] The present disclosure may be a system, method, and / or computer program product. The computer program product may include one or more computer-readable storage media having computer-readable program instructions thereon for causing a processor to implement aspects of the present disclosure.
[0070] The computer-readable program instructions described herein may be downloaded from a computer-readable storage medium to a corresponding computing / processing device, or downloaded to an external computer or external storage device via a network (e.g., the Internet, a local area network, a wide area network, and / or a wireless network). The network may include copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers.
[0071] Aspects of the present disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present disclosure. It should be understood that each block in the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0072] These computer-readable program instructions may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions executed via the processor of the computer or other programmable data processing apparatus create means for implementing the functions / acts specified in one or more blocks of the flowchart and / or block diagram.
[0073] The flowchart and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagrams may represent a module, segment, or portion of instructions, which includes one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently, or the blocks may sometimes be executed in the reverse order, depending upon the functionality involved. It should also be noted that each block of the block diagrams and / or flowchart illustrations, and combinations of blocks in the block diagrams and / or flowchart illustrations, can be implemented by a system based on dedicated hardware that performs the specified functions or acts or a combination of dedicated hardware and computer instructions.
[0074] Although the foregoing relates to embodiments of the present disclosure, other and additional embodiments of the present disclosure may be envisioned without departing from the basic scope of the present disclosure, and the scope of the present disclosure is determined by the appended claims.
Claims
1. An automatic audio calibration method for an audio system in a room, comprising: Capturing a video of the room by a camera; Retrieving environmental information and listener information from the video; Estimating environmental effects in the sound field at the listener based on the environmental information and the listener information; And Generating a compensation filter for the audio system to compensate for the estimated environmental effects.
2. The method according to claim 1, wherein retrieving the environmental information and the listener information from the video comprises: Identifying objects from the video; Selecting at least one main reflector and one listener; And Obtaining position information of the at least one main reflector and position information of the listener.
3. The method according to claim 2, wherein estimating the environmental effects in the sound field at the listener comprises estimating at least one reflected sound pressure of at least one reflected sound based on the environmental information and the listener information, wherein the at least one reflected sound is caused by at least one reflecting plane of the at least one main reflector.
4. The method according to claim 1, wherein generating the compensation filter for the audio system comprises generating filter coefficients based on the estimated environmental effects.
5. The method according to claim 4, further comprising applying the generated filter coefficients to an EQ filter in the audio system.
6. The method according to claim 3, wherein estimating at least one reflected sound pressure comprises: Obtaining a first distance indicating a distance from a speaker in the audio system to the head or ear of the listener; And For each reflector: Obtaining a second distance indicating a perpendicular distance from the head or ear of the listener to a plane where the reflecting plane of the reflector is located; Obtaining a third distance indicating a perpendicular distance from the speaker to the reflecting plane; And Estimating the reflected sound pressure based on the first distance, the second distance, and the third distance.
7. The method according to claim 4, wherein generating the filter coefficients based on the estimated environmental effects comprises generating the filter coefficients such that an amplitude response of the compensation filter with the generated filter coefficients compensates for a difference between an amplitude response of the total sound and an amplitude response of the direct sound.
8. The method according to claim 7, wherein the total sound includes a superposition of the direct sound and at least one reflected sound, and the direct sound indicates a sound wave that is emitted from the speaker and reaches the listener directly without any reflection.
9. The method according to claim 2, wherein selecting at least one main reflector comprises selecting at least one reflector whose size is greater than a size threshold as the at least one main reflector.
10. The method according to claim 1, wherein the environmental information includes position information of at least one main reflector, and the listener information includes position information of the head or ear of the listener.
11. An automatic audio calibration system for an audio system in a room, comprising: A camera configured to capture a video of the room by the camera; And A processor coupled to the camera and configured to: Retrieve environmental information and listener information from the video; Estimate the environmental impact in the sound field at the listener based on the environmental information and the listener information; And Generate a compensation filter for the audio system to compensate for the estimated environmental impact.
12. The system according to claim 11, wherein the processor is further configured to: Identify an object from the video; Pick out at least one main reflector and one listener; and Obtain the position information of the at least one main reflector and the position information of the listener.
13. The system according to claim 12, wherein the processor is further configured to estimate at least one reflected sound pressure of at least one reflected sound based on the environmental information and the listener information, wherein the at least one reflected sound is caused by at least one reflection plane of the at least one main reflector.
14. The system according to claim 11, wherein the processor is further configured to generate filter coefficients based on the estimated environmental impact.
15. The system according to claim 14, wherein the processor is further configured to apply the generated filter coefficients to an EQ filter in the audio system.
16. The system according to claim 13, wherein the processor is further configured to: Obtain a first distance indicating a distance from a speaker in the audio system to the head or ear of the listener; And For each reflector: Obtain a second distance indicating the perpendicular distance from the head or ear of the listener to the plane of the reflection plane of the reflector; Obtain a third distance indicating the perpendicular distance from the speaker to the reflection plane; And Estimate the reflected sound pressure based on the first distance, the second distance, and the third distance.
17. The system according to claim 14, wherein the processor is further configured to generate the filter coefficients such that the amplitude response of the compensation filter with the generated filter coefficients compensates for the difference between the amplitude response of the total sound and the amplitude response of the direct sound.
18. The system according to claim 17, wherein the total sound includes the superposition of the direct sound and at least one reflected sound, and the direct sound indicates a sound wave that is emitted from the speaker and reaches the listener directly without any reflection.
19. The system according to claim 12, wherein the processor is further configured to select at least one reflector having a size greater than a size threshold as the at least one main reflector.
20. A computer-readable storage medium comprising computer-executable instructions that, when executed by a computer, cause the computer to perform the method according to any one of claims 1 to 10.