Headphones with a digital audio processor system utilizing 3D mapping to modify the original audio stream and a method of modifying the original audio stream

Headphones with 3D spatial mapping sensors dynamically adjust sound effects based on the listener's environment, addressing the static nature of existing technologies and providing enhanced spatial awareness and immersion.

WO2025186780A1PCT designated stage Publication Date: 2025-09-11ZIOBRO PAWEL
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/IB2025/052484
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-03-08
Filing Date
2025-03-07
Publication Date
2025-09-11

AI Technical Summary

Technical Problem

Existing audio technologies fail to dynamically adapt sound effects to the actual spatial environment of the listener, providing a static and inaccurate auditory experience.

Method used

Headphones equipped with 3D spatial mapping sensors, such as LIDAR, cameras, or ultrasonic sensors, modify the audio stream in real-time based on the listener's surroundings, calculating spatial parameters to dynamically adjust sound effects like reverb and echo.

Benefits of technology

The headphones provide a dynamic and accurate auditory experience by adapting sound effects to the listener's environment, enhancing spatial awareness and immersion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IB2025052484_12092025_PF_FP_ABST
    Figure IB2025052484_12092025_PF_FP_ABST
Patent Text Reader

Abstract

The subject of the invention is a set of headphones (1) in which the original audio stream (11) is processed by a digital audio processor system (2) that utilizes 3D spatial mapping for modification, performed by 3D spatial mapping sensors (3), characterized in that the left speaker (4) and right speaker (5) of the headphones (1) are connected to the audio processor (2), and the audio processor (2) is connected to a left 3D spatial mapping sensor (6), located in the left headphone (7) and oriented toward the left side of the listener, as well as to a right 3D spatial mapping sensor (8), located in the right headphone (9) and oriented toward the right side of the listener, and the audio processor (2) is further connected to a source (10) of the original audio stream (11). The subject of the invention also includes a method of processing the original audio stream (11) by the digital audio processor (2), which adapts to changing spatial parameters of the environment surrounding the listener.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Headphones with a digital audio processor system utilizing 3D mapping to modify the original audio stream and a method of modifying the original audio stream

[0002] The subject of the invention is headphones with a digital audio processor system utilizing 3D spatial mapping to modify the original audio stream and a method of modifying the original audio stream. The digital processor system modifies the original sound, taking into account the parameters of the 3D space in which the listener is located at a given moment. The intended result of such modification is to evoke in the listener a sensation of auditory contact with the space in which they are located and moving at a given moment.

[0003] In the prior art document US2011299707 Al, a method and computer program related to spatial sound used in over-ear headphones is disclosed. The solution comprises a set of headphones with an accelerometer and tilt sensor for tracking the position and orientation of the headphone set, as well as a computing device including a headphone position processor to receive information on the position and orientation of the headphones. The position of the virtual speaker is determined by processors (VSLP), which receive a digital signal containing audio information from the digital audio stream. A digital signal containing information on the location and orientation of the headphones from the headphone position processor is sent, and a digital signal containing audio information is sent to a summing processor (receiving digital output signals from VSLP), then the signals are summed and sent to a digital-to-analog converter (DAC). The digital-to-analog converter converts the summed digital output signals received from the VSLP into an analog signal and outputs the analog signal to the headphone set.

[0004] In contrast, patent document US2008031475 Al describes a device and method of a personal audio assistant. The solution consists of an ambient microphone, canal microphone, ear canal receiver, sealing section, logic circuit, communication module, memory unit, and user interaction element. The user interaction element is configured to send playback commands to the logic circuit upon activation by the user, whereupon the logic circuit reads recording parameters stored in the memory unit and sends audio content to the ear canal receiver according to the recording parameters. Furthermore, patent document US2017099539 Al presents a headphone device comprising a first headphone cup, a sensor contained within the first headphone cup, and a controller. The first earcup is connected to a headband and contains a speaker system and a thermal control subsystem. The controller is configured to detect output data generated by the sensor and, based on the output data, to transmit a signal to the thermal control subsystem contained within the first earcup to modify the temperature associated with at least the first earcup.

[0005] The invention provides headphones with a digital audio processor system utilizing 3D mapping to modify the original audio stream and a method of modifying the original audio stream, by which the auditory impressions received by the user reflect the environment in which the user is located, more precisely imitating the spatial conditions of the room in which the listener is located.

[0006] Essence of the invention

[0007] The subject of the invention is headphones that modify the original audio stream through a digital audio processor system that uses 3D spatial mapping by 3D mapping sensors, characterized in that the left and right headphone speakers are connected to the digital audio processor, and the digital audio processor is connected to a left 3D spatial mapping sensor located on the left side of the listener, placed in the left headphone, and a right 3D spatial mapping sensor located on the right side of the listener, placed in the right headphone, and the digital audio processor is connected to the source of the original audio stream.

[0008] Preferably, the 3D spatial mapping sensor is an optical distance sensor.

[0009] Preferably, the optical sensors are laser sensors.

[0010] Preferably, the optical sensors are LIDAR laser sensors.

[0011] Preferably, the optical sensors are cameras.

[0012] Preferably, the 3D spatial mapping sensor is an ultrasonic distance sensor. Preferably, the 3D spatial mapping sensor is any combination of LIDAR sensors, cameras, and ultrasonic distance sensors.

[0013] Preferably, the source of the original audio stream is connected to the headphones by wire or wirelessly.

[0014] Preferably, the source of the original audio stream is connected to the headphones wirelessly via Bluetooth.

[0015] The subject of the invention is also a method of modifying the original audio stream by an digital audio processor adapting to the changing spatial parameters of the environment surrounding the listener on their left and right side, characterized by the following steps: a) the digital audio processor receives the original audio stream from the source, and from the sensors receives information about the space surrounding the listener, then b) the digital audio processor, based on information about the space surrounding the listener, determines the parameters of this space, then c) the digital audio processor, based on the determined spatial parameters, calculates coefficients for calculating reverb, echo, and the direction of the original audio stream source, then d) the digital audio processor, based on the calculated coefficients, modifies the parameters of the original audio stream by adding reverb, echo, and direction of the source, then e) the modified audio stream is emitted through the headphones.

[0016] Preferably, the spatial information comprises data on the distance to various objects and spaces on the left side of the listener for the left headphone and on the right side of the listener for the right headphone.

[0017] Preferably, the spatial parameters are volume, distance to objects, the surface area on which the listener is located, acoustic absorption of the room, the average sound absorption coefficient, the area of the planes delimiting the room, including the floor, the average distance from all objects and surfaces, the farthest measured distance to a surface for the space on the left side of the listener for the left headphone and the space on the right side of the listener for the right headphone.

[0018] Preferably, steps from a) to d) are performed cyclically.

[0019] The essence of the invention can also be illustrated by home Hi-Fi systems, especially the sound effects they generate. Some of them offer the possibility of modifying the sound so that the listener has the impression that they are, for example, "in a small room", "in a large room", "in a small hall", or "in a large hall". These effects are applied statically at the listener's request, without any examination of the space in which the listener is located.

[0020] The essence of the invention described in this document is that these sound effects will be applied dynamically based on the examined space in which the listener is located. If the listener is in a small room, the sound in the headphones will be modified to match the parameters of sound in a "small room," while if the listener moves from a small room to a large room, the digital audio processor, using 3D space scanning by 3D sensors, will detect this and dynamically modify the sound to match the parameters of sound in a "large room".

[0021] The subject of the invention is shown in the drawing, in which:

[0022] Fig. 1 shows a view of the headphones with individual components,

[0023] Fig. 2 shows a general view of the headphones,

[0024] Fig. 3 and Fig. 4 show the listener in a room,

[0025] Fig. 4A shows the view of pyramids mapped in space,

[0026] Fig. 4B shows the delay line of a FIR echo filter.

[0027] Example of embodiment - headphones

[0028] An embodiment is shown in the drawing (Fig. 4) below (to keep the drawing legible, only the operation of the right headphone is shown. The left headphone operates the same way, but if it were also shown in the drawing, it would become unreadable). Headphones (1) modifying the original audio stream (11) through a digital audio processor (2) connected to the source (10) of the original audio stream (11), os well os to the left speaker (4) and right speaker (5) of the headphones (1), characterized in that the diaitalaudio processor (2) uses 3D spatial mapping for modification through 3D mapping sensors (3) located in the headphones, in particular the digital audio processor (2) is connected to the left sensor (6) mapping the 3D space on the left side of the listener, placed in the left headphone (7), and the right sensor (8) mapping the 3D space on the right side of the listener, placed in the right headphone (9).

[0029] The digital audio processor is a key component of modern audio systems, constituting the heart of acoustic signal modification technology. Its main task is to optimize and modify sound to adapt it to specific listening requirements and conditions. Thanks to advanced DSP (Digital Signal Processing) algorithms, the digital audio processor analyzes and transforms the input signal in real time, providing additional auditory effects.

[0030] In the basic embodiment, the 3D spatial mapping sensors (3) are optical distance sensors such as LIDAR sensors. An optical distance sensor is a device that uses light to accurately measure the distance to an object without physical contact. Its operation is based on principles of optics and electronics, allowing for fast and precise distance measurements. The sensor emits a beam of light, often laser, which, after reflecting off an object, returns to the device's detector. Based on the time that elapsed between emission and reception of the light, and knowing the speed of light, the device is able to calculate the distance to the object. There are several measurement technologies and methods used in optical distance sensors, including Time of Flight (ToF), which directly measures the time of flight of light, laser triangulation, using geometric phenomena to calculate distance, and interferometry, which uses the interference of light for highly precise measurements at short distances.

[0031] In another embodiment of the invention, the optical sensors are cameras. Ca m e ra s mounted on the headphones track characteristic points in the environment, allowing precise determination of the user's position in space. This technology, known as "inside-out tracking," allows for free movement without the need for additional stationary external sensors. In "inside-out tracking" technology, cameras are placed on the headphones and actively scan the environment, recording characteristic points or spatial features such as edges, walls, furniture, and other distinct objects. Based on changes in the observed features, image processing and data analysis algorithms calculate the device's orientation in space, allowing precise tracking of spatial parameters. Monochromatic or RGB cameras are most commonly used, as well as depth cameras (e.g. ToF - Time of Flight), which record detailed information about the surroundings. "Inside-out tracking" uses a variety of algorithms that allow precise determination of the device's position and orientation in space. These algorithms are based on advanced image processing, sensor data analysis, and artificial intelligence techniques. SLAM is one of the main algorithms used in "inside-out tracking," which enables simultaneous mapping of the environment and determination of the device's own position within it. The SLAM algorithm uses camera data to create a realtime map of the environment. There are different SLAM variants, including visual SLAM (V- SLAM), which mainly uses visual data, and SLAM based on depth sensors.

[0032] In another embodiment of the invention, the 3D spatial mapping sensor (3) is an ultrasonic distance sensor. An ultrasonic distance sensor is a device that uses high-frequency sound waves, inaudible to the human ear, to measure the distance to an object. It works based on the principle of echolocation. The sensor emits short ultrasonic pulses that travel through the air and reflect off encountered obstacles. The echo reflected from the object returns to the sensor, where it is recorded by the built-in receiver. Based on the time that elapsed between sending and receiving the ultrasound, and knowing the average speed of sound in the air, the device is able to calculate the distance to the object.

[0033] The source (10) of the original audio stream (11) is connected to the headphones either bv wire or wirelessly. The source (10) of the original audio stream, in the context of music playback, may be various devices and platforms that deliver an audio stream to headphones, speakers, or other playback systems. Modern technologies allow the use of many types of audio streams, both analog and digital. Smartphones and tablets are among the most popular audio stream sources due to their versatility and easy access to a wide range of music applications and streaming services. They allow for wireless transmission of the audio stream to headphones via Bluetooth, as well as wired transmission of the audio stream through audio connectors. Computers and laptops also play an important role as audio stream sources, especially in a home work or entertainment environment. They can play music / sound from various applications, games, or streaming services, as well as from local files stored on the hard drive. Connection to playback devices can be done both wired and wirelessly.

[0034] Example of embodiment - method

[0035] The method of sound modification will be presented in a specific situation in which the listener is located in a room shown in Fig. 4 with the following assumptions:

[0036] 1. The listener is located in a room with dimensions 5m x 6m x 2.9m,

[0037] 2. The distance between the LIDAR sensors is 0.1m,

[0038] 3. The listener is wearing headphones shown in Fig. 1 and Fig. 2, with the sound processor enabled that uses 3D spatial mapping, the subject of this patent,

[0039] 4. LIDAR sensors are used, transmitting information about the measured distances to the sound processor,

[0040] 5. The LIDAR sensor in the right headphone scans the space by shooting laser beams in different directions on the right side of the listener and measures the distances from the listener to objects and surfaces located on their right side. The left headphone operates in the same way for the left side of the listener; however, the visualization of this operation was omitted in Fig. 4 to improve the readability of the drawing,

[0041] 6. The source of the audio signal is the listener's smartphone, which transmits it via Bluetooth to the headphones,

[0042] 7. The center of the listener's head is located in the geometric center of the room.

[0043] The digital audio processor, according to the procedure described in claim 10), modifies the original audio stream as described in steps a) to e), which will form the basis of the following description of the method of sound modification.

[0044] In step a) of the method according to the invention: „the digital audio processor (2) receives the original audio stream (11) from the source (10), and information about the space surrounding the listener from the sensors (6)(8)." The digital audio processor receives the original audio stream via Bluetooth, as well as data concerning measured distances from the LIDAR laser sensors.

[0045] In step b) of the method according to the invention: „the digital audio processor (2), based on the spatial information, determines the parameters of the space"

[0046] This step is performed by the digital audio processor based on the transmitted distances and calculates the parameters of this space in continuous (real-time) processing, in particular:

[0047] 1. Volume of space - for each side of the listener the following is calculated: VL (approximate volume of the space on the left side of the listener) and VP (approximate volume of the space on the right side of the listener), and based on their sum V is determined - (approximate volume of the space around the listener),

[0048] To calculate the volume of the space, it must be divided into triangular pyramids. To do this, it is sufficient to group all laser beams by nearest neighbors. This way, the space will be divided into adjacent triangular pyramids. Fig. 4 shows one of the triangular pyramids, determined in space based on three neighboring laser beams. Each such pyramid has the characteristics shown in Fig. 4a with labels:

[0049] C - point from which the laser beams are emitted - LIDAR laser sensor,

[0050] A, B, C, D - vertices of the triangular pyramid, a, b, c - vectors defined by LIDAR laser beams (based on the distance measured by the laser and the angle of the laser beam emission) from the sensor to each vertex of the pyramid [m],

[0051] VIDS,.- volume of the i-th pyramid in the space on the left side of the listener [m3].

[0052] The volume will be calculated using the formula for the volume of a triangular pyramid:

[0053] VLOs^^^ aXb }* cl

[0054] 6 Then, all pyramid volumes will be summed and presented as an approximate value of the actual volume on the left side of the listener:

[0055] VLOs,. - volume of the i-th triangular pyramid in the space on the left side of the listener [m3],

[0056] VL - approximate volume of the space on the left side of the listener [m3],

[0057] I - number of triangular pyramids into which the space was divided, i - the i-th pyramid in the space.

[0058] Measurement errors:

[0059] It should be noted that the volume VL is an approximate value of the actual volume of space on the left side of the listener, because:

[0060] (1) its value is subject to error resulting from measurement accuracy,

[0061] (2) the placement of LIDAR sensors on the headphones makes it impossible to measure the space located in front of and behind the listener with a width equal to the distance between the LIDAR sensors of the left and right headphones.

[0062] The measurement accuracy problem described in (1) arises from the fact that LIDAR laser beams define pyramids in space whose base, determined by vertices A, B, and D, is by definition flat and does not fully reflect the actual curvature of the measured surface, but is only its approximation, which directly affects the accuracy of the VL calculation.

[0063] This error can be minimized by aiming to divide the space into the smallest possible adjacent pyramids. Then their bases will better represent the measured surfaces. However, such action will generate a larger number of LIDAR laser beam emissions and will negatively impact the energy consumption required for headphone operation. Therefore, during implementation of the invention, the number of emitted LIDAR laser beams to determine volume will be selected experimentally to achieve satisfactory VL calculation results while saving energy.

[0064] The problem of the impossibility to measure space described in (2) arises from the fact that the vectors defined by LIDAR laser beams from the left (right) headphone must be directed at most toward the ideal left (right) hemisphere of the listener with the center at the LIDAR sensor point of the left (right) headphone.

[0065] More formally, for an assumed space with coordinates X, Y, Z, if we assume that:

[0066] 1. the center of the listener's head with headphones is at the coordinates 0,0,0

[0067] 2. the LIDAR of the left headphone is at the coordinates -1,0,0

[0068] 3. the LIDAR of the right headphone is at the coordinates 1,0,0

[0069] 4. the LIDAR sensors measure disjoint spaces, meaning they have no common parts,

[0070] Then these assumptions imply the following conditions:

[0071] 1. the left headphone sensor measures at most the space X,Y,Z where X < -1

[0072] 2. the right headphone sensor measures at most the space X,Y,Z where X > 1

[0073] 3. no sensor measures the space X,Y,Z where X is in the range (-1,1)

[0074] The measurement ranges of the space for each sensor cannot be larger, because in such a case it would be impossible to guarantee that the measured spaces by the LIDAR sensors would be disjoint (the laser beams from both sensors would intersect in space). Therefore, in the actual implementation of the invention, the measurement errors described above are disregarded, and the adopted approximate value of VL is sufficient to achieve the goal.

[0075] The same calculations are performed for the volume VP for the right side of the listener. The final volume V is calculated as the sum:

[0076] V = VL+ VP

[0077] V - approximate volume of the space around the listener [m3],

[0078] VL - approximate volume of the space on the left side of the listener [m3], VP - approximate volume of the space on the right side of the listener [m3]. Waviness of the space (using the definition of "surface waviness") - based on the distances to objects and their distribution in space, we determine: a a_.v,g„ (average sound absorption coefficient [real number in the range [0,1]]). This coefficient is determined experimentally based on different types of surface waviness and the distribution of objects in space. For example, if the space surrounding the listener is a uniform solid, e.g. composed of bare walls, then the waviness coefficient will be low. However, if the space contains objects making it an uneven solid, e.g. furniture (as shown by highly irregular distance measurements), then the waviness coefficient will be high. Based on the waviness of the space, the digital audio processor adopts appropriate average sound absorption coefficients. Acoustic absorption of the room - The acoustic absorption of the room A will be calculated using the formula:

[0079] A - acoustic absorption of the room [m2sabins], a a_.v,g„ - average sound absorption coefficient [real number in the range [0,1]],

[0080] S - approximate surface area of the spatial boundaries of the room, including the floor [m2].

[0081] To determine the surface area S of the spatial boundaries of the room, the method of dividing the space into adjacent triangular pyramids used in the volume calculation in point 1 will be applied. The area of the triangle between vertices A, B, and D will be calculated using the formula: a, b, c - vectors defined by LIDAR laser beams (based on measured distance and laser beam emission angle) from the sensor to each vertex of the pyramid [m], f*ABDi- su rface area°f he triangle spanned by vertices A, B, D for the i-th pyramid in the space [m2].

[0082] All triangle surface areas spanned on vertices A, B, D of the pyramids dividing the space will be summed to obtain the parameter S, according to the formula:

[0083] I - number of triangular pyramids on the left and right side of the listener into which the space is divided [natural number], i - the i-th pyramid in the space [natural number],

[0084] S - approximate surface area of the spatial boundaries of the room, including the floor [m2].

[0085] The calculated surface area S is an approximate value of the actual surface area due to the measurement error described in point 1. For our purposes, this error is disregarded and the approximate value is accepted. Surface area of the space enclosing the room - For each side of the listener, we respectively calculate: PL - approximate surface area of the space enclosing the room on the left side of the listener, PP - approximate surface area of the space enclosing the room on the right side of the listener, and based on these, we determine P - approximate surface area enclosing the room around the listener.

[0086] To calculate the surface area enclosing the space — in other words, the surface area the listener is located on — we will use the calculations obtained in point 1 and point 3. In point 1, the space was divided into triangular pyramids with vertices A, B, C, and D, where vertex C is located at the center of the LIDAR sensor from which the laser beam is emitted. In point 3, the surface areas of triangles spanned between vertices A, B, and D for all pyramids were summed. Let's assume that the set of all such triangles on the left side of the listener will be named ZL and will have cardinality I, and the surface area of the triangle ABD for the i-th triangular pyramid in the space for set ZL will be named PZL^gpj.

[0087] To calculate the surface area enclosing the room, the set ZL should be filtered to form a subset XL, which contains only those triangles whose vertices A, B, and D are located at the level of the listener's ankles or lower — that is, the vectors a, b, and c have a vertical component greater than the listener's height. (The listener's height W will be set by the user in the headphone configuration menu.) Let's assume that this subset of triangles satisfying the above condition will be named XL, with cardinality K, and the surface area of the triangle ABD for the k-th triangular pyramid in space from the set XL will be named PXL^gp^. Thus, the dependencies are:

[0088] XL^ZL

[0089] XL - set of triangles whose vertices A, B, and D are at ankle level or below, i.e. vectors a, b, c are longer than the height of the listener on the left side (from the ankles downward),

[0090] ZL - set of triangles formed by vertices A, B, and D of all triangular pyramids into which the left-side space is divided, as well as: k<K oraz i<I oraz K<I

[0091] K - number of triangles in set XL,

[0092] I - number of triangles in set ZL, k - k-th triangle in set XL, i - i-th triangle in set ZL. Therefore, the desired parameter PL will be calculated as:

[0093] - surface area of the k-th triangle belonging to the set XL [m2],

[0094] PL - approximate surface area enclosing the room on the left side of the listener [m2].

[0095] The calculated surface area PL is an approximate value of the actual surface area enclosing the room on the left side of the listener due to the measurement error described in point 1. For our purposes, this error is disregarded and the approximate value is accepted.

[0096] The same calculations are performed to determine the surface area enclosing the room on the right side of the listener.

[0097] Finally, the total surface area enclosing the room around the listener is:

[0098] P = PL + PP

[0099] P - approximate surface area enclosing the room around the listener [m2]. Average distance from all objects and surfaces - The approximate value of the average distance to all objectsawill be calculated using the formula:

[0100] DL^ - distance measured by the z-th laser beam on the left side of the listener [m],

[0101] DP^ - distance measured by the z-th laser beam on the right side of the listener

[0102] [m], ZL - number of laser beams emitted on the left side of the listener [natural number],

[0103] ZP - number of laser beams emitted on the right side of the listener [natural number],

[0104] D - approximate average distance from all objects [m], avg

[0105] The calculated average distance from all objects is an approximate value of the actual average distance due to the measurement error described in point 1. For our purposes, this error is disregarded and the approximate value is accepted. Echo delay - Based on the average distance to all objects calculated in point 5, and knowing the average speed of sound in air, the echo delay is calculated:

[0106] DE - echo delay, expressed as a number of delay samples for a given sampling frequency fs [unitless number],

[0107] V , - average speed of sound at 15°C = 340.3 [m / sec], sound f - sampling frequency [Hz] = [s-1], s

[0108] D - average distance from all objects [m], avg Maximum distance from all objects and surfaces - The value D is calculated max based on the following formula: Dj. - distance measured by the i-th laser beam [m],

[0109] Z - number of laser beams emitted on the left and right sides of the listener [natural number],

[0110] D - maximum distance from all objects to the listener [m], max

[0111] In step c) of the method according to the invention: „the digital audio processor (2), based on the determined spatial parameters, calculates coefficients for computing reverb, echo, and the direction of the original audio stream source (11)"

[0112] 1. Reverberation time

[0113] Based on the spatial parameters calculated by the digital audio processor as described in step b) in points 1), 3), and 4), the reverberation time of the room can be determined using several possible formulas, namely::

[0114] 1. Sabine's formula for non-dampened rooms (with low acoustic absorption), i.e. AL (from point 3.) < 0.2, with evenly distributed absorption and long reverberation time,

[0115] 2. For highly dampened rooms, i.e. AL > 0.2, with evenly distributed absorption and short reverberation time,

[0116] 3. For rooms with a volume greater than 1000 m3,

[0117] 4. Norris-Eyring formula,

[0118] 5. Millington-Sette formula,

[0119] 6. Knudsen formula.

[0120] For our room from Fig. 4, we assume the first one: Sabine's formula for nondampened rooms (with low acoustic absorption), i.e. A < 0.2, that is:

[0121] T - reverberation time of the room [s],

[0122] V - volume of space [m3],

[0123] A - acoustic absorption of the room [m2sabins = m2],

[0124] SWS = 0.161 s / m - constant derived from Wallace Sabine's experiments (assuming standard air at ~20°C and 50% humidity) Echo delay

[0125] Based on the spatial parameters calculated by the digital audio processor according to step b) points 2) and 6), a digital FIR echo filter with a single delay line as shown in Fig. 4b may be used: y[n\ = x[n\ + a* x[n — DE] y[n] - output signal with applied echo, x[n] - input signal before echo is added, a - echo attenuation coefficient [real number in the range [0,1]],

[0126] DE - echo delay expressed as a number of samples for the given sampling frequency fs [unitless number],

[0127] The attenuation coefficient a for the above equation can be calculated based on the average sound absorption coefficient a_avg using the formula: a - echo attenuation coefficient [real number in the range [0,1]],

[0128] - average sound absorption coefficient [real number in the range [0,1]], avg

[0129] N - average number of reflections before decay [natural number].

[0130] The average number of reflections N before decay for the above formula can be estimated using the data described in step b) in points 1) and 3), using the formula:

[0131] N - average number of reflections before decay [natural number],

[0132] L- average free path length of the sound wave before the next reflection [m],

[0133] A - acoustic absorption of the room [m2sabins],

[0134] V - volume of space [m3].

[0135] The approximate average free path L of the sound wave before the next reflection for the above formula is determined using data from step b) points 1) and 3), based on the empirical formula used in acoustics:

[0136] S - surface area of the spatial boundaries of the room including the floor [m2],

[0137] V - volume of space on the left side of the listener [m3],

[0138] L- approximate average free path of the sound wave before the next reflection [m], 3. Source direction vector

[0139] Based on the spatial parameters calculated by the digital audio processor described in step b), particularly in point 7), the direction of sound arrival is determined by the longest calculated vector (based on the maximum distance from all objects and surfaces).

[0140] In step d) of the method according to the invention: „the digital audio processor (2), based on the calculated coefficients, modifies the parameters of the original audio stream (11) by adding reverb, echo, and source direction (10)"

[0141] Based on the formulas described in step c) and the usage example shown in Fig. 4, we calculate:

[0142] 1. Reverberation time:

[0143] Room volume: V = (5m -0.1m) * 6m * 2.9m = 85.26m-

[0144] Approximate surface area of the spatial boundaries of the room, including the floor:

[0145] S = 2 * (6m * 2.9m + 2.9m * (5m - 0.1m) + 6m * (5m -0.1m)) = 2 * (17.4m2+ 14.21m2+ 29.4m2) = 122.02m2

[0146] Average sound absorption coefficient (assumed experimentally): ^avg=0-1

[0147] Acoustic absorption of the room: A = a * S = 0.1 * 122.02m2= 12.202m2ovq

[0148] Using Sabine's formula for non-dampened rooms: 2. Echo delay

[0149] Assumed standard sampling frequency: f = 44100 Hz s_

[0150] Average speed of sound at 15°C: V . = 340,3 m / sec sound - - -

[0151] Average distance from all objects: D ~ 4.17m ova

[0152] The above result cannot be calculated directly because we do not have a physical set of LIDAR distance measurements. Therefore, for the purposes of this exemplary embodiment in the document, we apply a geometric approximation using the 3D formula for the average distance from the center of a rectangular cuboid, based on the room size data. Calculated echo delay as a number of samples at the given sampling frequency:

[0153] DE = (D * f ) / V _.= (4.17 m * 44100 Hz) / 340,3 m / sec = 540 (rounded down) ava s sound

[0154] 3. Source direction vector

[0155] Based on the room size data and assuming: X - room width, , Y - room depth, Z - room height, origin is at the geometric center of the room then, one of the possible source direction vectors selected by the digital audio processor may be the vector pointing to a chosen corner, e.g.: v = ((5m-0.1m) / 2, 6m / 2, 2.9m / 2) = (2.45, 3, 1.45) [ml

[0156] Vector length: D„ = D ~ 4.14m . max

[0157] In step e) of the method according to the invention: „the modified audio stream is emitted through the headphones (1)."

[0158] Steps from a) to d) are executed cyclically every 0.5 seconds. In this way, if the listener moves through space or changes orientation, then after a maximum of 0.3 seconds, the sound properties will be updated based on newly calculated LIDAR distance measurements.

Claims

Patent Claims1. Headphones (1) modifying the original audio stream (11) through a digital audio processor (2) connected to the source (10) of the original audio stream (11), as well as to the left speaker (4) and right speaker (5) of the headphones (1), characterized in that the digital audio processor (2) uses 3D spatial mapping for modification through 3D mapping sensors (3) located in the headphones, in particular the digital audio processor (2) is connected to the left sensor (6) mapping the 3D space on the left side of the listener, placed in the left headphone (7), and the right sensor (8) mapping the 3D space on the right side of the listener, placed in the right headphone (9).

2. The system according to claim 1, wherein the 3D spatial mapping sensor (6)(8) is an optical distance sensor.

3. The system according to claim 1 or 2, wherein the optical distance sensors are laser sensors.

4. The system according to any of claims 1 to 3, wherein the optical distance sensors are LIDAR laser sensors.

5. The system according to claim 1 or 2, wherein the optical sensors are cameras.

6. The system according to claim 1 or 2, wherein the 3D spatial mapping sensor (6)(8) is an ultrasonic distance sensor.

7. The system according to claim 1 or 2, wherein the 3D spatial mapping sensor (6)(8) comprises any combination of LIDAR sensors, cameras, and ultrasonic distance sensors.

8. The system according to claim 1, wherein the source (10) of the original audio stream is connected to the headphones either via a wired or a wireless connection.

9. The system according to claim 8, wherein the source (10) of the original audio stream is connected to the headphones wirelessly via Bluetooth.

10. A method of modifying the original audio stream (11) by means of an digital audio processor (2) adapting to changing spatial parameters of the space surrounding the listener, characterized in that it comprises the steps: a) the digital audio processor (2) receives the original audio stream (11) from the source (10), and information about the space surrounding the listener from the sensors (6)(8), then, b) the digital audio processor (2), based on the spatial information, determines the parameters of the space, then c) the digital audio processor (2), based on the determined spatial parameters, calculates coefficients for computing reverb, echo, and the direction of the original audio stream source (11), then d) the digital audio processor (2), based on the calculated coefficients, modifies the parameters of the original audio stream (11) by adding reverb, echo, and source direction (10), then e) the modified audio stream is emitted through the headphones (1).

11. The method according to claim 10, wherein the spatial information comprises data about the distance to various objects and spaces on the left side of the listener for the left headphone, and to various objects on the right side of the listener for the right headphone.

12. The method according to claim 10, wherein the spatial parameters include: volume, distance to objects, surface area under the listener, acoustic absorption of the room, average sound absorption coefficient, area of spatial boundaries including the floor, average distance from all objects and surfaces, the farthest measured distance to a surface on the left side of the listener for the left headphone, and on the right side of the listener for the right headphone.

13. The method according to claim 10, wherein steps a) to d) are executed cyclically.

Citation Information

Patent Citations

  • Audio system and method of operation therefor

    US20130272527A1

  • Mixed reality spatial audio

    US20190116448A1

  • Extrapolation of acoustic parameters from mapping server

    US20210377690A1