Audio signal distribution of virtual sound sources

By determining the location of speakers and objects in the audio system, calculating distances and modifying the audio signal using the distance attenuation function, the problem that existing audio systems are difficult to provide immersive audio in multi-speaker and mobile object scenes is solved, real-time response to mobile objects and accuracy of tone positioning is achieved.

CN120050590APending Publication Date: 2025-05-27HARMAN INT IND INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411630056.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-11-27
Filing Date
2024-11-15
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

Existing audio systems struggle to provide an immersive audio experience when processing multiple speakers and moving objects, and have limited responsiveness to speakers and object locations.

Method used

By determining the position of the speaker and the object in the listening environment, calculating the distance between the two, and modifying the amplitude of the audio signal using the distance attenuation function, the modified audio signal is generated for transmission to the speaker.

Benefits of technology

Real-time response to mobile objects is achieved, the perceived accuracy of tone and positioning is improved, and a realistic immersive audio experience is provided without the need for a large amount of processing resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120050590A_ABST
    Figure CN120050590A_ABST
Patent Text Reader

Abstract

In various embodiments, a computer-implemented method includes: determining a speaker position of a speaker in a listening environment; tracking an object position of the object within the listening environment; retrieving the audio signal; calculating the distance between the object position and the loudspeaker position; generating a modified audio signal, wherein the amplitude of the modified audio signal is based on (i) the audio signal, (ii) the distance, and (iii) a distance attenuation function; and transmitting the modified audio signal to a speaker.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Various embodiments relate generally to audio output devices and, more particularly, to audio signal distribution of virtual sound sources. Background Art

[0002] Various consumer devices output sound to enhance the user's experience when interacting with the consumer device. For example, various products generate sound to entertain users. In such products, the sound generation circuit stores pre-recorded sound files or generates the sound to be output. When the product receives an input (such as pressing a button), the sound generation circuit loads the pre-recorded sound file or generates the sound and drives the speaker to output the corresponding audio.

[0003] At least one disadvantage of conventional sound-generating devices is that such devices have limited ability to play immersive sound. For example, some devices use low-power microcontrollers or storage systems to minimize cost; however, the limited storage and processing capabilities of such systems restrict the devices to using speakers with a limited set of parameters to reproduce sound. As a result, the sound produced by conventional sound-generating devices has difficulty reproducing the timbre of pre-recorded or generated sounds. In addition, many devices cannot produce sound or lack the ability to be updated to output new or different sounds.

[0004] In response to the above limitations of such devices, it is often desirable to output the audio associated with the device through a separate sound system (such as a speaker group). The speakers are typically located at certain locations within the physical space. For example, a given room includes a sound bar and additional satellite speakers located near the sound bar. In another example, the room may include speakers organized as a home theater, wherein the center speaker is located near the center of the front wall of the room, and the front left speaker, front right speaker, rear left speaker, and rear right speaker are located in corresponding corners of the room. The audio playback device transmits a signal to each speaker so that a listener in the physical space hears the combined output of all speakers, thereby hearing the sound associated with the sound-emitting device.

[0005] At least one disadvantage of conventional sound systems is that such systems are not responsive to the position or movement of the device within the listening environment associated with the sound being produced. For example, the speakers within a given sound system generate a sound field that includes one or more sweet spots corresponding to the target position of the listener in the listening environment. The sweet spots are typically tuned in the sound field to produce ideal sound quality. However, because the sound system does not take into account the position of the speakers and the device associated with the sound. As a result, the listener may perceive the apparent position of the sound produced by the sound system to be different from the device that "generates" the sound. The inconsistency between the actual position of the device and the apparent position of the sound produced reduces the immersive experience experienced by the listener.

[0006] As previously indicated, there is a need in the art for more efficient techniques for providing audio to multiple speakers from varying locations. Summary of the invention

[0007] In various embodiments, a computer-implemented method includes: determining a speaker position of a speaker in a listening environment; tracking an object position of an object within the listening environment; retrieving an audio signal; calculating a distance between the object position and the speaker position; generating a modified audio signal, wherein an amplitude of the modified audio signal is based on (i) the audio signal, (ii) the distance, and (iii) a distance attenuation function; and transmitting the modified audio signal to the speaker.

[0008] Additional embodiments provide, among other things, non-transitory computer-readable storage media storing instructions for implementing the above methods, as well as interactive toys, devices, and systems configured to implement the above methods.

[0009] At least one technical advantage of the disclosed technology relative to the prior art is that using the disclosed technology, a sound system can distribute audio signals to one or more speakers in a physical listening area in a manner that indicates the apparent position of an object within a listening environment and improves the perceived accuracy in terms of timbre and positioning. Specifically, by determining the position of an object within a listening environment and attenuating the sound signal based on the distance between the object and the corresponding one or more speakers, the sound system provides perceptually accurate sound to the tracked object in real time, thereby effectively providing realistic immersive audio in the listening environment that responds to the movement of the object within the listening environment without requiring large and expensive processing resources. In addition, by using techniques that operate with various numbers of speakers, a sound system using the disclosed technology can provide perceptually accurate audio that responds to different numbers of speakers located at various locations within the listening environment. These technical advantages provide one or more technical improvements over prior art methods. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] In order that the manner in which the above-mentioned features of various embodiments are used can be understood in detail, a more particular description of the inventive concept briefly outlined above can be obtained by reference to various embodiments, some of which are shown in the accompanying drawings. However, it should be noted that the drawings only show typical embodiments of the inventive concept and are therefore not to be considered as limiting the scope in any way, and that there are other equally effective embodiments.

[0011] Figure 1 is a schematic diagram illustrating an audio processing system according to various embodiments;

[0012] Figure 2 According to various embodiments, Figure 1an example physical listening environment and a corresponding virtual listening environment modeled by an audio processing system; and

[0013] Figure 3 A flow chart illustrating method steps for generating an audio signal based on the position of a tracked object according to various embodiments is set forth. DETAILED DESCRIPTION

[0014] In the following description, numerous specific details are set forth to provide a more thorough understanding of various embodiments. However, it will be apparent to one skilled in the art that the concepts of the present invention may be practiced without one or more of these specific details.

[0015] Figure 1 1 is a schematic diagram illustrating an audio processing system 100 according to various embodiments. As shown, the audio processing system 100 includes, but is not limited to, a computing device 110, one or more sensors 150, and one or more speakers 160. The computing device 110 includes, but is not limited to, a processing unit 112 and a memory 114. The memory 114 includes, but is not limited to, an audio processing application 120, a virtual environment 130, one or more audio signals 140, and one or more sound profiles 182. The virtual environment 130 includes, but is not limited to, a virtual object 132, one or more virtual microphones 134, and position data 136. The object 180 includes, but is not limited to, a sound profile 182(1) and an audio signal 140(1).

[0016] The audio processing system 100 can be implemented in various forms, such as an interactive device including a processor and local memory, a personal computer, and the like. For example, the audio processing system 100 can be incorporated into one or more interactive toys (e.g., a bird toy including a voice box). Additionally or alternatively, in some embodiments, the audio processing system 100 can be incorporated into other types of non-toy consumer devices. The audio processing system 100 can use a dedicated processing device and / or a separate computing device (such as a user's mobile computing device or a cloud computing system) to perform processing functions. The audio processing system 100 can use any number of various types of sensors to detect various environmental values, which can be attached to other system components, integrated with other system components, or provided separately.

[0017] The computing device 110 is a device that generates audio signals to drive one or more speakers 160 to partially generate a sound field. In various embodiments, the computing device 110 transmits a set of modified audio signals to the set of speakers 160 in the audio processing system 100. In various embodiments, the computing device 110 can be a central unit in a home theater system, a sound bar, and / or another device that communicates with one or more speakers 160. The computing device 110 is included in one or more devices, such as consumer products (e.g., interactive toys, portable speakers, gaming devices, other products, etc.), smart home devices (e.g., smart lighting systems, security systems, digital assistants, etc.), communication systems (e.g., teleconferencing systems, video conferencing systems, speaker amplification systems, etc.), and the like. In various embodiments, the computing device 110 is located in various environments, including but not limited to indoor environments (e.g., living rooms, conference rooms, meeting rooms, home offices, etc.) and / or outdoor environments (e.g., terraces, rooftops, gardens, etc.). In some embodiments, the computing device 110 is a low-power, limited processing, and / or limited memory device that implements lightweight processing of incoming data. For example, computing device 110 may be a Raspberry Pi (e.g., Pi Pi Pi or Pi ), which includes a processor (such as a digital signal processor), a memory (e.g., 1MB to 4MB RAM), and a storage device (e.g., a flash memory card). For example, computing device 110 may be a development board such as 4.0 microcontroller development board, or any other board containing a processor used as a digital signal processor, such as Cortex M4, or other lightweight computing devices.

[0018] Processing unit 112 may be any suitable processor, such as a central processing unit (CPU), a graphics processing unit (GPU), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), a digital signal processor (DSP), a multi-core processor, and / or any other type of processing unit or a combination of two or more of the same type and / or different types of processing units, such as a system on a chip (SoC) or a CPU configured to operate in conjunction with a GPU. In general, processing unit 112 may be any technically feasible hardware unit capable of processing data and / or executing software applications.

[0019] The memory 114 may include a random access memory (RAM) module, a flash memory unit, or any other type of memory unit or combination thereof. The processing unit 112 is configured to read data from and write data to the memory 114. In various embodiments, the memory 114 includes a non-volatile memory, such as an optical drive, a magnetic drive, a flash drive, or other storage device. In some embodiments, a separate data storage device (such as an external device including a network ("cloud storage device")) supplements the memory 114. The audio processing application 120 within the memory 114 can be executed by the processing unit 112 to implement the overall functionality of the computing device 110, including the audio processing application 120 and / or running simulations and solvers associated with the virtual environment 130, and thus coordinate the operation of the entire computing device 110. In various embodiments, an interconnect bus (not shown) connects the processing unit 112, the memory 114, and any other components of the computing device 110.

[0020] The audio processing application 120 determines the relative distance of the object 180 from one or more speakers 160 and generates audio signals for the set of speakers 160 for reproduction. The audio processing application 120 generates the audio signals by first determining the relative position (e.g., location and / or orientation) of the object 180 and the set of speakers 160 and calculating the distance between the object and the set of speakers 160. The audio processing application 120 uses the corresponding calculated distances to generate a set of modified audio signals, the set of modified audio signals being adjusted at least according to the calculated distances. The audio processing application 120 then transmits the set of modified audio signals to the set of speakers 160 for reproduction.

[0021] In various embodiments, the audio processing application 120 determines the current position of each speaker in the set of speakers 160 within the physical listening environment. Additionally or alternatively, the audio processing application 120 tracks the movement of one or more speakers in the set of speakers 160. For example, the audio processing application 120 receives sensor data from one or more sensors 150 (e.g., tracking data for a given speaker as a series of optical data, and / or a series of auditory data received in response to a test signal generated by the computing device 110). In such instances, the sensor data indicates the position and / or orientation of each speaker 160 at a given time. In some embodiments, the sensor data indicates that at least one speaker in the set of speakers 160 (e.g., speaker 160(4)) is moving. The audio processing application processes the sensor data to determine the respective positions of the set of speakers 160, and causes the computing device 110 to store the determined positions in a common coordinate system as part of the position data 136. In some embodiments, the audio processing application 120 receives sensor data generated by one or more sensors on the speaker 160(4). For example, the speaker 160 may include one or more sensors (not shown), such as a position sensor and / or an IMU that acquires various sensor data (e.g., acceleration measurements, magnetic field measurements, angular velocity, etc.). In such instances, the speaker acquires sensor data while moving and transmits a series of messages containing the acquired sensor data. In such instances, the audio processing application 120 receives and aggregates the sensor data included in the messages and determines the trajectory and / or current position of the speaker 160.

[0022] In various embodiments, the audio processing application 120 tracks the position of the object 180. In some embodiments, the object 180 is a physical object within the physical listening environment 210. In such instances, the audio processing application 120 determines the position of the physical object 180 within the physical listening environment and / or tracks the trajectory of the physical object 180 moving within the physical listening environment. In one example, the audio processing application 120 obtains sensor data from one or more sensors 150 and determines the current position of the physical object 180 by processing the obtained sensor data. Then, the computing device 110 stores each determined position as part of the position data 136 in a common coordinate system. Alternatively, in some embodiments, the object 180 is a virtual object 132 within the virtual environment 130. In such instances, the audio processing application 120 tracks the virtual object 132 within the virtual environment 130 based on the position data 136 generated for the virtual object 132. For example, the audio processing application 120 and / or a separate application (not shown) generates a virtual environment 130 including the virtual object 132. In such instances, the application managing the virtual environment 130 generates position data for the virtual object 132. In such instances, the audio processing application 120 tracks the virtual object 132 as it moves in the virtual environment 130 by retrieving portions of the position data 136 corresponding to the virtual object 132.

[0023] In various embodiments, the audio processing application 120 calculates a set of distances from the object 180 to the set of speakers 160. In some embodiments, the audio processing application 120 calculates a physical distance (e.g., a Euclidean distance) from the object 180 to each speaker 160(1) to 160(5) in the physical listening environment to generate the calculated distances. Additionally or alternatively, in some embodiments, the audio processing application 120 calculates a distance from the virtual object 132 to a set of virtual microphones 134 within the virtual environment 130 to generate the calculated distances. For example, the virtual environment 130 includes a set of virtual microphones 134 located at locations within the virtual environment 130 that correspond to locations of the set of speakers 160 within the physical listening environment. The audio processing application 120 calculates a distance between the virtual object 132 and the virtual microphone 134 that represents the Euclidean distance between the object 180 and the set of speakers 160 within the physical listening environment.

[0024] In various embodiments, the audio processing application 120 generates audio signals for the speakers 160 based on a set of calculated distances and one or more distance decay functions. In various embodiments, the audio processing application 120 uses one or more distance decay functions to modify the amplitude and / or phase of the input audio signal 140(1) to generate audio signals for the set of speakers based on the corresponding calculated distances between the corresponding speakers 160 and the object 180. Additionally or alternatively, the audio processing application 120 uses other functions to modify the input audio signal 140(1) based on the orientation of the corresponding speakers 160 relative to the object 180. In some embodiments, the audio processing application 120 also applies various panning techniques to the input audio signal 140(1) to simulate the sound of the object 180 as the object 180 moves along the trajectory 222.

[0025] In some embodiments, the physical listening environment does not include physical objects. In such instances, the audio processing application 120 tracks the trajectory of the virtual object 132 within the virtual environment 130 based on the position data 136 generated for the virtual object 132. In such instances, the position of the virtual object 132 within the virtual environment 130 represents the position of the physical object (e.g., object 180) within the physical listening environment. For example, the audio processing application 120 and / or a separate application (not shown) generates a virtual object 132 (e.g., a virtual ball in an AR game) within the virtual environment 130. In such instances, the application that manages the virtual environment 130 also generates position data for the virtual object 132. In such instances, the audio processing application 120 uses the position data corresponding to the virtual object 132 to track the virtual object 132 and calculates the distance between the position of the virtual object 132 and the position of the virtual microphone 134.

[0026] In various embodiments, the audio processing application 120 selects the distance decay function from a set of candidate distance decay functions. For example, the computing device 110 may store a set of candidate distance decay functions, such as a linear function, a linear square function, or an inverse function, that attenuates the gain or changes the phase of the input audio signal 140 (1) based on the distance between the object 180 and the given speaker 160. In such instances, the audio processing application 120 uses the selected distance decay function to modify the amplitude and / or phase of the input audio signal for the given speaker 160 based on the calculated distance between the given speaker 160 and the object 180. In some embodiments, the distance decay function attenuates the audio signal between a minimum distance (Dmin) and a maximum distance (Dmax). In such instances, the audio processing application 120 compares the calculated distance to a minimum distance threshold (Dmin) based on the minimum distance (e.g., zero or some other minimum distance) and / or a maximum distance threshold based on the maximum distance (Dmax). When the audio processing application 120 determines that the calculated distance satisfies the minimum threshold and / or the maximum threshold, the audio processing application 120 applies the selected distance decay function.

[0027] In one example, the audio processing application 120 uses a linear function that attenuates the amplitude of a given signal based on the distance (D) between the object 180 and the speaker 160 (or the virtual object 132 and the virtual microphone 134). The linear function can also modify the amplitude outside the minimum threshold and the maximum threshold. For example, Formula 1 calculates the amplitude based on the calculated value of the distance 262 compared to the minimum threshold and the maximum threshold:

[0028]

[0029] In another example, the audio processing application 120 uses a linear square function that attenuates the amplitude of a given signal according to the square of the distance (D), where the amplitude decays as the distance increases. The linear square function can also modify the amplitude outside of the minimum threshold and the maximum threshold. For example, piecemeal formula 2 calculates the amplitude based on the calculated value of the distance 262 compared to the minimum threshold and the maximum threshold:

[0030]

[0031] In a further example, the audio processing application 120 uses an inverse function, i.e., the amplitude of a given signal is a function of the inverse of the distance (D), which decays as the distance increases. The inverse function can also modify the amplitude outside of the minimum and maximum thresholds. For example, piecemeal formula 3 calculates the amplitude based on the calculated value of the distance 262 compared to the minimum and maximum thresholds:

[0032]

[0033] In some embodiments, in addition to the inverse function, the audio processing application 120 also uses a taper that gradually attenuates the amplitude of a given signal when the distance exceeds a maximum threshold to the taper point (T). For example, piecemeal formula 4 is based on the calculated value of distance 262 and the minimum and maximum thresholds and the taper point ( For example , 4*(D max -D min )) to calculate the amplitude:

[0034]

[0035] Additionally or alternatively, in some embodiments, the audio processing application 120 uses a distance decay function that further modifies the amplitude and / or phase based on the calculated orientation difference. In such instances, the calculated amplitude is a function of the distance (D) and one or more angles representing the orientation difference, as shown in Formula 5:

[0036]

[0037] In various embodiments, the audio processing application 120 drives the computing device 110 to transmit the set of modified audio signals to the set of speakers 160. In some embodiments, each respective speaker in the set of speakers 160 receives one of the modified audio signals from the computing device 110 via a wire, a wireless stream, or via a network. Upon receiving the modified audio signal, each speaker 160 in the set of speakers 160 reproduces the corresponding modified audio signal to generate sound waves within the physical listening environment. In various embodiments, the sound waves generated by the set of speakers 160 combine to generate a sound field that provides a perceptually accurate location of the physical object 180 within the physical listening environment.

[0038] The virtual environment 130 is a computer model that simulates the operation and physical quantities within the virtual acoustic environment and the operation of one or more virtual devices in the virtual acoustic environment. In some embodiments, the application (e.g., audio processing application 120, a separate application, etc.) that manages the virtual environment 130 is trained using data that simulates measurement data recorded in the test acoustic environment.

[0039] The virtual object 132 and the one or more virtual microphones 134 represent objects and / or devices within the physical listening environment. For example, the virtual object 132 represents the object 180, and the one or more virtual microphones 134 represent the one or more speakers 160. In various embodiments, the audio processor uses the virtual object 132 and / or the virtual microphones 134 to calculate the distance used to generate the set of modified audio signals. For example, the audio processing application 120 may initially determine the positions of the speakers 160(1)-160(5) within the physical listening environment 210. The audio processing application 120 may use the reciprocity principle of sound to exchange the positions of the audio transmitter and the audio receiver within the virtual environment 130. In such an example, the audio processing application 120 places a set of virtual microphones 134(1)-134(5) at positions within the virtual environment 130 that correspond to the positions of the speakers 160(1)-160(5) within the physical listening environment 210. Additionally or alternatively, the audio processing application 120 simulates the physical object 180 as a virtual object 132 that acts as an audio emitter. In such an example, the audio processing application 120 places the virtual object 132 at a location within the virtual environment 130 that corresponds to the location of the physical object 180 within the physical listening environment 210. When the audio processing application 120 calculates the distances between the virtual object 132 and the virtual microphones 134(1)-134(5), the calculated distances 262(1)-262(5) within the virtual environment 130 correspond to the calculated distances 262(1)-262(5) between the physical object 180 and the speakers 160(1)-160(5) within the physical listening environment 210.

[0040] In various embodiments, the position of the virtual object 132 and / or the set of virtual microphones 134 is represented in the form of a combination of position and orientation. For example, the position data 136 for a given virtual microphone 134 includes the coordinates of the position of the virtual microphone within the virtual environment 130, as well as orientation information, such as a set of angles relative to a normal orientation within the virtual environment 130 (e.g., ).

[0041] The memory 114 stores one or more audio signals 140 and one or more sound profiles 182. For example, the computing device 110 receives a sound profile 182(1) including an input audio signal 140(1) from an object 180 and stores the sound profile 182(1) in the memory 114. In some embodiments, the audio processing application 120 receives the input audio signal 140(1) and the sound profile 182(1) separately. Additionally or alternatively, in some embodiments, the computing device 110 stores one or more sound profiles 182 and / or one or more audio signals 140 associated with multiple objects. In such instances, the audio processing application 120 identifies the object 180, identifies the sound profile 182(1) corresponding to the object 180, and retrieves the audio signal 140(1) associated with the sound profile 182(1).

[0042] The object 180 is a physical object within the physical listening environment or an object representing the location of the virtual object 132 within the physical listening environment. In various embodiments, the audio processing application 120 tracks the object 180 in the physical listening environment and generates a set of audio signals associated with the object 180. For example, the object 180 can be an interactive toy (e.g., an ambulance) that stores an audio signal (e.g., a siren). In such an example, the audio processing application 120 tracks the current location of the interactive toy within the physical listening environment and generates a set of audio signals for the set of speakers 160 to reproduce the sounds of the interactive toy.

[0043] In some embodiments, object 180 includes a set of tracking sensors (not shown) that can be used to determine the position and / or movement of object 180 within the physical listening environment. For example, object 180 may include various types of tracking sensors that acquire sensor data, such as optical sensors, position sensors, IMUs, audio sensors, and the like. In such instances, the object sends the sensor data in one or more messages to audio processing application 120 for processing to determine the position of object 180. Additionally or alternatively, object 180 stores sound profile 182(1) and / or audio signal 140(1). In various embodiments, computing device 110 acquires sound profile 182(1) and / or audio signal 140(1) from object 180 and stores sound profile 182(1) and / or audio signal 140(1). In such instances, audio processing application 120 may then identify object 180 and retrieve sound profile 182(1) and / or audio signal 140(1) to generate the set of audio signals for reproduction by the set of speakers 160.

[0044] One or more sensors 150 include various types of sensors that acquire sensor data from a physical listening environment. For example, the sensor 150 may include an auditory sensor, such as a microphone, to receive various types of sounds (e.g., subsonic pulses, ultrasonic waves, voice commands, etc.). In some embodiments, the sensor 150 includes an optical sensor (such as an RGB camera, a time-of-flight camera, an infrared camera, a depth camera, a quick response (QR) code tracking system, a potentiometer, a proximity or presence sensor), a motion sensor (such as an accelerometer or an inertial measurement unit (IMU) (e.g., a three-axis accelerometer, a gyroscope sensor, and / or a magnetometer)), a pressure sensor, and the like. In addition, in some embodiments, the sensor 150 may include a wireless sensor (including a radio frequency (RF) sensor (e.g., sonar and radar)) and / or a wireless communication protocol (including Bluetooth, Bluetooth Low Energy (BLE), a cellular protocol, and / or near field communication (NFC)).

[0045] One or more speakers 160 each provide sound output by reproducing the corresponding received audio signal. For example, one or more speakers 160 can be components of a wired or wireless speaker system, or any other device that generates sound output. In various embodiments, two or more speakers 160 can be incorporated into a speaker array and / or a single device (e.g., arranged in a body of a form factor including multiple speakers) and share a common location. In various embodiments, one or more speakers are implemented using any number of different conventional form factors, such as a single consumer product, a discrete speaker device, a personal speaker, a body-worn (head-worn, shoulder-worn, hand-worn, etc.) speaker device, etc. In some embodiments, one or more speakers 160 may be connected to an output device that provides other forms of output in addition, such as a display device that provides visual output.

[0046] Figure 2 An example physical listening environment 210 and a Figure 1 The corresponding virtual listening environment 130 modeled by the audio processing system 100 of the embodiment of the present invention. As shown in the figure, the physical listening environment 210 includes but is not limited to a set of speakers 160 and a physical object 180. The virtual environment 130 includes but is not limited to a set of virtual microphones 134 and a virtual object 132.

[0047] In operation, the audio processing application 120 determines the locations of the speakers 160(1)-160(5) in the physical listening environment 210. The audio processing application 120 uses the determined locations to calculate the distances 262(1)-262(5) between the physical object 180 and the speakers 160(1)-160(5). The audio processing application 120 uses the set of calculated distances 262 to modify a set of audio signals reproduced by the speakers 160(1)-160(5) within the physical listening environment 210. When generating an audio signal for a given speaker (e.g., speaker 160(2)), the audio processing application 120 modifies the input audio signal 140(1) using a distance decay function that modifies the amplitude and / or phase of the input audio signal 140(1) based on the calculated distances 262(2). In this way, the audio processing application 120 drives the speakers 160(1)-160(2) to generate a sound field including perceptually accurate sounds of the physical object 180 in real time, which responds to the movement of the physical object 180 without requiring large and expensive processing resources.

[0048] The physical listening environment 210 is a portion of a real-world environment that includes one or more speakers 160 for reproducing audio signals heard by a listener. In various embodiments, the physical listening environment 210 may include various numbers of speakers 160. In such instances, the audio processing application 120 tracks each of the speakers 160 in the physical listening environment 210 and distributes audio signals to each of the speakers 160.

[0049] In various embodiments, the physical listening environment 210 includes at least one physical object 180. In such instances, the audio processing application 120 tracks the physical object 180 and generates a set of audio signals associated with the physical object 180, wherein the speaker 160 reproduces the set of audio signals to generate a sound field including sounds corresponding to the physical object 180. For example, the physical object 180 can be an interactive toy (e.g., an ambulance) that stores audio signals (e.g., a siren). In such instances, the audio processing application 120 tracks the current position of the interactive toy within the physical listening environment 210 and generates a set of audio signals for the speaker 160 to reproduce. The speaker 160 reproduces the set of audio signals to generate a sound field that provides audio signals to the interactive toy in a manner that provides a perceptually accurate representation of the interactive toy's position within the physical listening environment 210 (e.g., accurate timbre, positioning, etc.).

[0050] In various embodiments, the audio processing application 120 tracks the movement of one or more of the speakers 160(1)-160(5) within the physical listening environment 210. In such instances, the audio processing application 120 receives sensor data from one or more sensors 150 indicating the position and / or orientation of each speaker 160(1)-160(5) at a given time. In some embodiments, the sensor data indicates that at least one (e.g., speaker 160(2)) is moving. In one example, the audio processing application 120 acquires the sensor data in the form of tracking data that includes a series of optical data acquired by an optical sensor and / or a series of auditory data received by one or more microphones in response to a test signal generated by the computing device 110. The audio processing application 120 processes the tracking data to determine a current position of each speaker 160(1)-160(5), wherein the position includes a position and an orientation. For example, the position of speaker 160(2) includes coordinates of the position of speaker 160(2) within physical listening environment 210, as well as orientation information, such as a set of angles relative to a normal orientation within physical listening environment 210 (e.g., ). Additionally or alternatively, in some embodiments, audio processing application 120 receives sensor data (e.g., acceleration measurements, magnetic field measurements, angular velocity, etc.) generated by a position sensor and / or IMU on speaker 160(2). For example, speaker 160(2) transmits a series of messages containing sensor data as it moves. In such instances, audio processing application 120 receives and aggregates the sensor data included in the messages and determines a trajectory and / or current position of speaker 160(2).

[0051] In various embodiments, the audio processing application 120 tracks the trajectory 222 of the physical object 180 within the physical listening environment 210. For example, the audio processing application 120 processes sensor data received from one or more sensors 150 to detect the presence of the physical object 180 within the physical listening environment 210. In such instances, the audio processing application 120 determines the current position of the physical object 180 and / or tracks the trajectory 222 of the physical object 180 within the physical listening environment 210. The computing device 110 then stores each determined position as part of the position data 136 in the form of a combination of position and orientation.

[0052] In various embodiments, the audio processing application 120 uses the virtual environment 130 to track the positions of the speakers 160(1)-160(5) and / or the physical objects 180. In some embodiments, the audio processing application 120 generates the virtual environment 130 as a virtual simulation of the physical listening environment 210. Alternatively, in some embodiments, a separate application (not shown) such as an augmented reality (AR), virtual reality (VR), and / or extended reality (XR) application generates the virtual environment 130. In such instances, the audio processing application 120 uses the virtual environment 130 to calculate the distance 262 within the virtual environment 130 and uses the calculated distance 262 when generating audio signals for the speakers 160(1)-160(5).

[0053] For example, the audio processing application 120 may initially determine the locations of the speakers 160(1)-160(5) within the physical listening environment 210. The audio processing application 120 may use the reciprocity principle of sound to exchange the locations of the audio transmitters and audio receivers within the virtual environment 130. In such an example, the audio processing application 120 places a set of virtual microphones 134(1)-134(5) within the virtual environment 130 at locations corresponding to the locations of the speakers 160(1)-160(5) within the physical listening environment 210. Additionally or alternatively, the audio processing application 120 simulates the physical object 180 as a virtual object 132 that acts as an audio transmitter. In such an example, the audio processing application 120 places the virtual object 132 at a location within the virtual environment 130 that corresponds to the location of the physical object 180 within the physical listening environment 210. When the audio processing application 120 calculates the distances between the virtual objects 132 and the virtual microphones 134(1) to 134(5), the calculated distances 262(1) to 262(5) within the virtual environment 130 correspond to the calculated distances 262(1) to 262(5) between the physical objects 180 and the speakers 160(1) to 160(5) within the physical listening environment 210.

[0054] In various embodiments, the audio processing application 120 generates an audio signal for the speaker 160 based on the calculated distances 262(1)-262(5) and one or more distance decay functions. In various embodiments, the audio processing application 120 uses one or more distance decay functions to modify the amplitude and / or phase of the input audio signal 140(1) to generate an audio signal for each speaker 160(1)-160(5) based on the corresponding calculated distances 262(1)-262(5) between the speaker 160 and the tracked object. Additionally or alternatively, the audio processing application 120 uses other functions to modify the input audio signal 140(1) based on the orientation of the speaker 160 relative to the physical object 180. In some embodiments, the audio processing application 120 also applies various panning techniques to the audio signal 140(1) corresponding to the physical object 180 to simulate the sound of the physical object 180 as the physical object 180 moves along the trajectory 222.

[0055] Alternatively, in some embodiments, the physical listening environment 210 does not include the physical object 180. In such instances, the audio processing application 120 tracks a trajectory 252 of the virtual object 132 within the virtual environment 130 based on the position data 136 generated for the virtual object 132. For example, the audio processing application 120 and / or a separate application (not shown) generates the virtual object 132 (e.g., a virtual ball in an AR game). In such instances, the application that manages the virtual environment 130 also generates position data for the virtual object 132 as the virtual object 132 traverses along the trajectory 252. In such instances, the audio processing application 120 uses the position data corresponding to the virtual object 132 to track the virtual object 132 and calculates the distances 262(1) to 262(5) between the position of the virtual object 132 and the positions of the virtual microphones 134(1) to 134(5).

[0056] Figure 3 A flow chart illustrating method steps for generating an audio signal based on the position of a tracked object according to various embodiments. Figure 1 to Figure 2 The method steps are described herein as a system with a plurality of method steps, but those skilled in the art will understand that any system configured to perform the method steps in any order falls within the scope of the present disclosure.

[0057] As shown, method 300 begins at step 302, where an audio processing application 120 tracks a set of one or more speakers 160. In various embodiments, the audio processing application 120 executed on a computing device 110 tracks movement of a set of speakers 160 (e.g., speakers 160(1) through 160(5)) within a physical listening environment 210. Additionally or alternatively, in various embodiments, the audio processing application 120 determines a current position of each speaker in the set of speakers 160. In various embodiments, the audio processing application 120 receives sensor data from one or more sensors 150 coupled to the computing device 110, wherein the sensor data indicates a position and / or orientation of each speaker 160 at a given time. In some embodiments, the sensor data indicates that at least one speaker in the set of speakers 160 (e.g., speaker 160(4)) is moving.

[0058] In one example, the audio processing application 120 obtains sensor data (e.g., tracking data as a series of optical data for a given speaker 160, and / or a series of auditory data received in response to a test signal generated by the computing device 110) from one or more sensors 150 coupled to the computing device 110. In some embodiments, the computing device 110 determines a current position of each speaker in the set of speakers 160 based on the sensor data. The computing device 110 then stores each determined position as part of the position data 136 in the form of a combination of a position and an orientation. For example, the position data 136 for a given speaker 160(4) includes coordinates of the position of the speaker 160(4) within the physical listening environment 210, as well as orientation information, such as a set of angles relative to a normal orientation within the physical listening environment 210 (e.g., ).

[0059] Additionally or alternatively, in some embodiments, audio processing application 120 receives sensor data (e.g., acceleration measurements, magnetic field measurements, angular velocity, etc.) generated by a position sensor and / or IMU on speaker 160 (4). For example, speaker 160 (4) transmits a series of messages containing sensor data as it moves. In such instances, audio processing application 120 receives and aggregates the sensor data included in the messages and determines a trajectory and / or current position of speaker 160.

[0060] At step 304, the audio processing application 120 tracks the position of the object. In various embodiments, the audio processing application 120 tracks the position of the object 180 within the listening environment. In some embodiments, the object is a physical object 180 within the physical listening environment 210. In such instances, the audio processing application 120 determines the position of the physical object 180 within the physical listening environment 210 and / or tracks the trajectory 222 of the physical object 180. In one example, the audio processing application 120 obtains sensor data from one or more sensors 150 coupled to the computing device 110 and determines the current position of the object 180 based on the sensor data. The computing device 110 then stores each determined position as part of the position data 136 in the form of a combination of position and orientation.

[0061] Alternatively, in some embodiments, the object is a virtual object 132 within a virtual environment 130 corresponding to a physical listening environment 210 including a speaker 160. In such instances, the audio processing application 120 can track a trajectory 252 of the virtual object 132 within the virtual environment 130 based on the position data 136 generated for the virtual object 132. For example, the audio processing application 120 and / or an XR application (not shown) generates a virtual environment 130 including the virtual object 132. In such instances, the application that manages the virtual environment 130 generates the position data of the virtual object 132 as the virtual object 132 traverses along the trajectory 252. In such instances, the audio processing application 120 tracks the virtual object 132 by retrieving the portion of the position data 136 corresponding to the virtual object 132.

[0062] At step 306, the audio processing application 120 calculates the distance from the tracked object to the speaker 160. In various embodiments, the audio processing application 120 calculates a set of distances between the position of each speaker in the set of speakers 160 and the tracked object. In some embodiments, the audio processing application 120 calculates a physical distance (e.g., a Euclidean distance) from the physical object 180 to each speaker 160(1)-160(5) in the physical listening environment 210 to generate the calculated distances 262(1)-262(5).

[0063] Additionally or alternatively, in some embodiments, the audio processing application 120 calculates a distance from the virtual object 132 to a set of virtual microphones 134(1)-134(5) within the virtual environment 130 to generate the calculated distances 262(1)-262(5). For example, the virtual environment 130 includes a set of virtual microphones 134 located at locations within the virtual environment 130 that correspond to locations of the set of speakers 160 within the physical listening environment 210. For example, the location of the virtual microphone 134(4) within the virtual environment 130 corresponds to the location of the speaker 160(4) within the physical listening environment 210. Thus, the calculated distance 262(4) between the virtual object 132 and the virtual microphone 134(4) represents the Euclidean distance between an object (virtual object 132 or physical object 180) within the physical listening environment 210 and the speaker 160(4).

[0064] At step 308, the audio processing application 120 selects a distance decay function. In various embodiments, the audio processing application 120 selects a distance decay function from a set of candidate distance decay functions to use when generating a set of audio signals for the set of speakers 160. In various embodiments, the audio processing application 120 uses the distance decay function to apply the calculated distance to the given speaker to modify the amplitude and / or phase of the input audio signal when generating the audio signal for the given speaker for reproduction. For example, the computing device 110 may store a set of candidate distance decay functions, such as a linear function, a linear square function, or an inverse function, that attenuate the gain or change the phase of the input audio signal based on the distance between the tracked object and the given speaker.

[0065] At step 310, the audio processing application 120 generates an audio signal for the speaker 160 based on the calculated distance 262 and the selected distance decay function. In various embodiments, the audio processing application 120 uses the selected distance decay function to modify the amplitude and / or phase of the input audio signal of each speaker 160(1)-160(5) based on the corresponding calculated distance 262(1)-262(5) between the speaker 160 and the tracked object.

[0066] In some embodiments, the audio processing application 120 receives an input audio signal 140(1) from an object 180 for generating an audio signal. For example, the computing device 110 receives a sound profile 182(1) including the input audio signal 140(1) from the object 180 and stores the sound profile 182(1) in the memory 114. In some embodiments, the audio processing application 120 receives the input audio signal 140(1) and the sound profile 182(1) separately. Additionally or alternatively, in some embodiments, the computing device 110 stores one or more sound profiles 182 and / or one or more audio signals 140 associated with multiple objects. In such instances, the audio processing application 120 identifies the object 180, identifies the sound profile 182(1) corresponding to the object 180, and retrieves the audio signal 140(1) associated with the sound profile 182(1).

[0067] In various embodiments, upon retrieving the input audio signal 140(1), the audio processing application 120 generates a set of modified audio signals for the set of speakers 160 by modifying the input audio signal 140(1) using the selected distance decay function. For example, the audio processing application 120 generates a modified audio signal for the speaker 160(4) by modifying the amplitude of the input audio signal 140(1) by applying the selected distance decay function. The distance decay function modifies the amplitude of the input audio signal 140(1) based on the calculated distance 262(4) such that the amplitude of the modified audio signal decreases as the calculated distance 262(4) between the object and the speaker 160(4) increases.

[0068] In some embodiments, the distance decay function attenuates the audio signal between a minimum distance and a maximum distance. In such instances, the audio processing application 120 compares the calculated distance 262(4) to a minimum distance threshold and / or a maximum distance threshold. When the audio processing application 120 determines that the calculated distance 262(4) satisfies the threshold, the audio processing application 120 applies the selected distance decay function.

[0069] At step 312, the audio processing application 120 transmits the modified audio signal to the speaker 160. In various embodiments, the audio processing application 120 drives the computing device 110 to transmit the set of modified audio signals to the set of speakers 160. In some embodiments, each respective speaker in the set of speakers 160 receives one of the modified audio signals from the computing device 110 via a wire, a wireless stream, or via a network. Upon receiving the modified audio signal, the speaker 160 reproduces the modified audio signal to generate sound waves within the physical listening environment 210. In various embodiments, the sound waves generated by the set of speakers 160 combine to generate a sound field that provides a perceptually accurate location of the physical object 180 within the physical listening environment 210.

[0070] While transmitting the audio signal to the set of speakers 160, the audio processing application 120 returns to step 302 or 304 to optionally track any additional movement of the object 180 and / or one or more of the speakers in the set of speakers 160. For example, the audio processing application 120 returns to step 302 to detect movement of the speaker 160(1) to a new location within the physical listening environment 210. In such an instance, the audio processing application 120 repeats at least a portion of the method 300 to calculate the distance between the object 180 and the speaker 160(1) at the new location. Alternatively, the audio processing application 120 proceeds to step 304 to track the trajectory 222 of the object 180 and repeat the method 300 to calculate the distance between the object 180 and the set of speakers 160 at the new location.

[0071] In summary, the audio processing application tracks the position of one or more speakers in the physical listening environment. The audio processing application identifies the position of the speaker in the corresponding virtual listening environment and places the virtual microphone at the identified position. The audio processing application also tracks one or more objects in the physical listening environment. The audio processing application identifies the position of one or more objects in the corresponding virtual listening environment and places the virtual sound source at the identified position. In some embodiments, the audio processing application determines the distance between the position of the object and the position of the speaker for each speaker. In some embodiments, the virtual sound source is a virtual object in the virtual listening environment. In some embodiments, the audio processing application determines the distance between the position of the virtual sound source and the position of the virtual microphone corresponding to the speaker in the virtual listening environment for each speaker.

[0072] When determining the distance, the audio processing application then generates an audio signal for each speaker based on the corresponding distance. When generating the audio signal, the audio processing application determines the amplitude of a given audio signal for the speaker based on the determined distance to the virtual sound source and the distance attenuation function. The audio processing application then distributes the audio signal to the corresponding speakers for reproduction in the physical listening environment.

[0073] At least one technical advantage of the disclosed technology relative to the prior art is that using the disclosed technology, a sound system can distribute audio signals to one or more speakers in a physical listening area in a manner that indicates the apparent position of an object within a listening environment and improves the perceived accuracy in terms of timbre and positioning. Specifically, by determining the position of an object within a listening environment and attenuating the sound signal based on the distance between the object and the corresponding one or more speakers, the sound system provides perceptually accurate sound to the tracked object in real time, thereby effectively providing realistic immersive audio in the listening environment that responds to the movement of the object within the listening environment without requiring large and expensive processing resources. In addition, by using techniques that operate with various numbers of speakers, a sound system using the disclosed technology can provide perceptually accurate audio that responds to different numbers of speakers located at various locations within the listening environment. These technical advantages provide one or more technical improvements over prior art methods.

[0074] 1. In various embodiments, a computer-implemented method includes: determining a speaker position of a speaker in a listening environment; tracking an object position of an object within the listening environment; retrieving an audio signal; calculating a distance between the object position and the speaker position; generating a modified audio signal, wherein an amplitude of the modified audio signal is based on (i) the audio signal, (ii) the distance, and (iii) a distance attenuation function; and transmitting the modified audio signal to the speaker.

[0075] 2. The computer-implemented method of clause 1, wherein the object is a physical object within the listening environment.

[0076] 3. A computer-implemented method as described in clause 1 or 2, wherein the object is an interactive toy.

[0077] 4. The computer-implemented method of any one of clauses 1 to 3, wherein determining the speaker position of the speaker comprises tracking the speaker.

[0078] 5. The computer-implemented method of any of clauses 1 to 4, wherein the distance decay function comprises at least one of: a linear function, a linear square function, or an inverse function.

[0079] 6. A computer-implemented method as described in any of clauses 1 to 5, wherein the speaker position includes a position and an orientation, and the amplitude of the modified audio signal is further based on the orientation of the speaker relative to the object.

[0080] 7. A computer-implemented method as described in any of clauses 1 to 6, further comprising: determining an additional speaker position for each of one or more additional speakers in the listening environment, and calculating an additional distance between the object position and the additional speaker position for each of the one or more additional speakers; generating an additional modified audio signal, wherein an amplitude of the additional modified audio signal is based on (i) the audio signal, (ii) the additional distance, and (iii) the distance attenuation function; and transmitting the additional modified audio signal to the additional speaker.

[0081] 8. The computer-implemented method of any of clauses 1 to 7, wherein the speaker is a speaker array.

[0082] 9. The computer-implemented method of any of clauses 1 to 8, wherein retrieving the audio signal comprises receiving, from the object, a sound profile comprising the audio signal.

[0083] 10. The computer-implemented method of any of clauses 1 to 9, wherein retrieving the audio signal comprises identifying the object and loading the audio signal from a sound profile corresponding to the object.

[0084] 11. In various embodiments, one or more non-transitory computer-readable media storing instructions that, when executed by one or more processors, cause the one or more processors to perform the following steps: determining a speaker position of a speaker in a listening environment; tracking an object position of an object within the listening environment; retrieving an audio signal; calculating a distance between the object position and the speaker position; generating a modified audio signal, wherein an amplitude of the modified audio signal is based on (i) the audio signal, (ii) the distance, and (iii) a distance attenuation function; and transmitting the modified audio signal to the speaker.

[0085] 12. One or more non-transitory computer-readable media as described in clause 11, further comprising generating one or more virtual microphones in a virtual environment corresponding to the listening environment, wherein each virtual microphone corresponds to a speaker among the one or more speakers, and each virtual microphone is located at a microphone position within the virtual environment corresponding to the speaker position within the listening environment.

[0086] 13. The one or more non-transitory computer-readable media of clause 11 or 12, wherein the object is a virtual object within the virtual environment.

[0087] 14. The one or more non-transitory computer-readable media of any one of clauses 11 to 13, wherein determining the speaker position of the speaker comprises tracking the speaker.

[0088] 15. One or more non-transitory computer-readable media as described in any of clauses 11 to 14, wherein the speaker position includes a position and an orientation, and the amplitude of the modified audio signal is further based on the orientation of the speaker relative to the object.

[0089] 16. The one or more non-transitory computer-readable media of any of clauses 11 to 15, wherein retrieving the audio signal comprises identifying the object and loading the audio signal from a sound profile corresponding to the object.

[0090] 17. In various embodiments, an interactive toy that generates an audio signal for reproduction comprises: at least one sensor that acquires sensor data; and a computing device that determines a speaker position of a speaker in a listening environment; tracks an object position of the interactive toy within the listening environment based on the sensor data; retrieves an audio signal; calculates a distance between the object position and the speaker position; generates a modified audio signal, wherein an amplitude of the modified audio signal is based on (i) the audio signal, (ii) the distance, and (iii) a distance decay function; and transmits the modified audio signal to the speaker.

[0091] 18. The interactive toy of clause 17, wherein determining the speaker position of each of the one or more speakers comprises tracking each of the one or more speakers to the speaker position.

[0092] 19. An interactive toy as described in claim 17 or 18, wherein the computing device further determines an additional speaker position for each of one or more additional speakers in the listening environment, and calculates an additional distance between the object position and the additional speaker position for each of the one or more additional speakers; generates an additional modified audio signal, wherein the amplitude of the additional modified audio signal is based on (i) the audio signal, (ii) the additional distance and (iii) the distance attenuation function; and transmits the additional modified audio signal to the additional speaker.

[0093] 20. The interactive toy of any of clauses 17 to 19, wherein the distance decay function comprises at least one of: a linear function, a linear square function, or an inverse function.

[0094] Any and all combinations of any claim elements recited in any claim and / or any elements described in this application are within the intended scope of the invention and protection in any manner.

[0095] Descriptions of various embodiments have been presented for purposes of illustration, but are not intended to be exhaustive or limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments.

[0096] Aspects of the present embodiment may be embodied as a system, method or computer program product. Therefore, aspects of the present disclosure may take the form of a complete hardware implementation, a complete software implementation (including firmware, resident software, microcode, etc.) or a combination of software and hardware implementations, which may be collectively referred to herein as "modules", "systems" or "computers". In addition, any hardware and / or software techniques, processes, functions, components, engines, modules or systems described in the present disclosure may be implemented as circuits or circuit sets. In addition, aspects of the present disclosure may take the form of a computer program product embodied in one or more computer-readable media on which computer-readable program codes are embodied.

[0097] Any combination of one or more computer-readable media may be utilized. A computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium. A computer-readable storage medium may be, for example, but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or apparatus, or any suitable combination of the foregoing. More specific examples of computer-readable storage media (a non-exhaustive list) would include the following media: an electrical connection with one or more conductors, a portable computer floppy disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing. In the context of this document, a computer-readable storage medium may be any tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or apparatus.

[0098] Aspects of the present disclosure are described above with reference to the flowchart and / or block diagram of the method, device (system) and computer program product according to the embodiments of the present disclosure. It should be understood that each frame in the flowchart and / or block diagram and the frame combination in the flowchart and / or block diagram can be realized by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer or other programmable data processing device to produce a machine. The instruction enables the function / action specified in one or more frames of the flowchart and / or block diagram to be realized when the processor of the computer or other programmable data processing device is executed. This processor can be, but is not limited to, a general-purpose processor, a special-purpose processor, a specific application processor or a field programmable gate array.

[0099] The flowcharts and block diagrams in the accompanying drawings illustrate the possible specific implementation architectures, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram may represent a module, section or portion of a code, and the code includes one or more executable instructions for implementing the specified one or more logical functions. It should also be noted that in some alternative implementation schemes, the functions mentioned in the box may not appear in the order mentioned in the accompanying drawings. For example, two boxes shown in succession can actually be executed substantially simultaneously, or these boxes can sometimes be executed in reverse order, depending on the functionality involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of boxes in the block diagram and / or flowchart can be implemented by a dedicated hardware-based system that performs a specified function or action, or a combination of dedicated hardware and computer instructions.

[0100] While the foregoing is directed to embodiments of the present disclosure, other and further embodiments of the present disclosure may be devised without departing from the basic scope thereof, which is determined by the following claims.

Claims

1. A computer-implemented method comprising: Determine the speaker positions of the speakers in the listening environment; tracking an object position of an object within the listening environment; Retrieving audio signals; calculating the distance between the object position and the speaker position, generating a modified audio signal, wherein an amplitude of the modified audio signal is based on (i) the audio signal, (ii) the distance, and (iii) a distance decay function; and The modified audio signal is transmitted to the speaker.

2. The computer-implemented method of claim 1, wherein the object is a physical object within the listening environment.

3. The computer-implemented method of claim 1, wherein the object is an interactive toy. The computer-implemented method of claim 1 , wherein determining the speaker position of the speaker comprises tracking the speaker.

5. The computer-implemented method of claim 1, wherein the distance decay function comprises at least one of: a linear function, a linear-square function, or an inverse function.

6. The computer-implemented method of claim 1, wherein: The speaker position includes a position and an orientation; and The amplitude of the modified audio signal is further based on the orientation of the speaker relative to the object.

7. The computer-implemented method of claim 1 , further comprising: determining, for each of the one or more additional speakers in the listening environment, an additional speaker position; and For each additional speaker of the one or more additional speakers: calculating additional distances between the object position and the additional speaker positions, generating an additional modified audio signal, wherein an amplitude of the additional modified audio signal is based on (i) the audio signal, (ii) the additional distance, and (iii) the distance decay function; and The additional modified audio signal is transmitted to the additional speaker.

8. The computer implemented method of claim 1, wherein the speaker is a speaker array.

9. The computer-implemented method of claim 1, wherein retrieving the audio signal comprises receiving, from the object, a sound profile that includes the audio signal.

10. The computer-implemented method of claim 1, wherein retrieving the audio signal comprises: Identify the object; as well as The audio signal is loaded from a sound configuration file corresponding to the object.

11. One or more non-transitory computer readable media storing instructions that, when executed by one or more processors, cause the one or more processors to perform the following steps: Determine the speaker positions of the speakers in the listening environment; tracking an object position of an object within the listening environment; Retrieving audio signals; calculating the distance between the object position and the speaker position, generating a modified audio signal, wherein an amplitude of the modified audio signal is based on (i) the audio signal, (ii) the distance, and (iii) a distance decay function; and The modified audio signal is transmitted to the speaker.

12. The one or more non-transitory computer-readable media of claim 11, further comprising: One or more virtual microphones are generated in a virtual environment corresponding to the listening environment, wherein: Each virtual microphone corresponds to a speaker among the one or more speakers, and Each virtual microphone is located at a microphone position within the virtual environment that corresponds to the speaker position within the listening environment.

13. The one or more non-transitory computer-readable media of claim 12, wherein the object is a virtual object within the virtual environment.

14. The one or more non-transitory computer-readable media of claim 11, wherein determining the speaker position of the speaker comprises tracking the speaker.

15. The one or more non-transitory computer-readable media of claim 11, wherein: The speaker position includes a location and an orientation, and The amplitude of the modified audio signal is further based on the orientation of the speaker relative to the object.

16. The one or more non-transitory computer-readable media of claim 11, wherein retrieving the audio signal comprises: Identify the object; as well as The audio signal is loaded from a sound configuration file corresponding to the object.

17. An interactive toy for generating an audio signal for reproduction, comprising: at least one sensor, the at least one sensor acquiring sensor data; as well as A computing device, wherein: Determine the speaker positions of the speakers in the listening environment; tracking an object position of the interactive toy within the listening environment based on the sensor data; Retrieving audio signals; calculating the distance between the object position and the speaker position, generating a modified audio signal, wherein an amplitude of the modified audio signal is based on (i) the audio signal, (ii) the distance, and (iii) a distance decay function; and The modified audio signal is transmitted to the speaker.

18. The interactive toy of claim 17, wherein determining the speaker position of each of the one or more speakers comprises tracking each of the one or more speakers to the speaker position.

19. The interactive toy of claim 17, wherein the computing device further: for each of the one or more additional speakers in the listening environment, determining an additional speaker position; and For each additional speaker of the one or more additional speakers: calculating additional distances between the object position and the additional speaker positions, generating an additional modified audio signal, wherein an amplitude of the additional modified audio signal is based on (i) the audio signal, (ii) the additional distance, and (iii) the distance decay function; and The additional modified audio signal is transmitted to the additional speaker.

20. The interactive toy of claim 17, wherein the distance decay function comprises at least one of: a linear function, a linear square function, or an inverse function.