Virtual sound source audio translation based on proximity

By using proximity sensors in the audio system to measure the distance between the virtual sound source and the speaker and generate a gain value to control the audio signal, the problem of hardware and software complexity when generating immersive audio scenes in the prior art is solved, and more flexible and efficient audio processing is achieved.

CN120128871APending Publication Date: 2025-06-10HARMAN INT IND INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411784523.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-12-04
Filing Date
2024-12-06
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

The prior art requires complex hardware and software systems to track the precise location of speakers when generating immersive audio scenarios, resulting in increased costs and complexity and difficult to apply to small and low computing power products.

Method used

By determining the distance between the virtual sound source and the speaker, using a proximity sensor to generate a gain value, control the audio signal output by the speaker, and implementing a proximity-based translation of the audio object.

Benefits of technology

This technology enables generation of immersive spatial audio scenarios with fewer hardware components and reduced software processing, improving the flexibility of the audio system and easy real-time audio processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120128871A_ABST
    Figure CN120128871A_ABST
Patent Text Reader

Abstract

Audio translation techniques for virtual sound sources are described. In some embodiments, the techniques include determining a first distance between a first speaker and a virtual sound source; determining a second distance between the second loudspeaker and the virtual sound source; generating a first audio output signal of the first speaker based on the input audio signal, the first distance, and the second distance; generating a second audio output signal of a second speaker based on the input audio signal, the first distance, and the second distance; transmitting the first audio output signal to a first speaker for output; and transmitting the second audio output signal to a second speaker for output.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross - Reference to Related Applications

[0002] This application claims the benefit of U.S. Provisional Patent Application No. 63 / 607,819, filed on Dec. 8, 2023, titled "Proximity - Based Audio Object Panning". The subject matter of this related application is hereby incorporated herein by reference. Technical Field

[0003] The contemplated embodiments generally relate to audio systems and, more particularly, to proximity - based virtual sound source audio panning. Background Art

[0004] The use of virtual reality (VR) and augmented reality (AR) is becoming increasingly popular in many applications. VR is an interactive experience that typically replaces the user's real - world environment simulated via auditory and visual feedback, while AR is an interactive experience characterized by the combination of the real world and the virtual world, real - time interaction, and precise 3D registration of virtual and real objects. VR and AR are designed to provide users with an immersive experience of a virtual world or a real world enhanced with virtual sensing information. VR and AR applications now include entertainment ( For example , video games and multimedia entertainment), education ( For example , medical and military training), and business ( For example , virtual meetings or other interactions).

[0005] A significant challenge in VR and AR applications is how to map virtual sound sources to real speakers within a listening area to create an immersive and realistic audio experience. Creating such a precise spatial audio scene using multiple speakers requires information about the positions of the speakers in the listening room so that each speaker appropriately processes the input audio signal for the correct physical location of the simulated virtual sound source. Tracking the precise positions of the speakers can be complex and requires expensive sensors such as IR transmitters / receivers, UWB transmitters / receivers, and / or software systems similar to game engines. These methods significantly increase the cost and complexity of VR or AR audio systems, making them impractical for small and low - computing - power products (such as low - cost toys) to generate an immersive audio scene that accurately simulates the positions of toy or other virtual sounds.

[0006] As previously mentioned, there is a need in the art for improved techniques to generate immersive audio scenes. Summary of the Invention

[0007] One embodiment of the present disclosure describes a computer - implemented method that includes determining a first distance between a first speaker and a virtual sound source; determining a second distance between a second speaker and the virtual sound source; generating a first audio output signal for the first speaker based on an input audio signal, the first distance, and the second distance; generating a second audio output signal for the second speaker based on the input audio signal, the first distance, and the second distance; transmitting the first audio output signal to the first speaker for output; and transmitting the second audio output signal to the second speaker for output.

[0008] At least one technical advantage of the disclosed technology over the prior art is that, using the disclosed technology, a spatial audio scene including a virtual sound source such as a toy or other object can be created, where the audio signal representing the virtual sound source is generated by speakers physically separated from the virtual sound source. Using the disclosed technology, proximity sensors can be used to generate a spatial audio scene, where the proximity sensors measure the distance from the virtual sound source to each speaker that generates the audio signal representing the virtual sound source. Thus, compared to other immersive spatial audio methods, the disclosed technology can produce an immersive spatial audio mix with fewer hardware components and reduced software processing. Another advantage of the disclosed technology is that the technology is flexible with respect to various characteristics of the sound system that generates the spatial audio scene, such as the number of speakers included in the sound system, the position of the speakers within the listening area, and the number and position of virtual sound sources within the spatial audio scene. Another advantage of the disclosed technology is that the technology is easily implemented using real - time audio processing, which is important for immersive applications that require fast processing and distribution of audio to maintain a certain audio experience. These technical advantages represent one or more technical improvements over prior art methods. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] To enable a more specific understanding of the manner in which the above - described features of various embodiments can be obtained, the inventive concept briefly summarized above may be described in more detail by reference to various embodiments, some of which are illustrated in the accompanying drawings. However, it should be noted that the drawings only illustrate typical embodiments of the inventive concept and should not be considered in any way to limit the scope, and there are other equivalent embodiments.

[0010] Figure 1 is a schematic diagram showing an audio system according to various embodiments;

[0011] Figure 2 is a conceptual diagram of an audio system and a listening area according to one embodiment;

[0012] Figure 3 is a conceptual diagram of an audio system and a listening area according to another embodiment;

[0013] Figure 4 shows a cross-fade curve for determining a gain value according to various embodiments;

[0014] Figures 5A to 5E shows various cross-fade curves for determining gain values according to various other embodiments; and

[0015] Figure 6 is a flowchart of method steps for generating a spatial audio scene according to various embodiments. DETAILED DESCRIPTION

[0016] In the following description, numerous specific details are set forth to provide a more thorough understanding of the various embodiments. However, it will be apparent to one of ordinary skill in the art that the inventive concept may be practiced without one or more of these specific details.

[0017] INTRODUCTION

[0018] According to various embodiments, a spatial audio scene is generated by an audio system using proximity-based audio object panning. In an embodiment, the spatial audio scene includes at least one virtual sound source, such as a toy or other object, which operates as a virtual sound source within the spatial audio scene. To generate the spatial audio scene, an audio signal representing the virtual sound source is generated by a speaker of the audio system, the speaker being physically separated from the virtual sound source and thus not included within the virtual sound source. In an embodiment, the volume of the audio signal generated by each speaker varies based on the position change of the virtual sound source relative to the speaker. Thus, the spatial audio scene provides an immersive audio experience for a user near the audio system, where the movement of the virtual sound source is reflected by proximity-based audio object panning. In an embodiment, a proximity sensor may be used to generate the spatial audio scene, the proximity sensor measuring the distance from the virtual sound source to each speaker, wherein the gain value of each speaker is determined based on the measured distance from the virtual sound source to that speaker. Since the gain value of each speaker is based on the measured distance between the virtual sound source and the speaker, a complex two-dimensional or three-dimensional map of the relative positions of the speaker and the virtual sound source is not required to generate the spatial audio scene. Thus, compared to other immersive spatial audio methods, the audio system can generate an immersive spatial audio mix with fewer hardware components and reduced software processing.

[0019] SYSTEM OVERVIEW

[0020] Figure 1is a block diagram of an audio system 100 configured to implement one or more embodiments of the present disclosure. As shown, the audio system 100 includes, but is not limited to, a computing device 110, one or more distance sensors 150, a plurality of speakers 160, and a virtual sound source 102. The computing device 110 includes, but is not limited to, a processing unit 112 and a memory 114. The memory 114 stores (but is not limited to) an audio panning application 120. The virtual sound source 102 can be an interactive toy or other object, where an audio signal representing the virtual sound source 102 is generated by the speakers 160. Generally, the audio signal can be a sound effect or other sound nominally generated by the virtual sound source 102 but actually generated by the speakers 160. As shown, the speakers 160 are not arranged within the virtual sound source 102 but are physically separated from the virtual sound source 102.

[0021] In operation, when outputting an audio signal corresponding to the virtual sound source 102, the audio system 100 uses the measured distances between the virtual sound source 102 corresponding to a physical device ( For example , a toy or other object that may not have audio output capabilities) and each of the plurality of speakers 160 to generate a gain setting for each of the plurality of speakers 160. Specifically, the audio panning application 120 uses the distance sensors 150 to determine the distances between the virtual sound source 102 and each of the speakers 160. Then, the audio panning application 120 uses the measured distances and one or more gain curves to determine the gain values for each of the plurality of speakers 160. Then, the gain values are used to control the volume of the audio signals output by each speaker 160 to produce a sound associated with the virtual sound source 102.

[0022] The distance sensor 150 includes various types of sensors for measuring the distance between each of the virtual sound sources 102 and the speakers 160. In some embodiments, the distance sensor 150 is a proximity sensor. The distance sensor 150 can use any technically feasible distance measurement technique, including but not limited to using ultrasonic waves, infrared light, computer imaging mode, and / or the like. In some embodiments, one or more distance sensors 150 are disposed within the virtual sound source 102. For example, in some embodiments, the distance sensor 150 includes a sensor disposed within the virtual sound source 102 that is capable of detecting inaudible audio signals generated by each speaker 160 to determine the distance between the virtual sound source 102 and each speaker. In another example, in some embodiments, the distance sensor 150 includes an ultrasonic sensor that measures the distance to the speaker 160 by transmitting sound waves toward the speaker 160 and measuring the time interval required for a portion of the transmitted sound waves to be reflected back to the ultrasonic sensor. Additionally or alternatively, in some embodiments, one or more distance sensors 150 are disposed within each speaker 160. Additionally or alternatively, in some embodiments, one or more distance sensors 150 of the audio system 100 are physically separated from both the virtual sound source 102 and the speakers 160.

[0023] In some embodiments, in addition to the distance sensor 150, the audio system 100 further includes other types of sensors to obtain information about the acoustic environment. Other types of sensors include cameras, quick response (QR) code tracking systems, motion sensors such as accelerometers or inertial measurement units (IMUs) ( For example , triaxial accelerometers, gyroscopic sensors, and / or magnetometers), pressure sensors, and the like. Additionally, in some embodiments, the distance sensor 150 can include wireless sensors (including radio frequency (RF) sensors ( For example , sonar, and radar)) and / or wireless communication protocols (including Bluetooth, Bluetooth low energy (BLE), cellular protocols, and / or near field communication (NFC)).

[0024] Each of the plurality of speakers 160 can be any technically feasible type of audio output device. For example, in some embodiments, the plurality of speakers 160 includes one or more digital speakers that receive an audio output signal in digital form and convert the audio output signal into a change in air pressure or sound energy via a transduction process. According to various embodiments, each of the plurality of speakers 160 generates an audio signal (output sound) for the virtual sound source 102 at a volume or gain level determined by the audio translation application 120.

[0025] The computing device 110 enables the implementation of the various embodiments described herein. In Figure 1In the illustrated embodiment, computing device 110 includes a processing unit 112 and a memory 114.

[0026] The processing unit 112 can be any suitable processor, such as a central processing unit (CPU), a graphics processing unit (GPU), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), a digital signal processor (DSP), and / or any other type of processing unit or combination of different processing units, such as a CPU and / or DSP configured to operate in conjunction with a GPU. Generally, the processing unit 112 can be any technically feasible hardware unit capable of processing data and / or executing software applications, such as the audio panning application 120.

[0027] The memory 114 can include random access memory (RAM) modules, flash memory cells, or any other type of memory cell or combination thereof. The processing unit 112 is configured to read data from and write data to the memory 114. In various embodiments, the memory 114 includes non-volatile memory, such as an optical drive, a magnetic drive, a flash drive, or other storage devices. In some embodiments, a separate data repository, such as an external data repository, included in a network (“cloud storage device”) can supplement the memory 114. The audio panning application 120 within the memory 114 can be executed by the processing unit 112 to implement the overall functionality of the computing device 110 and thus, overall, coordinate the operation of the audio system 100. In various embodiments, an interconnect bus (not shown) connects the processing unit 112, the memory 114, the speakers 160, the distance sensor 150, and any other components of the computing device 110.

[0028] In various embodiments, the computing device 110 is included within the virtual sound source 102, thereby allowing the virtual sound source 102 to use the audio panning application 120 to generate appropriate sound levels for each speaker 160 that produces an audio output representative of the virtual sound source 102. Alternatively, in some embodiments, the computing device 110 is separate from the virtual sound source 102, and the distance sensor 150 is disposed within the virtual sound source 102. One such embodiment is described below in connection with Figure 2 In such an embodiment, the computing device 110 can be included in a home theater system, a soundbar, a vehicle system, another computing device ( For example , a desktop computer, a laptop computer, a tablet computer, a mobile device, etc.) and / or the like. Similarly, in such an embodiment, the computing device 110 can be included in one or more devices separate from the virtual sound source 102, such one or more devices such as consumer products ( For example , a portable speaker, a gaming device, a modular toy component, etc.), a vehicle (For example , the main unit of an automobile, a truck, a van, etc.), a smart home device ( For example , a smart lighting system, a security system, a digital assistant, etc.), a communication system ( For example , a teleconference system, a video conference system, a speaker amplification system, etc.), etc. In various embodiments, the virtual sound source 102 and the computing device 110 are located in various environments, including but not limited to indoor environments ( For example , a living room, a meeting room, a conference hall, a home office, etc.)

[0029] and / or outdoor environments ( For example , a patio, a roof, a garden, etc.).

[0030] Figure 2 is a conceptual diagram of the audio system 100 and the listening area 200 according to one embodiment. As shown, the computing device 110 of the audio system 100 is separated from the virtual sound source 102, and the distance sensor 150 is disposed within the virtual sound source 102. For example, the computing device 110 may be a modular sound unit of an interactive toy system including the virtual sound source 102, and the virtual sound source 102 may be an interactive toy included in the interactive toy system. In such an embodiment, the virtual sound source 102 may be an interactive toy or other object in which a suitable sound effect or other audio output is generated by the speaker 160. In Figure 2 the embodiment shown, the audio system 100 includes but is not limited to a first speaker 262, a second speaker 264, and a third speaker 266, each of which can be consistent with Figure 1 the speaker 160. In addition, the virtual sound source 102, the first speaker 262, the second speaker 264, and the third speaker 266 are disposed within and / or near the listening area 200 that can be occupied by a user (not shown).

[0031] In operation, the distance sensor 150 determines the distance D1 between the virtual sound source 102 and the first speaker 262, the distance D2 between the virtual sound source 102 and the second speaker 264, and the distance D3 between the virtual sound source 102 and the third speaker 266. Then, the distance sensor 150 transmits the distances D1, D2, and D3 to the computing device 110, so that the audio panning application 120 can determine the suitable gain value G1 of the first speaker 262, the suitable gain value G2 of the second speaker 264, and the suitable gain value G3 of the third speaker 266. The following is combined with Figure 4Describe techniques for determining a gain value G1, a gain value G2, and a gain value G3. Once the computing device 110 determines the gain value G1, the gain value G2, and the gain value G3, the computing device 110 generates a first audio output signal 202 to the first speaker 262, a second audio output signal 204 to the second speaker 264, and a third audio output signal 206 to the third speaker 266. In some embodiments, the computing device 110 generates the first audio output signal 202, the second audio output signal 204, and the third audio output signal 206 based on an input audio signal 210 and specific gain values, where the input audio signal 210 represents sound associated with the virtual sound source 102. In such an embodiment, the computing device 110 generates the first audio output signal 202 by modifying the input audio signal 210 with the gain value G1, generates the second audio output signal 204 by modifying the input audio signal 210 with the gain value G2, and generates the third audio output signal 206 by modifying the input audio signal 210 with the gain value G3. The computing device 110 then transmits the first audio output signal 202 to the first speaker 262, the second audio output signal 204 to the second speaker 264, and the third audio output signal 206 to the third speaker 266. The first speaker 262 then generates an audio output (not shown) based on the first audio output signal 202. Similarly, the second speaker 264 generates an audio output (not shown) based on the second audio output signal 202, and the third speaker 266 generates an audio output (not shown) based on the third audio output signal 206. In this way, audio from one or more virtual sound sources ( For example , the virtual sound source 102) is distributed to any number of physical speakers ( For example , the first speaker 262, the second speaker 264, and the third speaker 266) such that the localization of the virtual sound source within the listening area 200 is perceptually accurate.

[0032] Figure 3 is a conceptual diagram of an audio system 100 and a listening area 300 according to another embodiment. As shown in the figure, the audio system 100 includes a first speaker 362, a second speaker 364, and a third speaker 366, each of which can be consistent with the Figure 1 speaker 160. In addition, the virtual sound source 102, the first speaker 362, the second speaker 364, and the third speaker 366 are disposed within and / or near a listening area 300 that can be occupied by a user (not shown). Compared to the embodiment of the audio system 100 shown in Figure 2 , in Figure 3In this case, different distance sensors are provided within each speaker of the audio system 100. Accordingly, distance sensor 352 is provided within the first speaker 362, distance sensor 354 is provided within the second speaker 364, and distance sensor 356 is provided within the third speaker 366.

[0033] In operation, distance sensor 352 determines the distance D1 between the virtual sound source 102 and the first speaker 362, distance sensor 354 determines the distance D2 between the virtual sound source 102 and the second speaker 364, and distance sensor 356 determines the distance D3 between the virtual sound source 102 and the third speaker 366. Distance sensor 352 transmits the distance D1 to the computing device 110, distance sensor 354 transmits the distance D2 to the computing device 110, and distance sensor 356 transmits the distance D3 to the computing device 110. Then, the audio panning application 120 can determine a suitable gain value G1 for the first speaker 362 based on the distance D1, determine a suitable gain value G2 for the second speaker 364 based on the distance D2, and determine a suitable gain value G3 for the third speaker 366 based on the distance D3. Once the computing device 110 has determined the gain values G1, G2, and G3, the computing device 110 generates a first audio output signal 302 to the first speaker 362, a second audio output signal 304 to the second speaker 364, and a third audio output signal 306 to the third speaker 366. In some embodiments, the computing device 110 generates the first audio output signal 302, the second audio output signal 304, and the third audio output signal 306 based on the input audio signal 310 and specific gain values, where the input audio signal 310 represents the sound associated with the virtual sound source 102. In such embodiments, the computing device 110 generates the first audio output signal 302 by modifying the input audio signal 310 with the gain value G1, generates the second audio output signal 304 by modifying the input audio signal 310 with the gain value G2, and generates the third audio output signal 306 by modifying the input audio signal 310 with the gain value G3. Then, the computing device 110 transmits the first audio output signal 302 to the first speaker 362, transmits the second audio output signal 304 to the second speaker 364, and transmits the third audio output signal 306 to the third speaker 366. Then, the first speaker 362 generates an audio output (not shown) based on the first audio output signal 302. Similarly, the second speaker 364 generates an audio output (not shown) based on the second audio output signal 304, and the third speaker 366 generates an audio output (not shown) based on the third audio output signal 306. In this way, the audio from one or more virtual sound sources ( For example , the virtual sound source 102) is distributed to any number of physical speakers ( For example, the first speaker 362, the second speaker 364, and the third speaker 366), such that the localization of the virtual sound source within the listening area 200 is perceptually accurate.

[0034] Proximity-based Audio Object Panning Process

[0035] According to various embodiments, an audio system uses proximity-based audio object panning to generate a spatial audio scene. Specifically, in an audio system including multiple speakers and a virtual sound source, as the virtual sound source moves away from the first speaker and closer to the second speaker, the gain value of the audio input signal modified by the first speaker decreases, while the gain value of the audio input signal modified by the second speaker increases. The resulting effect is that the perceived sound source position (sometimes referred to as the "phantom panning") moves away from the first speaker and closer to the second speaker, which produces the desired result: when the virtual sound source moves within the listening area, the phantom panning follows the virtual sound source. Thus, when the virtual sound source moves between and / or around the speakers of the audio system, the audio system produces different audio outputs from each speaker, and the perceived position of the virtual sound source matches the physical position of the virtual sound source.

[0036] In some embodiments, the audio panning application 120 employs gain calculations to determine the gain value for each speaker. In such embodiments, the gain value for each speaker is based on the distance between the virtual sound source and the speaker, where the gain value of a particular speaker is functionally equivalent to the volume knob setting of that particular speaker. In some embodiments, it is assumed that the audio signal representing the sound generated by the virtual sound source remains approximately constant, and it is assumed that each speaker of the audio system has approximately the same sensitivity to changes in the gain value associated with that speaker ( For example , the change in output volume). In such embodiments, the sum of the squares of each gain value equals a constant value. Alternatively, when it is known that one or more speakers have different sensitivities, appropriate offset gain values can be applied to one or more speakers to compensate for the sensitivity differences to changes in the gain values determined by the audio panning application 120.

[0037] In some embodiments, the gain calculations employed by the audio panning application 120 are based on a simple panning algorithm to determine the gain value for each speaker. In such embodiments, the panning algorithm can be represented by a crossfade curve. For the sake of clarity in description, the use of a crossfade curve to determine the gain value is described herein with respect to an audio system including two speakers ( For example , Figure 2 the first speaker 262 and the second speaker 264 in For example , Figure 2 ), and a single virtual sound source ( Figure 4Describe an embodiment of such a cross-fade curve.

[0038] Figure 4 The cross-fade curve 400 for determining gain values according to various embodiments is shown. The cross-fade curve 400 includes a set of two gain curves that vary according to the position of the virtual sound source 102. Specifically, the cross-fade curve 400 includes a first gain curve 410 that represents a set of gain values for a first speaker ( For example , any one of speakers 262, 264, 266, 362, 364, or 366), and a second gain curve 420 (dashed line) that represents a set of gain values for a second speaker ( For example , the other of speakers 262, 264, 266, 362, 364, or 366). The left side of the cross-fade curve 400 indicates the gain values of the first and second speakers when the virtual sound source 102 is closer to the first speaker. Conversely, the left side of the cross-fade curve 400 indicates the gain values of the first and second speakers when the virtual sound source 102 is closer to the second speaker. Thus, the first gain curve 410 has a higher value on the left side of the cross-fade curve 400 and a lower value on the right side of the cross-fade curve 400, while the second gain curve 420 has a higher value on the right side of the cross-fade curve 400 and a lower value on the left side of the cross-fade curve 400.

[0039] Based on the distance D1 (between the virtual sound source 102 and the first speaker) and the distance D2 (between the virtual sound source 102 and the second speaker), the gain values of the first and second speakers are determined using the cross-fade curve 400. In some embodiments, the distance ratio of D1 and D2 is calculated and used as an input to the first gain curve 410 and the second gain curve 420. For example, in Figure 4 the embodiment shown, the distance ratio on the left side of the cross-fade curve 400 is the minimum value ( For example , 0.01), the distance ratio on the right side of the cross-fade curve 400 is the maximum value ( For example , 10), and the distance ratio at the center of the graph is 1. Additionally, in Figure 4 the embodiment shown, the possible gain values of the first and second speakers vary from 0 (occurring at the minimum distance ratio of the speakers) to 1 (occurring at the maximum distance ratio of the speakers). Thus, in such an embodiment, when the current position of the virtual sound source 102 is midway between the first and second speakers, the distance ratio of the current position of the virtual sound source 102 is 1 / 1, i.e., 1.0. As shown, for Figure 4An embodiment of the first gain curve 410 and the second gain curve 420 shown, when the current position of the virtual sound source 102 is in the middle of the first speaker and the second speaker and the distance ratio of the current position of the virtual sound source 102 is 1, the gain value indicated by the first gain curve 410 is 0.707, and the gain value indicated by the second gain curve 420 is 0.707. Therefore, in the above embodiment, the gain values of the first speaker and the second speaker can be determined based on the distance D1 and the distance D2 without performing two-dimensional or three-dimensional mapping of the relative positions of the first speaker, the second speaker, and the virtual sound source 102 within the listening area 200.

[0040] In the embodiment described in conjunction with Figure 4 the cross-fade curve 400 is implemented as two gain curves for the first speaker and the second speaker respectively, which vary linearly according to the distance ratio of the distance D1 and the distance D2. In other embodiments, the cross-fade curve 400 can be implemented as any other set of technically feasible gain curves for the first speaker and the second speaker. An example embodiment of such a gain curve is described below in conjunction with Figures 5A to 5E

[0041] Figures 5A to 5E Various cross-fade curves for determining gain values according to various other embodiments are shown. Figure 5A A cross-fade curve 510 including a first gain curve 512 and a second gain curve 514 is shown. As shown, in the first gain curve 512 and the second gain curve 514, the gain values of the first speaker and the second speaker vary according to the distance ratio of the distance D1 and the distance D2 such that the total sound power generated by the first speaker and the second speaker is constant.

[0042] Figure 5B A cross-fade curve 520 including a first gain curve 522 and a second gain curve 524 is shown. As shown, in the first gain curve 522 and the second gain curve 524, the gain values of the first speaker and the second speaker vary according to the distance ratio of the distance D1 and the distance D2 to produce a slow fade effect. Therefore, in the Figure 5B embodiment shown, the gain values of the first speaker and the second speaker change such that when the distance ratio reaches or exceeds the first value 526, the sound generated by the first speaker undergoes a slow fade, and when the distance ratio drops below the second value 528, the sound generated by the second speaker undergoes a slow fade.

[0043] Figure 5C ​Shows a cross-fade curve 530 including a first gain curve 532 and a second gain curve 534. As shown, in the first gain curve 532 and the second gain curve 534, the gain values of the first speaker and the second speaker vary according to the distance ratio of the distance D1 and the distance D2 to produce a slow cut-off effect. Therefore, in Figure 5C the embodiment shown, the gain values of the first speaker and the second speaker change such that when the distance ratio reaches or exceeds the first value 536, the sound generated by the first speaker will undergo a slow cut-off, and when the distance ratio drops below the second value 538, the sound generated by the second speaker will undergo a slow cut-off.

[0044] Figure 5D Shows a cross-fade curve 540 including a first gain curve 542 and a second gain curve 544. As shown, in the first gain curve 542 and the second gain curve 544, the gain values of the first speaker and the second speaker vary according to the distance ratio of the distance D1 and the distance D2 to produce a fast cut-off effect. Therefore, in Figure 5D the embodiment shown, the gain values of the first speaker and the second speaker change such that when the distance ratio reaches or exceeds the first value 546, the sound generated by the first speaker will undergo a fast cut-off, and when the distance ratio drops below the second value 548, the sound generated by the second speaker will undergo a fast cut-off.

[0045] Figure 5E Shows a cross-fade curve 550 including a first gain curve 552 and a second gain curve 554. As shown, in the first gain curve 552 and the second gain curve 554, the gain values of the first speaker and the second speaker vary according to the distance ratio of the distance D1 and the distance D2 to produce a transition effect. Therefore, in Figure 5E the embodiment shown, the gain values of the first speaker and the second speaker change such that when the distance ratio reaches or exceeds the transition value 556, the sound generated by the first speaker transitions to the sound generated by the second speaker, and when the distance ratio drops below the transition value 556, the sound generated by the second speaker transitions to the sound generated by the first speaker 264.

[0046] In the above embodiments, a relatively simple panning algorithm employing a cross-fade curve is used to describe the determination of the gain values for an audio system including two speakers. In other embodiments, the gain values for an audio system including any number of speakers can be determined. In such embodiments, a surround panning algorithm can be employed to describe the appropriate gain curve functions for each of the any number of speakers. The analytical expressions for the surround panning algorithm can be easily obtained from Figure 4 and Figures 5A to 5Ederived from the cross-fade curve, and those skilled in the art can implement these analytical expressions to implement such an embodiment for determining the gain value.

[0047] Figure 6 is a flowchart of method steps for generating a spatial audio scene according to various embodiments. Although the method steps are shown in sequence, those skilled in the art will understand that some method steps can be executed in a different order, repeated, omitted, and / or performed by Figure 6 components other than those described in Figures 1 to 5E the system of

[0048] As shown, method 600 begins at step 602, where the audio panning application 120 determines the distances between the speakers and the virtual sound source, such as distances D1, D2, and D3. Generally, different distances are determined for each speaker of the audio system 100. In some embodiments, one or more distance sensors 150 disposed within the virtual sound source 102 are used to determine distances D1, D2, and D3. Alternatively, in some embodiments, different distance sensors disposed within each speaker of the audio system 100 are used to determine distances D1, D2, and D3.

[0049] At step 604, the audio panning application 120 determines the distance ratio for the current position of the virtual sound source 102. For example, in some embodiments, the distance ratio for the current position of the virtual sound source 102 is the ratio of a first distance between a speaker of the audio system 100 and the virtual sound source 102 to a second distance between a second speaker of the audio system 100 and the virtual sound source 102.

[0050] In step 606, the audio panning application 120 determines the gain value for each speaker of the audio system 100. In embodiments where the audio system 100 includes two speakers, the gain value for each speaker can be determined based on the distance ratio determined in step 604. In embodiments where the audio system 100 includes three or more speakers, a more complex algorithm can be employed to determine the gain value rather than using the distance ratio. For example, in such embodiments, the gain value for each speaker can be determined based on distances D1, D2, and D3 and a suitable surround sound panning algorithm.

[0051] At step 608, the audio panning application 120 based on the gain value of the speaker and the input audio signal representing the sound associated with the virtual sound source 102 ( For example, the input audio signal 210) generates an audio signal for each speaker. Typically, the audio signal for a particular speaker is generated by modifying the input audio signal with a gain value associated with that particular speaker.

[0052] At step 610, the audio panning application 120 transmits the audio output signal for each speaker to the corresponding speaker. In some embodiments, the audio output signal is transmitted wirelessly, while in other embodiments, the audio output signal is transmitted via Bluetooth, WiFi, or any other technically feasible wireless protocol.

[0053] In summary, techniques for generating a spatial audio scene using proximity-based audio object panning are disclosed. In an embodiment, the spatial audio scene includes at least one virtual sound source, such as a toy or other object, which operates as a virtual sound source within the spatial audio scene. To generate the spatial audio scene, the audio signals representing the virtual sound sources are generated by the speakers of the audio system, which are physically separated from the virtual sound sources and thus not included within the virtual sound sources. In an embodiment, the volume of the audio signal generated by each speaker varies based on the position change of the virtual sound source relative to the speaker. In an embodiment, a proximity sensor can be used to generate the spatial audio scene, the proximity sensor measuring the distance from the virtual sound source to each speaker, wherein the gain value for each speaker is determined based on the measured distance from the virtual sound source to that speaker.

[0054] At least one technical advantage of the disclosed techniques over the prior art is that, using the disclosed techniques, a spatial audio scene including virtual sound sources such as toys or other objects can be created, wherein the audio signals representing the virtual sound sources are generated by speakers physically separated from the virtual sound sources. Using the disclosed techniques, a proximity sensor can be used to generate the spatial audio scene, the proximity sensor measuring the distance from the virtual sound source to each speaker that generates the audio signal representing the virtual sound source. Thus, compared to other immersive spatial audio methods, the disclosed techniques can generate an immersive spatial audio mix using fewer hardware components and reduced software processing. Another advantage of the disclosed techniques is that the techniques are flexible with respect to various characteristics of the sound system that generates the spatial audio scene, such as the number of speakers included in the sound system, the position of the speakers within the listening area, and the number and position of virtual sound sources within the spatial audio scene. Another advantage of the disclosed techniques is that the techniques are easily implemented using real-time audio processing, which is important for immersive applications that require fast processing and distribution of audio to maintain a certain audio experience.

[0055] Aspects of the present disclosure are also described in accordance with the following clauses.

[0056] 1. In some embodiments, a computer-implemented method for generating sound for a virtual sound source includes: determining a first distance between a first speaker and the virtual sound source; determining a second distance between a second speaker and the virtual sound source; generating a first audio output signal for the first speaker based on an input audio signal, the first distance, and the second distance; generating a second audio output signal for the second speaker based on the input audio signal, the first distance, and the second distance; transmitting the first audio output signal to the first speaker for output; and transmitting the second audio output signal to the second speaker for output.

[0057] 2. The computer-implemented method of clause 1, wherein generating the first audio output signal for the first speaker includes determining a first gain value for the first speaker based on the first distance and the second distance.

[0058] 3. The computer-implemented method of clause 1 or 2, wherein generating the first audio output signal for the first speaker further includes modifying the input audio signal using the first gain value to produce the first audio output signal.

[0059] 4. The computer-implemented method of any one of clauses 1 to 3, wherein determining the first gain value for the first speaker includes: calculating a distance ratio of the first distance and the second distance; and selecting the first gain value based on the distance ratio.

[0060] 5. The computer-implemented method of any one of clauses 1 to 4, wherein selecting the first gain value based on the distance ratio includes selecting the first gain value from a cross-fade curve.

[0061] 6. The computer-implemented method of any one of clauses 1 to 5, wherein the cross-fade curve includes one of the following: a set of two gain curves that vary linearly according to the distance ratio; a set of two gain curves that vary according to the distance ratio such that the total sound power produced by the first speaker and the second speaker is constant; a set of two gain curves that vary according to the distance ratio to produce a slow fade effect; a set of two gain curves that vary according to the distance ratio to produce a slow cut-off effect; a set of two gain curves that vary according to the distance ratio to produce a fast cut-off effect; or a set of two gain curves that vary according to the distance ratio to produce a transition effect.

[0062] 7. The computer-implemented method of any one of clauses 1 to 6, wherein generating the second audio output signal for the second speaker includes determining a second gain value for the second speaker based on the first distance and the second distance.

[0063] 8. A computer-implemented method as described in any one of clauses 1 to 7, further comprising: determining a third distance between the first speaker and the virtual sound source; determining a fourth distance between the second speaker and the virtual sound source; determining a third gain value for the first speaker based on the third distance and the fourth distance; and determining a fourth gain value for the second speaker based on the third distance and the fourth distance, wherein the sum of the squares of the first gain value and the second gain value is equal to a specific value, and the sum of the squares of the third gain value and the fourth gain value is equal to the specific value.

[0064] 9. A computer-implemented method as described in any one of clauses 1 to 8, wherein determining the first gain value for the first speaker comprises: calculating a distance ratio of the first distance and the second distance; and selecting the first gain value based on the distance ratio.

[0065] 10. A computer-implemented method as described in any one of clauses 1 to 9, wherein determining the first distance comprises receiving a distance from a distance sensor disposed within the virtual sound source.

[0066] 11. A computer-implemented method as described in any one of clauses 1 to 10, wherein determining the first distance comprises receiving a distance from a distance sensor disposed within the first speaker.

[0067] 12. A computer-implemented method as described in any one of clauses 1 to 11, wherein the virtual sound source comprises an interactive toy.

[0068] 13. A computer-implemented method as described in any one of clauses 1 to 12, wherein the input audio signal corresponds to the virtual sound source.

[0069] 14. In some embodiments, one or more non-transitory computer-readable media storing instructions that, when executed by one or more processors, cause the one or more processors to perform the following steps: determining a first distance between a first speaker and a virtual sound source; determining a second distance between a second speaker and the virtual sound source; generating a first audio output signal for the first speaker based on an input audio signal, the first distance, and the second distance; generating a second audio output signal for the second speaker based on the input audio signal, the first distance, and the second distance; transmitting the first audio output signal to the first speaker for output; and transmitting the second audio output signal to the second speaker for output.

[0070] 15. One or more non-transitory computer-readable media as described in clause 14, wherein generating the first audio output signal of the first speaker includes determining a first gain value of the first speaker based on the first distance and the second distance.

[0071] 16. One or more non-transitory computer-readable media as described in clause 14 or 15, wherein generating the first audio output signal of the first speaker further includes modifying the input audio signal using the first gain value to produce the first audio output signal.

[0072] 17. One or more non-transitory computer-readable media as described in any one of clauses 1 to 16, wherein determining the first gain value of the first speaker includes: calculating a distance ratio of the first distance and the second distance; and selecting the first gain value based on the distance ratio.

[0073] 18. One or more non-transitory computer-readable media as described in any one of clauses 1 to 17, wherein generating the second audio output signal of the second speaker includes determining a second gain value of the second speaker based on the first distance and the second distance.

[0074] 19. One or more non-transitory computer-readable media as described in any one of clauses 1 to 18, wherein determining the first distance includes receiving a distance from a distance sensor disposed within the first speaker.

[0075] 20. In some embodiments, a system includes: a first speaker; a second speaker; one or more distance sensors operable to determine a first distance between the first speaker and a virtual sound source and a second distance between the second speaker and the virtual sound source; a memory storing instructions; and one or more processors configured to perform the following steps when executing the instructions; determining the first distance; determining the second distance; generating a first audio output signal of the first speaker based on an input audio signal, the first distance, and the second distance; generating a second audio output signal of the second speaker based on the input audio signal, the first distance, and the second distance; transmitting the first audio output signal to the first speaker for output; and transmitting the second audio output signal to the second speaker for output.

[0076] The descriptions of the various embodiments have been presented for purposes of illustration, but these descriptions are not intended to be exhaustive or limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments.

[0077] Aspects of the present implementation can be embodied as a system, a method, or a computer program product. Accordingly, aspects of the present disclosure may take the form of an entirely hardware implementation, an entirely software implementation (including firmware, resident software, microcode, etc.), or an implementation combining software aspects with hardware aspects, all of which may generally be referred to herein as a "module" or "system". In addition, aspects of the present disclosure may take the form of a computer program product embodied in one or more computer-readable media having computer-readable program code embodied thereon.

[0078] Any combination of one or more computer-readable media may be utilized. The computer-readable media may be a computer-readable signal medium or a computer-readable storage medium. A computer-readable storage medium may be, for example but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer-readable storage medium would include the following media: an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing media. In the context of this document, a computer-readable storage medium may be any tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device.

[0079] Aspects of the present disclosure have been described above with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present disclosure. It should be understood that each block in the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions may be provided to a processor of a general purpose computer, a special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, enable the performance of the functions / acts specified in one or more blocks of the flowchart and / or block diagram. Such a processor may be, but is not limited to, a general purpose processor, a special purpose processor, an application specific processor, or a field programmable gate array or arrays.

[0080] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of the possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, segment, or portion of code that includes one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions noted in the blocks may occur out of the order noted in the accompanying drawings. For example, two blocks shown in succession may, in fact, be executed substantially concurrently, depending on the functionality involved, or may sometimes be executed in the reverse order. It should also be noted that each block of the block diagrams and / or flowchart illustrations, and combinations of blocks in the block diagrams and / or flowchart illustrations, can be implemented by special-purpose hardware-based systems that perform the specified functions or acts, or by combinations of special-purpose hardware and computer instructions.

[0081] While the foregoing is directed to embodiments of the present disclosure, other and further embodiments of the present disclosure may be devised without departing from the basic scope thereof, and the scope thereof is determined by the appended claims.

Claims

1. A computer-implemented method for generating sounds for a virtual sound source, the computer-implemented method comprising: determining a first distance between the first speaker and the virtual sound source; determining a second distance between a second speaker and the virtual sound source; generating a first audio output signal of the first speaker based on an input audio signal, the first distance, and the second distance; generating a second audio output signal of the second speaker based on the input audio signal, the first distance, and the second distance; transmitting the first audio output signal to the first speaker for output; as well as The second audio output signal is transmitted to the second speaker for output. 2 . The computer-implemented method of claim 1 , wherein generating the first audio output signal of the first speaker comprises determining a first gain value for the first speaker based on the first distance and the second distance. 3 . The computer-implemented method of claim 2 , wherein generating the first audio output signal for the first speaker further comprises modifying the input audio signal using the first gain value to produce the first audio output signal.

4. The computer-implemented method of claim 2, wherein determining the first gain value for the first speaker comprises: calculating a distance ratio between the first distance and the second distance; as well as The first gain value is selected based on the distance ratio. 5 . The computer-implemented method of claim 4 , wherein selecting the first gain value based on the distance ratio comprises selecting the first gain value from a cross-fade curve.

6. The computer-implemented method of claim 5, wherein the cross-fade curve comprises one of: a set of two gain curves that vary linearly according to the distance ratio; a set of two gain curves that vary according to the distance ratio so that the total sound power produced by the first speaker and the second speaker is constant; a set of two gain curves that vary according to the distance ratio to produce a slow fade effect; a set of two gain curves that vary according to the distance ratio to produce a slow cut-off effect; a set of two gain curves that vary according to the distance ratio to produce a fast cut-off effect; or a set of two gain curves that vary according to the distance ratio to produce a transition effect. 7 . The computer-implemented method of claim 2 , wherein generating the second audio output signal of the second speaker comprises determining a second gain value for the second speaker based on the first distance and the second distance.

8. The computer-implemented method of claim 7, further comprising: determining a third distance between the first speaker and the virtual sound source; determining a fourth distance between the second speaker and the virtual sound source; determining a third gain value of the first speaker based on the third distance and the fourth distance; as well as determining a fourth gain value of the second speaker based on the third distance and the fourth distance, The sum of the squares of the first gain value and the second gain value is equal to a specific value, and the sum of the squares of the third gain value and the fourth gain value is equal to the specific value.

9. The computer-implemented method of claim 7, wherein determining the first gain value for the first speaker comprises: calculating a distance ratio between the first distance and the second distance; as well as The first gain value is selected based on the distance ratio.

10. The computer-implemented method of claim 1, wherein determining the first distance comprises receiving the distance from a distance sensor disposed within the virtual sound source.

11. The computer-implemented method of claim 1 , wherein determining the first distance comprises receiving the distance from a distance sensor disposed within the first speaker.

12. The computer-implemented method of claim 1, wherein the virtual sound source comprises an interactive toy.

13. The computer-implemented method of claim 1, wherein the input audio signal corresponds to the virtual sound source.

14. One or more non-transitory computer-readable media storing instructions that, when executed by one or more processors, cause the one or more processors to perform the following steps: determining a first distance between the first speaker and the virtual sound source; determining a second distance between a second speaker and the virtual sound source; generating a first audio output signal of the first speaker based on an input audio signal, the first distance, and the second distance; generating a second audio output signal of the second speaker based on the input audio signal, the first distance, and the second distance; transmitting the first audio output signal to the first speaker for output; as well as The second audio output signal is transmitted to the second speaker for output. 15 . The one or more non-transitory computer-readable media of claim 14 , wherein generating the first audio output signal of the first speaker comprises determining a first gain value for the first speaker based on the first distance and the second distance.

16. The one or more non-transitory computer-readable media of claim 15, wherein generating the first audio output signal for the first speaker further comprises modifying the input audio signal using the first gain value to produce the first audio output signal.

17. The one or more non-transitory computer-readable media of claim 15, wherein determining the first gain value for the first speaker comprises: calculating a distance ratio between the first distance and the second distance; as well as The first gain value is selected based on the distance ratio.

18. The one or more non-transitory computer-readable media of claim 17, wherein generating the second audio output signal of the second speaker comprises determining a second gain value for the second speaker based on the first distance and the second distance.

19. The one or more non-transitory computer-readable media of claim 14, wherein determining the first distance comprises receiving the distance from a distance sensor disposed within the first speaker.

20. A system comprising: First speaker; Second speaker; one or more distance sensors operable to determine a first distance between a first speaker and a virtual sound source and a second distance between a second speaker and the virtual sound source; a memory storing instructions; as well as One or more processors, wherein the one or more processors are configured to perform the following steps when executing the instructions: determining the first distance; determining the second distance; generating a first audio output signal of the first speaker based on an input audio signal, the first distance, and the second distance; generating a second audio output signal of the second speaker based on the input audio signal, the first distance, and the second distance; transmitting the first audio output signal to the first speaker for output; as well as The second audio output signal is transmitted to the second speaker for output.