Dual-purpose sound system with audio rendering and ultrasonic beacon signaling
By combining an audio speaker system with an ultrasonic beacon system, a combined audio/ultrasonic signal is generated for positioning, solving the problem of high cost of ultrasonic positioning systems and achieving the dual functions of high-precision object tracking and positioning and audio playback.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- NOKIA NETWORKS OY
- Filing Date
- 2025-10-10
- Publication Date
- 2026-04-10
AI Technical Summary
Existing ultrasonic positioning systems have high infrastructure costs and are complex to install and maintain for indoor positioning.
By combining an audio speaker system with an ultrasonic beacon system, the speaker emits audible sound waves and ultrasonic signals. The audio and ultrasonic signals are then mixed by a dual-audio system to generate a combined audio/ultrasonic signal for positioning.
It enables both audio signal playback and high-precision object tracking and positioning without increasing hardware costs, reducing the installation and maintenance costs of dedicated beacon systems.
Smart Images

Figure CN121831685A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the field of indoor / local positioning, and more particularly, to ultrasonic or acoustic positioning. BACKGROUND
[0002] Indoor or local positioning systems for tasks such as robot positioning, object navigation, inventory tracking, etc. are expected to play an increasing role in residential and commercial buildings. These tasks can be performed with an ultrasonic positioning system that is analogous to an indoor global positioning system (GPS) that facilitates positioning and tracking with high accuracy. One problem with ultrasonic positioning systems is the cost of the infrastructure. SUMMARY
[0003] A dual-purpose sound system and associated methods of implementing the dual-purpose sound system are described herein. The dual-purpose sound system combines an audio or sound system with an ultrasonic beacon system. The sound system includes a plurality of speakers that are configured to emit audible sounds of music, voice (e.g., paging or announcements), etc. For example, a residence can have a smart speaker, a stereo system, or a surround sound system (e.g., 5.1 surround sound, 5.1.2 surround sound, 5.1.4 surround sound, Dolby Atmos®, etc.). Commercial buildings such as retail stores, malls, hotels, distribution centers, warehouses, manufacturing plants, etc. can have speakers installed, such as on a ceiling, to play music, voice announcements, etc. The dual-purpose sound system described herein uses the speakers as an ultrasonic beacon system for positioning or object tracking. Thus, the dual-purpose sound system acts as a two-in-one system, i.e., an audio system for music / voice and an ultrasonic beacon system that emits beacon signaling for positioning or object tracking. One technical benefit is the avoidance of the costs, setup, and maintenance associated with a separate beacon system.
[0004] In one embodiment (also referred to as an aspect), an apparatus includes a first stage including a master volume controller configured to perform volume control of audio signals of audio channels of a sound system to generate variable volume audio signals of the audio channels, and a second stage including an audio mixer configured to mix the variable volume audio signals of a target set of the audio channels with fixed volume ultrasonic signals for positioning to generate at least two combined audio / ultrasonic signals configured for output to speakers of the sound system corresponding to the target set of the audio channels.
[0005] In one embodiment, the audio mixer of the second stage includes a second audio mixer, and the first stage further includes a first audio mixer configured to mix the audio signals to generate mixed audio signals provided to the master volume controller.
[0006] Other embodiments can include computer readable media, other systems or apparatus, or other methods or means for performing the acts described herein. Additionally, one or more embodiments as described above can be combined with one or more of the other embodiments described herein.
[0007] The above summary presents a basic understanding of some aspects of the specification. This summary is not an extensive overview of the specification. It is intended to neither identify key or critical elements of the specification nor delineate any of the specific embodiments of the specification, or any scope of the claims. Its sole purpose is to present some concepts of the specification in a simplified form as a prelude to the more detailed description that is presented later. BRIEF DESCRIPTION OF DRAWINGS
[0008] Some embodiments of the present application will now be described by way of example only, and with reference to the accompanying drawings. In all the drawings, like reference numerals refer to like elements or elements of the same type.
[0009] Figure 1 An ultrasonic positioning system in an illustrative embodiment is illustrated.
[0010] Figure 2 A dual use sound system in an illustrative embodiment is illustrated.
[0011] Figures 3A-3D The structure of a speaker system in some examples is illustrated.
[0012] Figure 4 5.1 surround sound is illustrated as an example.
[0013] Figure 5A is a block diagram illustrating a dual audio system in an illustrative embodiment.
[0014] Figure 5B is a block diagram of a first stage of a dual audio system in an illustrative embodiment.
[0015] Figure 6 is a flow diagram illustrating a method of operating a dual audio system in an illustrative embodiment.
[0016] Figure 7 A dual audio system in operation in an illustrative embodiment is illustrated.
[0017] Figure 8 is a block diagram illustrating a dual audio system in an illustrative embodiment.
[0018] Figure 9 is a flow diagram illustrating a method of operating a dual audio system in an illustrative embodiment.
[0019] Figure 10 is a block diagram illustrating a dual audio system in an illustrative embodiment.
[0020] Figure 11 is a block diagram illustrating a dual audio system in an illustrative embodiment.
[0021] Figure 12 is a block diagram illustrating a dual audio system in an illustrative embodiment.
[0022] Figure 13 is a block diagram illustrating a dual audio system in an illustrative embodiment.
[0023] Figure 14 illustrates an upward-firing speaker in an illustrative embodiment.
[0024] Figure 15 illustrates an upward-firing speaker and a forward-firing speaker in an illustrative embodiment.
[0025] Figure 16 illustrates a dual audio system of a distributed sound system in an illustrative embodiment.
[0026] Figure 17 illustrates a user application for controlling a dual-purpose sound system in an illustrative embodiment.
[0027] Figures 18A-18B illustrates how audio signals and ultrasonic signals are frequency multiplexed in an illustrative embodiment.
[0028] Figure 19 is a flowchart illustrating a method of mixing audio signals and ultrasonic signals in an illustrative embodiment. DETAILED DESCRIPTION
[0029] The accompanying drawings and following description are illustrative of specific examples. It should be understood, therefore, that the application is not limited to the specific details, examples, and conditions described and / or illustrated herein, but rather, the intention is to cover all modifications, equivalents, and alternatives falling within the scope and spirit of the application. Any example described herein is intended to be illustrative of the principles of the application and should not be construed as limiting the application to only the specifically enumerated examples and conditions. Consequently, the application concept(s) is (are) limited only by the claims, and equivalents thereof, as they can vary from time to time.
[0030] Figure 1An ultrasonic positioning system 100 in an illustrative embodiment is shown. The ultrasonic positioning system 100 is configured to track or determine a position of a mobile acoustic receiver 102. The ultrasonic positioning system 100 uses a beacon system 106 that includes a plurality of fixed transmitters 104, each configured to emit or issue ultrasonic signals or pulses, referred to as beacons. To track the position, the acoustic receiver 102 includes a microphone or the like that receives the ultrasonic signals. For example, the acoustic receiver 102 can include an acoustic tag (e.g., a flat tag) that includes a circuit with an attached microphone or microphone array configured to receive the ultrasonic signals. The acoustic receiver 102 or a positioning server 130 communicatively coupled to the acoustic receiver 102 is configured to compute measurements of the ultrasonic signals (e.g., time of arrival (ToA), angle of arrival (AoA), etc.) to estimate a distance between the acoustic receiver 102 and one or more transmitters 104. In Figure 1 an example, the ultrasonic positioning system 100 is configured or implemented within an enclosed or indoor space 150, such as within a (residential or commercial) building 152 or some other enclosed structure (e.g., a manufacturing facility, a warehouse, a distribution center, etc.). However, the ultrasonic positioning system 100 can also be configured or implemented in an outdoor space. The transmitters 104 (also referred to as transmitter nodes) are installed at fixed known positions within the indoor space 150. For example, the transmitters 104 can be installed at or towards a ceiling 154 of the indoor space 150, or near another boundary within the indoor space 150. Any other beacon positions can also work as long as the geometric accuracy of the AoA or ToA is sufficient at the mobile location. Each transmitter 104 issues ultrasonic signals / pulses / beacons in the ultrasonic spectrum. The acoustic receiver 102 (also referred to as a receiver / acoustic node, an acoustic tracker tag, an ultrasonic tag, etc.) is associated with a target or tracked object 110 that moves or is in a non-fixed position within the indoor space 150. The acoustic receiver 102 can be carried by or disposed in the tracked object 110 to be positioned, or can have a predefined positional relationship with the tracked object 110. As such, by determining the position of the acoustic receiver 102, the position of the tracked object 110 can be determined accordingly, such as by co-locating or being in a predefined positional relationship with respect to the acoustic receiver 102. For example, the acoustic receiver 102 can be mounted on a mobile device or mobile robot 112, as shown in Figure 1 However, the acoustic receiver 102 can be worn or held by a human user (e.g., a smartphone, a smartwatch, a wearable tracker, etc.), attached to an inventory item, or otherwise attached to the tracked object 110.
[0031] In operation, transmitter 104 emits ultrasonic waves 120 (also referred to as ultrasonic beacon signals or ultrasonic beacons) within indoor space 150, and acoustic receiver 102 performs measurements based on the received ultrasonic waves 120. Acoustic receiver 102 may report measurements 122 to a centralized system / server (e.g., via a wireless interface), such as positioning server 130. Positioning server 130 may calculate or determine the location information of acoustic receiver 102 based on the measurements, such as performing multilateral positioning by using an estimated distance to transmitter 104. Positioning server 130 can thus track the orientation or location of acoustic receiver 102 (e.g., in the coordinate system of indoor space 150). It should be noted that in some embodiments, acoustic receiver 102 may process measurements 122 to calculate or determine location information and report the location information to positioning server 130. Positioning server 130 may also provide path planning or driving direction to mobile robot 112 (e.g., based on the location of mobile robot 112). In other cases, a server is not required (e.g., when only mobile robot 112 uses the location information).
[0032] like Figure 1 The ultrasonic positioning system 100 shown facilitates high-precision positioning and tracking. A potential problem with ultrasonic positioning systems is generally the cost of infrastructure, installation, etc. For example, it may be necessary to purchase, install, and calibrate a dedicated transmitter 104 to implement some ultrasonic positioning systems.
[0033] In the embodiments described herein, an ultrasonic beacon system for positioning is implemented using a sound system comprising multiple speakers. Figure 2 A dual-purpose sound system 200 is illustrated in an illustrative embodiment. The dual-purpose sound system 200 is a system that combines features of an audio / sound system with features of an ultrasonic beacon system. The dual-purpose sound system 200 can be configured or implemented within an indoor space 150, such as in a (residential or commercial) building 152 or some other enclosed or partially enclosed structure. The dual-purpose sound system 200 includes a plurality of loudspeakers 204 mounted in fixed, known locations (such as within indoor space 150). Loudspeakers 204 are devices configured to convert electrical signals into sound waves (i.e., acoustic energy) and are also referred to herein as loudspeakers. Although... Figure 2 The dual-purpose sound system 200 includes four speakers 204, but the dual-purpose sound system 200 may include more or fewer speakers 204, such as N The number of speakers 204 can depend on the environment in which the dual-purpose sound system 200 is implemented. Regarding the room space 150, the number of speakers 204 can depend on the size of the room space 150, where larger rooms sometimes use more speakers 204 than smaller rooms.
[0034] The loudspeakers 204 can be positioned at any of a variety of orientations within the indoor space 150. For example, a residence can have a smart speaker, a stereo system, or a surround sound system (e.g., 5.1 surround sound, 5.1.2 surround sound, 5.1.4 surround sound, Dolby Atmos, etc.). Generally, these loudspeakers 204 can be placed on a shelf, on a stand, on the floor, attached to a wall, installed in or on a ceiling, etc. Commercial buildings, such as retail stores, malls, hotels, distribution centers, warehouses, manufacturing facilities, etc., can have loudspeakers installed, such as on a ceiling. In one embodiment, the loudspeakers 204 can be placed near, such as adjacent to, a boundary 260 (e.g., a wall, a ceiling, a floor, etc.) of the indoor space 150. The positions of the loudspeakers 204 are known a priori, that is, the orientations of the loudspeakers 204 are predefined such that the orientations of the acoustic wireless receivers 102 can be determined using the orientations of the loudspeakers 204, and in turn, the orientations of the tracked objects 110 to be located.
[0035] In one embodiment, at least a subset (e.g., two or more) of the loudspeakers 204 are configured to emit audible sound waves 220 and ultrasonic sound waves 120. The audible sound waves 220 are sound waves having frequencies in the audible spectrum (e.g., about 20 Hz to 20 kHz) that are detectable by a human 212. For example, the loudspeakers 204 can emit audible sound waves 220 for (background) music, paging messages / notifications, etc. The loudspeakers 204 can emit ultrasonic sound waves 120 such that the loudspeakers 204 act as emitters for localization. Thus, the dual-purpose sound system 200 acts as an audio / sound system and an ultrasonic beacon system to track or determine the orientations of the acoustic receivers 102. As described above, the acoustic receivers 102 are associated with a target or tracked object 110 that is mobile or in a non-fixed position, such as within the indoor space 150. The acoustic receivers 102 can be carried by or disposed in the tracked object 110 to be located, or can have a predefined positional relationship with the tracked object 110. As such, by determining the orientations of the acoustic receivers 102, the orientations of the tracked object 110 can be determined accordingly, such as by co-locating or being in a predefined positional relationship with respect to the acoustic receivers 102.
[0036] The dual audio system 202 is communicatively coupled to the loudspeakers 204, such as through a wired or wireless connection. The dual audio system 202 is configured to provide drive or output signals 222 to the loudspeakers 204. As will be described in further detail below, the output signal(s) 222 generated by the dual audio system 202 include a combined audio / ultrasonic signal 224. The combined audio / ultrasonic signal 224 includes an ultrasonic signal mixed with an audio signal. The audio signal is an electrical representation of audible sound waves that can be captured, transmitted, stored, and / or processed by an electronic device. The ultrasonic signal is an electrical representation of ultrasonic sound waves that can be captured, transmitted, stored, and / or processed by an electronic device.
[0037] In operation, the dual audio system 202 provides the combined audio / ultrasonic signal 224 to the two or more loudspeakers 204. The loudspeakers 204 that receive the combined audio / ultrasonic signal 224 from the dual audio system 202 emit ultrasonic sound waves 120 in response to the ultrasonic signal of the combined audio / ultrasonic signal 224 and audible sound waves 220 in response to the audio signal of the combined audio / ultrasonic signal 224. The ultrasonic sound waves 120 are used for localization of the tracked object 110. As described above, the acoustic receiver 102 performs measurements based on the received ultrasonic sound waves 120 to determine or compute location information of the acoustic receiver 102. At the same time, the audible sound waves 220 are used to play music, voice notifications (e.g., pages), etc., which can be heard by the person 212 located within the indoor space 150. One technical benefit is that the dual-purpose sound system 200 acts as a two-in-one system, i.e., an audio system for music / voice and an ultrasonic beacon system for object tracking. The dual-purpose sound system 200 can replace or augment existing art audio or sound systems, such as a distributed sound system for background music, a paging system, a surround sound system, a single-unit Bluetooth speaker or a smart speaker, etc., by simultaneously adding ultrasonic signals to the audio signals output to the loudspeakers 204. With such a combined system, any tracked object 110 can be tracked while music or other audible sound is played. The acoustic receiver 102 can also be connected to a server in the ultrasonic localization system 100 in order to offload the computation of locations, receive path planning information or driving directions, or receive any other control information.
[0038] Figures 3A-3D FIG. 1 illustrates the structure of a loudspeaker system in some examples. In Figure 3A The loudspeaker system 300 includes a forward-firing loudspeaker 204-1 that includes a loudspeaker driver 302 disposed within an enclosure 304. The loudspeaker driver 302 is an electro-acoustic transducer that converts an audio signal into a corresponding sound. In Figure 3BIn some embodiments, the speaker system 300 includes a pair of speakers 204, such as a forward- firing speaker 204-2 and a tweeter 204-3 disposed within the enclosure 304. Figure 3C In some embodiments, the speaker system 300 includes a pair of speakers 204, such as a forward- firing speaker 204-4 and an upward-firing speaker 204-5 disposed within the enclosure 304. Figure 3D In some embodiments, the speaker system 300 includes four speakers 204, such as a forward- firing speaker 204-6, a left speaker 204-7, a right speaker 204-8, and a surround speaker 204-9 disposed within the enclosure 304. Figures 3A-3D The speaker system 300 in FIGS. 1-3 is provided as an example, and other speaker systems are also contemplated herein.
[0039] Generally, a sound system (e.g., the dual-purpose sound system 200) is configured with multiple (e.g., two) audio channels. For example, mono sound has one audio channel, while stereo sound has two audio channels (left and right channels). 5.1 surround sound has six channels, including a center channel (CNT), a front left (FL) channel, a front right (FR) channel, two surround channels (left surround (SL) and right surround (SR)), and a low frequency effects (LFE) channel designed for a subwoofer. 5.1.2 surround sound has eight channels, including channels similar to 5.1 surround sound, plus a high front left (HFL) channel and a high front right (HFR) channel. N
[0040] Figure 4 FIG. 4 illustrates 5.1 surround sound 400 as an example. A sound system, such as a 5.1 surround sound system, is configured with audio channels 402. The audio channels 402 are representations of sound. In the 5.1 surround sound 400, the audio channels 402 include a center channel 402-1 (CNT), a front left (FL) channel 402-2, a front right (FR) channel 402-3, a left surround (SL) channel 402-4, a right surround (SR) channel 402-5, and an LFE channel 402-6.
[0041] In embodiments described herein, the dual-audio system 202 is configured to superimpose an ultrasonic signal / channel on the multiple audio channels 402 of the sound system. Figure 5A is a block diagram illustrating a dual audio system 202 in an illustrative embodiment. The dual audio system 202 is an apparatus, data processing element, circuitry, etc. configured to generate a combined audio / ultrasonic wave signal 224 provided to a set of loudspeakers 204. The dual audio system 202 includes a two-stage mixer 506 configured to combine a variable volume audio signal with a constant or fixed volume ultrasonic (beacon) signal. In a first stage 501 (i.e., stage-1), the two-stage mixer 506 is configured to render and / or mix the audio signal, and in a second stage 502 (i.e., stage-2), the two-stage mixer 506 is configured to mix the ultrasonic signal with the audio signal. In the two-stage mixer 506, a master volume control is placed between the audio signal mixer and the ultrasonic signal mixer. One technical advantage is that a variable volume audio signal can be mixed with a constant level ultrasonic signal. In other words, the ultrasonic signal level is not affected by the audio signal volume variations. In this way, a constant and sufficiently high beacon signal level ensures reliable tracking even if the tracked object 110 is far away from the loudspeakers 204 (e.g., 30 meters).
[0042] The first stage 501 of the two-stage mixer 506 includes an audio mixer 512 (also referred to as a first audio mixer, a first stage audio mixer, etc.) that is a system, apparatus, circuitry, equipment, etc. configured to mix an audio signal 520 to generate a mixed audio signal 522 for an audio channel 402 of a sound system. The first stage 501 also includes a volume controller 514 (also referred to as a master volume controller) that is a system, apparatus, circuitry, equipment, etc. configured to adjust or set the volume of an audio signal (e.g., the mixed audio signal 522) to generate a variable volume audio signal 524 for the audio channel 402. The volume controller 514 can individually control / adjust the volume of the N audio channel 402 as needed. The second stage 502 includes an audio mixer 516 (also referred to as a second audio mixer, a second stage audio mixer, etc.) that is a system, apparatus, circuitry, equipment, etc. configured to mix the variable volume audio signals 524 of a target set of audio channels 402 with a fixed volume ultrasonic signal 526 (also referred to as an ultrasonic beacon signal) to generate a combined audio / ultrasonic wave signal 224.
[0043] Figure 5Bis a block diagram of a first stage 501 of the dual-audio system 202 in an illustrative embodiment. In one embodiment, the audio mixer 512 can include a renderer 518 and a signal type audio mixer 519. The renderer 518 is configured to render source audio signals 520 to generate rendered audio signals 521. An encoded audio stream (i.e., an audio bitstream) received by the dual-audio system 202 can contain various signal types, such as channel signals, object signals, or higher-order ambisonic (HOA) coefficient signals. The encoded audio stream can also contain an arbitrary combination of signal types. Accordingly, the renderer 518 can perform a combination of, for example, channel signal rendering (e.g., upmixing or downmixing), object signal rendering, or HOA coefficient signal rendering. When the encoded audio stream contains multiple signal types, the signal type audio mixer 519 mixes the rendered audio signals 521 to generate mixed audio signals 522. The encoded audio stream can contain channel signals, object signals, HOA coefficient signals, or any other encoded audio signals. Various audio encoding standards support channel signals, object signals, and HOA coefficient signals, such as MPEG-H 3D Audio, developed by ISO / IEC Moving Picture Experts Group (MPEG), designated as ISO / IEC 23008-3, and MPEG-I Immersive Audio, designated as ISO / IEC 23090-4.
[0044] In Figure 5A , one or more subsystems of the dual-audio system 202 can be implemented on a hardware platform composed of analog and / or digital circuitry. One or more subsystems of the dual-audio system 202 can be implemented on a processor 530 that executes instructions 534 stored in a memory 532. The processor 530 includes integrated hardware circuitry configured to execute the instructions 534 to provide the functionality of the dual-audio system 202. Depending on the particular implementation, the processor 530 can include a set of one or more processors, or can include multiple processor cores. The memory 532 is a non-transitory computer-readable medium for data, instructions, applications, etc., and is accessible by the processor 530. The memory 532 is a hardware storage device capable of temporarily and / or permanently storing information. The memory 532 can include random access memory or any other volatile or non-volatile storage device.
[0045] The dual-audio system 202 can include various other components not specifically illustrated in Figure 5A , such as additional audio mixers, amplifiers, digital-to-analog (D / A) converters, analog-to-digital (A / D) converters, etc.
[0046] Figure 6 is a flowchart illustrating a method 600 of operating the dual-audio system 202 in an illustrative embodiment. Reference will be made to Figure 5AThe steps of the method 600 are described with respect to the dual audio system 202 in FIG. 2, but those skilled in the art will appreciate that the method 600 can be performed in other systems or devices. Moreover, the steps of the flowcharts described herein are not all inclusive and can include other steps not shown and can be performed in alternative orders.
[0047] In the first stage 501, the audio mixer 512 can render and / or mix the audio signals 520 to generate mixed audio signals 522 for the audio channels 402 of the sound system 200 (step 602). However, it should be noted that rendering / mixing at the first stage 501 can not be required depending on the content of the input audio streams. The volume controller 514 performs volume control of the audio signals (e.g., the mixed audio signals 522) to generate variable volume audio signals 524 for the audio channels 402 (step 604). Thus, the volume level of the audio signals for each audio channel 402 is set in the first stage 501 before mixing with any ultrasonic signals 526. In the second stage 502, the audio mixer 516 mixes, superimposes, or frequency multiplexes the variable volume audio signals 524 of the target set of audio channels 402 (i.e., two or more) with the fixed volume ultrasonic signals 526 to generate combined audio / ultrasonic signals 224 (step 606). The combined audio / ultrasonic signals 224 (i.e., two or more) are configured for output to the speakers 204 of the sound system 200 corresponding to the target set of audio channels 402. Thus, the dual audio system 202 provides the combined audio / ultrasonic signals 224 to the speakers 204 corresponding to the target set of audio channels 402. The dual audio system 202 can also provide output signals 222 (audio only) to one or more speakers 204 that are not within the target set of audio channels 402. One technical benefit is that the variable volume audio signals 524 are mixed with the fixed volume ultrasonic signals 526 to provide a dual purpose sound system.
[0048] Figure 7 A dual audio system 202 operating in an illustrative embodiment is illustrated. In this example, the dual audio system 202 includes a set of speakers 204 mounted in a building ceiling. A target set 702 of audio channels 402 is designated for localization in the dual audio system 202. Thus, the speakers 204 receiving the combined audio / ultrasonic signals 224 will emit audible sound waves 220 (indicated with “A”) detectable by a person 212, such as music, speech, etc., and will also emit ultrasonic waves 120 (indicated with “U”) that can be used to localize one or more tracked objects 110. One technical benefit is that the dual audio system 202 allows the sound system to operate as a dual purpose sound system 200. Thus, there is no need to install dedicated emitters 104, which can reduce the material and / or installation costs to implement localization, such as within an indoor space.
[0049] Traditionally, two separate systems are used for object tracking and audio playback, one system renders the music / voice signals and the other system renders the beacon signals. One reason for the two separate systems is that volume control is essential in the sound system for music / voice playback, but is detrimental in the beacon system, as the beacon system needs to provide a constant beacon signal level at all times when tracking or locating an object. A sound system with master volume control would affect all input signals, whether audio signals or beacon signals. In other words, when the volume changes, the beacon signal level also changes. However, to ensure reliable tracking, it is important that the beacon signal remains constant and relatively high at all times when tracking or locating an object. Furthermore, rendering audible signals and inaudible ultrasonic beacon signals on a single system is a challenge due to the frequency band limitations in electronics and loudspeakers.
[0050] The dual audio system 202 described herein effectively fuses the sound system with the ultrasonic beacon system. To this end, volume control is performed on the audio signals 522 before mixing with the ultrasonic signals 526. Therefore, any adjustment to the volume of the audio signals 522 does not affect the volume level of the ultrasonic signals 526. This enables variable volume audio signals to be mixed with constant level beacon signals. In other words, the beacon signal level is not affected by any volume changes of the audio signals. In this way, a constant and sufficiently high beacon signal level ensures reliable tracking even when the tracked object 110 is far away from the loudspeakers 204 (e.g. 30 meters or more).
[0051] Certain embodiments describing the configuration of the dual audio system 202 are provided below. The processes, systems and methods described in the following embodiments can be incorporated into the above-described embodiments as desired. In an embodiment, K is the number of audio channels of the encoded audio stream, N is the number of audio channels of the sound system, and M is the desired number of ultrasonic signals / channels. With N audio channels, the dual-purpose sound system 200 can support up to N ultrasonic signals, that is, the number of ultrasonic signals is limited by M ≤ N .
[0052] In an embodiment, a two-stage mixer concept is utilized in the dual audio system 202, which has a first stage audio mixer 512 that maps K source audio signals to N system audio signals 524, and a second stage audio mixer 514 that maps M ultrasonic signals 526 toN a second stage audio mixer 516 on the variable volume audio signals 524, where a volume control is placed between the first stage audio mixer 512 and the second stage audio mixer 516. Figure 8 is a block diagram illustrating the dual audio system 202 in the illustrative embodiment. In this embodiment, the dual audio system 202 includes a first stage 501 (i.e., first stage) and a second stage 502 (i.e., second stage) for generating a combined audio / ultrasonic signal 224. The first stage 501 includes a volume controller 514, where an audio decoder 802 and a first stage audio mixer 512 are disposed upstream of the volume controller 514. The audio decoder 802 is configured to decode an encoded audio stream 832 into a plurality of audio signals, referred to as source audio signals 520 (e.g., D1, D2, …, D K ). A renderer 518 is configured to render the source audio signals 520 to generate rendered audio signals 521 (e.g., R1, R2, …, R N ). For example, the renderer 518 can render the source audio signals 520 using metadata 840, where the metadata 840 can include information 842 about a listener’s position and / or orientation, position information 844 of the loudspeakers 204, etc. A signal type audio mixer 519 is configured to mix the rendered audio signals 521 to generate mixed audio signals 522 (e.g., P1, P2, …, P N ). For example, the signal type audio mixer 519 is configured to mix channel-based signals, object-based signals, HOA-based signals, etc. The volume controller 514 is configured to adjust or set the volume of the mixed audio signals 522 to generate variable volume audio signals 524 (e.g., A1, A2, …, A N ).
[0053] The second stage 502 includes a second stage audio mixer 516, where a D / A converter 808 and an amplifier 810 are disposed downstream of the second stage audio mixer 516. The second stage audio mixer 516 is configured to mix the plurality of variable volume audio signals 524 with fixed volume ultrasonic signals 526 (e.g., U1, U2, …, U M ) to generate digital output signals 826 (e.g., B1, B2, …, B N ). The D / A converter 808 is configured to convert the digital output signals 826 into analog output signals 828 (e.g., C1, C2, …, C N ). The amplifier 810 is configured to amplify the analog output signals 828 to generate amplified output signals 830 (e.g., L1, L2, …, L N). The set of amplified output signals 830 mixed with the fixed volume ultrasonic signals 526 is an example of the above-described combined audio / ultrasonic signals 224. The dual audio system 202 is configured to provide or supply the amplified output signals 830 (including the combined audio / ultrasonic signals 224) to a set of speakers 204 (e.g., S1, S2,..., S N ) that receive the combined audio / ultrasonic signals 224 will emit both signals simultaneously, one for the audio signal for the listener (i.e., the person 212) and one for the ultrasonic beacon signal for the tracked object 110. The listener can hear the audio signal, but not the ultrasonic beacon signal, while the acoustic receiver 102 can "hear" the ultrasonic beacon signal and use it for tracking or localization, and can also "hear" the audio signal and use it for equalization (of the speakers 204 or speech), and / or for speech communication (e.g., with a chatbot or a remote conferencing party), or for other interactions with the user. In one embodiment, each speaker 204 emits a different set of signals to enable the listener to hear spatial audio, and for the acoustic receiver 102 to calculate its position and to perform equalization.
[0054] It should be noted that while Figure 8 A digital audio input (audio bitstream) is shown, analog input signals can be processed with additional analog-to-digital (A / D) converters. Further, when processing analog audio inputs, the basic concepts of the dual audio system 202 can be implemented in analog hardware, eliminating the need for A / D and D / A conversion.
[0055] Figure 9 is a flowchart illustrating a method 900 of operating a dual audio system 202 in an illustrative embodiment. The steps of the method 900 will be described with reference to the dual audio system 202 in Figure 8 , but those skilled in the art will appreciate that the method 900 can be performed in other systems or devices.
[0056] In the first stage of the method 900, the audio decoder 802 receives an encoded audio stream 832 (step 902) and extracts the K source audio signals 520 contained in the encoded audio stream 832 (step 904). K The source audio signals 520 are then processed to accommodate the number of audio channels 402 of the dual-purpose sound system 200 (step 906). NThis task is performed by the first-stage audio mixer 512. The renderer 518 renders the source audio signal 520 to generate the rendered audio signal 521 (step 906). The encoded audio stream 832 (i.e., the audio bitstream) can use channel encoding, object encoding, HOA encoding, or another encoding scheme or channel signal type. Therefore, the renderer 518 can use the signal... S Additional auxiliary information or metadata 840 (e.g., location information) provided in the middle are used to perform channel signal rendering (e.g., K (520 source audio signals up-mixing or down-mixing), object signal rendering, or HOA coefficient signal rendering. For channel-based audio, when N < K hour, K The source audio signal 520 is downmixed into the dual-purpose sound system 200. N 402 audio channels. When N > K hour, K The source audio signal 520 is upmixed into the dual-purpose sound system 200. N 402 audio channels. When N = K At that time, no processing is required, and K The source audio signal 520 can be used as the audio signal 521 for rendering. The renderer 518 can use the speaker configuration (i.e., speaker position 844, and possible listener positions) to render. K A source audio signal 520. An audio mixer 519 mixes the rendered audio signal 521 to generate a mixed audio signal 522 (step 908). The encoded audio stream 832 can contain multiple audio streams, such as... Figure 8 "in j As indicated. The encoded audio stream 832 may contain channel signals, object signals, and HOA coefficient signals, which can be summed in the signal type audio mixer 519 to generate a mixed audio signal 522.
[0057] Following the first-stage audio mixer 512 but before the second-stage audio mixer 516, the volume controller 514 adjusts, controls, or sets the volume of the mixed audio signal 522 to generate a variable-volume audio signal 524 (step 910). This allows the user to set the volume of the dual-purpose sound system 200 (i.e., the volume level of music, voice notifications, etc.).
[0058] In the second stage of method 900, a second-stage audio mixer 516 mixes multiple variable-volume audio signals 524 with a fixed-volume ultrasonic signal 526 to generate a digital output signal 826 (step 912). The volume level of the ultrasonic signal 526 is preset and remains fixed, such as in the range of approximately -12 dB (where 0 dB is the maximum level) or any other level, depending on variables such as amplifier output power, speaker efficiency, and the expected sound pressure level at the listener's and receiver's locations. When setting up the dual-purpose sound system 200, the volume of the ultrasonic signal 526 can be calibrated to ensure the system is not overloaded. Other methods for preventing signal level overload are known to those skilled in the art, such as applying a dynamic range limiter. The second-stage audio mixer 516 is provided in digital format. N A digital output signal 826 is converted to an analog output signal 828 by the D / A converter 808 (step 914). The amplifier 810 amplifies the analog output signal 828 to generate an amplified output signal 830 (step 916). The amplified output signal group 830, mixed with the fixed-volume ultrasonic signal 526, is an example of a combined audio / ultrasonic signal 224. The amplifier 810 then provides, supplies, or feeds the amplified output signal 830 (including the combined audio / ultrasonic signal 224) to a group of speakers 204 (step 918). The speakers 204 may be in a single housing (e.g., a Bluetooth speaker, a smart speaker, etc.), in a separate housing (a stereo system or a surround system), in a ceiling / wall mount (e.g., a 70 / 100V distributed sound system), etc. One technical benefit is that the dual audio system 202 allows the dual-purpose sound system 200 to simultaneously emit an audio signal for listening and an ultrasonic beacon signal for positioning, while also allowing volume control of the audio signal independently of the ultrasonic beacon signal. Therefore, controlling the volume of the audio signal will not negatively affect the ultrasonic beacon signal.
[0059] The following examples illustrate K = N = M =2. In this example, the dual-purpose sound system 200 provides stereo audio to the listener and has two audio channels 402 (L and R). The dual-audio system 202 is configured to superimpose two ultrasonic beacon signals onto the two audio channels 402 for positioning.
[0060] Figure 10is a block diagram illustrating a dual audio system 202 in an illustrative embodiment. The dual audio system 202 includes a first stage 501 and a second stage 502 as described above. In the first stage 501, an audio decoder 802 is configured to decode an encoded audio stream 832 into two source audio signals 520 (e.g., D1 and D2). Because the two source audio signals 520 are contained in the encoded audio stream 832 and the two audio channels 402, the first stage audio mixer 512 does not need to down / mix the source audio signals 520. Thus, the source audio signals 520 represent the mixed audio signals 522 as discussed above. A volume controller 514 is configured to adjust or set the volume of the source audio signals 520 to generate two variable volume audio signals 524 (e.g., A1 and A2).
[0061] In the second stage 502, the second stage audio mixer 516 is configured to mix the two variable volume audio signals 524 with fixed volume ultrasound signals 526 (e.g., U1 and U2) to generate digital output signals 826 (e.g., B1 and B2). A D / A converter 808 is configured to convert the digital output signals 826 to analog output signals 828 (e.g., C1 and C2). An amplifier 810 is configured to amplify the analog output signals 828 to generate amplified output signals 830 (e.g., L1 and L2) that include the combined audio / ultrasound signals 224. The dual audio system 202 is configured to provide or supply the combined audio / ultrasound signals 224 to the pair of speakers 204 (e.g., S1 and S2). The speakers 204 that receive the combined audio / ultrasound signals 224 will emit both audio sounds and ultrasound beacon signals for tracking the tracked object 110 simultaneously for a listener (i.e., a person 212). One technical benefit is that a stereo system can be used as a dual-purpose sound system 200 that provides both audio and ultrasound beacon signals simultaneously.
[0062] The following embodiment illustrates an example of a dual-purpose sound system 200 with four speakers S1, S2, S3, and S4 (i.e., four speaker drivers) in a single enclosure. The speaker S1 can be mounted on the front side, the speaker S2 is mounted on the left side, the speaker S3 is mounted on the right side, and the speaker S4 is mounted on the top side. Thus, the dual-purpose sound system 200 has four audio channels 402 with corresponding labels, F (front), L (left), R (right), and S (surround). The dual audio system 202 is configured to superimpose four ultrasound beacon signals on the four audio channels 402 for localization.
[0063] Figure 11is a block diagram illustrating a dual audio system 202 in an illustrative embodiment. The dual audio system 202 includes a first stage 501 and a second stage 502 as described above. In the first stage 501, an audio decoder 802 is configured to decode an encoded audio stream 832 into K source audio signals 520 (e.g., D1, D2,..., D K ). The renderer 518 is configured to render the source audio signals 520 to generate four rendered audio signals 521 (e.g., R1, R2, R3, and R4). The encoded audio stream 832 can be mono, stereo, surround, or another encoding format. Thus, the renderer 518 is configured to render the source audio signals 520 for the available audio channels 402 (i.e., front, left, right, surround) and multiple sets of signals. The signal type audio mixer 519 is configured to mix the rendered audio signals 521 to generate four system audio signals 522 (e.g., P1, P2, P3, and P4). The volume controller 514 is configured to adjust or set the volume of the system audio signals 522 to generate four variable volume audio signals 524 (e.g., A1, A2, A3, and A4).
[0064] In the second stage 502, the second stage audio mixer 516 is configured to mix the four variable volume audio signals 524 with fixed volume ultrasound signals 526 (e.g., U1, U2, U3, and U4) to generate four digital output signals 826 (e.g., B1, B2, B3, and B4). The D / A converter 808 is configured to convert the four digital output signals 826 into four analog output signals 828 (e.g., C1, C2, C3, and C4). The amplifier 810 is configured to amplify the analog output signals 828 to generate amplified output signals 830 (e.g., L1, L2, L3, and L4) that include the combined audio / ultrasound signals 224. The dual audio system 202 is configured to provide or supply the combined audio / ultrasound signals 224 to a set of four speakers 204 (e.g., S1, S2, S3, and S4). The speakers 204 that receive the combined audio / ultrasound signals 224 will simultaneously emit audio sounds and ultrasound beacon signals for tracking the tracked object 110 for a listener (i.e., a person 212). Depending on the orientation of the speaker system, the front channel is a direct path signal, the left and right channels arrive at the listener or sound receiver 102 as reflections from the left and right walls, respectively, and the surround channel arrives at the listener or sound receiver 102 as a reflection from the ceiling. One technical benefit is that a smart speaker, for example, can be used as a dual-purpose sound system 200 that simultaneously provides audio and ultrasound beacon signals.
[0065] The following embodiment illustrates an example of a dual-purpose sound system 200, which includes a 5.1.2 surround sound system having a left front (FL) speaker S1, a center channel (C) speaker S2, a right front (FR) speaker S3, a left surround (SL) speaker S4, a right surround (SR) speaker S5, a high left front (HFL) speaker S6, and a high right front (HFR) speaker S7. Therefore, the dual-purpose sound system 200 has eight audio channels 402, correspondingly labeled (note that the LFE channel is not shown because it cannot be used to carry ultrasonic beacon signals). The dual-audio system 202 is configured to superimpose ultrasonic beacon signals onto both audio channels 402 for positioning.
[0066] Figure 12 This is a block diagram of a dual audio system 202 in an illustrative embodiment. The dual audio system 202 includes a first stage 501 and a second stage 502 as described above. In the first stage 501, an audio decoder 802 is configured to decode the encoded audio stream 832 into... K Individual audio signals 520 (e.g., D1, D2, ..., D... K Renderer 518 is configured to render source audio signal 520 to generate seven rendered audio signals 521 (e.g., R1, R2, ..., R7). Signal type audio mixer 519 is configured to mix rendered audio signals 521 to generate seven system audio signals 522 (e.g., P1, P2, ..., and P7). Volume controller 514 is configured to adjust or set the volume of system audio signals 522 to generate seven variable volume audio signals 524 (e.g., A1, A2, ..., and A7).
[0067] In the second stage 502, the second stage audio mixer 516 is configured to mix the plurality of variable volume audio signals 524 with the fixed volume ultrasonic signals 526 to generate seven digital output signals 826 (e.g., B1, B2, …, B7). In this embodiment, the second stage audio mixer 516 mixes two variable volume audio signals 524 (e.g., A6 and A7) with fixed volume ultrasonic signals 526 (e.g., U1 and U2). In this example, the fixed volume ultrasonic signals 526 are mixed with the high set channels, which are the HFL channel 1224-6 for the HFL speaker S6 and the HFR channel 1224-7 for the HFR speaker S7. For ultrasonic beacon signals, the high set channels can be preferred because they are more likely to provide line of sight (LOS), i.e., a direct path from the high set speaker to the acoustic receiver 102, which is less likely to be blocked by objects, people, etc. However, the fixed volume ultrasonic signals 526 can be otherwise superimposed on the audio channels 402. The D / A converters 808 are configured to convert the digital output signals 826 to analog output signals 828 (e.g., C1, C2, …, C7). The amplifiers 810 are configured to amplify the analog output signals 828 to generate amplified output signals 830 (e.g., L1, L2, …, L7), where the amplified output signals L6 and L7 include the combined audio / ultrasonic signals 224. The dual audio system 202 is configured to provide or supply the amplified output signals L1 to L5 to the speakers S1 to S5, respectively, and to provide the combined audio / ultrasonic signals 224 (L6 and L7) to the speakers S6 and S7. The speakers 204 that receive the combined audio / ultrasonic signals 224 will simultaneously emit audio sounds and ultrasonic beacon signals for tracking the tracked object 110 to the listener (i.e., the person 212). One technical benefit is that a 5.1.2 surround sound system can be used as a dual-purpose sound system 200 that simultaneously provides audio and ultrasonic beacon signals.
[0068] The following embodiment illustrates an example of a dual-purpose sound system 200 that includes a 5.1.4 surround sound system with a front left (FL) speaker S1, a center channel (C) speaker S2, a front right (FR) speaker S3, a surround left (SL) speaker S4, a surround right (SR) speaker S5, a high set front left (HFL) speaker S6, a high set front right (HFR) speaker S7, a high set surround left (HSL) speaker S8, and a high set surround right (HSR) speaker S9. Thus, the dual-purpose sound system 200 has ten audio channels 402 with corresponding labels (note that the LFE channel is not shown because it cannot be used to carry ultrasonic beacon signals). The dual audio system 202 is configured to superimpose ultrasonic beacon signals on three audio channels 402 for localization.
[0069] Figure 13is a block diagram illustrating a dual audio system 202 in an illustrative embodiment. The dual audio system 202 includes a first stage 501 and a second stage 502 as described above. In the first stage 501, an audio decoder 802 is configured to decode an encoded audio stream 832 into K source audio signals 520 (e.g., D1, D2,..., D K ). A renderer 518 is configured to render the source audio signals 520 to generate nine rendered audio signals 521 (e.g., R1, R2,..., R9). A signal type audio mixer 519 is configured to mix the rendered audio signals 521 to generate nine system audio signals 522 (e.g., P1, P2,..., and P9). A volume controller 514 is configured to adjust or set the volume of the system audio signals 522 to generate nine variable volume audio signals 524 (e.g., A1, A2,..., and A9).
[0070] In the second stage 502, a second stage audio mixer 516 is configured to mix the plurality of variable volume audio signals 524 with fixed volume ultrasound signals 526 to generate digital output signals 826 (e.g., B1, B2,..., B9). In this embodiment, the second stage audio mixer 516 mixes three variable volume audio signals 524 (e.g., A7, A8, and A9) with fixed volume ultrasound signals 526 (e.g., U1, U2, and U3). In this example, the fixed volume ultrasound signals 526 are mixed with the high set channels, which are the HFR channel 1324-7 for the HFR speaker S7, the HSL channel 1324-8 for the HSL speaker S8, and the HSR channel 1324-9 for the HSR speaker S9. However, the fixed volume ultrasound signals 526 can be superimposed on the audio channels 402 in other ways. A D / A converter 808 is configured to convert the digital output signals 826 into analog output signals 828 (e.g., C1, C2,..., C9). An amplifier 810 is configured to amplify the analog output signals 828 to generate amplified output signals 830 (e.g., L1, L2,..., L9), where the amplified output signals L7, L8, and L9 include the combined audio / ultrasound signals 224. The dual audio system 202 is configured to provide or supply the amplified output signals L1 to L6 to the speakers S1 to S6, respectively, and to provide or supply the combined audio / ultrasound signals 224 (L7, L8, and L9) to the speakers S7, S8, and S9. The speakers 204 that receive the combined audio / ultrasound signals 224 will simultaneously emit audio sounds and ultrasound beacon signals for tracking the tracked object 110 to the listener (i.e., the person 212). One technical benefit is that a 5.1.4 surround sound system can be used as a dual-purpose sound system 200 that simultaneously provides audio and ultrasound beacon signals.
[0071] While high placement of speakers can have advantages as described previously, it is sometimes not practical or possible to mount speakers 204 in or near the ceiling. Figures 14-15 An alternative to ceiling mounted or high placement of speakers is illustrated. Figure 14 An upward firing speaker 204-5 in an illustrative embodiment is illustrated. The upward firing speaker 204-5 is typically designed for a surround sound system such as Dolby Atmos and DTS:X, and implemented in a speaker system as small as a smart speaker. Like the forward firing speaker, the upward firing speaker 204-5 carries audible sound (Al) and can also carry an ultrasonic beacon signal (Ul). The audible sound (Al) and the ultrasonic beacon signal (Ul) reflect off the ceiling 154. Thus, the robot 112, etc. can "hear" the ultrasonic beacon signal (Ul) reflected off the ceiling 154, while the listener (i.e., the person 212) can hear the audible sound (Al) reflected off the ceiling 154. Since the location of the upward firing speaker 204-5 is known by the robot 112 through setup, the robot 112 knows the location of the "virtual speaker." For a beacon based positioning algorithm (e.g., ToA or AoA), the location of the virtual speaker is related to the calculation of the robot position and orientation. In system design, the selection of the upward firing speaker 204-5 can be guided by the desired beam width 1406, such that any unwanted direct sound from it to the robot 112 and the listener are sufficiently attenuated.
[0072] Figure 15 An upward firing speaker 204-5 and a forward firing speaker 204-4 in an illustrative embodiment are illustrated. As Figure 15 The single speaker system in the room 102 provides four unique sounds, that is, the high placement audio signal (Al) from the upward firing speaker 204-5 and the direct path audio signal (A2) from the forward firing speaker 204-4, and the high placement ultrasonic signal (Ul) and the direct path ultrasonic signal (U2). The robot 112 "hears" two ultrasonic beacon signals, such as the real beacon and the virtual beacon. The robot 112 thus has a sufficient number of ultrasonic beacon signals to perform positioning. In the positioning, the beacon coordinates are the location of the "real" speaker (i.e., the forward firing speaker 204-4) and the location of the virtual speaker (i.e., the upward firing speaker 204-5). The location of the virtual speaker can be determined based on ceiling symmetry, a line of symmetry (i.e., a symmetric surface in 3-D space) provided by the ceiling 154, etc.
[0073] Figure 15The loudspeaker system in FIG. 8 includes an upward-firing driver and a forward-firing driver. Examples of these drivers are full-range drivers, mid-range drivers, or coaxial drivers, where the low / mid-range driver and tweeter are coaxially arranged. They can also be arranged as two separate drivers, including a low / mid-range driver and a tweeter. In some cases, it can also be desirable to use more than two drivers in the upward-firing section or the forward-firing section. In Figures 14-15 In FIG. 9, the upward-firing loudspeaker 204-5 is shown as forward-tilted. The driver can be mounted in any other way, or without any tilt. One design variable is the beamwidth 1406 of the upward-firing loudspeaker 204-5 when mounted in the enclosure, where it is desirable for the sound of the upward-firing loudspeaker 204-5 to be mostly received by the listener and the robot 112 from the ceiling 154.
[0074] Although for simplicity of illustration, any number of listeners and any number of robots 112 or other tracked objects 110 can receive the signals in FIG. 10, the example includes one listener and one robot 112. Likewise, Figures 14-15 The example in FIG. 11 shows one loudspeaker system. However, the same concepts can be used for any desired number of loudspeaker systems. Figures 14-15 The example in FIG. 11 shows one loudspeaker system. However, the same concepts can be used for any desired number of loudspeaker systems.
[0075] The encoded audio stream 832, as discussed above, can be encoded as an MPEG-H 3D audio stream. MPEG-H supports high or elevated loudspeakers (i.e., loudspeakers above ear height (such as 5.1+4H), but also supports below ear height (such as 22.2). MPEG-H supports channel encoding, object encoding, and HOA encoding, and supports interactive three degrees of freedom (3DoF) rendering based on the orientation (yaw, pitch, and roll) of the listener.
[0076] Sound systems are often installed in commercial buildings, such as distribution centers, warehouses, manufacturing plants, and retail stores. Such sound systems are often designed for multiple purposes, but primarily as background music systems, paging systems, public address systems, etc. With the concepts described herein, these sound systems can also be used as ultrasonic beacon systems to locate vehicles, robots, appliances, people, etc.
[0077] In these commercial or corporate buildings, 70V / 100V distributed speaker systems are often used. These systems are also known as high impedance, high voltage, or constant voltage speaker systems, and have desirable properties for large installations. Speakers can be connected by their transformer "daisy chains," avoiding the need for individual cables from the amplifier to each speaker. In addition, due to the high input impedance of the transformer, the cable impedance is relatively small, allowing long cable runs without the need for large gauge cables. Such distributed sound systems typically render only a single mono audio signal, in which case the same mono audio is emitted on each speaker of the distributed system. Even though only a mono signal is emitted, the distributed sound system can still be organized as a multi-channel system, for example, to address a limited number of high impedance speakers that can be serviced by a single amplifier channel due to its output power limitations.
[0078] Figure 16 A two-audio system 1602 of the distributed sound system 1601 in the illustrative embodiment is illustrated. Audio input signals El, E2, through Enare provided to the two-audio system 1602. The audio input signals El, E2, through Enmay be provided in digital form, such as from a Wi-Fi, Ethernet, or USB port, or in analog form, such as from an XLR microphone input or line level input. The audio input signals El, E2, through Enare mixed by the two-audio system 1602 to provide a first audio signal 1620 and a second audio signal 1622. The first audio signal 1620 and the second audio signal 1622 are provided to the distributed sound system 1601, which provides a first audio output signal 1630 and a second audio output signal 1632. The first audio output signal 1630 and the second audio output signal 1632 are provided to the speakers 1603, which emit the first audio output signal 1630 and the second audio output signal 1632, respectively. K The audio input signals (e.g., encoded digital audio signals from a Wi-Fi, Ethernet, or USB port, and analog signals from an XLR microphone input or line level input) are first decoded in the case of digital signals, or otherwise first converted to digital signals by the A / D converter 1604. The two-stage mixer 506 can be implemented or carried out with a digital signal processor (DSP) 1610. The first stage audio mixer 512 can run common signal processing functions such as equalizers, dynamic range processors, effects processors, level adjustments, real-time spectral analysis, and signal mixing. A user can control the two-stage mixer 506 through a two-stage mixer controller 1612, such as a tablet, smartphone, or Mac / Win / Linux computer, etc. The two-stage mixer controller 1612 can be connected to the two-stage mixer through Ethernet, USB, Wi-Fi, etc.
[0079] The first stage audio mixer 512 provides K input channels and N output channels. For each channel, the audio signals 522 can be significantly different, or they can be identical (i.e., each of the audio signals contains the same mono audio mix). The latter is the case in which a mono mix is created in the first stage audio mixer 512. The second stage audio mixer 516 mixes or adds the ultrasonic signals 526 to the audio signals 524. In this embodiment and other embodiments described above, the beacon signals can be provided in digital form, for example, as Ma unique pre-computed data set of chirps, each chirp using a specified or different frequency band (e.g., with a bandwidth of 1 to 2 kHz). Thus, each speaker 204 can emit different chirps at different frequencies. Moreover, the duration of the chirps can be short, e.g., 20 milliseconds (ms), that is, 960 samples for a sampling frequency of 48 kHz, and repeated continuously at the desired rate. In another example, the frequency bands of the chirps can overlap, with the chirps differing in duration, in frequency increase (chirp up), in frequency decrease (chirp down), etc. Because the ultrasonic signals 526 can each be specified by a short sequence of signals (e.g., chirps) and a repetition interval, there is no need for physical input channels for these signals. This means that, in some cases, a regular audio system can be converted into a dual-purpose system 202 by adding the second stage audio mixer 516 through a software modification only, provided the hardware requirements are met. The first hardware requirement is that the D / A conversion provides an upper passband frequency limit close to (e.g., within 1 kHz of) the Nyquist frequency (e.g., > 23 kHz for a Nyquist frequency of 24 kHz) to enable the ultrasonic signals 526 to pass through. The second hardware requirement is that the frequency range of the speakers 204 extends into the ultrasonic range. The third hardware requirement is that the software performing the second stage mixer can be installed. In Figure 16 the example of FIG. 18, the first hardware requirement can be met by an oversampled sigma-delta D / A converter 1620 with a steep roll-off interpolation filter setting. The second hardware requirement can be met by a band, piezo, dome tweeter, or any other tweeter that extends into the ultrasonic range with a beam width sufficient to cover the area. The third hardware requirement can be met if the software of the system can be modified (e.g., as in a DSP implementation). If the hardware requirements are met or the system can be modified accordingly, the two-stage mixer can be implemented such that the ultrasonic signals 526 are mixed into the audio signals 524 after the master volume control and such that a constant ultrasonic signal level can be set to ensure good localization performance regardless of the audio volume setting.
[0080] In Figure 16 the example of FIG. 18, the first hardware requirement can be met by an oversampled sigma-delta D / A converter 1620 with a steep roll-off interpolation filter setting. The second hardware requirement can be met by a band, piezo, dome tweeter, or any other tweeter that extends into the ultrasonic range with a beam width sufficient to cover the area. The third hardware requirement can be met if the software of the system can be modified (e.g., as in a DSP implementation). If the hardware requirements are met or the system can be modified accordingly, the two-stage mixer can be implemented such that the ultrasonic signals 526 are mixed into the audio signals 524 after the master volume control and such that a constant ultrasonic signal level can be set to ensure good localization performance regardless of the audio volume setting. N N ) are fed into an amplifier 810 that amplifies the signals and converts the low impedance output into a high impedance output using a step-up transformer or electronic circuitry. The number of output channels of the amplifier 810 can be, for example, N = 4. The high impedance output enables the speakers 204 to be daisy-chained, as Figure 16 The impedance needs to be matched to the loudspeaker 204 at the loudspeaker 204. Therefore, a step-down transformer 1622 is inserted in front of each loudspeaker 204.
[0081] Distributed sound systems are commonly used in shopping malls, retail stores, restaurants, hotels and public buildings. In conventional distributed sound systems, mono signals are typically distributed in these spaces, in which case the same mono audio signal is played on all loudspeakers. Although if desired, Figure 16 The distributed audio system components in the system of FIG. 1 can still play the same mono audio on all loudspeakers 204, but the distributed beacon components typically provide a unique beacon signal on each loudspeaker channel of the active beacon channel.
[0082] For localization purposes, Figure 16 The loudspeakers 204 in the system of FIG. 1 are arranged in localization zones (i.e., Zone 1 through Zone H). The localization zones include loudspeakers 204 carrying different beacon signals, while they can carry the same audio signal (mono) or different audio signals, which is determined by the mixer. The localization zones provide the necessary beacons for localization. When adjacent localization zones are adjacent or overlapping, they can interfere with each other. To avoid such interference, adjacent localization zones can have different frequency channels. For example, in a 4-channel system (N = 4), localization zone 1 can include one loudspeaker of channel L1 and one loudspeaker of channel L2. Localization zone 2 can include one loudspeaker of channel L3 and one loudspeaker of channel L4. In this way, the beacon signal frequency of localization zone 1 will not interfere with the frequency of localization zone 2. N
[0083] The desired spacing of loudspeakers in a distributed sound system and the desired spacing of beacons in a distributed localization system are typically similar, mainly determined by the beam width or coverage angle of the loudspeakers and the ceiling height. However, different spacings of the two distributed systems can also be accommodated. For example, if the distributed audio components each require a smaller spacing, the mixer can provide only audio output channels, and in turn the distributed audio components can provide only audio loudspeaker channels. In this way, the localization zones can have the desired number of only audio channels. Figure 16
[0084] The transformer 1622 in a high impedance distributed system can reduce the frequency response at high frequencies. This means that the frequency band of the ultrasonic beacon signal can be attenuated. Such high frequency attenuation can be due to stray magnetic fields or core losses in the transformer. However, if the transformer 1622 is well designed, excellent high frequency performance can be achieved.
[0085] Figure 17 A user application 1704 for controlling a dual-purpose sound system in an illustrative embodiment is shown. A user's user device 1702 (e.g., a smartphone, tablet, PC, etc.) can implement the user application 1704. The user application 1704 enables the user to control the audio player and the beacon signal player. The user application 1704 can also add control functions and status indicators for the robot 112 or other tracked object 110. For the audio player component, the user application 1704 enables the user to set the source of the audio signal (e.g., a website, a server, an online radio, a personal computer, a smartphone, etc.). The user application 1704 allows typical controls on the source playback mode (if available) (e.g., start / end, fast forward / rewind, and play / pause of music or podcasts). Although the source of the beacon signal can also be set by the user, the source file name and directory can be pre-set during installation and can not require any intervention by the user. Likewise, the beacon signal playback function can be coupled to the robot on / off status or robot task. For example, if the robot is powered on, it can trigger the beacon signal to automatically turn on, and in the same way, turn off if the robot is powered off. The user application 1704 also enables the user to determine the task of the robot, and it can show the status of the robot. The user application 1704 can implement a dual-audio system as described above to frequency multiplex the audio signal and the beacon signal.
[0086] Instead of a single user application 1704 that acts as both an audio player and a beacon signal player and can perform as a robot controller, two separate applications can also be used. In this case, a first application (i.e., the audio player) can be set to redirect the audio output to a second application (i.e., the beacon signal player and robot control). The second application uses the audio signal as input and frequency multiplexes the audio signal with the beacon signal (typically a stereo or multi-channel signal). The combined audio / ultrasonic signal can then be emitted to the sound system.
[0087] Figures 18A-18B How the audio signal and the ultrasonic signal are frequency multiplexed in an illustrative embodiment is shown. Figure 18AFig. 1 illustrates the frequency band allocation for a sampling frequency of 48 kHz. In this case, the lower part of the frequency spectrum from about 0 Hz to 20 kHz, including the audible frequency spectrum, is used for audio signals (e.g. mono, stereo, surround etc.), while the upper part of the frequency spectrum from about 20 kHz to 24 kHz is allocated to ultrasonic signals. Since modern D / A devices use high order anti-aliasing filters with an internal oversampling function, computer generated digital signals with frequency components close to the critical Nyquist frequency of 24 kHz can generally be played out with little attenuation in the filtering process, leaving a large available frequency range of about 20 kHz to at least 23 kHz for the ultrasonic signals.
[0088] In many cases, the frequency multiplexing is a simple mixing process of two signal sets, including a set of 48 kHz sampled bandwidth limited audio signals and a set of 48 kHz sampled bandwidth limited ultrasonic signals. For a stereo system with two loudspeaker channels L1 and L2, two audio signals A1 and A2 and two beacon signals U1 and U2 can be mixed in various ways. One simple approach is to add the signals, i.e. L1 = A1 + U1, L2 = A2 + U2. This mixing assumes that the audio signals have little energy in the range from 20 kHz to 24 kHz and that the beacon signals have little energy in the frequency range from 0 Hz to 20 kHz.
[0089] Figure 18B Fig. 2 illustrates the frequency band allocation for a sampling frequency of 96 kHz. Here, the frequency band of the ultrasonic beacon signals is much wider, about 26 kHz of the available frequency bandwidth from 20 kHz to 46 kHz. This means that in this case many beacon channels can be allocated. The frequency multiplexing can be performed in the same way as in the 48 kHz sampling rate case.
[0090] Any of the examples as described above can use the spectral allocation for audio signals and beacon signals as described in Figures 18A-18B Fig. 1.
[0091] Figure 19 Fig. 1 is a flow chart illustrating a method 1900 of mixing audio signals and ultrasonic signals in an illustrative embodiment. The steps of the method 1900 will be described with reference to the user application 1704 in Figure 17 Fig. 1. However, those skilled in the art will appreciate that the method 1900 can be performed in other systems or devices. Figure 19 The examples in Fig. 1 assume a frame-wise processing, where a frame is a block of samples. Such a frame can include samples corresponding to a 20 ms audio interval. For a sampling frequency of 48 kHz, this means that a 20 ms audio interval will include 960 samples. The ultrasonic signal can be contained in a separate frame, which can be played repeatedly. The frame can containM channels, wherein M is the number of beacon signals. For a sampling rate of 48 kHz, the ultrasonic frequency band from 20 kHz to 24 kHz can be used. Typically, the computer-generated frame is used with M channels containing ultrasonic signal samples. The ultrasonic signal can include a sweep (chirp) or a pseudo-random noise sequence (PRNS).
[0092] The user application 1704 reads the ultrasonic frame from the file (step 1902). Since the same frame is played repeatedly, the ultrasonic frame is read once at the beginning. The user application 1704 reads the audio player status (step 1904). The audio player status can be a stop / pause mode or a play mode. If the audio player is in the play mode, the user application 1704 reads the next audio frame (step 1906). Otherwise, there is no need to read a frame. The user application 1704 reads the robot status (step 1908). If the robot is on or activated, the ultrasonic frame is mixed with the audio frame (step 1910). In other words, the second-level audio mixer 516 is activated to mix the variable volume audio signal 524 and the fixed volume ultrasonic signal 526. However, if the audio player is not in the play mode, the ultrasonic frame is mixed with a null frame or not mixed at all. The user application 1704 then determines whether to play out the frame. When the robot is off and the audio player is off (i.e., not in the play mode), the play out decision is NO. When the robot is on, the play out decision is YES, and the user application 1704 plays out the frame (step 1912). One technical benefit is that the beacon signal is turned on when the robot is activated. Otherwise, the ultrasonic beacon signal is not needed. Other criteria for turning on the beacon signal can also be used. It can be better to leave the beacon signal on when multiple robots are using the beacon signal. Note again in the above description that the robot can be replaced with an ultrasonic tag or any other tracked object 110. One technical benefit is that the user application 1704 is able to control / implement a dual audio system to play audio and provide an ultrasonic beacon system at the same time.
[0093] Any of the various elements or modules described in the drawings or described herein can be implemented as hardware, software, firmware, or some combination thereof. For example, an element can be implemented as dedicated hardware. Dedicated hardware elements can be referred to as "processors," "controllers," or some similar terminology. When provided by a processor, these functions can be provided by a single dedicated processor, by a single shared processor, or by a plurality of individual processors, some of which can be shared. Moreover, explicit use of the term "processor" or "controller" should not be construed to refer exclusively to hardware capable of executing software, and can implicitly include, without limitation, digital signal processor (DSP) hardware, network processor, application specific integrated circuit (ASIC), or other circuitry, field programmable gate array (FPGA), read only memory (ROM) for storing software, random access memory (RAM), nonvolatile storage medium (e.g., flash drive), logic or some other physical hardware component or module.
[0094] Furthermore, an element can be implemented as instructions executable by a processor or a computer to perform the functions of that element. Some examples of instructions are software, programs, code, and firmware. The instructions are operational when provided to the processor to direct the processor to perform the functions of the element. The instructions can be stored in the memory of the processor. Some examples of processor-readable storage media include digital or solid state memories, magnetic storage media such as a magnetic disks and magnetic tapes, hard drives, or optical storage media.
[0095] As used in this application, the term "circuitry" can refer to one or more or all of the following:
[0096] (a) hardware-only circuit implementations (such as implementations in only analog and / or digital circuitry) ;
[0097] (b) combinations of hardware circuits and software, such as (as applicable):
[0098] (i) a combination of analog and / or digital hardware circuit(s) with software / firmware;
[0099] (ii) any portions of hardware processor(s) with software (including digital signal processor(s)), software, and memory(ies) that work together to cause an apparatus, such as a mobile phone or server, to perform various functions); and
[0100] (c) hardware circuit(s) and or processor(s), such as a microprocessor(s) or a portion of microprocessor(s), that requires software (e.g., firmware) for operation, but it is not required that software be present for the hardware circuit(s) and or processor(s) to operate.
[0101] This definition of circuitry applies to all uses of this term in this application, including in any claims. As a further example, as used in this application, the term "circuitry" would also cover an implementation that includes one or more processors and / or a processor(s) working in conjunction with a software module and / or other interioating software and / or firmware. For example, if a claim recites a process being performed by, or instrumentation being used by, circuitry, such a claim would cover any implementation of that process or instrumentation by circuitry, including one or more processors working in conjunction with a software module and / or other interioating software and / or firmware.
[0102] Although specific embodiments were described herein, the scope of the disclosure is not limited to these specific embodiments. The scope of the disclosure is defined by the following claims and any equivalents thereof.
Claims
1. An apparatus for processing audio signals, the apparatus comprising: a first stage comprising a master volume controller configured to perform volume control of audio signals of audio channels of a sound system to generate variable volume audio signals of the audio channels; and a second stage comprising an audio mixer configured to mix the variable volume audio signals of a target set of the audio channels with a fixed volume ultrasonic signal for localization to generate at least two combined audio / ultrasonic signals configured for output to speakers of the sound system corresponding to the target set of the audio channels.
2. The apparatus of claim 1, wherein: the audio mixer of the second stage comprises a second audio mixer; and the first stage further comprises a first audio mixer configured to mix the audio signals to generate a mixed audio signal provided to the master volume controller.
3. The apparatus of claim 2, wherein: the first stage further comprises an audio decoder configured to decode encoded audio streams into source audio signals; the first audio mixer comprises a renderer and a signal type audio mixer; the renderer is configured to render the source audio signals to generate rendered audio signals for the audio channels of the sound system; and the signal type audio mixer is configured to mix the rendered audio signals to generate the mixed audio signal.
4. The apparatus of claim 2, wherein: the first audio mixer and the second audio mixer are implemented with digital signal processors.
5. The apparatus of claim 1, wherein: at least one of the first stage and the second stage comprises a digital-to-analog converter comprising an oversampling digital-to-analog converter.
6. The apparatus of claim 1, wherein: the audio channels of the sound system comprise front channels, left channels, right channels, and surround channels, and the audio mixer is configured to mix the variable volume audio signals of at least two of the audio channels with the fixed volume ultrasonic signal to generate the combined audio / ultrasonic signals.
7. The apparatus of claim 6, wherein: the audio mixer is configured to mix the variable volume audio signals of the front channels and the surround channels with the fixed volume ultrasonic signal to generate the combined audio / ultrasonic signals.
8. The apparatus of claim 1, wherein: the audio channels of the sound system comprise center channels, left channels, right channels, surround channels, and height channels, and the audio mixer is configured to mix the variable volume audio signals of at least two of the height channels with the fixed volume ultrasonic signal to generate the combined audio / ultrasonic signals.
9. The apparatus of claim 1, wherein: the sound system is a distributed sound system; The audio mixer is configured to mix the variable volume audio signals of at least two of the audio channels with the fixed volume ultrasonic signals to generate the combined audio / ultrasonic signals. Each of the audio signals of the at least two of the audio channels contains the same mono audio mix.
10. The apparatus of claim 1, further comprising: an application configured to determine whether a tracked object is activated and, when the tracked object is activated, activate the audio mixer to mix the variable volume audio signals of the target set of the audio channels with the fixed volume ultrasonic signals to generate the combined audio / ultrasonic signals.
11. The apparatus of claim 1, wherein: the fixed volume ultrasonic signals comprise chirps, each chirp using a different frequency band.
12. A method for processing audio signals, the method comprising: at a first stage, performing volume control of audio signals of audio channels of a sound system to generate variable volume audio signals of the audio channels; at a second stage, mixing the variable volume audio signals of a target set of the audio channels with fixed volume ultrasonic signals for localization to generate at least two combined audio / ultrasonic signals configured to be output to speakers of the sound system corresponding to the target set of the audio channels; mixing the audio signals at the first stage to generate mixed audio signals, wherein performing volume control of the audio signals comprises performing volume control of the mixed audio signals.
13. The method of claim 12, further comprising: at the first stage, decoding the encoded audio stream into source audio signals; and at the first stage, rendering the source audio signals to generate rendered audio signals for the audio channels of the sound system, wherein the mixing at the first stage comprises mixing the rendered audio signals to generate the mixed audio signals.
14. The method of claim 13, wherein: the fixed volume ultrasonic signals comprise chirps, each chirp using a different frequency band.
15. A computer readable medium having programming instructions embodied therewith, the programming instructions, when executed by a processor, are operable to perform a method of operating a dual audio system, the method comprising: at a first stage, performing volume control of audio signals of audio channels of a sound system to generate variable volume audio signals of the audio channels; and at a second stage, mixing the variable volume audio signals of a target set of the audio channels with fixed volume ultrasonic signals for localization to generate at least two combined audio / ultrasonic signals configured to be output to speakers of the sound system corresponding to the target set of the audio channels.