Methods and systems for imperceptible acoustic data transfer

MCLT audio watermarking with pilot symbols and phase modulation addresses perceptibility and efficiency issues in audio systems, enabling secure and efficient acoustic data transfer.

WO2026005776A1PCT designated stage Publication Date: 2026-01-02HARMAN INT IND INC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/US2024/035666
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-06-26
Publication Date
2026-01-02

AI Technical Summary

Technical Problem

Existing audio watermarking techniques face issues such as perceptibility, susceptibility to corruption, security vulnerabilities, limited data capacity, and computational inefficiency, making them unsuitable for secure and efficient acoustic data transfer in audio systems.

Method used

The implementation of Modulated Complex Lapped Transform (MCLT) audio watermarking with pilot symbols embedded using intra-frame and inter-frame pseudo-noise sequences and double phase modulation, enabling imperceptible, robust, and secure data transfer through loudspeakers and microphones.

Benefits of technology

This approach achieves imperceptible, high-capacity, secure, and computationally efficient acoustic data transfer, supporting features like automatic device onboarding and zero-cost channel synchronization, even on low-powered devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2024035666_02012026_PF_FP_ABST
    Figure US2024035666_02012026_PF_FP_ABST
Patent Text Reader

Abstract

In methods and systems for Acoustic Data Transfer, on an encode side, a message may be channel encoded, pilot symbols may be embedded therein based on intra-frame and inter-frame pseudo-noise (PN) sequences in accordance with double phase modulation, and a two-dimensional (2D) time-frequency data grid may be generated based on the pilot-embedded channel-encoded message. Based on a time-domain audio signal and the 2D time-frequency data grid, data-embedded Modulated Complex Lapped Transform (MCLT) frames may be generated, and a watermarked audio signal based on the data-embedded MCLT frames may be transmitted via a loudspeaker. On a decode side, a microphone may receive the audio signal, which may be processed by a Sliding MCLT and / or a dynamic-programming-based 2D cross-correlator to detect a watermarked data packet and estimate its timestamp. A segment of the received audio signal containing the watermarked data packet may then be extracted, channel-equalized, and channel-decoded to recover the message.
Need to check novelty before this filing date? Find Prior Art

Description

Docket No. P230134WO METHODS AND SYSTEMS FOR IMPERCEPTIBLE ACOUSTIC DATA TRANSFER FIELD

[0001] The disclosure relates generally to methods and systems for acoustic data transfer, including among devices of a loudspeaker-based audio system. BACKGROUND

[0002] Audio systems may have components including one or more loudspeakers and other components. Some audio systems may support communication (such as the provision of audio signaling and / or control signaling) among the components of the systems through wired connections. Other audio systems may support communication among the components of the systems through wireless connections, such as wireless communication links compliant with various revisions of a Bluetooth® specification and / or various revisions of a Wi-Fi® specification. (Bluetooth® is a registered trademark of Bluetooth SIG, Inc., Kirkland, WA. Wi-Fi® is a registered trademark of Wi-Fi Alliance, Austin, Texas.) Some audio systems and / or audio-system capabilities may also benefit from wireless communication techniques (including point-to-point wireless communication techniques) that avoid additional technologies, such as technologies compliant with additional specifications involving additional sensors, circuitries, and / or other hardware.

[0003] Acoustic Data Transmission (ADT) techniques send messages through aerial space by modifying an audio signal to include non-audio-based information, transmitting the audio signal via a loudspeaker, receiving the audio signal via a microphone, and recovering the non-audio- based information from the audio signal. ADT techniques may thereby turn a loudspeaker and a microphone into components of a wireless communication technique, and may thus be good alternatives to wireless communication technologies involving additional hardware (such as additional sensors and / or additional circuitries).

[0004] Various uses of ADT techniques may also be enabled and / or enhanced with audio watermarking techniques, by which a watermark including the data to be transmitted is inserted into an audio signal prior to transmission, and the data is recovered from the received watermarked audio signal, ideally without assuming knowledge of the original host audio as in blind audio watermarking. However, audio watermarking techniques may have various problems, such asDocket No. P230134WO perceptibility of the embedded watermark, susceptibility of the embedded watermark to damage from unintentional and / or intentional corruption during transmission, lack of security of the embedded watermark, unsuitability of the data capacity of the embedded watermark for the application thereof, and / or computational inefficiency of incorporating the data into the audio signal via the embedded watermark. SUMMARY

[0005] Disclosed herein are methods and systems for ADT techniques using Modulated Complex Lapped Transform (MCLT) audio watermarking. The embedded watermarks of the audio watermarking techniques may advantageously minimize or eliminate perceptible changes to the host audio signal, be successfully extracted with minimal damage following transmission, reduce and / or eliminate insecure extraction, provide a data capacity suitable for its applications, and be computationally efficient to implement. Accordingly, the audio watermarking of the methods and systems disclosed herein may be blind, imperceptible, robust, secure, of high capacity, and of low computational complexity.

[0006] Moreover, the methods and systems disclosed herein may advantageously enable and / or facilitate compute-efficient and imperceptible ADT techniques that may be run in real-time on low-powered devices (e.g., loudspeakers). Accordingly, these methods and systems may support useful features such as automatic new-device onboarding for Bluetooth® and Wi-Fi® devices, proximity detection between two devices, and / or zero-cost channel synchronization (for devices without Bluetooth® and / or Wi-Fi® hardware).

[0007] In some embodiments, the advantageous ADT techniques described herein above may be enabled by methods including receiving a decode-side digital signal from a microphone, the decode-side signal being based on an audio signal, and the audio signal being generated by a loudspeaker based on an encode-side digital signal from a first device. The decode-side signal may be processed by a second device using one or more of a Sliding MCLT (SMCLT) and a dynamic programming (DP)-based cross-correlator, to detect that a watermarked data packet is present in the decode-side digital signal, and to also estimate a timestamp of the watermarked data packet within the decode-side digital signal. Based on the detection that the watermarked data packet is present in the decode-side data signal, a segment of the decode-side digital signal containing the watermarked data packet may be extracted from the decode-side signal, the segment being basedDocket No. P230134WO on the estimated timestamp. The segment may then be processed to generate a recovered message. The use of the SMCLT to detect the presence and temporal position of the watermarked data packet within the decode-side signal may advantageously support an imperceptible, robust, secure, high capacity, and low computational complexity ADT technique.

[0008] For some embodiments, the advantageous ADT techniques described herein above may be enabled and / or facilitated by methods including channel-encoding a message, followed by embedding pilot symbols into the channel-encoded message. The plurality of pilot symbols may be based on an intra-frame pseudo-noise (PN) sequence (which may be random) and an inter-frame PN sequence (which may be optimized for autocorrelation), and the embedding may be in accordance with a double phase modulation, e.g., in which the two PN sequences may be multiplied to form a two-dimensional (2D) grid. Such embedding may advantageously facilitate and / or enable computationally feasible watermark detection in low-powered devices. Based on the pilot-embedded channel-encoded message, a 2D time-frequency data grid may then be generated. Next, based on a time-domain audio signal and the 2D time-frequency data grid, a plurality of data-embedded MCLT frames may be generated in accordance with MCLT-based watermarking. Data-embedded MCLT frames may then be used to generate a watermarked time-domain audio signal (e.g., using inverse MCLT). Finally, an audio signal based on the watermarked time-domain audio signal may be generated by a loudspeaker. The embedding of pilot symbols based on intra- frame and inter-frame PN sequences in accordance with double phase modulation and MCLT- based watermarking may advantageously support a blind, imperceptible, robust, secure, high capacity, and low computational complexity ADT technique.

[0009] In MCLT-based audio watermarking, data may be embedded by modifying the phase of select MCLT coefficients at select MCLT frames. MCLT-based audio watermarking may advantageously have a relatively high bitrate of ADT bandwidth with relatively high imperceptibility properties.

[0010] In additional embodiments, the advantageous ADT techniques described herein above may be enabled and / or facilitated by systems comprising decode-side components and / or encode- side components. The decode-side components may comprise a decode-side interface for an output of a microphone and a decode-side device coupled to the decode-side interface. The decode-side device may process a decode-side digital signal using one or more of an SMCLT and a DP-based 2D cross-correlator to detect that a watermarked data packet is present in the decode-side digitalDocket No. P230134WO signal provided by the microphone, and to estimate the timestamp of the watermarked data packet within the decode-side digital signal. Based on detecting that the watermarked data packet is present, a segment of the decode-side digital signal containing the watermarked data packet may be extracted, the segment being based on the estimated timestamp. The segment may be channel- equalized using a frequency-selective linear filter to generate a plurality of channel-distortion- compensated MCLT coefficients for a 2D time-frequency data grid. The plurality of channel- distortion-compensated MCLT coefficients may be channel-decoded to generate a recovered message.

[0011] It should be understood that the summary above is provided to introduce in simplified form a selection of concepts that are further described in the detailed description. It is not meant to identify key or essential features of the claimed subject matter, the scope of which is defined uniquely by the claims that follow the detailed description. Furthermore, the claimed subject matter is not limited to implementations that solve any disadvantages noted above or in any part of this disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] The disclosure may be better understood from reading the following description of non-limiting embodiments, with reference to the attached drawings, wherein below:

[0013] FIG.1 shows a schematic diagram of portions of an audio system supporting Acoustic Data Transfer (ADT), in accordance with one or more embodiments of the present disclosure;

[0014] FIG. 2 shows an encode-side pipeline of an audio system supporting ADT, in accordance with one or more embodiments of the present disclosure;

[0015] FIG. 3 shows a decode-side pipeline of an audio system supporting ADT, in accordance with one or more embodiments of the present disclosure;

[0016] FIG. 4 shows portions of a channel encoding stage of an encode-side pipeline, in accordance with one or more embodiments of the present disclosure;

[0017] FIGS. 5A and 5B depict scenarios of pilot symbol embedding in an MCLT-based audio watermarking technique, in accordance with one or more embodiments of the present disclosure;Docket No. P230134WO

[0018] FIG.6 shows a scenario of encode-side pilot embedding for efficient decode-side data packet detection and / or synchronization, in accordance with one or more embodiments of the present disclosure;

[0019] FIG.7 shows a scenario of MCLT every-other-frame and every-other-frequency data embedding, in accordance with one or more embodiments of the present disclosure;

[0020] FIGS.8A and 8B depict scenarios of a 2D ring buffer state, in accordance with one or more embodiments of the present disclosure;

[0021] FIG. 9 shows portions of a channel decoding stage of an encode-side pipeline, in accordance with one or more embodiments of the present disclosure; and

[0022] FIGS.10A-10B and 11A-11B show methods for ADT, in accordance with one or more embodiments of the present disclosure. DETAILED DESCRIPTION

[0023] FIGS.1 through 9 depict audio systems, and components thereof, for encode-side and decode-side ADT techniques. FIGS.10A-10B and 11A-11B depict methods for encode-side and decode-side ADT techniques. The systems, components thereof, and associated methods may advantageously permit ADT techniques which may enable and / or facilitate compute-efficient and imperceptible ADT techniques that may be run in real-time on low-powered devices (e.g., loudspeakers), thereby supporting useful features such as automatic new-device onboarding for Bluetooth® and Wi-Fi® devices, proximity detection between two devices, and / or zero-cost channel synchronization (for devices without Bluetooth® and / or Wi-Fi® hardware

[0024] FIG.1 shows an audio system 100 supporting ADT. Audio system 100 includes a first audio assembly 110 and a second audio assembly 160. First audio assembly 110 includes a first device 120 and a loudspeaker 130. First device 120 may have a first interface 112 for receiving a message to be transmitted to second audio assembly 160 via an ADT technique, and for receiving an audio signal in which to embed an audio watermark containing the message. First device 120 may also have a second interface 122 for an input of loudspeaker 130, which may generate a watermarked audio signal 132. First audio assembly 110, first device 120, and loudspeaker 130 may accordingly form parts of an encode-side portion of audio system 100.

[0025] Meanwhile, second audio assembly 160 includes a second device 170 and a microphone 180. Second device 170 may have a first interface 162 for providing a messageDocket No. P230134WO recovered from an embedded watermark via an ADT technique. Second device 170 may also have a second interface 172 for an output of microphone 180, which may process an audio signal 182. Audio signal 182 may be watermarked audio signal 132 as received at second audio assembly 160, e.g., as impacted by the environment between first audio assembly 110 and second audio assembly 160. Second audio assembly 160, second device 170, and microphone 180 (and portions thereof) may accordingly form parts of a decode-side portion of audio system 100.

[0026] In some embodiments, audio system 100 may include one or more components forming parts of encode-side portions of audio system 100, in a manner similar to first audio assembly 110 as described herein. Audio system 100 may also include one or more components forming parts of decode-side portions of audio system 100, in a manner similar to second audio assembly 160 as described herein.

[0027] For example, a loudspeaker component of audio system 100 may include a component substantially similar to first audio assembly 110 which forms at least part of an encode-side portion of audio system 100, while an audio / video (A / V) receiver of audio system 100 may include a component substantially similar to second audio assembly 160 which forms at least part of a decode-side portion of audio system 100. In that example, the loudspeaker may transmit a message to the A / V receiver via an ADT technique by embedding an audio watermark containing the message in an audio signal. The A / V receiver may then receive the audio signal and recover the message from the embedded watermark using an ADT technique.

[0028] Moreover, in various embodiments, audio system 100 may have a variety of components, each of which may form at least part of an encode-side portion of audio system 100, at least part of a decode-side portion of audio system 100, or even both. Thus, in a variety of embodiments, some components of audio system 100 may both transmit messages in audio signals via an ADT technique and recover messages from audio signals via an ADT technique, as disclosed herein. In a further example, a first loudspeaker component of audio system 100 may have an audio assembly that is substantially similar to first audio assembly 110, and a second loudspeaker component of audio system 100 may have an audio assembly that is substantially similar to second audio assembly 160. In that example, the first loudspeaker component may transmit a message (e.g., via the audio assembly substantially similar to first audio assembly 110) to the second loudspeaker component through an ADT technique by embedding an audio watermark containing the message in an audio signal. The second loudspeaker component may then receive the audioDocket No. P230134WO signal (e.g., via the audio assembly substantially similar to second audio assembly 160) and recover the message from the embedded watermark through an ADT technique. Thus, in various embodiments, loudspeaker components of audio systems may have both an audio assembly for transmitting messages and an audio assembly for receiving messages, to promote the passing of messages between loudspeaker components. Similarly, for some embodiments, a first audio assembly 110 may be a loudspeaker assembly that includes both a loudspeaker substantially similar to loudspeaker 130 and a microphone substantially similar to microphone 180, and a second audio assembly 160 may be another loudspeaker assembly that also includes both a loudspeaker substantially similar to loudspeaker 130 and a microphone substantially similar to microphone 180.

[0029] FIG.2 shows an encode-side pipeline 200 of an audio system supporting ADT (such as audio system 100). Encode-side pipeline 200 includes a channel encoding stage 210, a pilot embedding stage 220, a time-frequency data grid stage 230, and an MCLT data encoding stage 240. A message to be transmitted by one component of the audio system through an MCLT-based watermarking ADT technique (as disclosed herein) may be provided to encode-side pipeline 200, and after proceeding through the stages thereof, encode-side pipeline 200 may generate an audio signal having an audio watermark containing the message. The audio signal may in turn be played by, e.g., a loudspeaker, in accordance with the ADT technique. Various portions of encode-side pipeline 200 may accordingly relate to various components forming parts of an encode-side portion of an audio system, such as first audio assembly 110, first device 120, loudspeaker 130, and / or associated interfaces of audio system 100.

[0030] Turning briefly to the next drawing, FIG. 3 shows a decode-side pipeline 300 of an audio system supporting ADT (such as audio system 100). Decode-side pipeline 300 includes an MCLT watermark detector stage 310, an MCLT data extraction stage 320, a channel equalization stage 330, and a channel decoding stage 340. The played audio signal may be received by, e.g., a microphone, and processed by decode-side pipeline 300, which may generate a recovered message, which may match the message sent through the MCLT-based watermarking ADT technique. Various portions of decode-side pipeline 300 may accordingly relate to various components forming parts of a decode-side portion of an audio system, such as second audio assembly 160, second device 170, microphone 180, and / or associated interfaces of audio system 100.Docket No. P230134WO

[0031] Returning to FIG.2, the input to channel encoding stage 210 may be a message to be transmitted, which might not be known at the receiving end, and which may be and / or may include a set of bits (e.g., numerals having a binary range of values, such as 0 and 1). Channel encoding stage 210 may accordingly encode message data to be transmitted. The output of channel encoding stage 210 may be an expanded channel-encoded message, which may be a longer message including a set of bits.

[0032] The expanded channel-encoded message may be provided by channel encoding stage 210 as an input to pilot embedding stage 220. Pilot embedding stage 220 may also have a record of a predetermined set of pilot symbols (each of which may also be a set of bits), which is therefore known at encode-side pipeline 200 and is also assumed to be known at decode-side pipeline 300 (e.g., one or more potions of decode-side pipeline 300 have a record of the predetermined set of pilot symbols). Pilot embedding stage 220 may embed pilot symbols for watermark detection and / or synchronization, as well as to aid channel equalization at the receiver (e.g., on the decode side). The output of pilot embedding stage 220 is a pilot-embedded channel-encoded message, which may be a combined sequence of bits which includes both the expanded channel-encoded message and the pilot symbols.

[0033] The pilot-embedded channel-encoded message may be provided by pilot embedding stage 220 as an input to time-frequency data grid stage 230. The pilot-embedded channel-encoded message data may be converted to a 2D time-frequency data grid. The output of time-frequency data grid stage 230 may be a representation of a 2D matrix of bits to be embedded across a set of carrier frequencies and a set of time indices (e.g., across a set of MCLT frames).

[0034] The 2D matrix of bits may be provided by time-frequency data grid stage 230 as an input to MCLT data encoding stage 240, and a time-domain host audio signal to be watermarked may also be provided to MCLT data encoding stage 240. The data of the 2D matrix of bits may be embedded into the host audio signal using an MCLT-based embedding technique. The output of MCLT data encoding stage 240 may be a time-domain watermarked audio signal, which may contain the expanded channel-encoded message and the pilot symbols (e.g., the pilot-embedded channel-encoded message). In various embodiments, the watermarked audio signal might be sent for further processing and may be sent to a loudspeaker for transmission.

[0035] An advantage of the design of encode-side pipeline 200 and / or the stages therein may be that data embedding can happen in substantially real time on a frame-by-frame basis. As aDocket No. P230134WO result, MCLT data encoding stage 240 might advantageously avoid buffering significant portions of the host audio signal (and / or an entirety of the host audio signal).

[0036] Turning again to FIG. 3, the input to MCLT watermark detector stage 310 may be a received audio signal (e.g., received by a microphone), which may be a watermarked audio signal transmitted by a loudspeaker (of the sort discussed herein). The watermarked audio signal may have been distorted in various ways by virtue of travelling through the acoustic channel (e.g., by propagating from the loudspeaker, through the air, to the microphone). MCLT watermark detector stage 310 may use a local record of the predetermined set of pilot symbols to detect the presence of an encoded message in the audio signal (e.g., by detecting the watermark). The output of MCLT watermark detector stage 310 may be a binary decision regarding whether an audio signal bearing an encoded message has been detected (e.g., in a watermarked audio signal), along with an estimated timestamp (e.g., an approximate location in time) of the encoded message within the received audio signal. The MCLT watermark detection as disclosed herein may advantageously operate in substantially real-time on a sample-by-sample basis.

[0037] The binary decision regarding the presence of a watermarked audio signal, along with the estimated timestamp, may be provided by MCLT watermark detector stage 310 as inputs to MCLT data extraction stage 320, which may also be provided with the received audio signal. Using the estimated timestamp, a time-synchronized segment containing the watermarked audio segment in the MCLT domain may be extracted. The output of MCLT data extraction stage 320 may be a 2D time-frequency grid containing (e.g., consisting of) the complex MCLT coefficients at which the watermark, e.g., channel-encoded message, is embedded.

[0038] The 2D time-frequency grid containing the complex MCLT coefficients at which the watermark is embedded may be provided as input to channel equalization stage 330. The output of channel equalization stage 330 may be a channel-distortion-compensated 2D time-frequency grid of complex MCLT coefficients containing merely the channel-encoded message symbols.

[0039] The channel-distortion-compensated 2D time-frequency grid of complex MCLT coefficients may be provided as input to channel decoding stage 340. The output of channel decoding stage 340 may be the original message prior to channel encoding (e.g., on the encoding side). Accordingly, the output of channel decoding stage 340 may be the same message to be transmitted that was provided as an input to channel encoding stage 210.Docket No. P230134WO

[0040] Portions of the disclosed audio watermarking methods and systems pertaining to pilot embedding stage 220, MCLT data encoding stage 240, MCLT watermark detector stage 310, and / or MCLT data extraction stage 320 may advantageously enable and / or facilitate a more compute-efficient and imperceptible ADT technique which can run in substantially real time, and on relatively low-powered devices. Descriptions herein of channel encoding stage 210, channel equalization stage 330, and channel decoding stage 340 are not intended to be limiting. In various embodiments, any of a variety of channel encoding, channel equalization, and / or channel decoding methods and / or systems may be used.

[0041] In various embodiments, various methods and systems may accordingly implement the disclosed ADT techniques through the use of encode-side pipelines and decode-side pipelines as discussed herein.

[0042] On the encode side, turning to FIG.4, a message provided as input to channel encoding stage 210 may facilitate reliable ADT techniques by addressing distortion caused by channel distortion effects. The message may include a sequence of bits (e.g., numerals having a binary range of values, such as 0 and 1). Channel encoding stage 210 may be implemented in a variety of ways.

[0043] For example, in some embodiments, a channel encoding stage 400 may include a convolutional encoding portion 410 which may convolutionally encode the sequence of bits of the message, a Binary Phase Shift Keying (BPSK) portion 420 which may then map the values of the bits (for example by mapping values of “0” and “1” to “1” and “-1,” respectively), and a repetition portion 430, may repeat each symbol a number of times ^^. However, in various other embodiments, channel encoding stage 400 may include any of a variety of other portions pertinent to channel coding methods, e.g., polar coding portions, low-density parity-check (LDPC) portions, turbo coding portions, combinations of portions pertinent to various channel coding methods, and so on. Channel encoding stage 400 as disclosed herein is accordingly not intended to be limited to coding approaches incorporating convolutional encoding portion 410, BPSK portion 420, and / or repetition portion 430.

[0044] A random interleaving portion 440 of channel encoding stage 400 may then interleave and / or shuffle the data, for example by interleaving and / or shuffling the data with a predetermined key (which may be available for use on the decode side). The interleaving and / or shuffling may spread the data across frequency and across time to such an extent that decoding errors mayDocket No. P230134WO become as random as possible. A random phase alteration portion 450 of channel encoding stage 400 may then alternate the phase of the data with another predetermined key (which may, again, be available for use on the decode side). Phase alteration may advantageously promote watermark imperceptibility.

[0045] Following channel encoding stage 210, in pilot embedding stage 220, a set of pilot symbols may be embedded into the message (e.g., for watermark detection and / or synchronization and channel equalization on the decode side. Pseudo-noise (PN) may be used for pilot symbol generation. A well-selected PN architecture can provide advantageous autocorrelation and / or circular autocorrelation properties. For example, a PN sequence may include a set of pilot symbols having a minimum sidelobe level ratio (SLR) in the autocorrelation function for a given length of sequence. As another example, the PN architecture might incorporate a double phase modulation technique (as disclosed herein), which may include two PN sequences: a first PN sequence applied across frequency (e.g., an intra-frame PN sequence) and a second PN sequence applied across time (e.g., an inter-frame PN sequence). The intra-frame PN sequence may be random, and the inter- frame PN sequence may be optimized for low SLR in the autocorrelation function. The two PN sequences may be multiplied to form a 2D time-frequency grid. Along with SMCLT, such double phase modulation may advantageously enhance or enable computationally feasible watermark detection in low-powered devices. The embedding of the pilot symbols may be in accordance with double phase modulation combined with an MCLT-based audio watermarking method, which may advantageously facilitate low computational complexity.

[0046] FIGS. 5A and 5B depict scenarios of pilot symbol embedding in an MCLT-based audio watermarking technique. The use of pilot symbols may permit and / or facilitate estimates of distortions that may be introduced by the acoustic channel, such as reverberation. Pilot symbols may be embedded together with message symbols and padding symbols. Pilot symbols may be known at both the encoder and the decoder, while message symbols may be known at the encoder but not known at the decoder. Padding symbols may be known at the encoder, and may either be known or not known at the decoder. The padding symbols may aid the filtering operation at the edge of the data packet.

[0047] A data packet may be defined as a 2D grid with pilot symbols, message symbols, and padding symbols spread in frequency and time. Developed channel equalization techniques may support arbitrary data packet designs.Docket No. P230134WO

[0048] In FIG. 5A, in a first pilot-symbol embedding scenario 510, pilot symbols may be spread across carrier frequencies of a pilot frame, and pilot frames may then be spread across every few message frames. Padding frames may start and end the data packet, and may facilitate the filtering operation. In contrast, in FIG.5B, in a second pilot-symbol embedding scenario 520, pilot symbols may be spread across carrier frequencies of a pilot frame, and a block of pilot frames may be positioned between two blocks of message frames. Padding frames may start and end the data packet and may facilitate the filtering operation. Second pilot-symbol embedding scenario 520 may be advantageous since the pilot block may be efficiently reused for data packet detection and / or synchronization.

[0049] Once time-frequency data grid stage 230 has prepared the pilot-embedded channel- encoded message data converted to a 2D time-frequency data grid (e.g., the representation of the 2D matrix of bits to be embedded in the audio signal), at MCLT data encoding stage 240, a 2D frequency-and-time grid of pilot symbols may be embedded in the audio signal, along with message grids, using an MCLT-based watermark embedding technique.

[0050] FIG.6 shows a scenario 600 of encode-side pilot embedding for efficient decode-side data packet detection and / or synchronization. As discussed herein, pilot symbols may be known both on the encode side (e.g., the transmitter side) and at the decode side (e.g., the receiver side). The pilot symbols may be generated using two PN sequences. The first sequence, which may beknown as the intra-frame sequence, may be given by a vector ^^^ of size ^^^ ൈ 1, and may bemodulated with the host audio signal across frequency indices of eachThe second sequence,which may be known as the inter-frame sequence, may be given by a vector s2 of size 1 ൈ ^^ଶ,and may be modulated with the signal across frame indices. Scenario 600 accordingly depicts a double modulation technique. As a result of the double modulation technique, a 2D pilot grid of size ^^^ൈ ^^ଶmay be embedded between message blocks for efficient message detection at the receiver.

[0051] The MCLT technique employed by MCLT data encoding stage 240 may generate ^^ complex-values coefficients from a 2^^-length time-domain signal frame ^^^, where ^^ may denote a frame index and ^^^may have a 50% overlap with adjacent frames. The MCLT may be given by ^^^ ൌ ^^^,^ െ ^^^^^,^(1) ൌ^^^ െ ^^^^^^^^^^ ,(2)Docket No. P230134WO where ^^ and ^^ may be ^^×2^^ cosine and sine modulation matrices, respectively, and ^^ is a diagonal sine window matrix. The elements of ^^, ^^, and ^^ may be given by (3) ^^^^^, ^^^ ൌ ^2 cos ^ ^^ ^ 11 ^^ ൬^^ ^ ^ ൬^^ ^^ ൨ ^^ ^^ (4) (5)where ^^ = 0, 1, …,

[0052] The may 1் ்(6) ^^^ൌ 2^^൫^^ ^^^,^ ^ ^^ ^^^,^൯A final time-domain signalframes.

[0053] These implementations may be based on the concept that the MCLT may be efficiently computed from a Fast Fourier Transform (FFT). Computationally, a number of multiplicationsand / or additions of these implementations may be ^^^logଶ^^ ^ 1^ and ^^^3logଶ^^ ^ 3^ െ 2,respectively.

[0054] In the MCLT-based audio watermarking techniques discussed herein, data may be embedded by modifying the phase of select carriers as follows: ^^^^^^^^ ൌ |^^^^^^^|^^^^^^^ , (7)where d^^^^^ ∈ ^െ1,1^ is a datasignal is reconstructed and MCLT is applied again, the result may be a coefficient ^^^^^^^^ havingthe following relationship with ^^^^^^^^:^^1 ^^^^^^ ^^1 1 ^ (8) ^^^^^ ^ ^^ ^ ^^^^^^ 1^ ^^^^^^ ^ 1^ ^ ^^ ,where^^. In some embodiments, for ^^ not equal to 0 or ^^, the elements of ^^^,^may be derived as:Docket No. P230134WO ^െ ^^ା^ାௗìï^െ^^^1 ^^^2^^ െ 1^^2^^ ^ 1^ if |^^ െ ^^| ൌ 2^^(9) ^^^,^^^^^ ൌ^ା^í^െ1^ï 4else if |^^ െ ^^| ൌ 1î 0 otherwise. The embedding strategy in equation 7 may render accurate data extraction task at the receiver not always possible. Consequently, to ensure that ^^^^^^^^ ൌ |^^^^^^^|^^^^^^^ (10)the embedding strategy in equation 7 may be modified to ^^^ ^^^^ ൌ 2|^^ ^^^^|^^ ^^^^1 െ^^ ^^^ ^^ ^ ^1 ^^(11) ^^ ^ ି^,^ ^ି^ ^2^^^ ^^ െ 1 െ2^^^ ^^ ^ 1 ^ ^^^,^^^^ା^൨ ,may every

[0055] FIG.7 depicts a scenario 700 of an MCLT technique incorporating every-other-frame and every-other-frequency data embedding. The every-other-frame and every-other-frequency data embedding may advantageously address the inherent inter-frame and / or inter-frequency dependencies of MCLT.

[0056] On the decode side, MCLT watermark detector stage 310 may employ a sliding MCLT (SMCLT) method. SMCLT may be used for efficient audio watermark detection and / or synchronization. SMCLT may include computing an MCLT iteratively and efficiently for multiple incoming samples in time (e.g., every incoming sample in time), instead of performing a complete MCLT computation every incoming frame with one sample overlap as may be done by other MCLT-based methods, which may be computationally infeasible in a real-time application. An SMCLT method may also be used to extract the MCLT coefficients independently for specific frequency bins. Hence, SMCLT may also advantageously be computed merely for frequencies of interest.

[0057] MCLT may be related to the Discrete Fourier Transform (DFT) as follows: ^^^^^^ ൌ ^^^^^^^^ ^ ^^^^^ ^ 1^ (12)^^^^^^ ൌ ^^^^^^^^^^^^ (13)ଶ (14) 1 ^ା^ ^Docket No. P230134WO ଶெି^ (15) ^^^ ൌ ^ ^^ ^^ ^^ି^ଶగ^^ ^^ ^ ^ ^ଶெ, where ^^ may be an MCLT ^^ ൌ0, 1, ... , ^^ െ 1 may be abean MCLT for frequency bin ^^.

[0058] Notably, ^^^^^^ may be a DFT for a frame size of 2^^ samples. Hence, the SDFT algorithm can be easily combined with the fast MCLT algorithm above. The SDFT, described in a separate document, is here given by: ^^௧ା^^^^^ ൌ ^^^^^^൫^^௧^^^^ െ ^^^^^^ ^ ^^^^^ ^ 2^^^൯ (16)^where ^^ ൌ 0, 1, … , ∞ may be a time sample index.

[0059] Combining SDFT and fast MCLT may provide the following update rule: ^^௧ା^^^^^ ൌ ^^^^௧ା^^^^^ ^ ^^௧ା^^^^ ^ 1^ (18)^^௧ା^^^^^ ൌ ^^^^^^൫^^௧^^^^ െ ^^^^^^ ^ ^^^^^ ^ 2^^^൯ (19)

[0060] Notably, ^^^^^^ may be a time-invariant complex coefficient combining coefficients from MCLT and SDFT. For efficiency, ^^^^^^ may be saved in a table. Thus, for a given frequency index of interest, a per-sample time complexity of SMCLT, including the computational overhead of SDFT, may merely be one complex 90° rotation, two complex multiplications, one complex addition, two real additions, and two real subtractions. If ^^ is the number of MCLT frequency bins of interest, then in terms of memory complexity, 2^^ previous time samples of the time domain signal plus 2^^ previous DFT coefficients may be buffered. Additionally, another 2^^ precomputed coefficients for the term ^^^^^^ may be stored in a table.

[0061] For every incoming sample at the receiver, SMCLT may be applied to efficiently extract carriers of interest. Then, a DP-based 2D cross-correlator may be used to detect if a watermark is present in the receiver signal and to estimate its timestamp. The DP-based 2D cross- correlator may efficiently compute a cross-correlation function between a 2D segment of the 2D time-frequency grid corresponding to the known pilot symbols and the receiver signal. The processing pipeline of the DP-based 2D cross-correlator may be divided into two sequential steps: (1) correlating SMCLT-extracted coefficients with the intra-frame PN sequence, resulting in anDocket No. P230134WO intra-frame correlation coefficient; and (2) correlating the intra-frame correlation coefficient with an inter-frame PN sequence, resulting in the final cross-correlation coefficient corresponding to the given time sample index. The disclosed DP modulation technique, along with the disclosed SMCLT, may enable or enhance efficient packet detection at the decoder side using the disclosed DP-based cross-correlator.

[0062] An intra-frame correlation coefficient ^^^may be computed as follows: ேି^1భ^ ^^ ^ ^^^^^^^^^^^^^^^ ൌ ^ ^^ ^^, ^^ where ^^ may be a th carriercomputed using SMCLT a may range from 8 to 32 depending on the frame size used at the encoder (which may be defined in a range between 128 and 512 samples). The coefficient ^^^may be saved in a 2D ring buffer to avoid recomputation. In various embodiments, intra-frame correlation may be computed in any of a variety of ways which may incorporate a multiplication of an intra-frame PN sequence and carrier frequency coefficients, followed by a weighted summation of the resulting multiplied coefficients (wherein any of a wide variety of weighting may be performed).

[0063] As an example of saving the coefficient ^^^in the 2D ring buffer, FIGS. 8A and 8Bdepict scenarios of a 2D ring buffer state in which ^^ ൌ 2 and ^^ଶ ൌ 5. A first scenario 810 inFIG.8A depicts a 2D ring buffer state at a time ^^. In comparison, a second scenario 820 in FIG.8B depicts the 2D ring buffer state at a time ^^ ^ 1.

[0064] For a given row in the 2D ring buffer, a second and final correlation function may be computed as ேమି^1 ^ where ^^ଶ^^^^ may be a ^^-thAs disclosed, in MCLT-based audio watermarking, data may be embedded in every other frame, hence the term 2^^. Since ^^^is bounded to be between -1 and 1, regardless of the amount of noise, a simple thresholding decision may be employed to detect whether a data packet is present at time sample k.Docket No. P230134WO

[0065] Notably, the computational bottleneck of the DP-based 2D cross-correlator may be at the second correlation function computed using the inter-frame PN sequence, which might be as long as possible for desirable performance. However, to improve computational performance, the ^^^coefficients may be heavily quantized prior to buffering. For further computational efficiency, multiplications involving a PN sequence coefficient may be implemented as simple bitwise-AND operations merely at the sign bit.

[0066] In various embodiments, the intra-frame PN sequence may include any of a variety of random sequences. Its autocorrelation properties might not matter since it may be applied at non- overlapping frequency bins. A choice of an inter-frame PN sequence with good autocorrelation properties may have greater significance. Hence, the inter-frame PN sequence may be optimized to have low SLR for improved watermark detection performance.

[0067] Notably, by taking advantage of SMCLT combined with double phase modulation using intra-frame and inter-frame PN sequences, the asymptotic computational complexity of thewatermark detector, for a given time sample, may advantageously be merely ^^^^^^ ^ ^^ଶ^compared to ^^^^^ଶ^^^log^^ ^ ^^^^^ of a naïve brute search approach. In turn, this may enable thefeasible and real-time use of the disclosed ADT techniques.

[0068] Once the presence of the watermark is detected and its timestamp has been estimated, at the MCLT data extraction stage 320, inverse MCLT may be computed for the frames containing data. Next, soft data corresponding to a channel distorted version of ^^^^^^^^ (e.g., in equation 10) may be selected for further processing, e.g., channel equalization followed by channel decoding.

[0069] In some embodiments, the MCLT techniques disclosed herein may embed data in the 6.5 kilohertz (kHz) to 9.2 kHz band. For some embodiments, the band may extend to more than 9.2 kHz. In some embodiments, the MCLT techniques disclosed herein may have an MCLT frame size of ^^ set between 128 and 512. For some embodiments, the range of the MCLT frame size ^^ may extend to more than 512.

[0070] As disclosed herein, MCLT based audio watermarking may modify one or more phases of frequency carriers in one or more high frequency bands. MCLT based audio watermarking may also advantageously promote imperceptibility, such as by providing an average objective difference grade with Perceptual Evaluation of Audio Quality (PEAQ) above -1. In some embodiments, the MCLT based audio watermarking may provide an average objective differenceDocket No. P230134WO grade with PEAQ of about -0.5. Finally, MCLT based audio watermarking may advantageously facilitate high data rates.

[0071] Following MCLT data extraction, channel equalization may compensate distortion incurred to a signal following transmission (e.g., from loudspeaker to microphone). A distortion could be noise, frequency-selective impulse response of transmitter and / or receiver, reverberation, or lack of perfect synchronization, any and / or all of which may be present in ADT applications. The first type of distortion could be mitigated with low-rate channel coding, whereas the second, third, and fourth types of distortion might be more effectively addressed, without reducing the channel capacity, with a channel equalizer.

[0072] At channel equalization stage 330, a developed channel equalizer may take the form of a frequency-selective linear filter. At the encoder, a sequence of pilot symbols may be embedded along with the message symbols, and the pilot symbols may also be known at the decoder, where they may be used to estimate the linear filter coefficients. The filter coefficients may be estimated using a least squares (LS) estimator. Unlike traditional channel equalization, the decoder may merely know the phase of the pilot symbols, e.g., the magnitude of the pilot symbols will be proportional to the host audio signal, which might not be known in advance at the decoder side (as in blind audio watermarking). Hence, the extracted MCLT coefficients may be normalized to the unit circle to reduce the effect of varying energy in the host audio.

[0073] At channel decoding stage 340, the portions of channel encoding stage 210 may be reversed. First, a channel decoding stage 900 may include a phase de-alteration portion 910, at which the phase of the soft input may be de-alternated using the predetermined key of random phase alteration portion 450. Next, at a de-interleaving portion 920, the data may be de-interleaved using the predetermined key of random interleaving portion 440.

[0074] Then, in some embodiments, channel decoding stage 900 may also include a maximum likelihood (ML) estimate portion 930 which may de-alternate the phase of the soft input (e.g., by using the predetermined key of random phase alteration portion 450), and a soft input hard output (SIHO) Viterbi decoder portion 940 which may decode the convolutional codes. However, in various other embodiments, channel decoding stage 900 may include any of a variety of other portions pertinent to channel decoding methods. Channel decoding stage 900 as disclosed herein is accordingly not intended to be limited to decoding approaches incorporating ML estimate portion 930 and / or SIHO Viterbi decoder portion 940.Docket No. P230134WO

[0075] FIGS.10A and 10B show a first exemplary method for ADT. In FIGS.10A and 10B, a method 1000 comprises a receiving 1055, a processing 1060, an extracting 1070, and a processing 1080. In various embodiments, method 1000 may comprise a channel-encoding 1010, an embedding 1020, a generating 1030, a generating 1035, a generating 1040, a channel-equalizing 1085, and / or a channel-decoding 1090.

[0076] In receiving 1055, a decode-side digital signal may be generated by a microphone. The decode-side digital signal may be based on an audio signal that is generated by a loudspeaker based on an encode-side digital signal from a first device. In processing 1060, the decode-side digital signal may be processed, by a second device using one or more of an SMCLT and a 2D DP-based cross-correlator, to detect that a watermarked data packet is present in the decode-side digital signal, and also to estimate a timestamp of the watermarked data packet within the decode-side digital signal. In extracting 1070, based on detecting that the watermarked data packet is present in the decode-side digital signal, extracting a segment of the decode-side digital signal containing the watermarked data packet may be extracted, the segment of the decode-side digital signal being based on the estimated timestamp. In processing 1080, the segment may be processed to generate a recovered message.

[0077] In some embodiments, in channel-encoding 1010, a message may be channel-encoded. For some embodiments, in embedding 1020, pilot symbols may be embedded into the channel- encoded message. In some embodiments, in generating 1030, a 2D time-frequency data grid may be generated based on the pilot-embedded channel-encoded message. For some embodiments, in generating 1035, a plurality of data-embedded MCLT frames in accordance with MCLT-based watermarking may be generated based on a time-domain audio signal and the 2D time-frequency data grid. In generating 1040, the encode-side digital signal may be generated based on the plurality of data-embedded MCLT frames. The encode-side digital signal may be a watermarked time-domain audio signal.

[0078] In some embodiments, in channel-equalizing 1085, the processing of the segment to generate the recovered message may include channel-equalizing the segment using a frequency- selective linear filter to generate a plurality of channel-distortion-compensated MCLT coefficients for a 2D time-frequency data grid. For some embodiments, the plurality of channel-distortion- compensated MCLT coefficients may contain channel-encoded symbols for the recovered message. In some embodiments, in channel-decoding 1090, the processing of the segment toDocket No. P230134WO generate the recovered message may include channel-decoding the plurality of channel-distortion- compensated MCLT coefficients to generate the recovered message.

[0079] FIGS. 11A and 11B show a second exemplary method for ADT. In FIGS. 11A and 11B, a method 1100 comprises a channel-encoding 1110, an embedding 1120, a generating 1130, a generating 1135, a generating 1140, and a generating 1150. In various embodiments, method 1100 may comprise a generating 1155, a processing 1160, an extracting 1170, a channel-equalizing 1180, and / or a channel-decoding 1190.

[0080] In channel-encoding 1110, a message may be channel-encoded, using a first device. In embedding 1120, a plurality of pilot symbols may be embedded into the channel-encoded message. The plurality of pilot symbols may be based on an intra-frame PN sequence and inter- frame PN sequence, and the embedding may be in accordance with double phase modulation. In generating 1130, a 2D time-frequency data grid may be generated based on the pilot-embedded channel-encoded message. In generating 1135, a plurality of data-embedded MCLT frames may be generated in accordance with MCLT-based watermarking, based on a time-domain audio signal and the 2D time-frequency data grid. In generating 1140, a watermarked time-domain audio signal may be generated based on the plurality of data-embedded MCLT frames. In generating 1150, an audio signal based on the watermarked time-domain audio signal may be generated by a loudspeaker.

[0081] In some embodiments, the intra-frame PN sequence may be generated in a random manner and the inter-frame PN sequence may be optimized to have a low SLR for improved watermark detection performance. For some embodiments, data may be embedded every-other- frame and every-other-frequency in the MCLT domain.

[0082] In some embodiments, in generating 1155, a decode-side digital signal based on the watermarked time-domain audio signal may be generated by a microphone. For some embodiments, in processing 1160, the decode-side digital signal may be processed, by a second device using one or more of an SMCLT and a DP-based 2D cross-correlator, to detect that a watermarked data packet is present in the decode-side digital signal provided by the microphone, and / or to estimate a timestamp of the watermarked data packet within the decode-side digital signal.

[0083] In extracting 1170, a segment of the decode-side digital signal containing the watermarked data packet may be extracted based on detecting that the watermarked data packet isDocket No. P230134WO present in the decode-side digital signal, the segment of the decode-side digital signal being based on the estimated timestamp. For channel-equalizing 1180, the segment may be channel-equalized to generate a plurality of channel-distortion-compensated MCLT coefficients for the 2D time- frequency data grid. The plurality of channel-distortion-compensated MCLT coefficients may contain channel-encoded symbols for the message. In channel-decoding 1190, the plurality of channel-distortion-compensated MCLT coefficients may be channel-decoded to recover the message.

[0084] The methods disclosed herein (e.g., method 1000 and / or method 1100) may be configured for the operation of the systems disclosed herein (e.g., audio system 100 and / or components thereof, on an encoder side and / or a decoder side). Thus, the same advantages that apply to the systems may apply to the methods. In addition, in various embodiments, first audio assembly 110, first device 120, second audio assembly 160, and / or second device 170 may comprise one or more processors and a non-transitory memory having executable instructions that, when executed, cause the one or more corresponding processors to carry out various portions of the methods disclosed herein.

[0085] Note that the example control and estimation routines included herein can be used with various system configurations. The control methods and routines disclosed herein may be stored as executable instructions in non-transitory memory and may be carried out by the control system including the controller in combination with the various sensors, actuators, and other hardware. The specific routines described herein may represent one or more of any number of processing strategies such as event-driven, interrupt-driven, multi-tasking, multi-threading, and the like. As such, various actions, operations, and / or functions illustrated may be performed in the sequence illustrated, in parallel, or in some cases omitted. Likewise, the order of processing is not necessarily required to achieve the features and advantages of the example embodiments described herein, but is provided for ease of illustration and description. One or more of the illustrated actions, operations, and / or functions may be repeatedly performed depending on the particular strategy being used. Further, the described actions, operations, and / or functions may graphically represent code to be programmed into non-transitory memory of a computer readable storage medium, where the described actions are carried out by executing the instructions in a system including the various hardware components in combination with the electronic controller.Docket No. P230134WO

[0086] The description of embodiments has been presented for purposes of illustration and description. Suitable modifications and variations to the embodiments may be performed in light of the above description or may be acquired from practicing the methods. For example, unless otherwise noted, one or more of the described methods may be performed by a suitable device and / or combination of devices, such as the audio systems and components thereof described above with respect to FIGS.1-9. The methods may be performed by executing stored instructions with one or more logic devices (e.g., processors) in combination with one or more additional hardware elements, such as storage devices, memory, image sensors / lens systems, light sensors, hardware network interfaces / antennas, switches, actuators, clock circuits, and so on. The described methods and associated actions may also be performed in various orders in addition to the order described in this application, in parallel, and / or simultaneously. The described systems are exemplary in nature, and may include additional elements and / or omit elements. The subject matter of the present disclosure includes all novel and non-obvious combinations and sub-combinations of the various systems and configurations, and other features, functions, and / or properties disclosed.

[0087] The disclosure provides support for a method for ADT, comprising: receiving, from a microphone, a decode-side digital signal, the decode-side digital signal being generated by the microphone based on an audio signal, and the audio signal being generated by a loudspeaker based on an encode-side digital signal from a first device, processing the decode-side digital signal, by a second device using one or more of a SMCLT and a 2D DP-based cross-correlator, to detect that a watermarked data packet is present in the decode-side digital signal, and to estimate a timestamp of the watermarked data packet within the decode-side digital signal, and based on detecting that the watermarked data packet is present in the decode-side digital signal, extracting a segment of the decode-side digital signal containing the watermarked data packet, the segment of the decode- side digital signal being based on the estimated timestamp, and processing the segment to generate a recovered message. In a first example of the method, the processing of the segment to generate the recovered message further includes channel-equalizing the segment using a frequency- selective linear filter to generate a plurality of channel-distortion-compensated MCLT coefficients for a 2D time-frequency data grid. In a second example of the method, optionally including the first example, the plurality of channel-distortion-compensated MCLT coefficients contain channel-encoded symbols for the recovered message. In a third example of the method, optionally including one or both of the first and second examples, the processing of the segment to generateDocket No. P230134WO the recovered message further includes channel-decoding the plurality of channel-distortion- compensated MCLT coefficients to generate the recovered message. In a fourth example of the method, optionally including one or more or each of the first through third examples, the method further comprises: channel-encoding a message. In a fifth example of the method, optionally including one or more or each of the first through fourth examples, the method further comprises: embedding pilot symbols into the channel-encoded message. In a sixth example of the method, optionally including one or more or each of the first through fifth examples, the method further comprises: generating, based on the pilot-embedded channel-encoded message, a 2D time- frequency data grid, and generating, based on a time-domain audio signal and the 2D time- frequency data grid, a plurality of data-embedded MCLT frames in accordance with MCLT-based watermarking In a seventh example of the method, optionally including one or more or each of the first through sixth examples, the method further comprises: generating, based on the plurality of data-embedded MCLT frames, the encode-side digital signal, wherein the encode-side digital signal is a watermarked time-domain audio signal.

[0088] The disclosure also provides support for a method for ADT, the method comprising: channel-encoding a message, using a first device, embedding a plurality of pilot symbols into the channel-encoded message, the plurality of pilot symbols being based on an intra-frame PN sequence and inter-frame PN sequence, and the embedding being in accordance with double phase modulation, generating, based on the pilot-embedded channel-encoded message, a 2D time- frequency data grid, generating, based on a time-domain audio signal and the 2D time-frequency data grid, a plurality of data-embedded MCLT frames in accordance with MCLT-based watermarking, generating a watermarked time-domain audio signal based on the plurality of data- embedded MCLT frames, and generating, by a loudspeaker, an audio signal based on the watermarked time-domain audio signal. In a first example of the method, the intra-frame PN sequence is generated in a random manner and the inter-frame PN sequence is optimized to have a low SLR for improved watermark detection performance. In a second example of the method, optionally including the first example, data is embedded every-other-frame and every-other- frequency in the MCLT domain. In a third example of the method, optionally including one or both of the first and second examples, the method further comprises: generating, by a microphone, a decode-side digital signal based on the watermarked time-domain audio signal. In a fourth example of the method, optionally including one or more or each of the first through thirdDocket No. P230134WO examples, the method further comprises: processing the decode-side digital signal, by a second device using one or more of a SMCLT and a DP-based 2D cross-correlator, to detect that a watermarked data packet is present in the decode-side digital signal provided by the microphone, and to estimate a timestamp of the watermarked data packet within the decode-side digital signal. In a fifth example of the method, optionally including one or more or each of the first through fourth examples, the method further comprises: based on detecting that the watermarked data packet is present in the decode-side digital signal, extracting a segment of the decode-side digital signal containing the watermarked data packet, the segment of the decode-side digital signal being based on the estimated timestamp. In a sixth example of the method, optionally including one or more or each of the first through fifth examples, the method further comprises: channel-equalizing the segment to generate a plurality of channel-distortion-compensated MCLT coefficients for the 2D time-frequency data grid, wherein the plurality of channel-distortion-compensated MCLT coefficients contain channel-encoded symbols for the message. In a seventh example of the method, optionally including one or more or each of the first through sixth examples, the method further comprises: channel-decoding the plurality of channel-distortion-compensated MCLT coefficients to recover the message.

[0089] The disclosure also provides support for a system for ADT, comprising: a decode-side interface for an output of a microphone, and a decode-side device coupled to the decode-side interface, the decode-side device having one or more decode-side processors and a decode-side non-transitory memory having executable instructions that, when executed, cause the one or more decode-side processors to: process a decode-side digital signal provided by the microphone, using one or more of a SMCLT and a DP-based 2D cross-correlator, to detect that a watermarked data packet is present in the decode-side digital signal, and to estimate the timestamp of the watermarked data packet within the decode-side digital signal, based on detecting that the watermarked data packet is present in the decode-side digital signal, extracting a segment of the decode-side digital signal containing the watermarked data packet, the segment of the decode-side digital signal being based on the estimated timestamp, channel-equalize the segment using a frequency-selective linear filter to generate a plurality of channel-distortion-compensated MCLT coefficients for a 2D time-frequency data grid, and channel-decode the plurality of channel- distortion-compensated MCLT coefficients to generate a recovered message. In a first example of the system, the system further comprises: an encode-side interface for an input of a loudspeaker,Docket No. P230134WO and an encode-side device coupled to the encode-side interface, the encode-side device having one or more encode-side processors and an encode-side non-transitory memory having executable instructions that, when executed, cause the one or more encode-side processors to: channel-encode a message to generate a channel-encoded message, embed a plurality of pilot symbols into the channel-encoded message to generate a pilot-embedded channel-encoded message, the plurality of pilot symbols being based on an intra-frame PN sequence and inter-frame PN sequence, and the embedding being in accordance with double phase modulation, generate, based on the pilot- embedded channel-encoded message, a 2D time-frequency data grid, generate, based on a time- domain audio signal and the 2D time-frequency data grid, a plurality of data-embedded MCLT frames in accordance with MCLT-based watermarking, generate a watermarked time-domain audio signal based on the plurality of data-embedded MCLT frames, and generate, by the loudspeaker, an audio signal based on the watermarked time-domain audio signal. In a second example of the system, optionally including the first example, the intra-frame PN sequence is generated in a random manner and the inter-frame PN sequence is optimized to have a low SLR for improved watermark detection performance. In a third example of the system, optionally including one or both of the first and second examples, data is embedded every-other-frame and every-other-frequency in the MCLT domain.

[0090] As used in this application, an element or step recited in the singular and proceeded with the word “a” or “an” should be understood as not excluding plural of said elements or steps, unless such exclusion is stated. Furthermore, references to “one embodiment” or “one example” of the present disclosure are not intended to be interpreted as excluding the existence of additional embodiments that also incorporate the recited features. Terms such as “first,” “second,” “third,” and so on are used merely as labels, and are not intended to impose numerical requirements or a particular positional order on their objects. The following claims particularly point out subject matter from the above disclosure that is regarded as novel and non-obvious.

Claims

Docket No. P230134WO CLAIMS:

1. A method for acoustic data transfer (ADT), comprising: receiving, from a microphone, a decode-side digital signal, the decode-side digital signal being generated by the microphone based on an audio signal, and the audio signal being generated by a loudspeaker based on an encode-side digital signal from a first device; processing the decode-side digital signal, by a second device using one or more of a Sliding Modulated Complex Lapped Transform (SMCLT) and a two-dimensional (2D) dynamic programming (DP)-based cross-correlator, to detect that a watermarked data packet is present in the decode-side digital signal, and to estimate a timestamp of the watermarked data packet within the decode-side digital signal; and based on detecting that the watermarked data packet is present in the decode-side digital signal, extracting a segment of the decode-side digital signal containing the watermarked data packet, using the estimated timestamp; and processing the segment to generate a recovered message.

2. The method for ADT of claim 1, wherein the processing of the segment to generate the recovered message further includes channel-equalizing the segment using a frequency-selective linear filter to generate a plurality of channel-distortion-compensated Modulated Complex Lapped Transform (MCLT) coefficients for a 2D time-frequency data grid.

3. The method for ADT of claim 2, wherein the plurality of channel-distortion-compensated MCLT coefficients contain channel-encoded symbols for the recovered message.

4. The method for ADT of claim 2, wherein the processing of the segment to generate the recovered message further includes channel-decoding the plurality of channel-distortion-compensated MCLT coefficients to generate the recovered message.Docket No. P230134WO 5. The method for ADT of claim 1, further comprising: channel-encoding a message.

6. The method for ADT of claim 5, further comprising: embedding pilot symbols into the channel-encoded message.

7. The method for ADT of claim 6, further comprising: generating, based on the pilot-embedded channel-encoded message, a 2D time-frequency data grid; and generating, based on a time-domain audio signal and the 2D time-frequency data grid, a plurality of data-embedded Modulated Complex Lapped Transform (MCLT) frames in accordance with MCLT-based watermarking.

8. The method for ADT of claim 7, further comprising: generating, based on the plurality of data-embedded MCLT frames, the encode-side digital signal, wherein the encode-side digital signal is a watermarked time-domain audio signal.

9. A method for acoustic data transfer (ADT), the method comprising: channel-encoding a message, using a first device; embedding a plurality of pilot symbols into the channel-encoded message, the plurality of pilot symbols being based on an intra-frame pseudo-noise (PN) sequence and inter-frame PN sequence, and the embedding being in accordance with double phase modulation; generating, based on the pilot-embedded channel-encoded message, a two-dimensional (2D) time-frequency data grid; generating, based on a time-domain audio signal and the 2D time-frequency data grid, a plurality of data-embedded Modulated Complex Lapped Transform (MCLT) frames in accordance with MCLT-based watermarking; generating a watermarked time-domain audio signal based on the plurality of data- embedded MCLT frames; andDocket No. P230134WO generating, by a loudspeaker, an audio signal based on the watermarked time-domain audio signal.

10. The method for ADT of claim 9, wherein the intra-frame PN sequence is generated in a random manner and the inter- frame PN sequence is optimized to have low sidelobe level ratio (SLR).

11. The method for ADT of claim 9, wherein data is embedded every-other-frame and every-other-frequency in the MCLT domain.

12. The method for ADT of claim 9, further comprising: generating, by a microphone, a decode-side digital signal based on the watermarked time- domain audio signal.

13. The method for ADT of claim 12, further comprising: processing the decode-side digital signal, by a second device using one or more of a Sliding MCLT (SMCLT) and a dynamic programming (DP)-based 2D cross-correlator, to detect that a watermarked data packet is present in the decode-side digital signal provided by the microphone, and to estimate a timestamp of the watermarked data packet within the decode-side digital signal.

14. The method for ADT of claim 13, further comprising: based on detecting that the watermarked data packet is present in the decode-side digital signal, extracting a segment of the decode-side digital signal containing the watermarked data packet, using the estimated timestamp.

15. The method for ADT of claim 14, further comprising: channel-equalizing the segment to generate a plurality of channel-distortion-compensated MCLT coefficients for the 2D time-frequency data grid,Docket No. P230134WO wherein the plurality of channel-distortion-compensated MCLT coefficients contain channel-encoded symbols for the message.

16. The method for ADT of claim 15, further comprising: channel-decoding the plurality of channel-distortion-compensated MCLT coefficients to recover the message.

17. A system for acoustic data transfer (ADT), comprising: a decode-side interface for an output of a microphone; and a decode-side device coupled to the decode-side interface, the decode-side device having one or more decode-side processors and a decode-side non-transitory memory having executable instructions that, when executed, cause the one or more decode-side processors to: process a decode-side digital signal provided by the microphone, using one or more of a Sliding Modulated Complex Lapped Transform (SMCLT) and a dynamic programming (DP)-based two-dimensional (2D) cross-correlator, to detect that a watermarked data packet is present in the decode-side digital signal, and to and to estimate the timestamp of the watermarked data packet within the decode-side digital signal; based on detecting that the watermarked data packet is present in the decode-side digital signal, extracting a segment of the decode-side digital signal containing the watermarked data packet, the segment of the decode-side digital signal being based on the estimated timestamp; channel-equalize the segment using a frequency-selective linear filter to generate a plurality of channel-distortion-compensated Modulated Complex Lapped Transform (MCLT) coefficients for a 2D time-frequency data grid; and channel-decode the plurality of channel-distortion-compensated MCLT coefficients to generate a recovered message.

18. The system for ADT of claim 17, further comprising: an encode-side interface for an input of a loudspeaker; andDocket No. P230134WO an encode-side device coupled to the encode-side interface, the encode-side device having one or more encode-side processors and an encode-side non-transitory memory having executable instructions that, when executed, cause the one or more encode-side processors to: channel-encode a message to generate a channel-encoded message; embed a plurality of pilot symbols into the channel-encoded message to generate a pilot- embedded channel-encoded message, the plurality of pilot symbols being based on an intra- frame pseudo-noise (PN) sequence and inter-frame PN sequence, and the embedding being in accordance with double phase modulation; generate, based on the pilot-embedded channel-encoded message, a 2D time-frequency data grid; generate, based on a time-domain audio signal and the 2D time-frequency data grid, a plurality of data-embedded MCLT frames in accordance with MCLT-based watermarking; generate a watermarked time-domain audio signal based on the plurality of data- embedded MCLT frames; and generate, by the loudspeaker, an audio signal based on the watermarked time-domain audio signal.

19. The system for ADT of claim 18, wherein the intra-frame PN sequence is generated in a random manner and the inter-frame PN sequence is optimized to have low sidelobe level ratio (SLR).

20. The system for ADT of claim 18, wherein data is embedded every-other-frame and every-other-frequency in the MCLT domain.

Citation Information

Patent Citations

  • Audio watermarking via phase modification

    US10210875B2

  • Watermark decoder and method for providing binary message data

    US20130218313A1

  • Watermark signal provision and watermark embedding

    US8965547B2

  • Watermark generator, watermark decoder, method for providing a watermark signal in dependence on binary message data, method for providing binary message data in dependence on a watermarked signal and computer program using a differential encoding

    US9350700B2