Spectrogram-based temporal alignment for independent recording and playback systems
By combining spectrograms and cloud computing processors on smartphones, high-resolution impulse response measurements of multi-channel speaker systems were achieved, overcoming the limitations of accuracy and resolution in traditional methods and providing time alignment and deconvolution processing in complex environments.
Patent Information
- Application Number
- CN202480047688.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-07-31
- Filing Date
- 2024-08-05
- Publication Date
- 2026-02-13
AI Technical Summary
In multi-channel speaker systems, existing technologies struggle to accurately measure impulse response, especially when there are a large number of speakers, complex measurement microphone setups, and background noise. Traditional methods are limited by the fast Fourier transform calculations on smartphone DSPs, resulting in compromised accuracy and resolution of the impulse response.
A spectrogram-based time alignment method is adopted, which calculates the impulse response in real time on a smartphone, uses interval-by-interval matched filtering and statistical analysis of the spectrogram, and combines cloud computing processors to achieve high-resolution fast Fourier transform, overcoming hardware limitations and providing time alignment for independent recording and playback systems.
It achieves accurate time alignment of multi-channel speaker systems in complex environments, improves the accuracy and resolution of impulse response, reduces the impact of noise interference, and ensures the synchronous deconvolution processing effect of the speaker system.
Smart Images

Figure CN121533040A_ABST
Abstract
Description
Technical Field
[0001] One or more embodiments generally relate to time alignment for independent recording and playback, and more specifically, to providing spectrogram-based time alignment for independent recording and playback systems. Background Technology
[0002] Traditionally, equalization of indoor loudspeaker systems is achieved by exciting one loudspeaker at a time. However, with higher numbers of loudspeakers, limitations in measurement microphone setup, interference from conventional stimuli, and background noise, measuring the impulse response of a multichannel system in real time can be cumbersome. Furthermore, the accuracy and resolution of the impulse response are compromised due to the limitations of Fast Fourier Transform (FFT) calculations on smartphone digital signal processing (DSP) devices. Summary of the Invention
[0003] Technical solution One embodiment provides a computer-implemented method comprising sending a stimulus signal to a speaker; receiving a measurement signal via a microphone; converting the stimulus signal into a stimulus time-frequency representation; converting the measurement signal into a measurement time-frequency representation; selecting at least one frequency value between the stimulus time-frequency representation and the measurement time-frequency representation; using the selected at least one frequency value to perform a correlation (e.g., linear correlation) analysis; and determining a statistical pattern based on the correlation analysis to generate the start time of the stimulus signal. In some embodiments, a nonlinear analysis including a nonlinear model (e.g., a neural network) may be trained to identify the start or stop time of the stimulus signal in a recording.
[0004] Another embodiment includes a non-transitory processor-readable medium comprising a program that, when executed by a processor, provides the start time of a stimulus in a measurement, the providing step including the processor sending a stimulus signal to a speaker. The processor receives the measurement signal via a microphone. The processor transforms the stimulus signal into a stimulus time-frequency representation. The processor also transforms the measurement signal into a measurement time-frequency representation. The processor further selects at least one frequency value between the stimulus time-frequency representation and the measurement time-frequency representation. The processor also uses the selected at least one frequency value to perform a correlation (e.g., linear correlation) analysis. The processor further determines a statistical pattern based on the correlation analysis to generate the start time of the stimulus signal.
[0005] Another embodiment provides a device comprising: a memory storing instructions; and at least one processor executing the instructions, the instructions including processing configured to send a stimulus signal to a speaker. A measurement signal is received via a microphone. The stimulus signal is converted into a stimulus time-frequency representation. The measurement signal is converted into a measurement time-frequency representation. At least one frequency value is selected between the stimulus time-frequency representation and the measurement time-frequency representation. A correlation analysis is performed using the selected at least one frequency value. Based on the correlation analysis, a statistical pattern is determined to generate the start time of the stimulus signal.
[0006] These and other features, aspects and advantages of one or more embodiments will be understood with reference to the following description, the appended claims and the accompanying drawings. Attached Figure Description
[0007] To gain a more complete understanding of the nature and advantages of the embodiments and the preferred modes of use, reference should be made to the following detailed description, which should be read in conjunction with the accompanying drawings, wherein: Figure 1 This is an example architecture for measurement settings based on some embodiments; Figure 2A An example impulse response for deconvolution of multiple channels is shown using the disclosed techniques according to some embodiments; Figure 2B An example impulse response for deconvolution of multiple channels using threshold-based signal alignment is shown; Figure 3A Example graphs are shown, based on some embodiments, of interval matched filtering using cross-correlation based on a time alignment technique using monophonic scanning; Figure 3B Example graphs are shown, based on some embodiments, of interval matched filtering using cross-correlation based on a time alignment technique using monophonic scanning; Figure 3C The statistical pattern based on the start time point line plot is shown. Figures 3A-3B Example curves; Figure 4 An example audio recording spectrogram is shown according to some embodiments; Figure 5 An example input stimulus spectrogram is shown according to some embodiments; Figure 6 An example time-domain plot of a noisy recording is shown; Figure 7 It shows Figure 6 Example spectrogram of a noisy recording; Figure 8An example impulse response is shown in a smartphone application plotted using an incorrect start time under non-stationary noise, according to some embodiments; Figure 9 An example impulse response is shown in a smartphone application plotted using an incorrect start time under non-stationary noise, according to some embodiments; Figure 10 An example dot-line diagram of the impulse response of a smartphone application with a signal-to-noise ratio (SNR) of 0 dB (stable noise) is shown according to some embodiments; Figure 11 An example impulse response of the main channel of a smartphone application with 0dB SNR (stable noise) is shown according to some embodiments; Figure 12 An example spectrogram of a noisy (smooth) recording with 0dBSNR according to some embodiments is shown; Figure 13 An example is shown of the impulse response of twelve (12) channels of a soundbar measured using an external microphone according to some embodiments of the disclosed techniques; Figure 14 Illustrations are shown according to some embodiments Figure 13 A close-up view of the impulse response dot plot of the height channel; Figure 15 An example heatmap of relative channel delay error according to some embodiments is shown; Figure 16 An example level measurement heatmap is shown according to some embodiments for the accuracy of level measurements using a smartphone and a sound level meter; Figure 17 An example level measurement heatmap is shown, according to some embodiments, for the accuracy of levels measured using a smartphone application; Figure 18 An example screen display of a smartphone application according to some embodiments is shown; Figure 19 An example screen display of a smartphone application according to some embodiments is shown; Figure 20 Examples of screen displays of a smartphone application connected to a cloud computing environment using the disclosed techniques are shown according to some embodiments; Figure 21 Examples of screen displays of a smartphone application connected to a cloud computing environment using the disclosed techniques are shown according to some embodiments; Figure 22 Examples of screen displays of smartphone applications connected to a cloud computing environment using the disclosed techniques are shown according to some embodiments; and Figure 23 The process for time alignment for independent recording and playback, according to some embodiments, is illustrated. Detailed Implementation
[0008] The following description is for the purpose of illustrating the general principles of one or more embodiments and is not intended to limit the inventive concept claimed herein. Furthermore, specific features described herein may be used in combination with other described features in every possible combination and arrangement. Unless otherwise expressly defined herein, all terms shall be given their broadest interpretation, including the meaning implied in the specification and the meaning understood by those skilled in the art and / or the meaning as defined in dictionaries, papers, etc.
[0009] Descriptions of exemplary embodiments are provided on the following pages. The text and accompanying drawings are provided by way of example only to aid the reader in understanding the disclosed technology. They are not intended and will not be construed as limiting the scope of this disclosure in any way. Although certain embodiments and examples have been provided, it will be apparent to those skilled in the art, based on the disclosure herein, that changes can be made to the illustrated embodiments and examples without departing from the scope of the disclosed technology.
[0010] Some embodiments generally relate to time alignment for independent recording and playback, and more specifically, to providing spectrogram-based time alignment for independent recording and playback systems. One embodiment provides a computer-implemented method comprising: accessing a stimulus signal at each of a plurality of frequency bands; accessing a measured signal at each of the plurality of frequency bands; transforming the stimulus signal into a stimulus time-frequency representation; transforming the measured signal into a measured time-frequency representation; selecting at least one frequency value between the stimulus time-frequency representation and the measured time-frequency representation; performing a correlation analysis using the selected at least one frequency value; and determining a statistical pattern based on the correlation analysis to estimate the start time of the stimulus signal originating from the measured signal.
[0011] In one or more embodiments, the disclosed techniques enable synchronized deconvolution of multi-channel speaker systems. Some embodiments use a set of cyclically shifted sinusoidal scanning stimuli to excite the speakers and compute the impulse response in real time on smartphone (or other similar devices, such as tablets) applications via a cloud-based architecture. Independent recording and playback systems, as well as manual delays or system latency due to Bluetooth, Wi-Fi, or cloud-based communications, pose further challenges to the accuracy of measurements. To overcome these complexities, one or more embodiments implement a time alignment method that uses interval-by-interval matched filtering of the spectrogram, followed by statistical analysis of the results.
[0012] In some embodiments, as shown in Equation (Eq.) (1), the synchronous deconvolution process computes the impulse response estimate using a logarithmic scan of the inverse spectrum of autocorrelation and cross-correlation.
[0013] Eq. (1) in,
[0014] Input signal: ,in, From Eq. (1).
[0015]
[0016] .
[0017] Note that for an N-channel speaker system, the j-th channel input is... <m>One of the log scans is P-cyclically shifted.
[0018] Accurate impulse response calculations from deconvolution operations require high resolution fast Fourier transforms (FFTs). In some embodiments, on-device implementation of the synchronous deconvolution process as a smart phone application can be limited by the length of the FFT, which can be implemented on a processor. Applications attempting to implement high resolution FFTs often encounter overruns. To address the problem of hardware overruns, in one or more embodiments, processing is performed on a cloud computing based processor. In some embodiments, the application is implemented using MATLAB Mobile or other similar tools. In one example embodiment, the architecture implementing the synchronous deconvolution processing / algorithm application can be built using MATLAB Mobile's connection to MATHWORKS® cloud features. This enables the implementation of FFTs of lengths on the order of 219 (required for the example application) and above, which is sufficient to deconvolve scans of approximately the same length.
[0019] Figure 1 is an example architecture for a measurement setup according to some embodiments. The example architecture includes a phone 110 (e.g., a smart phone, tablet computer, etc.), a cloud computing environment 120 (e.g., one or more cloud computing devices, servers, etc.), and a speaker system 130 (e.g., a home theater system, a soundbar, etc.). In one or more embodiments, the synchronous deconvolution processing / algorithm application is launched from the phone 110. The cloud computing environment 120 plays the scan and the phone 110 records the measurement. In some embodiments, the measurement can be estimated from measurements taken with the phone 110 with its internal microphone and an external measurement microphone. In one or more embodiments, the recording from the phone 110 is sent to the cloud computing environment 120 where the processing / algorithm is implemented. The example architecture indicates the use of cloud computing based control to synchronize the start and stop times of the recording and playback from the speaker system 130.
[0020] Figure 2A An example impulse response 210 for deconvolution for multiple channels using the disclosed technology is shown according to some embodiments. In the example impulse response, the disclosed technology aligns the signals at 3.873 seconds. The channels include a left channel (L), a right channel (R), a center channel (C), a low frequency effect (LFE) or subwoofer channel, a left surround channel (Ls), a right surround channel (Rs), a left rear surround (Lrs), a right rear surround (Rrs), a left height channel (Lh), a right height channel (Rh), a left height surround channel (Lhs), and a right height surround channel (Rhs).
[0021] Figure 2B Example impulse responses 220 are shown that use threshold-based signal alignment for deconvolution of multiple channels. In the example impulse responses, threshold-based signal alignment aligns the signals at 3.027 seconds. The channels include L, R, C, LFE (or subwoofer channel), Ls, Rs, Lrs, Rrs, Lh, Rh, Lhs, and Rhs.
[0022] Figures 3A-3B Example plots 310 and 320 are shown that result from interval-by-interval matched filtering using cross-correlation based on time alignment techniques using single channel sweeps, according to some embodiments. In some embodiments, in a real-time smartphone application for measuring loudspeaker systems using N-channel deconvolution (N > 1) operations, it can be necessary to estimate the start time of the sweep in the recorded signal to accurately deconvolve the room loudspeaker impulse responses. If the start time is not correctly estimated from the recorded signal, the result is skewed and noisy impulse responses, thus inaccurately calculating the relative loudspeaker delays and levels at the microphone locations. Start time estimation methods using threshold-based approaches are severely affected by room reflections, background noise, and the buffer time between the input signal and the recorded signal, and can result in distorted pulses. Figure 3C Example plot 330 is shown that is based on Figures 3A-3B statistical patterns for start time point line plots.
[0023] Figure 4 Example recording spectrogram 400 is shown, according to some embodiments. Figure 5 Example input stimulus spectrogram 500 is shown, according to some embodiments. Time alignment based on high time resolution spectrograms (64 samples window and 50% overlap) is robust and reliable in real-time applications. Some embodiments use a frequency-by-frequency interval matched filtering of two spectrograms (e.g., as shown in Figure 4 and Figure 5 ). Figure 4 and Figure 5 represent the recording spectrogram and the input stimulus spectrogram, respectively. Note that this technique is independent of the number of channels being deconvolved, so the disclosed technique can be used to time align (N > 1) channels at a time. Spectrogram 400 shows graphically how the time alignment technique is implemented on a multi-channel exponential sine sweep (i.e., multiple exponential sweeps applied to multiple loudspeakers simultaneously). Statistical analysis of the start times derived for each frequency interval can yield the actual start time of the stimulus captured in the recording. In one or more embodiments, a statistical pattern is used.
[0024] In one or more embodiments, the following equations provide a clear implementation of the processing / algorithm. Cross-correlation of two frequency intervals SM and Sideal: Eq. (3) where SM[n] and Sideal[n] are the n-th elements of the measured and ideal stimulus spectrogram frequency bins. This technique can be easily extended to include joint analysis of multiple frequency bins from the spectrogram. The cross-correlation is performed over length N with shift m.
[0025] maxcorr = max(corr(SM,Sideal)) Eq. (4) Note: lag(Speak_index) = index(maxcorr) arrivalTime = lag(Speak_index) Eq. (5) Figure 6 An example time-domain dot plot 600 of a noisy recording (acquired from a synchronous deconvolution process / algorithm application) is shown. This recording is from a noisy environment with speech and impulsive noise (non-stationary). The environment replicates real-world scenarios in which an end user of a smartphone application can be located. The start and end regions represent pre-stimulus and post-stimulus buffer regions. The central region represents the recording of the stimulus that should be used for deconvolution with the original stimulus. If an amplitude threshold-based approach is used for start time detection, the false impulsive noise that is detected as the start time can be seen. In one or more embodiments, the disclosed technique is used to detect the actual stimulus with an accuracy of 64 samples (frequency resolution of the spectrogram).
[0026] Figure 7 An example spectrogram 700 of the noisy recording of Figure 6 is shown in accordance with some embodiments. In one or more embodiments, spectrograms such as spectrogram 700 are used to accurately compute the start and stop times of a noisy recording. A synchronous deconvolution process / algorithm application using the disclosed technique is used to compute a clean and accurate impulse response 900 ( Figure 9 ). If an alternative start time detection based on a threshold in the time or frequency domain is used, the start time is falsely detected, and thus a noisy, inaccurate impulse response 800 ( Figure 8 ) is obtained.
[0027] Figure 8 An example impulse response 800 plotted on a smartphone application using a false start time under non-stationary noise is shown in accordance with some embodiments. The impulse response 800 shown is plotted on a smartphone application using a false start time (computed using a threshold-based approach) under non-stationary noise.
[0028] Figure 9 An example impulse response 900 plotted on a smartphone application using the correctly identified start time under non-stationary noise is shown in accordance with some embodiments. The impulse response 900 is plotted on a smartphone application using the accurate start time under non-stationary noise, which is calculated using the spectral graph based method. The start time detected by the disclosed technology differs by about 0.8 seconds from the start time erroneously detected by the conventional method, and the effects caused by the erroneous start time are observed significantly in the impulse response 800 Figure 8
[0029] Figure 10 An example dot plot 1000 of the impulse response using a smartphone application with 0 dB signal-to-noise ratio (SNR) (stationary noise) is shown in accordance with some embodiments. The channels included in the example dot plot 1000 include L, R, C, LFE (or subwoofer channel), Ls, Rs, Lrs, Rrs, Lh, Rh, Lhs, and Rhs.
[0030] Figure 11 An example impulse response 1100 of the primary channel (L) using a smartphone application with 0 dB SNR (stationary noise) is shown in accordance with some embodiments. Figure 12 An example spectral graph 1200 of 0 dB SNR noisy recording (stationary) is shown in accordance with some embodiments. The start time and end time alignment using spectral graph was tested with vacuum cleaner noise (stationary) at 0 dB SNR. The impulse response 1100 generated for the speakers is clean and useful for delay and level correction except for the subwoofer (LFE) impulse response, which is corrupted due to the low frequency content of the noise. This can be seen in the spectral graph of the noisy recording in the spectral graph 1200, while the clean impulse can be seen in the impulse response 1100.
[0031] Figure 13 An example of the impulse response 1300 of twelve (12) channels measured for a soundbar using an external microphone using the disclosed technology is shown in accordance with some embodiments. The channels included in the example dot plot 1000 include L, R, C, LFE (or subwoofer channel), Ls, Rs, Lrs, Rrs, Lh, Rh, Lhs, and Rhs. Channel 4 shows the impulse response of the subwoofer labeled LFE. As expected, only the low frequency component present shows the impulse response. The relative delay of each channel is calculated by the prominence of the peaks. The function for the prominence is implemented using the "findpeaks" function, which returns a vector of local maxima (peaks) with the input signal vector data.
[0032] Eq. (6) In equation (6), P is the peak protrusion, H is the peak height or amplitude, and L and R are the left and right valleys. The first peak of the impulse response is found to be above a specific threshold of protrusion.
[0033] Eq. (7) In equation (7), P(x[i]) is the peak prominence at the i-th sample in the impulse response (IR), and Pt is the threshold for detecting the prominence of the effective peak. To calculate the relative channel delay, Eq. (8) is implemented as: Eq.(8) In Eq. (8), TOAi is the arrival time of the i-th channel, which is calculated based on the known sampling frequency and the peak_index in Eq. (7).
[0034] Eq. (9) In Eq. (9), This refers to the actual relative distance between the speakers in channels i and j, respectively. This actual relative distance is calculated using the difference between the measured distances between the microphone and the speakers in channels i and j. This measurement can be performed using a laser rangefinder or a similar device.
[0035] Figure 14 Illustrations are shown according to some embodiments Figure 13 A close-up view 1400 shows the impulse response dot plot 1300 of the height channels (Rh and Lh) as illustrated. In the example embodiment, the impulse responses of Rh and Lh are recorded from an example soundbar used for ceiling reflection analysis. The direct flight delay of the height channels can be calculated using Eq.(6) to Eq.(8) for the first impulse response, and the first reflection delay can be calculated using Eq.(6) to Eq.(8) for the second impulse response observed in close-up view 1400. Since the height channels on the soundbar point towards the ceiling, the first reflection impulse response is more prominent than the direct flight impulse response.
[0036] Eq. (10) Eq.(10) shows that Errori;j is the difference between the actual relative delay Delay(Relativei,j) and the delay Delay(Measuredi,j) measured using Equation (8).
[0037] Figure 15 Example heatmap 1500 of relative channel delay error according to some embodiments is shown. In one or more embodiments, heatmap 1500 is used to display an error matrix of relative channel delay, which shows that, using the disclosed techniques and their synchronous deconvolution processing / algorithm applications, there is a relative channel delay error of less than 1 ms for all channels. In some embodiments, the delay of the subwoofer / LFE channel is achieved based on maximizing the sum of the frequency responses of the subwoofer and the main channel. This maximization is estimated by iteratively delaying the impulse responses of the subwoofer and the main channel and then minimizing the standard deviation of the frequency response over the crossover region. Once the minimum standard deviation of the frequency response over the crossover region is found for an N-sample (tms) delay of the main channel or subwoofer, the required delay of the system and the timing alignment of all channels, which is best suited to the listener's position, can be determined.
[0038] Figure 16 An example level measurement heatmap 1600, according to some embodiments, is shown for the accuracy of levels measured using a smartphone and a sound level meter. In one or more embodiments, the sound power level (SPL) can be calculated using an impulse response derived from a synchronous deconvolution processing / algorithm application. Based on this, speaker level equalization can be performed for the location of the primary listener.
[0039] Eq. (11) In Eq.(11), H and X are the FFTs of the impulse response and powder noise (with reference SPL), respectively (derived for the i-th channel using synchronous deconvolution processing / algorithm application). The length of the FFT used is determined by the following equation: length(Impulse_Responsei) + length(Pink_Noise) - 1. Eq. (12) In Eq. (12), It is a temporal convolution of (measured) powder noise and (computed) impulse response using synchronous deconvolution processing / algorithm application.
[0040] Eq. (13) Figure 17 An example level measurement heatmap 1700 is shown, according to some embodiments, regarding the accuracy of levels measured using a smartphone application. Heatmap 1600 ( Figure 16 The graph 1700 shows a comparison between the level calculated using synchronous deconvolution processing / algorithm application and the level calculated using measurements performed using a sound level meter. The graph shows the level calculated by increasing the input of each channel by +6 dB (x-axis), and its reflection in the measurement data (y-axis) can be observed.
[0041] Real-time implementation of speaker measurement in smartphone applications (for level and latency) is a challenging process. One or more embodiments leverage methods that deploy computation on processors / servers in cloud-based environments to address the challenges of using high-resolution FFTs on devices. In some embodiments, cyclically shifted logarithmic sine scans are time-sensitive and suitable for use. Statistical features from high temporal resolution spectrograms are used to address another major challenge of playback and recording systems with asynchronous operation.
[0042] In some embodiments, the disclosed techniques use statistical features automatically derived from spectrograms to provide timing alignment between the audio measurement / recording system and a separate playback system. In one or more embodiments, the disclosed techniques also provide spectrogram-based, automatically derived start and end time synchronization for systems requiring alignment of stimulus and recorded signals. Audio playback and recording systems may require synchronization due to communication delays from manual, BLUETOOTH®, Wi-Fi, or cloud-based methods. Spectrogram-based start and end time synchronization can be effective for any system that may require stimulus and recorded signal alignment. In some embodiments, the disclosed techniques additionally provide speaker tuning independently of the stimulus (including at least one of pink noise, maximum length sequence (MLS), logarithmic sine sweep, multitone, or random white noise) and use automatically derived statistical features from the spectrogram to estimate start and stop times. One or more embodiments are independent of any stimulus used for speaker tuning (i.e., the disclosed techniques are applicable to pink noise, MLS, logarithmic sine sweep, and multitone and random white noise, etc.). In cases with unknown stimuli, some embodiments use statistical properties of the spectrogram to estimate start and stop times.
[0043] Figures 18-19 Example screen displays 1800 and 1900 are shown for a smartphone application according to some embodiments. A synchronous deconvolution processing / algorithm app implemented on the smartphone connects to a cloud computing environment (e.g., MathWorks® Cloud). Once the display is connected to the cloud processing environment, screen display 1800 shows the display. Screen display 1900 shows an exemplary settings screen.
[0044] Figures 20-22 Examples of screen displays illustrating a smartphone application connected to a cloud computing environment using the disclosed techniques are shown according to some embodiments. Screen display 2000 shows a sensor page with the smartphone's microphone activated. Screen display 2100 shows data and information derived from the use of the simultaneous deconvolution processing / algorithm application. Screen display 2200 shows the impulse response derived from the use of the simultaneous deconvolution processing / algorithm application.
[0045] Figure 23 A process 2300 for time alignment for independent recording and playback, according to some embodiments, is illustrated. In block 2310, process 2300 provides sending a stimulus signal to a speaker. In block 2320, process 2300 also provides receiving a measurement signal via a microphone. In block 2330, process 2300 also provides converting the stimulus signal into a stimulus time-frequency representation. In block 2340, process 2300 further provides converting the measurement signal into a measurement time-frequency representation. In block 2350, process 2300 also provides selecting at least one frequency value between the stimulus time-frequency representation and the measurement time-frequency representation. In block 2360, process 2300 also provides performing a correlation analysis using the selected at least one frequency value. In block 2370, process 2300 provides determining the start time of a statistical pattern to generate the stimulus signal based on the correlation analysis.
[0046] In some embodiments, process 2300 further includes the feature of playing a stimulus signal at one or more speakers.
[0047] In one or more embodiments, process 2300 further provides time alignment of one or more speakers based on the start time.
[0048] In some embodiments, process 2300 further includes the feature that the measurement signal is recorded on a computing device (e.g., a smartphone, tablet computer, etc.).
[0049] In one or more embodiments, process 2300 further includes the following features: the time alignment of one or more speakers is calibrated for an audio measurement or recording system with an independent playback system, and based on one or more statistical features automatically derived from a spectrogram.
[0050] In some embodiments, process 2300 further includes providing tuning of one or more loudspeakers independently of the stimulus signal (including at least one of powder noise, maximum length sequence, logarithmic sine sweep, multi-tone, or random white noise). Process 2300 also includes providing start and end time synchronizations for an audio measurement or recording system based on spectrogram auto-derived values, wherein the audio measurement or recording system requires alignment of the stimulus and measurement signals.
[0051] In one or more embodiments, process 2300 further includes the following feature: alignment of the stimulus signal and the measurement signal is required due to communication delays caused by manual, BLUETOOTH®, Wi-Fi, or cloud-based communication.
[0052] Embodiments have been described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products. Each block, or combination thereof, of such illustrations / diagrams can be implemented by computer program instructions. When provided to a processor, the computer program instructions create a machine, such that the instructions, executed via the processor, create means for implementing the functions / operations specified in the flowcharts and / or block diagrams. Each block in the flowcharts / block diagrams may represent a hardware and / or software module or logic. In alternative embodiments, the functions indicated in the blocks may occur out of order, simultaneously, etc., as shown in the figures.
[0053] The terms "computer program medium," "computer-usable medium," "computer-readable medium," and "computer program product" are generally used to refer to media such as main memory, secondary storage, removable storage drives, hard disks mounted in hard disk drives, and signals. These computer program products are means for providing software to a computer system. Computer-readable media allow a computer system to read data, instructions, messages or message packets, and other computer-readable information from the computer-readable medium. For example, computer-readable media may include non-volatile memory such as floppy disks, ROM, flash memory, disk drive memory, CD-ROM, and other permanent storage. For example, it is useful for transferring information such as data and computer instructions between computer systems. Computer program instructions may be stored in a computer-readable medium that can instruct a computer, other programmable data processing apparatus, or other means to function in a particular manner, causing the instructions stored in the computer-readable medium to produce an article of writing that includes instructions that implement the functions / actions specified in one or more boxes of a flowchart and / or block diagram.
[0054] As those skilled in the art will understand, aspects of the embodiments can be implemented as a system, method, or computer program product. Therefore, aspects of the embodiments can take the form of a completely hardware embodiment, a completely software embodiment (including firmware, resident software, microcode, etc.), or an embodiment combining software and hardware aspects, all of which are generally referred to herein as a "circuit," "module," or "system." Furthermore, aspects of the embodiments can take the form of a computer program product implemented in one or more computer-readable media having computer-readable program code implemented thereon.
[0055] Any combination of one or more computer-readable media may be used. A computer-readable medium may be a computer-readable storage medium. A computer-readable storage medium may be, for example, but not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or apparatuses, or any suitable combination of the foregoing. More specific examples (not an exhaustive list) of computer-readable storage media will include: electronic connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing. In the context of this document, a computer-readable storage medium may be any tangible medium that may include or store programs for use by or in conjunction with an instruction execution system, device, or apparatus.
[0056] Computer program code for performing aspects of one or more embodiments can be written in any combination of one or more programming languages, including object-oriented programming languages (such as Java, Smalltalk, C++, etc.) and conventional procedural programming languages (such as the "C" programming language or similar programming languages). The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter case, the remote computer can be connected to the user's computer via any type of network (including a local area network (LAN) or a wide area network (WAN)), or can be connected to an external computer (e.g., via the Internet through an Internet service provider).
[0057] The flowchart illustrations and / or block diagrams of the methods, apparatus (systems), and computer program products described above illustrate aspects of one or more embodiments. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a special-purpose computer or other programmable data processing apparatus to produce a machine, such that the instructions, executable via a processor of the computer or other programmable data processing apparatus, create means for implementing the functions / actions specified in one or more blocks of the flowchart illustrations and / or block diagrams.
[0058] These computer program instructions may also be stored in a computer-readable medium that can instruct a computer, other programmable data processing device or other means to function in a particular manner, such that the instructions stored in the computer-readable medium produce an article of writing, which includes instructions that implement the functions / actions specified in one or more boxes of a flowchart and / or block diagram.
[0059] Computer program instructions may also be loaded onto a computer, other programmable data processing apparatus or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer-implemented process, such that the instructions, which execute on the computer or other programmable device, provide a process for implementing the functions / actions specified in one or more boxes of a flowchart and / or block diagram.
[0060] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of instructions comprising one or more executable instructions for implementing a specified logical function. In some alternative embodiments, the functions indicated in the blocks may occur in a non-consecutive order. For example, two blocks shown consecutively may actually be executed substantially simultaneously, or these blocks may sometimes be executed in reverse order, depending on the functions involved. It will be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented by a dedicated hardware-based system that performs the specified function or action or executes a combination of dedicated hardware and computer instructions.
[0061] Unless expressly stated otherwise, references to elements in the singular form in the claims are not intended to mean "one and only one," but rather "one or more." All structural and functional equivalents of the elements of the exemplary embodiments described above, which are now known or hereafter known to those skilled in the art, are intended to be included in these claims. Unless an element is expressly described using the phrase "means for..." or "steps for...", the elements of the claims herein shall not be construed under the provisions of Section 112, paragraph 6, of Title 35, United States Code.
[0062] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the disclosed technology. As used herein, the singular forms "a," "an," and "described" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that, when used in this specification, the terms "comprising" and / or "including" specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0063] All means or steps plus functional elements in the appended claims are intended to include any structure, material, action, and equivalent for performing functions in conjunction with other claimed elements of the specific claims. Descriptions of embodiments have been presented for purposes of illustration and description, but are not intended to be exhaustive or limited to embodiments of the disclosed forms. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the disclosed technology.
[0064] Although embodiments have been described with reference to certain versions of the examples, other versions are also possible. Therefore, the spirit and scope of the appended claims should not be limited to the description of the preferred versions included herein.< / m>
Claims
1. A computer-implemented method for determining the onset time of a stimulus in a measurement, comprising: Send the stimulus signal to the speaker; Receive measurement signals via microphone; The stimulus signal is transformed into a stimulus time-frequency representation; The measurement signal is transformed into a time-frequency representation of the measurement; At least one frequency value is selected between the stimulation time-frequency representation and the measurement time-frequency representation; Perform correlation analysis using the selected at least one frequency value; and The correlation analysis is used to determine the statistical pattern that generates the stimulus signal at the start time.
2. The method according to claim 1, wherein, The stimulus signal is played at one or more speakers.
3. The method according to claim 2, further comprising: The timing alignment of the one or more speakers is calibrated based on the start time.
4. The method according to claim 1, wherein, The measurement signal is recorded at the computing device.
5. The method according to claim 3, wherein, The time alignment of the one or more loudspeakers is calibrated for an audio measurement or recording system with an independent playback system, and is based on one or more statistical features automatically derived from the spectrogram.
6. The method according to claim 5, further comprising: The stimulation signal, independent of at least one of powder noise, maximum length sequence, logarithmic sine sweep, multi-tone or random white noise, provides the tuning of the one or more loudspeakers; as well as Provides start and end time synchronization for the audio measurement or recording system based on automatically derived spectrograms, wherein the audio measurement or recording system requires alignment of the stimulus signal and the measurement signal.
7. The method according to claim 6, wherein, Alignment of the stimulus signal and the measurement signal is required due to communication delays caused by manual, Bluetooth, Wi-Fi, or cloud-based communication methods.
8. A non-transitory processor-readable medium including a program that, when executed by a processor, provides the start time of a stimulus in a measurement, the providing step comprising: The processor sends the stimulation signal to the speaker; The processor receives the measurement signal via a microphone; The processor converts the stimulus signal into a stimulus time-frequency representation; The processor transforms the measurement signal into a time-frequency representation of the measurement. The processor selects at least one frequency value between the stimulation time-frequency representation and the measured time-frequency representation; The processor performs a correlation analysis using the selected at least one frequency value; as well as The processor determines the start time for generating the stimulus signal based on the correlation analysis to establish a statistical pattern.
9. The non-transitory processor-readable medium according to claim 8, wherein, The stimulus signal is played at one or more speakers.
10. The non-transitory processor-readable medium according to claim 9, further comprising: The processor calibrates the timing alignment of the one or more speakers based on the start time.
11. The non-transitory processor-readable medium according to claim 8, wherein, The measurement signal is recorded at the computing device.
12. The non-transitory processor-readable medium according to claim 10, wherein, The time alignment of the one or more loudspeakers is calibrated for an audio measurement or recording system with an independent playback system, and is based on one or more statistical features automatically derived from the spectrogram.
13. The non-transitory processor-readable medium of claim 12, further comprising: The processor provides tuning for the one or more loudspeakers independently of the stimulus signal comprising at least one of powder noise, maximum length sequence, logarithmic sine sweep, multi-tone or random white noise; as well as The processor provides start and end time synchronization for the audio measurement or recording system, which is automatically derived from the spectrogram and requires alignment of the stimulus signal and the measurement signal.
14. The non-transitory processor-readable medium according to claim 13, wherein, Alignment of the stimulus signal and the measurement signal is required due to communication delays caused by manual, Bluetooth, Wi-Fi, or cloud-based communication methods.
15. An apparatus comprising: Memory, storing instructions; as well as At least one processor executes the instructions, the instructions including processing configured to perform the following operations: Send the stimulus signal to the speaker; Receive measurement signals via microphone; The stimulus signal is transformed into a stimulus time-frequency representation; The measurement signal is transformed into a time-frequency representation of the measurement; At least one frequency value is selected between the stimulation time-frequency representation and the measurement time-frequency representation; Perform correlation analysis using the selected at least one frequency value; and The correlation analysis is used to determine the statistical pattern that generates the stimulus signal at the start time.