Automatic detection of alignment between two audio signals

A prediction network system analyzes audio signals to automatically detect and adjust for synchronization issues in dubbed content, enhancing audio quality control by addressing inefficiencies in manual methods.

JP2025133686AInactive Publication Date: 2025-09-11DISNEY ENTERPRISES INC +1
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2024214122
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-07-18
Filing Date
2024-12-09
Publication Date
2025-09-11
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Audio quality control in dubbed media is largely manual and inefficient, failing to detect small, isolated errors in audio and video synchronization, which is a growing issue due to the increasing popularity of dubbed content.

Method used

A system using a prediction network with two branches to analyze reference and target audio signals, extracting features in a higher-dimensional space to detect and adjust for synchronization issues, including global, drifting, and intermittent offsets, trained with contrastive learning and data augmentation to enhance robustness.

Benefits of technology

Automatically detects and adjusts for various synchronization problems in dubbed audio, improving efficiency and accuracy beyond manual methods, ensuring precise audio-video alignment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025133686000001_ABST
    Figure 2025133686000001_ABST
Patent Text Reader

Abstract

To solve synchronization problems by detecting temporal offset between two audio signals.SOLUTION: A reference audio synchronization system 100 includes: a prediction network including a prediction network branch #1 that analyzes a first sample of a first audio signal and determines a first representation in space, and a prediction network branch #2 that analyzes a second sample of a second audio signal and determines multiple second representations in space; and an offset analyzer that compares the first representation and the multiple second representations and outputs an offset between the first sample and the second sample.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0001] Pursuant to 35 USC § 119(e), this application has the benefit of and claims the benefit of the filing date of U.S. Provisional Patent Application No. 63 / 560,419, filed March 1, 2024, entitled "Automatic Detection of Alignment Between Two Audio Signals," which is incorporated herein by reference in its entirety for all purposes. [Background technology]

[0002] Audio quality control is a largely manual process for companies. Audio errors can occur at various points in the content pipeline, extending from content mastering all the way to content distribution. Audio and video misalignment is one of the most distracting quality defects in media consumption today. Even slight offsets between audio and video can be noticeable to viewers. In the film industry, audio and video synchronization issues can be a major driver of viewer abandonment. Media synchronization issues arise for a myriad of reasons, but one area more prone to these errors is dubbing. In some instances, the visual components of dubbed media remain unchanged, but the audio track is replaced by a translated rendition of the original language. Creating and inserting a new audio track can result in desynchronization with the original audio track. As dubbed content consistently grows in popularity, the problem of unidentified synchronization errors is becoming increasingly prevalent.

[0003]

[0003] Quality control of dubbed media can be a largely manual process. Quality checks can involve specialized users listening for differences in the audio, or quality control technicians visually comparing the raw waveform for any errors. These subjective evaluations can be costly and inefficient, and cannot detect small, isolated errors. Summary of the Invention

[0004] The included drawings are for illustrative purposes only and are intended to serve to provide examples of possible structures and operations for the disclosed inventive systems, apparatus, methods, and computer program products. These drawings are in no way intended to limit any changes in form and detail that may be made by those skilled in the art without departing from the spirit and scope of the disclosed implementations. [Brief explanation of the drawings]

[0005] [Figure 1]

[0005] FIG. 1 illustrates a simplified system for analyzing an audio signal, according to some embodiments. [Figure 2]

[0006] FIG. 1 illustrates a simplified flowchart of a method for training a predictive network, according to some embodiments. [Figure 3]

[0007] FIG. 1 illustrates a simplified flowchart for predicting an offset, according to some embodiments. [Figure 4]

[0008] 10A-10C illustrate examples of offset analysis processes, according to some embodiments. [Figure 5]

[0009] FIG. 10 illustrates a graph of offsets that may be detected, according to some embodiments. [Figure 6]

[0010] FIG. 1 illustrates an example of a computing device, according to some embodiments. DETAILED DESCRIPTION OF THE INVENTION

[0006]

[0011] Described herein are techniques for audio signal analysis systems. In the following description, for purposes of explanation, numerous examples and specific details are set forth in order to provide a thorough understanding of some embodiments. Some embodiments as defined by the claims may include some or all of the features of these examples alone or in combination with other features described below, and may further include modifications and equivalents of the features and concepts described herein.

[0007]

[0012] System Overview

[0013] The system receives two audio signals for comparison. The audio signals can be different types of audio signals, such as, but not limited to, mono, stereo (2.0), surround (e.g., 5.1, 7.1), etc. The reference audio signal can be a first audio signal in a first language, such as the original language of the video. The target audio signal can be a second audio signal that can be related to the reference audio signal. For example, the target audio signal can be a dubbed audio signal in any language, such as a foreign language translated from the original language. The target audio signal can also be an encoded version of the reference audio signal, a narration track (a description mixed with the reference audio signal), or any derivative audio signal from the reference audio signal. The system compares the two audio signals to determine the offset between the signals. There may be multiple dubbed audio signals in multiple foreign languages ​​for a video. The system may perform the following analysis between the reference audio signal and each dubbed audio signal: Each reference audio signal and dubbed audio signal pair may have different offset issues that can be detected.

[0008]

[0014] In some embodiments, the prediction network includes two branches that extract features from two audio signals. For example, the first branch extracts features from a reference audio signal, and the second branch extracts features from a dubbed audio signal. Each branch receives samples of the respective audio signal corresponding to a time period. The time periods can be the same, such as raw audio samples from timestamps a to b, where b is t seconds greater than a, or from different time periods. Each branch outputs a representation of the respective sample in a higher-dimensional space. The system compares the outputs for the two audio signals to determine the temporal offset between the two audio signals. Different methods can be used to determine the offset, which are described below. The offset problem can be determined based on analyzing multiple samples across the two audio signals. For example, the offset is first calculated for timestamps 0s to 10s, then 10s to 20s, etc. In this description, certain time periods are used, such as 10 seconds, 200 milliseconds, and 20 milliseconds, but these time periods are merely examples and other time periods may be used. This allows for a more detailed look at the types of offsets that may exist in the reference audio signal and the target audio signal. The system can determine the offset between the audio signals and use that offset to adjust the dubbed audio signal to address synchronization issues.

[0009]

[0015] The system can be trained to recognize multiple problems that may occur in the dubbing process. For example, the system may be able to recognize global offsets, which may be further grouped into constant offsets and drifting offsets, and intermittent offsets, which are isolated / one-off offsets between portions of the audio signal (e.g., silence, uncorrelated audio from entirely different audio files, or temporal shifts). Global offsets may be offsets that occur throughout the entire video. Constant offsets are constant throughout the video. Drifting offsets may drift throughout the video in a certain direction. Intermittent offsets may be isolated offsets that may occur in isolated portions of the video. Audio problems may occur in sequences containing only dialogue compared to portions of audio containing only background cues.

[0010]

[0016] Some use cases in which quality control analytics can be used include:

[0011]

[0017] Dubbing: The two audio tracks may differ substantially during the speech segments, and may differ further in background noise (e.g., background laughter in a sitcom) and music.

[0012]

[0018] Encoding: For linear television (TV), common audio operations between input and output audio tracks include noise reduction, loudness reduction / normalization / amplification, filtering, and compression.

[0013]

[0019] Dialogue Mismatch: Lines of dialogue may be added to or deleted from the "dubbing standard" provided to a third party compared to the final original audio. This poses challenges for foreign language dubbing / subtitling teams, as the dub they produce may not reflect the dialogue present in the final version.

[0014]

[0020] Narration Track: When a description is added over the main audio track. The target audio track may become out of sync due to the added description.

[0015]

[0021] system

[0022] FIG. 1 illustrates a simplified system 100 for analyzing audio signals, according to some embodiments. A reference audio synchronization system 102 (synchronization system 102) receives a reference audio signal and a target audio signal. The audio signals may be associated with a timeline in a video, such as a two-hour video having a two-hour audio signal. As discussed above, the reference audio signal may be a first signal in a first language (e.g., the original language of the audio track). The target audio signal may be in a language different from the original language, such as a foreign language translated from the original language. In some examples, the original language may be English, and the dubbed language may be Spanish, French, etc. Although an original language is mentioned, the reference audio signal may be in a language other than the original language. Also, the two audio signals need not be in different languages, but may be the same language. For example, the two audio signals may have different background music or sounds.

[0016]

[0023] The synchronization system 102 may include a prediction network 103 that analyzes a reference audio signal and a target audio signal to generate a first representation #1 of the reference audio signal and a second representation #2 of the target audio signal. The representations may extract important features of the respective audio signals in a higher dimensional space. The representations may be embeddings of the respective audio signals.

[0017]

[0024] Different prediction networks may be used. The prediction network 103 includes a prediction network branch #1 104-1 and a prediction network branch #2 104-2. While this structure of the prediction network 103 is described, other structures capable of determining a representation of an audio signal may be used. In some embodiments, each prediction network branch may be based on a transformer architecture capable of extracting features from an audio signal and generating a representation of the audio signal. In some embodiments, the prediction network branch #1 104-1 analyzes a reference audio signal and outputs representation #1, and the prediction network branch #2 104-2 analyzes a target audio signal and outputs representation #2. In some embodiments, the logic of the prediction network branch #1 104-1 and the logic of the prediction network branch #2 104-2 may form a Siamese network, a transformer-based twin network, or other network in which the logic of both branches is similar. The values ​​of the parameters of prediction branch #1 104-1 and prediction network branch #2 104-2 may be adjusted based on a training process, as described below.

[0018]

[0025] In some embodiments, prediction network branch #1 104-1 and prediction network branch #2 104-2 analyze sequences of the reference audio signal and the target audio signal, respectively. Both signals may be segmented into shorter sequences (e.g., 10 seconds) within the video to collect multiple predictions along the signal. The sequences may include multiple samples. Each sample may be a window of the time range of the respective audio signal. The input to prediction network branch #1 104-1 and prediction network branch #2 104-2 may be a window, such as 200 milliseconds (ms). Representation #1 and Representation #2 may include one or more representations for a single sequence, such as a 200 ms sliding window with a 20 ms sliding increment generated for a 10 second sequence and input to prediction network branch #1 104-1 and prediction network branch #2 104-2. The use of short samples allows the system to take into account globally occurring problems such as constant offsets as well as intermittent offsets such as drift or isolated offsets that may be found in dubbing. Predictions for sequences can be determined based on an analysis of the offsets within the samples.

[0019]

[0026] The offset analyzer 106 may analyze representation #1 and representation #2 to predict an offset between the two representations. The offset may be a time, such as t milliseconds, between samples of representation #1 and representation #2. As described in more detail below, the offset analyzer 106 may predict the offset using different methods. For example, the offset analyzer 106 may select the first representation #1 for the first sample. Then, the offset analyzer 106 compares the distance between the first representation #1 and the second representation #2 for the second sample around the time of the first sample. The offset analyzer 106 selects the second representation #2 based on the distance (e.g., the representation closest in distance). Because these two samples are considered similar, the offset analyzer 106 determines an offset between the first sample and the second sample on the video timeline, such as 20 ms. Each pair of representations corresponds to a time period on the media timeline. The offset analyzer 106 may then output an offset for each sample of the reference audio signal.

[0020]

[0027] The audio synchronization analyzer 108 may analyze the offsets relative to the samples to determine whether the video and dubbed audio have any synchronization problems. As mentioned above, the audio synchronization analyzer 108 may be able to detect multiple synchronization problems that may occur in the target audio signal. Different types of synchronization problems are described in more detail below.

[0021]

[0028] training

[0029] The prediction network 103 may be trained using different methods. FIG. 2 shows a simplified flowchart 200 of a method for training the prediction network 103 according to some embodiments. At 202, the synchronization system 102 receives a training dataset including pairs of a first set of samples from a reference signal and a second set of samples from a target audio signal. Labels for the pairs are also provided. For example, the labels may indicate whether the samples are synchronized or unsynchronized, or the labels may list the offset between the samples. In some embodiments, a contrastive learning process may be used to distinguish between synchronized and unsynchronized samples. Synchronized samples may be labeled where the offset between the reference sample and the dubbed sample is below a threshold, such as an offset that is not noticeable to humans. For example, a range where some people may not be able to detect whether the signals are misaligned may be [-125, 45] ms, although other ranges may also be used. Unsynchronized samples may be labeled where the offset between the reference signal and the dubbed sample exceeds a threshold, such as an offset that is noticeable to humans. Synchronized and unsynchronized samples may be generated in different ways. In some embodiments, unsynchronized samples are generated by assuming all audio pairs are initially synchronized, and unsynchronized samples are generated by advancing or delaying the target audio signal relative to the reference audio signal. Portions of data may also be verified to be synchronized. The training dataset may contain multiple synchronization issues (e.g., uncorrelated samples, silence, etc.). Enhancements may be added as potential audio differences are found in the dubbed audio signal. In some embodiments, enhancements may include mixing with background laughter, adding Gaussian noise, random loudness gain, high-pass filtering, low-pass filtering, polarity inversion, and pitch shifting, although other enhancements may also be used.Training the prediction network with the above data augmentation improves the robustness of the prediction between the reference audio signal and the variations in the dubbed audio signal that may occur in production use cases.

[0022]

[0030] At 204, the synchronization system 102 inputs corresponding samples from the first set of samples and the second set of samples into branch #1 104-1 and branch #2 104-2, respectively. For example, for a first pair, the first sample and the second sample are input into branch #1 and branch #2, respectively. Subsequent pairs are input into branch #1 and branch #2, respectively.

[0023]

[0031] At 206, a first output is generated from branch #1 and a second output is generated from branch #2. The first output may be representation #1 for the first sample and the second output may be representation #2 for the second sample. The representations may be in a higher dimensional space and include important features of each sample.

[0024]

[0032] At 208, the offset analyzer 106 may compare the difference between the first output and the second output to a label. For example, the label may be in sync or out of sync. The label may also be the offset between the first sample and the second sample. For example, if the label is in sync, the synchronization system 102 may train the prediction network 103 to output a representation that is closer in distance in a higher-dimensional space. If the label is out of sync, the synchronization system 102 may train the prediction network 103 to output a representation that is farther away in distance in a higher-dimensional space. Pairs from multiple different offset problems may be used to train the prediction network 103. This trains the prediction network 103 to be more robust in detecting offsets for different problems.

[0025]

[0033] At 210, synchronization system 102 may optimize parameters of branch #1 or branch #2 based on a function, such as a loss function. In some embodiments, the parameters are adjusted using symmetric loss by minimizing the distance between synchronized samples and maximizing the distance between unsynchronized samples. The above process may be performed for all sample pairs.

[0026]

[0034] After training, the prediction network 103 can be used to generate representations that are used to determine offset predictions.

[0027]

[0035] Offset Prediction

[0036] FIG. 3 shows a simplified flowchart 300 for generating offsets according to some embodiments. At 302, the prediction network 103 determines a first set of samples from a reference signal sequence (e.g., a 10-second sequence) and a second set of samples from a target audio signal sequence (e.g., a 10-second sequence). The first and second samples may be a set of sequential windows of t milliseconds in duration, such as each sample being 200 milliseconds, although other time windows may also be used. Each window may be identified by a position indicator within the sequence, such as a video timestamp. For example, samples may be labeled by time ranges, such as 0-200 ms, 10-210 ms, 20-220 ms, etc. The 10-second sequence may be analyzed across the reference signal and the target audio signal.

[0028]

[0037] At 304, the synchronization system 102 inputs respective samples from the first set of samples and the second set of samples into branch #1 and branch #2. Branch #1 104-1 may generate a representation for the first set of samples independently of branch #2 104-2, which generates a representation for the second set of samples. This is because the output may be analyzed after the representations are generated. At 306, for each sample, a first output is generated from branch #1 104-1 and a second output is generated from branch #2 104-2. The output of branch #1 104-1 is a sequence of representations for the first set of samples, and the output of branch #2 104-2 is a sequence of representations for the second set of samples. The representations may be identified, such as by a position identifier, and are then analyzed by the offset analyzer 106.

[0029]

[0038] The offset analyzer 106 may generate the offset using different processes. In some embodiments, at 308, the offset analyzer 106 may compare a first representation for a first sample with multiple second representations from a second sample within a time range, such as plus or minus t seconds. The offset of the first sample may be based on analyzing the difference between the first representation for the first sample and multiple second representations for the second sample within the range. In some embodiments, the offset analyzer 106 determines a second representation within the range that is closest in distance to the first representation. Then, at 310, the offset analyzer 106 determines a second sample for the second representation that is determined to be most similar to the first sample for the first representation based on the shortest distance between the two samples. At 312, using a location identifier for the sample, the offset analyzer determines an offset between the two samples. For example, the second sample may be −10 ms from the first sample based on comparing times associated with location indicators. An example of calculating the offset is described in FIG. 4. An offset may be determined for every first sample. At 314, the offset analyzer 106 outputs the offsets for the first samples.

[0030]

[0039] 4 illustrates an example of an offset analysis process according to some embodiments. The following may be performed for each sequence (e.g., 10 seconds) for the reference signal and the target audio signal. Shown are first representations #1 406-1 through #5 406-5 for the first sample 402 and second representations #2 408-2 through #4 408-4 for the second sample 404. The first representations 406-1 through 406-5 relate to five sequential first samples included in the reference audio signal, and the second representations 408-2 through 408-4 relate to three sequential second samples included in the target audio signal within a time range. Other representations may exist for the audio but are not shown, and second representations for samples before the second representation #2 408-2 and after the second representation #4 408-4 are not shown. Each representation can be a set of sequential windows of 200 ms duration (e.g., 0 ms to 200 ms, 10 ms to 210 ms, 20 ms to 220 ms, ..., 9800 ms to 10000 ms) with a hop size of 10 ms. For each of these windows, the prediction network 103 generates an embedding (e.g., an embedding representing the time 0 ms to 200 ms, an embedding representing 10 ms to 210 ms, etc.).

[0031]

[0040] In offset generation, offset analyzer 106 may select instances of first representation 406 and then calculate an offset for each instance. In this example, first representation #3 406-3 may be selected by offset analyzer 106, but similar calculations are performed for first representations #1, #2, #4, #5, etc.

[0032]

[0041] In the process, the offset analyzer 106 calculates a range for selecting the second representation 408. The range can be plus or minus t seconds. For example, the offset analyzer 106 selects the second representation from samples around the time associated with the first sample #3, which is associated with the first representation #3 406-3. In this example, the embedding 406-3 (e.g., a representation of the reference audio window 20 ms to 220 ms) has the same timestamp as the embedding 408-3 (e.g., a representation of the target audio window 20 ms to 220 ms). In some embodiments, the offset analyzer 106 selects the second representation #2 408-2, the second representation #3 408-3, and the second representation #4 408-4, although other representations may also be selected. The offset analyzer 106 calculates the distance between the first representation #3 406-3 and the second representation #2 408-2, second representation #3 408-3, and second representation #4 408-4, respectively. The distance may be the L2 distance between the values ​​of the representations in a higher dimensional space. The offset analyzer 106 determines which of the second representations 408-2, 408-3, 408-4 is closest to the first representation 406-3.

[0033]

[0042] The offset analyzer 106 determines that the first representation 406-3 is most similar to the second representation 408-2 because it has the smallest calculated distance (e.g., 0.74). Because the first representation 406-3 and the second representation 408-2 are representations of different time intervals, the offset analyzer 106 determines that an offset exists (e.g., 10 ms in this case because they are adjacent samples). If the smallest distance is between corresponding representations (e.g., the first representation 406-3 and the second representation 408-3), the two samples are considered synchronized.

[0034]

[0043] The above process is repeated for every representation in the first sample 402. That is, each representation of the first sample 402 has a corresponding second sample (offset). The first sample 402 is the representation generated for just the first sequence (e.g., 0 s to 10 s). An offset is determined for every representation / sample of this sequence (e.g., a 200 ms window), resulting in multiple offset predictions for the sequence (e.g., 0 s to 10 s). In some embodiments, the offset analyzer 106 uses these predictions to determine a single offset for the segment 0 to 10 seconds by taking the median of all offsets predicted for the first sample 402. This process is repeated for every sequence of the reference signal (e.g., 0 s to 10 s, 10 s to 20 s, 20 s to 30 s, ..., N s to N+10 s). Other methods of determining the offset for a sequence, such as using an average of the offsets, may also be used.

[0035]

[0044] In one example, the offset analyzer 106 calculates the distance from each of the second representations #2, #3, and #4 to the first representation #3 406-3. This results in distances of 0.74, 1.4, and 6.8 for the second representations #2, #3, and #4, respectively. The offset analyzer 106 selects which second representation is most similar to the first representation #3 406-3 (e.g., the smallest distance). If 0.74 is the smallest distance between the second representation #2 408-2 and the first representation #3 406-3, this means that the two representations are most similar within this range. The offset between the second representation #2 408-2 and the first representation #3 406-3 is then determined. For example, the second sample #2 associated with the second representation #2 408-2 may be -10 ms in the timeline from the first sample #3 associated with the first representation #3 406-3. An offset of -10 ms is then determined. This method for determining the offset may be used, although other methods may be used.

[0036]

[0045] Therefore, distances in a higher dimensional space are used to determine pairs of samples in the timeline. An offset is determined from the samples. The use of distances in a higher dimensional space allows the prediction network 103 to be trained to generate representations based on multiple different problems. This allows offsets from different problems to be detected.

[0037]

[0046] Synchronization question type

[0047] The synchronization system 102 can detect different types of problems with the target audio signal. For example, the audio synchronization analyzer 108 can detect a global offset, an intermittent offset, and so on.

[0038]

[0048] The constant offset may be a constant offset that exists throughout the entire video. In some embodiments, the audio synchronization analyzer 108 produces an offset prediction for consecutive samples of the reference signal and the target audio signal, such as within a 10-second sequence. The offset prediction may be analyzed across all sequences to produce an overall prediction. In some embodiments, predictions with a confidence value greater than a threshold may be considered for the overall offset prediction. The audio synchronization analyzer 108 uses these sequences to calculate the overall offset prediction. The audio synchronization analyzer 108 may use the offset predictions for the sequences to accurately predict the overall offset.

[0039]

[0049] The audio synchronization analyzer 108 may also detect intermittent offsets in the video. Intermittent offsets may not be constant throughout the video, but may occur intermittently within the video. The audio synchronization analyzer 108 evaluates samples to determine when intermittent offsets may occur, such as by using clustering to determine clusters of offsets. FIG. 5 shows a graph 500 of offsets that may be detected by some embodiments. The Y-axis is the offset value, and the X-axis is time on the video timeline. At 502, an offset of approximately +45 ms is determined, at 504, an offset of approximately +1600 ms is determined, at 506, an offset of approximately −125 ms is determined, and then at 508, an offset of approximately +45 ms is determined. The audio synchronization analyzer 108 may analyze one or more 10 ms samples or sequences to determine when intermittent offsets occur. Now that the offset analyzer 106 has determined the offset for each sample or sequence in the reference audio signal, the audio synchronization analyzer 108 can determine intermittent offsets in the video. The offset of the sequence (e.g., 10 seconds) can be used to determine when intermittent offsets occur.

[0040]

[0050] Sequences containing only dialogue will generally be the most different in the target audio signal. This contrasts with portions of audio containing only background cues, where many similarities may exist, since background cues may not be subject to dubbing. The audio synchronization analyzer 108 can optimally detect offsets for background cues and audio. The predictive network 103 can be trained to output representations for background effects and audio that allow for offset detection in both situations in the video.

[0041]

[0051] The synchronization system 102 is also robust to potential audio differences in the target audio signal. The video may have different enhancements, such as mixing with background laughter, adding Gaussian noise, random loudness gain, high-pass filtering, low-pass filtering, polarity inversion, and pitch shifting. The prediction network 103 can be trained with these enhancements to generate representations of the audio that may contain slight differences. The offsets due to the representations can be found by the audio synchronization analyzer 108 to detect problems.

[0042]

[0052] The audio synchronization analyzer 108 can also detect actual production errors, such as when dialogue is not present in the target audio signal. The predictive network 103 can be trained to generate representations that detect these errors.

[0043]

[0053] Therefore, the prediction network 103 can be trained to generate representations to detect different causes of offset. This makes the prediction network 103 robust when different offsets are caused by different problems in the same video. For example, the timing of laughter in a laughter track may not be the same in the target audio signal. The audio synchronization analyzer 108 can detect this offset in addition to other intermittent offsets due to other causes.

[0044]

[0054] conclusion

[0055] Therefore, the synchronization system 102 provides a robust offset prediction that does not rely on similarity assumptions. The synchronization system 102 can detect different types of offsets caused by different types of causes. This results in an improved process for detecting offsets.

[0045]

[0056] system

[0057] FIG. 6 illustrates an example of a computing device according to some embodiments. According to various embodiments, a system 600 suitable for implementing embodiments described herein includes a processor 601, a memory 603, a storage device 605, an interface 611, and a bus 615 (e.g., a PCI bus or other interconnect fabric). The system 600 may operate as a variety of devices, such as any of the devices or services described herein. While a particular configuration is described, a variety of alternative configurations are possible. The processor 601 may perform operations such as those described herein. Instructions for performing such operations may be embodied in the memory 603, on one or more non-transitory computer-readable media, or on some other storage device. Various specially configured devices may also be used in place of, or in addition to, the processor 601. The memory 603 may be a random access memory (RAM) or other dynamic storage device. The storage device 605 may include a non-transitory computer-readable storage medium that retains information, instructions, or some combination thereof, for example, instructions that, when executed by the processor 601, configure or enable the processor 601 to perform one or more operations of the methods described herein. The bus 615 or other communication components may support communication of information within the system 600. The interface 611 may be connected to the bus 615 and configured to send and receive data packets over a network. Examples of supported interfaces include, but are not limited to, Ethernet, Fast Ethernet, Gigabit Ethernet, Frame Relay, Cable, Digital Subscriber Line (DSL), Token Ring, Asynchronous Transfer Mode (ATM), High Speed ​​Serial Interface (HSSI), and Fiber Distributed Data Interface (FDDI). These interfaces may include ports suitable for communication with an appropriate medium. They may also include an independent processor and / or volatile RAM.A computer system or computing device may include or be in communication with a monitor, printer, or other suitable display for producing any of the results described herein to a user.

[0046]

[0058] Any of the disclosed implementations may be embodied in various types of hardware, software, firmware, computer-readable media, and combinations thereof. For example, some techniques disclosed herein may be implemented, at least in part, by a non-transitory computer-readable medium containing program instructions, state information, etc. to configure a computing system to perform various services and operations described herein. Examples of program instructions include both machine code, such as produced by a compiler, and higher-level code that may be executed via an interpreter. The instructions may be embodied in any suitable language, such as, for example, Java, Python, C++, C, HTML, any other markup language, JavaScript, ActiveX, VBScript, or Perl. Examples of non-transitory computer-readable media include, but are not limited to, magnetic media such as hard disks and magnetic tape, flash memory, optical media such as compact discs (CDs) or digital versatile discs (DVDs), magneto-optical media, and other hardware devices such as read-only memory ("ROM") devices and random access memory ("RAM") devices. A non-transitory computer readable medium may be any combination of such storage devices.

[0047]

[0059] In the above specification, various techniques and mechanisms may be described in the singular for clarity. However, it should be noted that some embodiments include multiple iterations of a technique or multiple instantiations of a mechanism unless otherwise stated. For example, a system may use a processor in various contexts, but may use multiple processors while remaining within the scope of this disclosure unless otherwise stated. Similarly, various techniques and mechanisms may be described as including a connection between two entities. However, a connection does not necessarily imply a direct, unobstructed connection, as various other entities (e.g., bridges, controllers, gateways, etc.) may exist between the two entities.

[0048]

[0060] Some embodiments may be implemented in a non-transitory computer-readable storage medium for use by or in connection with an instruction execution system, apparatus, system, or machine. The computer-readable storage medium includes instructions for controlling a computer system to perform methods described by some embodiments. The computer system may include one or more computing devices. The instructions, when executed by one or more computer processors, may be configured or operable to perform those described in some embodiments.

[0049]

[0061] As used herein or throughout the claims that follow, the words "a," "an," and "the" include plural referents unless the context clearly indicates a different interpretation. Also, as used herein or throughout the claims that follow, the meaning of "in" includes "in" and "on," unless the context clearly indicates a different interpretation.

[0050]

[0062] The above description illustrates various embodiments, along with examples of how aspects of some embodiments may be implemented. The above examples and embodiments should not be considered merely embodiments, but are presented to illustrate the versatility and advantages of some embodiments, as defined by the following claims. Based on the above disclosure and the following claims, other configurations, embodiments, implementations, and equivalents may be used without departing from the scope of this specification, as defined by the claims.

Claims

1. analyzing a first sample of a first audio signal to determine a first representation in space; analyzing a plurality of second samples for a second audio signal to determine a plurality of second representations within the space; comparing distances in the space between the first representation and the plurality of second representations in the space to select a second representation; determining an offset between the first sample in the first audio signal and a second sample in the second audio signal associated with the second representation; outputting the offset; A method comprising:

2. Analyzing the first sample and analyzing the plurality of second samples includes: analyzing the first sample using a first branch of a model; analyzing the plurality of second samples using a second branch of the model; and The method of claim 1 , comprising:

3. The method of claim 2 , wherein the first branch and the second branch include the same logic for generating a first representation and a second representation, respectively, in the space.

4. the first branch comprises first parameters trained to generate a first representation; the second branch comprises second parameters trained to generate a second representation; The method of claim 3.

5. selecting a time period based on the first sample; selecting the plurality of second samples based on the time period; and The method of claim 1 further comprising:

6. Comparing the first representation with the plurality of second representations includes: comparing distances in the space between the first representation and a second representation in the plurality of second representations; selecting a second representation from the plurality of second representations based on each of the distances; and The method of claim 1 , comprising:

7. Selecting the second sample includes: selecting the second representation from the plurality of second representations based on the second representation having a smallest distance to the first representation in the space; The method of claim 6 , comprising:

8. the first sample is from a first sequence of first samples in the first audio signal; the second sample is from a second sequence of second samples in the second audio signal. The method of claim 1.

9. determining an offset relative to the first sample to each second sample; determining an offset for the first sequence based on the offset for the first sample; The method of claim 8 further comprising:

10. determining a training data set comprising pairs of first and second training audio samples; analyzing the pairs with the model to output a first training representation and a second training representation; adjusting parameters of the model based on the labels associated with the pairs and the distances between the respective first and second training representations; The method of claim 1 further comprising:

11. Adjusting the parameters includes: adjusting parameters to cause the model to output a first or second training representation that is closer in distance in the space when the pair has a label indicating that the first and second training samples are synchronized; adjusting the parameters to cause the model to output the first training representation or the second training representation that is further apart in space when the pair has a label indicating that the first training sample and the second training sample are not synchronized; The method of claim 10, comprising:

12. Determining the offset comprises: determining a first position identifier for the first sample within the first audio signal; determining a second position identifier for the second sample within the second audio signal; determining the offset based on the first location identifier and the second location identifier; The method of claim 1 , comprising:

13. Determining the offset comprises: determining as the offset an offset relative to a first time for the first sample in the first audio signal and a second time for the second sample in the second audio signal; The method of claim 1 , comprising:

14. the comparing of the first representation with the plurality of second representations in the space is based on distances in the space between the first representation and the plurality of second representations; the offset is determined based on a time difference between the first sample and the second sample in the first audio signal and the second audio signal. The method of claim 1.

15. analyzing the offset to adjust synchronization between the first audio signal and the second audio signal; The method of claim 1 further comprising:

16. analyzing the offset to identify synchronization problems between the first audio signal and the second audio signal. The method of claim 1 further comprising:

17. the first audio signal is a reference audio signal in the original language of the video; the second audio signal is a dubbed audio signal in a translation of the original language for the video; The method of claim 1.

18. A non-transitory computer-readable storage medium having stored thereon computer-executable instructions that, when executed by a computing device, cause the computing device to: analyzing a first sample of a first audio signal to determine a first representation in space; analyzing a plurality of second samples for a second audio signal to determine a plurality of second representations within the space; comparing the first representation to the plurality of second representations in the space to select a second representation; determining an offset between the first sample and a second sample associated with the second representation; outputting the offset; 1. A non-transitory computer-readable storage medium operable to perform the steps of:

19. Analyzing the first sample and analyzing the plurality of second samples includes: analyzing the first sample using a first branch of a model; analyzing the plurality of second samples using a second branch of the model; and 20. The non-transitory computer-readable storage medium of claim 17, comprising:

20. one or more computer processors; A computer-readable storage medium wherein the computer-readable storage medium comprises: analyzing a first sample of a first audio signal to determine a first representation in space; analyzing a plurality of second samples for a second audio signal to determine a plurality of second representations within the space; comparing the first representation to the plurality of second representations in the space to select a second representation; determining an offset between the first sample and a second sample associated with the second representation; outputting the offset; 20. An apparatus comprising: instructions for controlling said one or more computer processors to operate to:

Citation Information

Patent Citations

  • Apparatus and method for evaluating voice signal quality

    JP2003167596A

  • Synchronization of spatial audio parametric coding with externally supplied downmix

    JP2008522243A

  • Information processor, information processing method, program, recording medium, and information processing system

    JP2013135310A

  • Echo cancellation method and apparatus based on delay time estimation

    JP2021500778A

  • Hierarchy encoding apparatus and hierarchy encoding method

    WO2005106850A1