Acoustic synthesis system
The acoustic synthesis system uses video signals to resample and align acoustic signals, addressing clock deviations in remote collaborations, ensuring high-quality sound synthesis.
Patent Information
- Application Number
- PCT/JP2024/000985
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-01-16
- Publication Date
- 2025-07-24
AI Technical Summary
In remote collaborations such as remote ensemble performances, the deviation of sampling clocks for separately sampled sounds causes deterioration of sound quality due to the inaccuracy of crystal oscillators, and existing synchronization methods fail to synchronize acoustic signals embedded in video signals without clock information.
An acoustic synthesis system that utilizes video signals to reproduce a clock, resamples acoustic signals from both the video and local clocks, and synthesizes them without clock deviation by selecting signals close to the local clock.
Enables high-quality acoustic signal synthesis by aligning sampling clocks across multiple signals, even when clock information is absent, thereby improving sound quality in remote collaborations.
Smart Images

Figure JP2024000985_24072025_PF_FP_ABST
Abstract
Description
Sound Synthesis System
[0001] TECHNICAL FIELD This disclosure relates to the synthesis of an audio signal to be transmitted along with a video signal.
[0002] In recent years, with the trend toward remote collaboration in society, there has been a demand for remote collaboration using sounds other than human speech, such as remote ensembles. In remote ensembles, separate sounds sampled at different locations must be appropriately synthesized at each location and fed back to the performers. In this sound synthesis, deviations in the sampling clocks of the separately sampled sounds cause degradation in the sound quality after synthesis.
[0003] Crystal oscillators are commonly used as clock sources for electronic devices, with an accuracy of approximately ±20 to ±100 ppm. For example, even if sampling is set to 48 kHz, if a crystal oscillator with an accuracy of approximately ±100 ppm is used, the sampling frequency will be accurate to a minimum of 47.9952 kHz and a maximum of 48.0048 kHz. In other words, a deviation of up to approximately 9.6 samples per second will occur. In this case, approximately 9.6 samples of data will be left over or missing during audio synthesis, causing degradation of sound quality.
[0004] One known method for addressing this sampling clock discrepancy is to synchronize the sampling clock itself. Audio equipment that requires high sound quality has a word clock terminal and a clock synchronization mechanism that enables it to distribute its own clock or sample in accordance with a received clock. Furthermore, Dante (registered trademark), a network audio system for remotely transmitting sound, uses a time synchronization protocol called PTP (Precision Time Protocol) to synchronize sampling and playback clocks at remote locations.
[0005] Furthermore, in remote collaborations such as remote ensembles, video signals are often used in addition to audio signals. In such cases, a signal format in which the audio signal is embedded within the video signal is sometimes used. In this signal format, the audio signal is synchronized with the video signal, and it is not possible to incorporate a mechanism for synchronizing the sampling clock of the audio signal. Furthermore, if the packetized signal containing the video signal containing the audio signal does not contain clock information, it is difficult for the device receiving the signal to identify clock discrepancies with other signals.
[0006] Therefore, for high-quality remote collaboration, such as a remote ensemble performance, it is necessary to synthesize and generate an audio signal with no sampling clock deviation from packetized video and audio signals that are sampled with multiple different clocks and do not contain clock information.
[0007] Sheu, Jia-Shing, Ho-Nien Shou, and Wei-Jun Lin. “Realization of an Ethernet-based synchronous audio playback system.” Multimedia Tools and Applications 75.16 (2016): 9797-9818.
[0008] An object of the present disclosure is to enable synthesis of an audio signal from a plurality of packetized video and audio signals that do not contain clock information, without causing deviation in the sampling clock.
[0009] As mentioned above, in remote collaborations such as remote ensembles, not only audio signals but also video signals captured simultaneously with the audio signals are often used. Therefore, in the present disclosure, a video signal transmitted together with the audio signals is used instead of the audio signals themselves.
[0010] The audio synthesis system and audio synthesis method disclosed herein receive an audio signal together with a video signal, use the video signal to regenerate a clock, resample the audio signal from the clock and a local clock, and use the resampled audio signal to synthesize with another audio signal.
[0011] The acoustic signal may be upsampled from the clock and a local clock, an acoustic signal that is closest to the local clock may be selected from the upsampled acoustic signals, and the selected acoustic signal may be synthesized with another acoustic signal.
[0012] The clock may be recovered using at least one of: (i) a synchronization signal for recovering the video signal; or (ii) the arrival time of the video signal.
[0013] The above disclosures can be combined as much as possible.
[0014] According to the present disclosure, it is possible to synthesize an audio signal from a plurality of packetized video and audio signals that do not include clock information, without causing deviation in the sampling clock.
[0015] 1 illustrates an example embodiment of an audio synthesis system of the present disclosure; FIG. 2 illustrates an example embodiment of a resampling unit; FIG. 3 illustrates an example of processing by the resampling unit; FIG. 4 illustrates an example embodiment of an audio synthesis system of the present disclosure;
[0016] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the drawings. Note that the present disclosure is not limited to the embodiments shown below. These implementation examples are merely illustrative, and the present disclosure can be implemented in various forms with various modifications and improvements based on the knowledge of those skilled in the art. Note that components with the same reference numerals in this specification and drawings indicate the same components.
[0017] The audio synthesis system of the present disclosure receives an audio signal along with a video signal, uses the video signal to regenerate a clock, resamples the audio signal from the clock and a local clock, and uses the resampled audio signal to synthesize with another audio signal.
[0018] The clock can be recovered using at least one of (i) a synchronization signal for recovering the video signal, or (ii) the arrival time of the video signal. In the first embodiment, the method (i) is used, and in the second embodiment, the method (ii) is used.
[0019] In the following embodiment, an example is shown in which the acoustic signal is upsampled from the clock and a local clock, an acoustic signal that is closest to the local clock is selected from the upsampled acoustic signals, and the selected acoustic signal is synthesized with another acoustic signal. This will be described in detail below.
[0020] 1 shows an example of an audio synthesis system according to this embodiment. The audio synthesis system of this embodiment includes a packet receiving unit 11, a signal separating unit 12, an audio reproducing unit 13, a clock source 14, and a mixing unit 15. The clock source 14 generates a reference frequency f lo Output the local clock.
[0021] The packet receiving unit 11 receives packets transmitted from multiple transmission sources. The signal separating unit 12 outputs packets to different sound reproducing units 13, one for each transmission source. At this time, the signal separating unit 12 separates the video signal and the audio signal in the packet and outputs them.
[0022] In this way, a sound reproducing unit 13 is provided for each packet transmission source. In this embodiment, as an example, only a sound reproducing unit 13_k that processes a packet transmitted from the k-th transmission source is shown.
[0023] In this embodiment, the sound reproducing unit 13_k includes a clock reproducing unit 31, a frequency detecting unit 32, and a resampling unit 33. The clock reproducing unit 31 detects a clock from a video signal and outputs a clock signal. The frequency detecting unit 32 uses a local clock from the clock source 14 to detect the clock frequency f of the clock signal from the clock reproducing unit 31. k The resampling unit 33 detects the clock frequency f detected by the frequency detection unit 32. k , and the acoustic signal is converted to a reference frequency f lo Resample with .
[0024] The mixing unit 15 mixes the sound signals resampled by each sound reproducer 13. In this way, the sound synthesis system according to this embodiment is an sound mixer that has a plurality of sound reproducers 13 and a mixer that mixes the sound signals from the sound reproducers 13.
[0025] A video signal includes various signals that separate the video signal itself on the time axis, such as a vertical synchronization signal determined by the video frame rate. Therefore, in this embodiment, the vertical synchronization signal included in the video signal is used to recover a clock from a signal that does not have a synchronization mechanism such as clock transmission. The signal that separates the video signal itself on the time axis is not limited to the vertical synchronization signal. For example, information on the video frame rate (vertical synchronization), horizontal synchronization, and resolution from the kth transmission source may be acquired in advance, and the clock may be detected from the video signal using this information. This information may be received and exchanged in advance from the kth transmission source, or it may be received together with the video signal.
[0026] When the clock recovery unit 31 detects a vertical synchronization signal from the video signal, it outputs the signal as a clock signal. The clock recovery unit 31 may output the clock signal using information on the time of reception of the video signal. The timing at which the clock recovery unit 31 outputs the clock signal may be determined each time predetermined information is detected, or may be determined using multiple pieces of predetermined information. By observing multiple signal sections in this way, the accuracy can be improved.
[0027] 2 shows an example of the configuration of the resampling unit 33 provided in the sound reproducing unit 13. The resampling unit 33 includes an upsampling unit 51, an interpolation unit 52, and a downsampling unit 53.
[0028] The upsampling unit 51 operates at a clock frequency f k For example, the source signals S11 to S13 of the acoustic signal shown in FIG. 3A are upsampled by a factor of four, as indicated by the circles in FIG. 3B. The value of the upsampled interpolated sample can be any value, but can be set to 0, for example.
[0029] Here, the sampling frequency is, for example, f k and f lo The frequency of the least common multiple of f can be used. k and f lo If the difference between the two is small, the least common multiple becomes very large, and the amount of calculation becomes unrealistic. Therefore, the sampling frequency in the up-sampling unit 51 is set to f k It may be set to a positive multiple of f. k For example, multiples of f k This can be set to about 2 to 8 times the amount of calculation.
[0030] The interpolation unit 52 calculates the value of the interpolated sample. For this data value, an appropriate intermediate data value according to the input acoustic signal can be used. For example, the calculation of the intermediate data value is performed by f k A low pass filter (LPF) that cuts off frequencies equal to or greater than 1 / 2 can be used. This LPF can be configured as a finite impulse response (FIR) filter.
[0031] The downsampling unit 53 samples a sample point that coincides with the local clock from the clock source 14. For example, when the time indicated by the dashed dotted line in Figure 3(c) coincides with the local clock from the clock source 14, the downsampling unit 53 selects signals S21, S24, and S27 that coincide with the timing of the dashed dotted line from among the source signals S11 to S13 and the interpolated signals S21 to S27.
[0032] During downsampling, if there is no sample point that matches the local clock, a nearby sample point can be selected. The value of a nearby sample point that does not match the local clock will differ from the value when it is matched to the local clock. Therefore, data for the local sample clock point can be generated from multiple nearby sample points using a process with a low computational load, such as linear interpolation.
[0033] Here, in the upsampling, if it is 2 to 8 times, f lo There may be cases where there is no sample time where the sampling time satisfies the above condition. In this case, a nearby sample point can be used instead. Furthermore, the downsampling unit 53 outputs the selected signals S21, S24, and S27 at timings that match the local clock from the clock source 14. This allows the mixing unit 15 to combine the audio signals resampled by each audio playback unit 13 without any deviation in the sampling clock.
[0034] As described above, the sound reproducing unit 13_k uses the video signal stored in the packet transmitted from the kth transmission source to reproduce the sound signal stored in the packet transmitted from the kth transmission source at the reference frequency f lo The sound signals output from each sound reproducing unit 13 to the mixing unit 15 are resampled at a reference frequency f lo Since the signals are resampled at , they can be synthesized in the mixing unit 15 without any deviation in the sampling clock.
[0035] Second Embodiment Fig. 4 shows an example of an audio synthesis system according to this embodiment. The operations of the audio reproducing unit 13 and the mixing unit 15 are the same as those in the first embodiment. In this embodiment, the clock reproducing method in the clock reproducing unit 31 is different from that in the first embodiment.
[0036] In this embodiment, when the packet receiving unit 11 receives packets transmitted from multiple transmission sources, it outputs the packets to different sound reproducing units 13 for each packet transmission source. The signal separating unit 12 outputs the packets to different sound reproducing units 13 for each packet transmission source. At this time, the signal separating unit 12 separates the sound signals in the packets and outputs them to different sound reproducing units 13 for each packet transmission source.
[0037] The clock recovery unit 31 recovers a clock from the time of reception of a packet containing video data from the packet receiving unit 11. Since the amount of video data is much larger than the amount of audio data, the averaging effect is high, and it is possible to reduce the error in the clock recovered from the arrival rate of that data.
[0038] For example, when one frame of data is divided into n pieces on average at a frame rate R, the packet arrival period is 1 / (Rn). Therefore, the clock recovery unit 31 can recover the time clock from the packet arrival period.
[0039] The number of packets arriving in T seconds is TRn. Therefore, the clock recovery unit 31 may recover the clock from the number of packets arriving per unit time.
[0040] In this embodiment, as in the first embodiment, the clock may be recovered using the video signal contained in the packet.
[0041] (Other Embodiments) The packet receiving unit 11 may output signals by shortening the buffer length as much as possible and maintaining the burstiness of received packets. Also, the signal separating unit 12 may separate video signals and audio signals by shortening the buffer length as much as possible and maintaining the burstiness of data from the packet receiving unit 11. This allows the clock regenerating unit 31 to regenerate the clock with high accuracy.
[0042] Furthermore, in addition to or instead of the above method, the resampling unit 33 may apply a polyphase filter to reduce the amount of calculation. Furthermore, resampling can be performed in a manner other than the above method.
[0043] The packet receiving unit 11, the signal separating unit 12, and the clock regenerating unit 31 may be integrated into one unit. For example, the necessary data may be extracted directly from the received packet to regenerate the clock.
[0044] The device of the present invention can also be realized by a computer and a program, and the program can be recorded on a recording medium or provided via a network.
[0045] As described above, the present disclosure recovers a clock from a video signal, rather than from an audio signal itself. In addition to pixel data streams, video signals contain signals indicating divisions of the video signal, such as vertical and horizontal synchronization signals. Based on these synchronization signals, a clock can be recovered with high precision. Therefore, an audio signal with no sampling clock deviation can be synthesized and generated from packetized video and audio signals that are sampled with multiple different clocks and do not contain clock information.
[0046] 11: Packet receiving unit 12: Signal separating unit 13: Sound reproducing unit 14: Clock source 15: Mixing unit 31: Clock reproducing unit 32: Frequency detecting unit 33: Resampling unit 51: Upsampling unit 52: Interpolating unit 53: Downsampling unit
Claims
1. An audio synthesis system that receives an audio signal together with a video signal, reproduces a clock using the video signal, resamples the audio signal from the clock and a local clock, and synthesizes the resampled audio signal with another audio signal.
2. The audio synthesis system according to claim 1, wherein the reproduction of the clock is performed using at least one of (i) a synchronization signal for reproducing the video signal, or (ii) the arrival time of the video signal.
3. The audio synthesis system according to claim 1, which upsamples the audio signal from the clock and the local clock, selects an audio signal close to the local clock from the upsampled audio signals, and synthesizes the selected audio signal with another audio signal.
4. An audio synthesis method that receives an audio signal together with a video signal, reproduces a clock using the video signal, resamples the audio signal from the clock and a local clock, and synthesizes the resampled audio signal with another audio signal.
Citation Information
Patent Citations
Synchronization of digital audio to digital images
JP1996511373A
Video server
JP2000243036A
System and method for providing video conferencing synchronization
US7084898B1