Method for generating an output audio signal intended to accompany a slowed-down video signal, method for broadcasting a slowed-down audio-video signal, output audio signal, and slowed-down audio-video signal

The method of segmenting and synchronizing audio signals with timecodes addresses the challenge of providing timely and realistic sound for slow-motion video in live broadcasts, ensuring high-quality playback and viewer engagement.

WO2026058226A1PCT designated stage Publication Date: 2026-03-19THE FAKTORY
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-09-15
Publication Date
2026-03-19

AI Technical Summary

Technical Problem

Existing methods for generating audio to accompany slow-motion video sequences in live broadcasts fail to provide realistic and synchronized sound in a timely manner, especially when the audio is delayed by several seconds, leading to inappropriate mixes and viewer discomfort.

Method used

A method involving segmenting input audio signals, assigning timecodes, generating filler grains through interpolation, and synchronizing these signals with slow-motion video using metadata, allowing for rapid production and seamless integration into live broadcasts.

Benefits of technology

Enables the rapid generation and synchronization of realistic audio for slow-motion sequences, ensuring high-quality playback without operator intervention, enhancing viewer experience and broadcast coherence.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

The first aspect of the invention relates to a method for generating an output audio signal intended to accompany a slowed-down video signal, comprising the following steps: a) capturing an input audio signal; b) splitting the input audio signal into a plurality of segments; c) recording the plurality of segments of the input audio signal; d) assigning a time code to each segment of the input audio signal; e) recording the time code of a segment in metadata linked to the segment of the audio signal; f) generating an output audio signal of a predetermined duration corresponding to a speed slowed down by a predetermined factor for each of the segments of the audio signal, as follows: - dividing the segment of the input audio signal into a set of grains (G0, G1, G2,...) each having a length (L0, L1, L2,...) and a temporal position (P0, P1, P2,...), - inserting the grains (G0, G1, G2,...) into the segment of the output audio signal at a temporal position (P'0, P'1, P'2,...) that corresponds to the original temporal position of the grains (G0, G1, G2,...) corrected by a factor dependent on the speed of the slowed-down sequence, - generating filler grains (G'0, G'1,...) intended to fill the spaces between the grains (G0, G1, G2,...) of the segment of the output audio signal and - inserting the filler grains (G'0, G'1,...) into the spaces between the grains (G0, G1, G2,...) of the segment of the output audio signal, - the filler grains (G'0, G'1,...) being generated by interpolation; g) recording the output audio signals corresponding to each segment while keeping the time code in the metadata. Other aspects of the invention relate to a method for broadcasting a slowed-down audio-video signal based on the use of the output audio signal thus generated, and to the output audio signal as such.
Need to check novelty before this filing date? Find Prior Art

Description

T0560-BE-P 1 Method for generating an output audio signal intended to accompany a slow-motion video signal, method for broadcasting a slow-motion audio-video signal, output audio signal and slow-motion audio-video signal.

[0001] Description

[0002] The present invention relates to a method for generating an output audio signal to accompany a slow-motion video signal, as well as a method for broadcasting a slow-motion audio-video signal. Other aspects of the invention concern the output audio signal thus generated and the synchronized slow-motion audio-video signal.

[0003] Indication of prior art

[0004] When broadcasting an event, such as a sporting event, it is often desirable to show slow-motion sequences to allow viewers to better visualize and understand certain key moments or simply relive them. Examples include a penalty kick in a football match, a foul in a tennis match, a breakaway during a cycling race, or an overtaking maneuver in a motor race. To accompany these slow-motion sequences, it is not possible to use the original audio of the sequence, as slowing it down by the same factor would alter the audio signal's frequency due to the reduced speed, making it unrealistic and even unpleasant for the viewer.This is why such slow-motion sequences are often accompanied only by ambient sound (for example, from the stadium, the road, or the track in the cases mentioned above) and / or commentary from sports commentators. In the case of a live broadcast, this mix can sometimes become inappropriate due to timing issues such as a restart of play, a whistle, or a sudden change in atmosphere in the stadium, on the track, or at the circuit.

[0005] Application WO-A1-2025 / 186474, the contents of which are incorporated herein by reference, proposes a method for generating an output audio signal (A') to accompany a slow-motion video signal (Vr). This method comprises the following steps performed in order: - division of an input audio signal (A) into a set of grains (GO, Gl, G2, ...) each possessing a length (L0, LI, L2, ...) and a time position (PO, Pl, P2, ...) based on the start of the grains, - insertion of grains (GO, Gl, G2, ...) into the output audio signal (A 1 ) at a temporal position (P'0, P'1, P'2, ...) which corresponds to the original temporal position of the beginning of the grains (GO, Gl, G2, ...) corrected by a factor dependent on the speed of the slowed sequence, - generation of filler grains (G'0, G'1, ...) intended to fill the spaces between the grains (GO, Gl, G2, ...) of the output audio signal (A 1 ), the generation of the filling grains being achieved, for example, by diffusion-based audio inpainting, - insertion of the filler grains (G'0, G'1, ...) into the spaces between the grains (G0, G1, G2, ...) of the output audio signal (A 1 In this way, a very realistic audio output signal can be obtained, perfectly matching the slow-motion video signal.

[0006] This method already provides remarkable results. The inventors have now focused on developing a very fast solution to use the method described above during a live broadcast. It should be noted that a slow-motion replay is never truly shown live, as it depicts a completed action that occurred in the past. For example, when a goal is scored, the operator typically waits 3 to 5 seconds before requesting a slow-motion replay of the preceding action, meaning that the sound of the assist is already at least 10 seconds old.

[0007] It would therefore be desirable to provide a solution that allows for the very rapid delivery, for example in 10 seconds or even less, of realistic sound to accompany a slow-motion video sequence. It would also be desirable to propose a method for broadcasting a slow-motion video signal that incorporates the sound thus generated.

[0008] Description of the invention

[0009] The invention that is the subject of this patent application aims to solve this technical problem. To achieve this, an output audio signal must first be generated to accompany a slow-motion video signal. Secondly, a slow-motion audio-video signal using the generated output audio signal can be played back.

[0010] According to the invention, a method is provided for generating an output audio signal to accompany a video signal slowed down by a predetermined factor. This method comprises the following steps: a) capturing an input audio signal; b) dividing the input audio signal into a plurality of segments; c) recording the plurality of segments of the input audio signal; d) assigning a timecode to each segment of the input audio signal; e) recording the timecode of a segment in metadata associated with the segment of the input audio signal; f) generating an output audio signal of a predetermined duration corresponding to a slowed-down speed by the predetermined factor for each segment of the audio signal, as follows: - division of the input audio signal segment into a set of grains (GO, Gl, G2, ...) each possessing a length (L0, LI, L2, ...) and a time position (PO, Pl, P2, ...), - insertion of grains (GO, Gl, G2, ...) into the output audio signal segment at a time position (P'0, P'1, P'2, ...) which corresponds to the original time position of the grains (G0, Gl, G2, ...) corrected by a factor dependent on the speed of the slowed-down sequence, - generation of filler grains (G'0, G'1, ...) intended to fill the spaces between the grains (GO, Gl, G2, ...) of the output audio signal segment and - insertion of filler grains (G'O, G'1, ...) into the spaces between the grains (GO, Gl, G2, ...) of the output audio signal segment, - the generation of the filling grains (G'O, G'1, ...) is carried out by interpolation; g) recording of the output audio signals corresponding to each segment of the input audio signal while preserving the time code in the metadata.

[0011] It is therefore understood that in step a), the input audio signal is systematically recorded (this is the original sound recorded live during the event). The recording of the input audio-video signal is done in an appropriate format. For example, it could be an SDI signal recorded in PAL, NTSC, HD, or even 4K format.

[0012] In step b), the input audio signal is cut into segments with a duration of, for example, 0.5 to 3 seconds. Typically, the segments are 1 second long.

[0013] It is understood that the steps are carried out successively in order, with the exception of steps c), d), and e), which can be interchanged. Indeed, in one embodiment, a timecode is assigned to each segment of the audio signal before the segments are recorded. In this case, it is even advantageous to record the segments and their timecodes simultaneously. In another embodiment, the procedure is carried out following the sequence a) to g).

[0014] The timecode assigned in step d) is a well-known and widely used time reference in the audio and video industries for synchronizing and marking recorded material, for example, in an SDI signal. SDI (Serial Digital Interface) is a digital video signal transmission standard that can carry any type of video signal, such as PAL, NTSC, HD, or even 4K formats. In addition to the video signal, the SDI signal can carry up to 32 audio channels or more, making it a versatile standard for professional productions. The timecode allows for the synchronization of multiple video and audio signals. For example, in a OB van or in a fixed production center, such as the control rooms of broadcast stations that have virtualized OB van production, all cameras and recorders can be precisely aligned to facilitate editing.Timecode provides a unique time reference for each frame in the video stream, simplifying the search and precise editing of specific segments. Timecode is expressed in hours, minutes, seconds, and frames. For recordings in digital files, the timecode indicating the start time of the recording is included in the metadata. One specific timecode is the SMPTE timecode (defined by the Society of Motion Picture and Television Engineers in the SMPTE 12M specification). Adherence to this standard ensures that the timecode is compatible with professional equipment used in production and post-production. Several documents describe the insertion of a timecode; for example, US-A1-2019 / 208239 and US-A1-2021 / 297717.

[0015] In step f), an output audio signal is generated with a predetermined duration corresponding to the predetermined slowdown factor of the video signal. For example, if you want to play a 5-second sequence slowed down by a factor of 1, it will be played over 10 seconds, and for a factor of 3, it will be played over 15 seconds. Therefore, the output audio signal must also be 10 or 15 seconds long, respectively. For example, the predetermined slowdown factor can be 1.5, 2, 3, or 4. The preferred predetermined slowdown factors are 2 and 3 because they correspond to conventional standards in the live event broadcasting industry. Of course, the predetermined slowdown factor will be adjusted depending on the nature of the event. It is understood that this step (f) can be implemented continuously. According to the invention, step (f) is implemented as follows: - division of the input audio signal segment into a set of grains (GO, Gl, G2, ...) each possessing a length (L0, LI, L2, ...) and a time position (PO, Pl, P2, ...), - insertion of grains (GO, Gl, G2, ...) into the output audio signal segment at a time position (P'0, P'1, P'2, ...) which corresponds to the original time position of the grains (G0, Gl, G2, ...) corrected by a factor dependent on the speed of the slowed-down sequence, - generation of filler grains (G'0, G'1, ...) intended to fill the spaces between the grains (GO, Gl, G2, ...) of the output audio signal segment and - insertion of filler grains (G'0, G'1, ...) into the spaces between the grains (G0, G1, G2, ...) of the output audio signal segment, - The generation of the filling grains (G'0, G'1, ...) is carried out by interpolation. This method is indeed very fast and provides a very realistic result.

[0016] Advantageously, to avoid accumulating delays in the generation of these output signals, several successive segments of the input audio signal are processed in parallel. By selecting such segments of sufficiently short duration (for example, 1 to 10 seconds, or even less than one second), it can be ensured that, even over very long periods, slowed-down audio output signals at the appropriate speed are available within the required timeframe for the playback of a slowed-down sequence.

[0017] In an advantageous embodiment, for each segment of the input audio signal, several output audio signals are generated at slowed-down speeds with different predetermined slow-down factors. This allows the operator to subsequently slow down the action they wish to rebroadcast to a speed chosen from among the different speeds of the various output audio signals. Again, the operator will pre-select the two or more predetermined slow-down factors best suited to the event being broadcast. For example, for a football or tennis match, slow-down factors of 2 and 3 are perfectly appropriate.

[0018] Once one or more audio output signals intended to accompany a slow-motion video signal have been generated, it remains to ensure that this signal can be perfectly synchronized with the displayed images. This synchronization must be possible without operator intervention.

[0019] To this end, the inventors have devised a method for broadcasting a slowed-down audio-video signal. This method comprises the following steps: a) selecting a sequence from the video component of an input audio-video signal, comprising an audio component and a video component, for broadcast at a slowed-down speed by a predetermined slowdown factor; b) detecting the timecode of the audio component of the selected input audio-video signal; c) searching for the timecode identified in the previous step in the metadata of the segments of the generated output audio signal as described above and identifying the location of the audio signal segment corresponding to the slowed-down audio-video sequence; d) broadcasting the video component of the slowed-down audio-video signal combined with the generated output audio signal as described above, the two signals being synchronized by the timecode of the audio signal.

[0020] When implementing the broadcasting method of the invention, the operator selects a sequence of the input signal that they wish to broadcast at a slower speed and sends it to the system's input. Naturally, the predetermined slowdown factor must correspond to the slowdown factor of the generated output audio signal (or to that of one of the generated output audio signals). Next, the timecode associated with the audio component (which will not be broadcast) of the input audio-video signal (and whose video component will be broadcast at a slower speed) is detected. The system can then search for the timecode identified in the previous step in the metadata of the segments of the generated output audio signal as described above and identify the location of the audio signal segment corresponding to the slowed-down audio-video sequence.

[0021] Finally, the video component of the audio-video signal is played back in slow motion combined with the generated output audio signal as described above, the two signals being synchronized by the timecode of the audio signal.

[0022] The inventors were able to determine that these operations could be implemented in an extremely short timeframe, fully compatible with the requirements of live event broadcasting.

[0023] As we have seen above, it is possible to generate several output audio signals so that the operator can select the slowdown factor of the video sequence that best suits the situation he wishes to present.

[0024] In some cases, the event being broadcast is so important that it is covered by several cameras, each capturing an input audio-video signal. For a tennis match, for example, there might be fourteen or more cameras. Before broadcasting a slow-motion sequence of an audio-video signal, the first step is therefore to choose the best input audio-video signal from those available. For example, this could be the signal from the camera closest to the action or from a more distant camera that provides a better overview. This selection must ensure that the captured action is complete and reproduced with optimal quality. With the integration of a generated output audio signal into the slow-motion sequences, it is preferable to be able to verify the quality of the audio signal associated with the selected input audio-video signal, adding further complexity for directors and camera operators.

[0025] According to one embodiment of the invention, a sound preview function is proposed to facilitate this process for operators broadcasting a slow-motion audio-video signal. This function allows them to verify that the output audio signal is synchronized with the images before broadcasting the replay to a wide audience. This ensures audio-video coherence for an optimal viewing experience.

[0026] To this end, the operator is provided with an accelerated version (by a factor of two or three, for example) of the generated output audio signal. Even though this audio signal is inaudible at normal speed because it is accelerated, it becomes usable for the operator when they play back the video signal at a slowed-down speed corresponding to the acceleration of the provided audio signal. Indeed, when the operator combines the slow-motion video signal, the generated and accelerated audio signal then returns to its nominal speed (we explained above how synchronization can be achieved with timecode). This audio signal preview function thus allows slow-motion operators to have an audio preview of the generated output signal across all audio-video recordings before final broadcast.It should be noted that the broadcast signal does not include this accelerated audio signal, and any potential quality loss resulting from this acceleration / slowing process does not affect the quality of the broadcast signal. On the other hand, this audio signal, even if slightly degraded, is sufficient to allow effective control of the audio signal and thus enables the selection of the correct input audio-video signal. Based on this advantageous variant, we therefore propose a method for broadcasting a slowed-down audio-video signal as explained above, in which the input audio-video signal is selected after previewing a slowed-down audio-video signal. This signal is called the preview audio-video signal. This preview audio-video signal was obtained by combining the video component of the original audio-video signal with the output audio signal generated according to the method described above.The output audio signal was transmitted at an accelerated speed by a factor corresponding to the deceleration factor.

[0027] Advantageously, the broadcast method involves showing a sequence consisting of an initial still image, the slow-motion sequence itself, and a final still image. The operator proposes a still image to the director, who gives the go-ahead to start the slow-motion broadcast, often using a special effect to announce the slow motion to viewers. This slow motion is played at a predetermined speed. The slow motion concludes with a still image, summarizing the event and indicating to viewers that the slow motion has ended before returning to the live event or showing another sequence. For example, at the end of the slow-motion sequence, a still shot corresponding to the end of the broadcast sequence or a representative image from the sequence chosen by the operator or director can be shown.To maintain viewer attention, it's best to clearly signal that they are watching a past event and then return to the live event. That's why it's always best for a director to begin a slow-motion sequence with a still image and end it with a still image as well. This allows viewers to clearly distinguish what belongs to the past from what is happening live.

[0028] According to an advantageous embodiment, a smooth transition of the broadcast audio signal from the ambient sound of the live event to the generated output audio signal is provided at the beginning of the sequence, and a smooth transition of the broadcast audio signal from the generated output audio signal back to the ambient sound of the live event at the end of the sequence. This operation is achieved using the "fade out, fade in" technique. Fade in and fade out are commonly used audio techniques for creating smooth transitions. A fade in gradually increases the volume of a sound from silence to a normal level. Conversely, a fade out gradually decreases the volume from a normal level to silence. These transitions create a fluid introduction or conclusion, whether for a visual scene or a sound element.In this case, we can use a device (for example, a mixer) that will implement the playback of the slow-motion sequence and start the audio sequence (sometimes even before the image appears) with the ambient sound of the moment (for example, from the stadium). The volume of this sound gradually decreases (fading out) to be replaced by a generated sound whose volume gradually increases (fading in).

[0029] According to an advantageous embodiment, the broadcast of the still image at the beginning or end of the sequence is accompanied by an audio signal corresponding to that image. This audio signal can be generated as described above by taking as the input audio signal a segment at the beginning of the audio component of the input audio-video signal, this segment being chosen to correspond to the image. For example, it could be a 1- to 2-second segment whose midpoint corresponds to the timecode of the image. It should be understood that to generate such an audio signal, audio inpainting will be performed as described in Belgian patent application number 2024 / 5138, but this time using this The segment is treated as a series of preceding and following grains (inpainting with itself), creating a sort of 'infinite sound loop'. A sound-based audio signal corresponding to this still image can be generated for 10 to 20 seconds without the viewer realizing that the sound is not real but generated. The ability to add sound to a still image is valuable for ensuring a smooth transition and avoiding abrupt sound changes that the human ear finds difficult to accept. Indeed, when the operator waits for the director's signal to start playing the slow-motion sequence, some time may elapse, and it is only one or two seconds before the director's signal that the audio mix will begin. Therefore, it will only begin on the director's instruction, just a fraction of a second before the still image opens and the slow-motion sequence begins.This opening sequence is usually accompanied by special effects intended to signal to the viewer that a slow-motion sequence is being shown.

[0030] This ability to synchronize sound with still images and slow-motion replays is a true revolution for those in charge of audio mixing during sports broadcasts. It significantly improves the quality of the viewing experience, thereby increasing audience numbers and revenue.

[0031] The present invention has been described using a single audio channel, but it is understood that the invention is not limited to this embodiment and extends to all multichannel formats used in production. This includes not only stereo, but also formats such as Dolby 5.1, Dolby 7.1, and others. The present invention therefore supports all types of input audio signals in terms of the number of channels, bandwidth, and sampling rate (generally 48 kHz, but compatible with other speeds such as 44.1 kHz or 96 kHz). It also works with various resolutions, such as 16-bit, 24-bit, or 32-bit.

[0032] Another aspect of the invention relates to the output audio signal as it can be generated in the manner described above.

[0033] Finally, yet another aspect of the invention relates to the slowed-down audio-video signal as it can be broadcast in the manner described above.

Claims

Demands 1. A method for generating an output audio signal to accompany a slow-motion video signal, comprising the following steps: a) capturing an input audio signal; b) dividing the input audio signal into a plurality of segments; c) recording the plurality of segments of the input audio signal; d) assigning a timecode to each segment of the input audio signal; e) recording the timecode of a segment in metadata associated with the segment of the audio signal; f) generating an output audio signal of a predetermined duration corresponding to a slowed-down speed by a predetermined factor for each of the segments of the audio signal, as follows: - generation of a first graphical representation of the input audio signal segment; - division of the first graphical representation of the input audio signal segment into a set of grains (GO, Gl, G2, ...) each possessing a length (LO, LI, L2, ...) and a time position (PO, Pl, P2, ...), - insertion of the grains (GO, Gl, G2, ...) into a second graphical representation representing the segment of the output audio signal at a time position (P'O, P'1, P'2, ...) which corresponds to the original time position of the grains (GO, Gl, G2, ...) corrected by a factor dependent on the speed of the slowed-down sequence, - generation by inpainting audio based on the diffusion of filler grains (G'O, G'1, ...) intended to fill the spaces between the grains (GO, Gl, G2, ...) of the second graphical representation representing the segment of the output audio signal and - insertion of the filler grains (G'O, G'1, ...) into the spaces between the grains (GO, Gl, G2, ...) of the second graphical representation representing the output audio signal segment, - conversion of the second graphical representation representing the output audio signal by the inverse operation of the first sub-step of step f); g) recording of the output audio signals corresponding to each segment while preserving the timecode in the metadata.

2. Method for generating an output audio signal according to claim 1, wherein steps a), b), f) and g) are carried out sequentially and steps c) and d) are carried out in any order between steps b) and f), step e) necessarily following step d).

3. Method for generating an output audio signal according to claim 2, wherein steps c) and e) are carried out simultaneously.

4. Method for generating an output audio signal according to any one of claims 1 to 3, wherein several successive segments of the input audio signal are processed in parallel.

5. Method for generating an output audio signal according to any one of claims 1 to 4, wherein, for each segment of the audio signal of the input signal, several output audio signals are generated at different slowed-down speeds of different predetermined factors.

6. Method for generating an output audio signal according to any one of the preceding claims, wherein the generation of the filling grains (G'O, G'1, ...) is carried out by audio inpainting, preferably based on diffusion.

7. A method for broadcasting a slowed-down audio-video signal comprising the following steps: a) selecting a sequence of an input audio-video signal, comprising an audio component and a video component, for broadcasting at a slowed-down speed of a predetermined slowdown factor; b) detecting the timecodes of the audio component of the selected input audio-video signal; c) searching for the timecode identified in the previous step in the metadata of the segments of the output audio signal generated according to any one of claims 1 to 7 and identifying the location of the audio signal segment corresponding to the slowed-down audio-video sequence; d) broadcasting the video component of the slowed-down audio-video signal combined with the output audio signal generated according to any one of claims 1 to 6, the two signals being synchronized by the timecode of the audio signal.

8. Method of broadcasting a slowed-down audio-video signal according to claim 7, wherein the input audio-video signal is selected after previewing a slowed-down audio-video signal, referred to as the preview audio-video signal, obtained by combining the video component of the original audio-video signal with the output audio signal generated according to any one of claims 1 to 6, said output audio signal having been transmitted at an accelerated speed by a factor corresponding to the slowing-down factor.

9. A method for broadcasting a slow-motion audio-video signal according to claim 7 or 8, wherein the broadcast sequence comprises in order i) a fixed shot corresponding to the beginning of the sequence of the selected audio component; ii) broadcast of the slowed-down sequence; ii) a fixed shot corresponding to the end of the broadcast sequence or to a representative image of the sequence.

10. Audio output signal obtained by a method according to any one of claims 1 to 6.

11. Slow-motion audio-video signal obtained by combining the video component of the slow-motion audio-video signal with the output audio signal generated according to any one of claims 1 to 6, the two signals being synchronized by the timecode of the audio signal.

Citation Information

Patent Citations

  • Time-stretching of an audio signal

    EP2509073A1

  • Method and apparatus for synchronizing audio and video signals

    US20150380054A1

  • Systems and Methods for Providing Audio Content During Trick-Play Playback

    US20190208239A1

  • Method and system for transmitting alternative image content of a physical display to different viewers

    US20210297717A1

  • Method for generating an audio signal

    WO2025186474A1