Data processing apparatus, system and method

The data processing apparatus enhances audio-visual content by visually indicating sound sources in video frames using audio data processing techniques, addressing the challenge of sound source identification for individuals who are hard of hearing.

WO2026002730A1PCT designated stage Publication Date: 2026-01-02SONY SEMICON SOLUTIONS CORP +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2025/066961
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-06-25
Filing Date
2025-06-17
Publication Date
2026-01-02

AI Technical Summary

Technical Problem

People who are hard of hearing struggle to identify the source of sounds in audio-visual content, as conventional subtitles provide a less rich experience compared to hearing the sounds themselves.

Method used

A data processing apparatus that modifies video frames to visually indicate sound sources based on associated audio data, using techniques like image segmentation, sound separation, and multi-modal machine learning to enhance the visual representation of sound origins.

Benefits of technology

Enhances the audio-visual experience for individuals who are hard of hearing by providing a richer and more immersive indication of sound sources within video content, improving their ability to understand sound origins.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2025066961_02012026_PF_FP_ABST
    Figure EP2025066961_02012026_PF_FP_ABST
Patent Text Reader

Abstract

A data processing apparatus comprising circuitry configured to: receive input video and audio data; modify a region of a video frame of the input video data to generate a modified video frame, the region being modified based on an audio sample of the input audio data temporally associated with the video frame; and output modified video data representing the modified video frame.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] DATA PROCESSING APPARATUS, SYSTEM AND METHOD

[0002] BACKGROUND

[0003] Field of the Disclosure

[0004] The present disclosure relates to a data processing apparatus, system and method.

[0005] Description of the Related Art

[0006] The “background” description provided is for the purpose of generally presenting the context of the disclosure. Work of the presently named inventors, to the extent it is described in the background section, as well as aspects of the description which may not otherwise qualify as prior art at the time of filing, are neither expressly or impliedly admitted as prior art against the present disclosure.

[0007] People who are hard of hearing may struggle to fully experience audio-video (AV) content. For example, even if a video is provided with subtitles indicating both sounds (e.g. thunder, crashes, gunfire, etc.) and speech, it may still be unclear from where in the video frame a particular sound originates. The detailed nature of the sound may also be unclear, especially for rich sounds which can be difficult to meaningfully describe using text alone. This leads to a situation in which people who are not hard of hearing are able to experience rich and variable sounds while people who are hard of hearing see only text. This can detract from the experience of people who are hard of hearing.

[0008] There is therefore a desire to address this problem.

[0009] SUMMARY

[0010] The present technology is defined by the claims.

[0011] BRIEF DESCRIPTION OF THE DRAWINGS

[0012] Non-limiting embodiments and advantages of the present disclosure are explained with reference to the following detailed description taken in conjunction with the accompanying drawings, wherein:

[0013] Fig. 1 schematically shows an example data processing apparatus;

[0014] Fig. 2 schematically shows modification of video data;

[0015] Fig. 3 shows a first example video modification technique;

[0016] Fig. 4 shows a second example video modification technique; and

[0017] Fig. 5 shows an example method.

[0018] Like reference numerals designate identical or corresponding parts throughout the drawings. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0019] Fig. 1 shows an example data processing apparatus 100 for implementing the present technology. The data processing apparatus 100 comprises a processor 101 for executing electronic instructions, a memory 102 (e.g. volatile memory) for storing the electronic instructions to be executed and electronic input and output information associated with the electronic instructions, a storage medium 103 (e.g. non-volatile memory) for long term (persistent) storage of information, a communication interface 104 for sending information to and / or receiving information from one or more other apparatuses and a user interface 105 (e.g. a touch screen, a non-touch screen, buttons, a keyboard and / or a mouse) for receiving commands from and / or outputting information to a user. Each of the processor 101 , memory 102, storage medium 103, communication interface 104 and user interface 105 are implemented using appropriate circuitry, for example. The processor 101 controls the operation of each of the memory 102, storage medium 103, communication interface 104 and user interface 105.

[0020] The present technology derives information from the audio data provided with video data which helps provide a richer experience to people who are hard of hearing compared to the use of conventional subtitles. The present technology is also applicable to any video data (with its corresponding audio data) and requires no special treatment (e.g. the addition of specific metadata) to be applied to the video data. The video and audio data may be stored (e.g. on storage medium 104) and played back or streamed from another apparatus (e.g. received by communication interface 104 over a network). The output is a modified version of the video data which may then either be played back (e.g. via a display of user interface 105) or streamed to another apparatus (e.g. transmitted by communication interface 104 over a network).

[0021] With the present technology, the video data is modified according to the audio data so that visual information representing the audio data is provided as part of the video data. In particular, sound sources in the displayed video (that is, objects shown in the displayed video which generate sound represented by the audio data, such as cloud which generates a “thunder” sound or a lion which generates a “roar” sound) are visually marked while they are emitting a sound and visually unmarked once they stop emitting the sound. The marking of the sound source can indicate characteristics of the sound such as pitch, volume, frequency, timbre, intensity and others.

[0022] It is noted that, although the present technology may be particularly beneficial for people who are hard of hearing, it may also be beneficial for people who are not hard of hearing and / or have other condition(s) which may make identifying the source of a sound in a video more challenging (e.g. those with neurological conditions such as auditory agnosia or auditory processing disorder).

[0023] Fig. 2 shows a simplified example of input(s) and output(s) according to the present technology. The input(s) include video data 201 (representative of successive video images I frames) and audio data 202 associated with the video data 201 (e.g. audio data such as speech and other sounds captured at the same time and location as the video data). As an example, a video frame 206A of the input video data 201 is shown. The frame 206A includes an object 205 which is a sound source in the video. That is, the object 205 is a source of one of the sounds represented by the audio data 202. In this case, the sound is thunder and the object 205 is a cloud emitting lightning (lightning cloud). The object 205 defines a region of the frame 206A.

[0024] A person who is hard of hearing may not be able to hear the thunder when the audio data 202 is used to generate output audio. Conventionally, subtitles (that is, text appearing on a display with the output video indicating speech and / or describing sounds present in the audio data) may therefore be use to indicate the occurrence of the thunder. For example, the wording “#Thunder#” may be displayed at the same time as the occurrence of the thunder. However, compared to a person who is not hard of hearing and thus is able to hear the thunder in the output audio, this is a far less rich experience.

[0025] With the present technology, enhancement 203 of the video data is thus performed to generate output video data 204 (modified video data) in which the appearance of the sound source 205 itself is changed (modified) to indicate the sound of the thunder. This is shown in the output video frame (modified video frame) 206B (which is a modified version of the video frame 206A), where the visual appearance of the sound source object 205 has been modified to include an outline 207 indicating the occurrence of the thunder.

[0026] Other types of visual change(s) could also be implemented (instead of or in addition to the addition of outline 207). For example, the object 205 could be made larger (e.g. by 10%, 15% or 20% in area) and / or one or more colours of the object could be changed. The object 205 could also be rendered in higher resolution compared to other objects in the image and / or rendered with an apparent sharper focus compared to other objects it the image.

[0027] Animated effects could also be applied to the object 205 so the appearance of the object changes from frame to frame. For example, the object 205 may be made to alternately increase and decrease in size and / or brightness and / or changes in colour. As another example, the object 205 may be surrounded by a circle (or circle segment) in which the curved line of the circle changes to zigzags or waves of variable amplitude and / or frequency depending on the amplitude and / or frequency of the sound.

[0028] It will be appreciated that any appropriate visual change to indicate the generation of sound from a particular sound source could be used. To help provide processed image frames 206B which include the necessary visual change(s) but maintain context and quality, a method such as pixel- guided diffusion (as exemplified in [1]) or differential diffusion (as exemplified in [2]) may be used. These techniques involve generating an output image frame from an input image frame by reconstructing the input image frame and changing only changing the appearance of one or more selected portions of the input image frame (e.g. an image segment defining object 205 in the example of Fig. 2). The output video data 204 thus visually indicates the sound source of a particular sound represented by the audio data in a way which integrated with the content of the video data and in a way which is richer and more immersive than the existing simple text-based descriptions provided as subtitles. This helps improve the audio-visual experience viewers, in particular viewers who are hard of hearing.

[0029] There are a number of ways in which the enhancement 203 to generate the output video data 204 from the input video data 201 and input audio data 202 may be implemented. Some example technique(s) are described with reference to Figs. 3 and 4.

[0030] Fig. 3 shows a first example technique. The technique is executed by processor 101 , for example.

[0031] At step 301 , the input video data 201 and corresponding input audio data 202 are obtained. For example, they may be received via communication interface 104 or obtained from storage medium 103.

[0032] At step 302, each video frame represented by the input video data is segmented. For each segmented video frame, a number of detected objects and the respective outlines of those detected objects in the video frame are identified. Any suitable known image segmentation technique may be used. For example, the DINOv2 segmentation technique disclosed in [3] may be used.

[0033] At step 303, the input audio data 202 is processed to separate the different sounds (components) represented by the input audio data. For example, for each segmented frame of the input video data 201 , a sample (audio sample) of the associated audio data of a predetermined duration (e.g. 0.5, 1 or 2 seconds) either side of the output time of the frame is processed in this way. In this way, the audio sample is temporally associated with the video frame. This allows, for example, the voice of a speaker to be separated from the background noise produced by thunder. Any suitable known sound separation technique may be used. For example, the sound separation technique disclosed in [4] may be used.

[0034] At step 304, the separated audio components for each segmented frame are localised in the frame. This involves, for example, taking each separated audio component (e.g. one or more separate voices, thunder, etc.) and using one or more characteristics of that audio component to associate it with one or more of the segmented objects in the frame. In the example of Fig. 2, this allows the segmented lightning cloud 205 to be associated with a separated audio component representing thunder.

[0035] The association may involve taking, as inputs, data representing the audio component and data representing each segmented object in the frame. The output is then an identification of one or more of the segmented objects.

[0036] Alternatively, the association may involve taking, as inputs, data representing the audio component and data representing the entire frame. The output is then one or more pixel locations in the video frame indicating the estimated location(s) of the audio component. The pixel location(s) are then used to identify, as the determined sound source(s), the segmented object(s) at those pixel location(s).

[0037] Localisation of each separated audio component (that is, association one or more of the separated audio components with a respective one or more objects in a frame) may be implemented using any suitable known audio localisation technique. For example, the audio localisation techniques disclosed in [5] and / or [6] may be used.

[0038] [5] enables the person speaking a particular instance of dialogue in a video to be identified and is trained using the known Oxford-BBC Lip Reading Sentences 2 (LRS2) dataset.

[0039] [6] enables the source of each of a plurality of predetermined sound classifications in a video to be localised. It uses a MUSIC training data set consisting of 685 videos (with their corresponding sound tracks) of 11 predetermined musical instrument classifications. In an example, with the present technology, a new training data set of a similar size but with different videos (and their corresponding sound tracks) and corresponding predetermined sound classifications (e.g. “thunder”, “crowd”, “ocean”, etc.) may be used.

[0040] Techniques such as those disclosed in [5] and [6] may be used in combination to enable both the localisation of predetermined non-speech sounds and the determination of an individual speaker (e.g. from multiple potential speakers in any given video frame) for a given instance of dialogue.

[0041] The output of step 304 is an indication of which segmented object(s) in each frame are associated with a separated audio component. It will be appreciated that, depending on the separated audio component(s) and segmented object(s) for each frame, the output of step 304 may indicate no segmented objects (e.g. if the input data represents a silent part of a scene), all segmented objects (e.g. if each segmented object represents a sound source) or only a portion of the segmented objects (e.g. if only some of the segmented objects represent a sound source). Thus, more generally, each separated audio component may be associated with any number of segmented objects in the frame, the number ranging from no segmented object to all segmented objects (e.g. if the separated audio component is any form of speech and all segmented objects in the frame are people talking).

[0042] At step 305, each video frame is modified to visually indicate each segmented object associated with a separated sound component, thereby generating the frames of the output video data 204. Thus, in the example of Fig. 2, the input video frame 206A is modified to generate the output video frame 206B in which the visual appearance of lightning cloud 205 is adjusted to include the outline 207. The outline 207 indicates that a separated sound component (i.e. that representing thunder) is associated with the lightning cloud 205.

[0043] In an example, the change to the visual appearance of an object associated with a separated sound component may be different depending on the nature of that sound component. Thus, for example, an object segment corresponding to a person speaking may be modified to make that person appear subtly larger, in higher resolution and / or in sharper focus compared to other detected object segments. On the other hand, an object segment corresponding to a more dramatic sound such as thunder (as in the case of lightning cloud 205) may be modified in a less subtle way (e.g. through outline 207 which may or may not be animated). In this case, the sound component associated with an object may be classified with one of a plurality of predetermined classifications (e.g. “speech”, “thunder”, “crowd”, “ocean”, etc.). Each of the plurality of predetermined classifications is then associated with a different respective visual modification which is applied to the object determined to be the source of the sound. Any suitable classification technique may be used. For example, the classification technique(s) disclosed in [7] may be used.

[0044] Fig. 4 shows a second example technique. The technique (including the functions of encoders 403, 404 and 405, concatenator 406, multi-modal machine learning (ML) model 407 and decoder 408) is executed by processor 101 , for example.

[0045] The technique of Fig. 4 uses a trained multi-modal ML model 407. In an example, the ML model may be trained using image pairs generated by the technique of Fig. 3 together with temporally associated samples of audio data. In this case, each image pair comprises an input image (e.g. image frame 206A) and a target output image (e.g. image frame 206B). The image frames and associated audio data samples may be taken from existing AV content (with permission of the AV content owner) such as films and / or television shows. In an example, ~10 million such image pairs and associated audio data samples may be used to train the ML model.

[0046] Using the technique of Fig. 3 to generate the training data for training the ML model 407 for the technique of Fig. 4 allows the generation of the training data to be largely automated, thereby reducing time and labour required in generating the training data. Furthermore, once trained, the ML model 407 of Fig. 4 allows the generation of output video data 204 in a computationally timely and efficient manner (thereby enabling effective application of the present technology to real time AV content streaming and / or playback).

[0047] As shown in Fig 4, the input video data 201 is provided to a first encoder 403. The encoder 403 generates a video token (input video token) representative of each video frame of the video data. For example, the video token may be a vector. Any suitable known way of tokenising a video frame so it is represented as a vector may be used. For example, the encoder 403 may be a vector quantized autoencoder such as that described in [8],

[0048] The input audio data 202 is provided to a second encoder 404. The encoder 404 generates an audio token (input audio token) representative of the audio sample associated with each video frame. For example, the audio token may be a vector. Any suitable known way of tokenising an audio sample so it is represented as a vector may be used. For example, the encoder 404 may use the wav2vec 2.0 technique disclosed in [9],

[0049] Optionally, status data 401 and / or coordinate data 402 may also be used with the example of Fig. 4. The status data 401 indicates whether or not sounds are to be visualised or not.

[0050] For example, the status data may simply indicate “on” to enable sound visualisation (in which case, for example, input image frame 206A is converted to image frame 206B) or “off’ to disable sound visualisation (in which case, for example, input image frame 206A is output without modification).

[0051] The status data may also indicate particular sound components and / or classifications which are to be turned “on” or “off”. For example, a user may wish for speech to be visualised (e.g. by through visual indication of the person speaking) but not for background sounds such as “thunder” to be visualised. In this case, the status data my indicate “thunder off’ or “mute thunder” so that sounds classified as “thunder” are not visualised in the output video data 204. In this case, for example, the image frame 206B would not include the outline 207 of the lightning cloud 205 and thus the appearance of the lightning cloud 205 in output image frame 206B will remain unchanged from that of input image frame 206A.

[0052] More general status data, such as “background sounds off” or “mute background sounds” may also be use. Muting background sounds in this case will mean no sounds are visualised in the output video data 204 except speech. Similarly, status data such as “speech off” or “mute speech” will mean all sounds except speech are visualised in the output video data 204. This may be useful if, for example, a user is happy to use subtitles to follow the dialogue of a video but wishes for non-speech sounds to be visualised (rather than or in addition to a description of the sound via subtitles).

[0053] The coordinate data 402 indicates video frame pixel coordinates to which the status data 401 is to be applied. For example, a user may indicate video frame pixel coordinates corresponding to the location of an object in the output video data (e.g. lightning cloud 205) together with status data “on” or “off’ to indicate whether or not sounds associated with that object should be visualised.

[0054] The status data 401 may be provided as textual data by a viewer (e.g. by a physical or on-screen keyboard). The coordinate data 402 may be provided by a viewer selecting a location on a display showing the output video data 204 using a cursor (controlled by a mouse or remote control, for example) or touch screen. The coordinate data 402 is recorded as textual data indicative of the horizontal and vertical pixel coordinates of the location of the display selected by the viewer.

[0055] The status data 401 and / or coordinate data 402 is user input data representing user input. It is provided via user interface 105, for example. Alternatively, if the data processing apparatus 100 transmits the output video frames to another data processing apparatus (display apparatus, not shown) for display (e.g. via communication interface 105), the status data 401 and / or coordinate data 402 is provided via a user interface of this other apparatus (which forms a system with the data processing apparatus) and transmitted to the data processing apparatus 100 (e.g. via communication interface 105). In either case, the user interface (either user interface 105 or the user interface of the other data processing apparatus) comprises (or is configured to communicate with) the display on which the output video data is displayed and one or more appropriate input apparatuses (e.g. mouse, remote control and / or touch screen circuitry) allowing a user to provide the status data 401 and / or coordinate data 402. The display may be comprised as part of a television, monitor, smartphone, tablet computer or virtual reality (VR) and / or augmented reality (AR) headset, for example.

[0056] The status data 401 and / or coordinate data 402 are provided to a third encoder 405. The encoder 405 generates a text token representative of the textual data of the status data 401 and / or coordinate data 402 (the encoder 405 first concatenating the textual data of the status data 401 and coordinate data 402 if both are present). For example, the text token may be a vector. Any suitable known way of tokenising text so it is represented as a vector may be used. For example, the encoder 405 may use the “Wordpiece” technique disclosed in

[0010] ,

[0057] When status data 401 and / or coordinate data 402 is to be used, suitable status and / or coordinate training data is used when training the ML model 407. Thus, for example, for each image pair and associated audio data sample of the training data, status and / or coordinate data is also provided as training data. For each image pair, the modification of the input image to generate the target output image corresponds to the status and / or coordinate data.

[0058] For example, looking again at the example of Fig. 2, a first instance of training data may comprise the frame 206A as the input image, the frame 206B (with outline 207 applied to lightning cloud 205) as the target output image, status data “thunder on” and unspecified coordinate data. A second instance of training data may comprise the frame 206A as the input image, the same frame 206A (with no outline 207 applied to lightning cloud 205) as the target output image, status data “thunder off’ and unspecified coordinate data. A third instance of training data may comprise the frame 206A as the input image, the frame 206B (with outline 207 applied to lightning cloud 205) as the target output image, status data “on” and coordinate data at which the lightning cloud 205 is located. A fourth instance of training data may comprise the frame 206A as the input image, the same frame 206A (with no outline 207 applied to lightning cloud 205) as the target output image, status data “off’ and coordinate data at which the lightning cloud 205 is located.

[0059] It will be appreciated that the use (and tokenisation) of textual data for the status data 401 and the large number of potential adjustments that can be made to an input image to generate a target output image in the training data mean the use of the technique of Fig. 4 (in particular, the use of multi-modal ML model 407) provides improved flexibility and customisation in determining the appearance of sound sources in output video data.

[0060] For instance, as well as the status data simply being able to indicate whether a particular sound source or type of sound is “on” or “off’, the desired type and / or extent of the change in appearance of sound sources in the output video data may be indicated. For example, status data 401 “show thunder in a subtle way” may result in a less extreme modification to the lightning cloud 205 (e.g. by slightly increasing the size, brightness, resolution and / or relative focus of the lightning cloud). On the other hand, status data 401 “show thunder in a more obvious way” may result in a more extreme modification to the lightning cloud 205 (e.g. by inclusion of a thick outline 207 in a bold colour).

[0061] The textual data of the status data 401 and / or coordinate data 402 may thus indicate any desired characteristic of the appearance of sound sources (and thus of the modified video data), including activation or deactivation of modification of one or more regions of a video frame of the input video data (e.g. indicating whether visual representation of the sound associated with a particular sound source is “on” or “off’), a desired appearance of one or more modified regions of a video frame of the input video data and / or a location (e.g. as indicated by coordinate data 402) of one or more regions of a video frame of the input video data (to which modification and / or a given type of modification is to be applied).

[0062] Once the video, audio and text tokens have been generated by the encoders 403, 404 and 405, they are concatenated by the concatenator 406 to form a single, combined, token (e.g. a vector) which is input to the ML model 407.

[0063] The ML model 407 is a multi-modal ML model. A multi-modal ML model is configured to take, as an input, a plurality of different input modalities and / or output a plurality of different input modalities. Images, audio and text are examples of different modalities. An ML model which is configured to take, as an input, data representative of a plurality of such modalities and / or output data representative a plurality of such modalities is thus a multi-modal ML model.

[0064] The ML model 407 takes, as input data, the combined token generated by concatenator 406 representing a video frame of the input video data, the audio data sample temporally associated with that video frame and, optionally, the textual status data 401 and / or coordinate data 402. The ML model 407 outputs, as output data, an output video token (e.g. a vector) representing a video frame of output video data. The output video frame (e.g. frame 206B) is a modified version of the input video frame (e.g. frame 206A). The modification depends on the audio data sample associated with the input video frame and, if present, the status data 401 and / or coordinate data 402. Any suitable known multi-modal ML model 407 may be used. For example, the ML model 407 may be a multi-modal artificial neural network such as those described in

[0011] or

[0012] ,

[0065] The output video token (representative of the output video frame) is input to decoder 408 which detokenises the output video token to generate the output video frame. Any suitable known detokenisation technique may be used. For example, a video diffusion decoder such as that described in

[0013] may be used.

[0066] The output video frames thus form the output video data 204. This is output with the original audio data 202. Users who are not hard of hearing can thus hear the original audio while viewing the output video. On the other hand, users are who are hard of hearing and thus who cannot hear the original audio (or, at least, cannot hear it sufficiently well) are still able to experience the audio information in an enhanced and immersive way due to the modifications made to the output video. It is noted that status data 401 and / or coordinate data 402 may also be used with the technique of Fig. 3. For example, if coordinate data 402 is provided, the video frame segment located at the relevant pixel coordinates is selected. Visual modification of that segment can then be turned “on” or “off’ in accordance with the status data 401. For example, if a viewer indicates the coordinates of a pixel of the segment defining lightning cloud 205 and provides status data “on”, the outline 207 will appear in the output image frame 206B. However, if the status data “off’ is provided, the outline 207 will not appear in the output image frame 206B (and thus the appearance of the lightning cloud 205 in the output image frame 206B will remain the same as that of the input image frame 206A). In this case, the status data may not be provided as textual data but, rather, as an “on / off’ toggle button (not shown) or the like (e.g. a virtual button displayable with the output video data 204 together with playback controls such as “Play”, “Pause”, etc.).

[0067] Although the present technology provides a way of indicating audio information to users who are hard of hearing in an enhanced way compared to conventional subtitles, subtitles may nonetheless be provided with the modified output video data generated according to the present technology.

[0068] Data representing the provided subtitles may be present (e.g. as metadata) with the original video data 201 and / or audio data 202 (e.g. in the language of the audio data and / or one or more other languages). Alternatively, if no data representing subtitles is present, the audio data 202 may be processed using a suitable known speech-to-text technique (e.g. the Wave2Seq technique disclosed in

[0014] ) to generate the subtitle data (e.g. in real time as the video and audio data are streamed or played back).

[0069] The subtitle text may be placed in a predetermined region of the output video frames (as per conventional subtitles). Alternatively, the subtitle text may be placed at a variable location in output video frames depending on the location of the source of the sound indicated by the subtitles in each frame. In this case, a segment of the current image frame associated with a sound component represented by the subtitles (e.g. the segment representing the current person speaking) may be determined (using the image segmentation, sound separation and / or sound localisation techniques described above, for example). The subtitles may then be placed in the vicinity (e.g. overlaid on, underneath or next to) of the determined segment.

[0070] Fig. 5 shows an example method. The method is executed by processor 101 , for example.

[0071] The method starts at step 501 .

[0072] At step 502, input video and audio data are received.

[0073] At step 503, a region of a video frame of the input video data is modified to generate a modified video frame. The region is modified based on an audio sample of the input audio data. The audio sample is temporally associated with the video frame.

[0074] At step 504, modified video data representing the modified video frame is output. The method ends at step 505.

[0075] Example(s) of the present technology are defined by the following numbered clauses:

[0076] 1 . A data processing apparatus comprising circuitry configured to: receive input video and audio data; modify a region of a video frame of the input video data to generate a modified video frame, the region being modified based on an audio sample of the input audio data temporally associated with the video frame; and output modified video data representing the modified video frame.

[0077] 2. A data processing apparatus according to clause 1 , wherein the circuitry is configured to: segment the video frame into one or more segments; separate the audio sample into one or more audio components; perform audio localisation to associate a segment of the one or more segments with an audio component of the one or more audio components; and modify, as the region of the video frame, the segment associated with the audio component.

[0078] 3. A data processing apparatus according to clause 1 , wherein the circuitry is configured to generate the modified video data from the input video and audio data using a multi-modal machine learning model.

[0079] 4. A data processing apparatus according to clause 3, wherein the circuitry is configured to: tokenise the video frame of the input video data to generate an input video token; tokenise the audio sample of the input audio data to generate an input audio token; provide the input video and audio tokens as an input to the multi-modal machine learning model; obtain, as an output of the multi-modal machine learning model, an output video token representing the modified video frame; and detokenise the output video token to generate the modified video frame.

[0080] 5. A data processing apparatus according to clause 4, wherein training the multi-modal machine learning model comprises using a plurality of instances of training data, each instance of training data comprising: an input video frame; an audio sample temporally associated with the input video frame; and a target output video frame, the target output video frame being generated by modifying a region of the input video frame based on the audio sample.

[0081] 6. A data processing apparatus according to clause 5, wherein the target output video frame is generated by: segmenting the input video frame into one or more segments; separating the audio sample into one or more audio components; performing audio localisation to associate a segment of the one or more segments with an audio component of the one or more audio components; and modify, as the region of the input video frame, the segment associated with the audio component.

[0082] 7. A data processing apparatus according to any one of clauses 3 to 6, wherein the circuitry is configured to: receive textual data indicating a desired characteristic of the modified video data; and generate the modified video data from the input video, input audio data and the textual data using the multi-modal machine learning model.

[0083] 8. A data processing apparatus according to clause 7, wherein the circuitry is configured to: tokenise the video frame of the input video data to generate an input video token; tokenise the audio sample of the input audio data to generate an input audio token; tokenise the textual data to generate a text token; provide the input video token, input audio token and text token as an input to the multimodal machine learning model; obtain, as an output of the multi-modal machine learning model, an output video token representing the modified video frame; and detokenise the output video token to generate the modified video frame.

[0084] 9. A data processing apparatus according to clause 8, wherein training the multi-modal machine learning model comprises using a plurality of instances of training data, each instance of training data comprising: an input video frame; an audio sample temporally associated with the input video frame; textual data indicating a desired characteristic of the modified video data; and a target output video frame, the target output video frame being generated by modifying a region of the input video frame based on the audio sample and according to the desired characteristic indicated by the textual data.

[0085] 10. A data processing apparatus according to any preceding clause, wherein the circuitry is configured to: receive user input data representing user input; determine, based on the user input data, one or more of: activation or deactivation of modification of one or more regions of a video frame of the input video data; a desired appearance of one or more modified regions of a video frame of the input video data; and a location of one or more regions of a video frame of the input video data.

[0086] 11. A system comprising: a data processing apparatus according to any preceding clause; and a display apparatus configured to output the modified video frame.

[0087] 12. A computer-implemented data processing method comprising: receiving input video and audio data; modifying a region of a video frame of the input video data to generate a modified video frame, the region being modified based on an audio sample of the input audio data temporally associated with the video frame; and outputting modified video data representing the modified video frame.

[0088] 13. A program for controlling a computer to perform a method according to clause 12.

[0089] 14. A computer-readable storage medium storing a program according to clause 13.

[0090] Numerous modifications and variations of the present disclosure are possible in light of the above teachings. It is therefore to be understood that, within the scope of the claims, the disclosure may be practiced otherwise than as specifically described herein.

[0091] In so far as embodiments of the disclosure have been described as being implemented, at least in part, by one or more software-controlled information processing apparatuses, it will be appreciated that a machine-readable medium (in particular, a non-transitory machine-readable medium) carrying such software, such as an optical disk, a magnetic disk, semiconductor memory or the like, is also considered to represent an embodiment of the present disclosure. In particular, the present disclosure should be understood to include a non-transitory storage medium comprising code components which cause a computer to perform any of the disclosed method(s).

[0092] It will be appreciated that the above description for clarity has described embodiments with reference to different functional units, circuitry and / or processors. However, it will be apparent that any suitable distribution of functionality between different functional units, circuitry and / or processors may be used without detracting from the embodiments.

[0093] Described embodiments may be implemented in any suitable form including hardware, software, firmware or any combination of these. Described embodiments may optionally be implemented at least partly as computer software running on one or more computer processors (e.g. data processors and / or digital signal processors). The elements and components of any embodiment may be physically, functionally and logically implemented in any suitable way. Indeed, the functionality may be implemented in a single unit, in a plurality of units or as part of other functional units. As such, the disclosed embodiments may be implemented in a single unit or may be physically and functionally distributed between different units, circuitry and / or processors.

[0094] Although the present disclosure has been described in connection with some embodiments, it is not intended to be limited to these embodiments. Additionally, although a feature may appear to be described in connection with particular embodiments, one skilled in the art would recognize that various features of the described embodiments may be combined in any manner suitable to implement the present disclosure.

[0095] REFERENCES

[0096]

[0001] Matsunaga et al, “Fine-grained Image Editing by Pixel-wise Guidance Using Diffusion

[0097] Models”, 2023, https: / / arxiv.org / pdf / 2212.02024.pdf

[0098] [2] Levin et al, “Differential Diffusion: Giving Each Pixel Its Strength”, 2024, https : / .00950.

[0099] [3] Oquab et al, “DINOv2: Learning Robust Visual Features without Supervision”, 2024, https: / / arxiv.org / pdf / 2304.07193.pdf

[0100] [4] Wisdom et al, “What's All the FUSS About Free Universal Sound Separation Data?”, 2020, https: / / arxiv.org / abs / 201 1 .00803

[0101] [5] Wu et al, “Time Domain Audio Visual Speech Separation”, 2019, https: / / arxiv.Org / abs / 1904.03760

[0102] [6] Zhao et al, “The Sound of Pixels”, 2018, https: / / arxiv.org / abs / 1804.03160

[0103] [7] Srivastava et al, “OmniVec: Learning robust representations with cross modal sharing”, 2023, https: / / arxiv.org / pdf / 231 1 ,05709v1

[0104] [8] van den Oord et al, “Neural Discrete Representation Learning”, 2018, https: / / arxiv.0rg / pdf / 171 1 .00937

[0105] [9] Baevski et al, “wav2vec 2.0: A Framework for Self-Supervised

[0106] Learning of Speech Representations”, 2020, https: / / arxiv.Org / pdf / 2006.1 1477

[0107]

[0010] Wu et al, “Google’s Neural Machine Translation System: Bridging the Gap between Human and Machine Translation”, 2016, https: / / arxiv.org / pdf / 1609.08144v2

[0108]

[0011] Mizrahi et al, “4M: Massively Multimodal Masked Modeling”, 2023,

[0109] : / / arxiv.org / pdf / 2312.06647

[0110]

[0012] Girdhar et al, “IMAGEBIND: One Embedding Space To Bind Them All”, 2023, http s . / / a rxi v.0 rg / pdf / 2305.05665

[0111]

[0013] Blattmann et al, “Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large

[0112] Datasets”, 2023, https: / / arxiv.org / pdf / 231 1 .15127

[0113]

[0014] Wu et al, “Wav2Seq: Pre-training Speech-to-Text Encoder-Decoder Models Using

[0114] Pseudo Languages”, 2022, https: / / arxiv.org / abs / 2205.01 Q86

Claims

CLAIMS1 . A data processing apparatus comprising circuitry configured to: receive input video and audio data; modify a region of a video frame of the input video data to generate a modified video frame, the region being modified based on an audio sample of the input audio data temporally associated with the video frame; and output modified video data representing the modified video frame.

2. A data processing apparatus according to claim 1 , wherein the circuitry is configured to: segment the video frame into one or more segments; separate the audio sample into one or more audio components; perform audio localisation to associate a segment of the one or more segments with an audio component of the one or more audio components; and modify, as the region of the video frame, the segment associated with the audio component.

3. A data processing apparatus according to claim 1 , wherein the circuitry is configured to generate the modified video data from the input video and audio data using a multi-modal machine learning model.

4. A data processing apparatus according to claim 3, wherein the circuitry is configured to: tokenise the video frame of the input video data to generate an input video token; tokenise the audio sample of the input audio data to generate an input audio token; provide the input video and audio tokens as an input to the multi-modal machine learning model; obtain, as an output of the multi-modal machine learning model, an output video token representing the modified video frame; and detokenise the output video token to generate the modified video frame.

5. A data processing apparatus according to claim 4, wherein training the multi-modal machine learning model comprises using a plurality of instances of training data, each instance of training data comprising: an input video frame; an audio sample temporally associated with the input video frame; and a target output video frame, the target output video frame being generated by modifying a region of the input video frame based on the audio sample.

6. A data processing apparatus according to claim 5, wherein the target output video frame is generated by: segmenting the input video frame into one or more segments; separating the audio sample into one or more audio components; performing audio localisation to associate a segment of the one or more segments with an audio component of the one or more audio components; andmodify, as the region of the input video frame, the segment associated with the audio component.

7. A data processing apparatus according to claim 3, wherein the circuitry is configured to: receive textual data indicating a desired characteristic of the modified video data; and generate the modified video data from the input video, input audio data and the textual data using the multi-modal machine learning model.

8. A data processing apparatus according to claim 7, wherein the circuitry is configured to: tokenise the video frame of the input video data to generate an input video token; tokenise the audio sample of the input audio data to generate an input audio token; tokenise the textual data to generate a text token; provide the input video token, input audio token and text token as an input to the multimodal machine learning model; obtain, as an output of the multi-modal machine learning model, an output video token representing the modified video frame; and detokenise the output video token to generate the modified video frame.

9. A data processing apparatus according to claim 8, wherein training the multi-modal machine learning model comprises using a plurality of instances of training data, each instance of training data comprising: an input video frame; an audio sample temporally associated with the input video frame; textual data indicating a desired characteristic of the modified video data; and a target output video frame, the target output video frame being generated by modifying a region of the input video frame based on the audio sample and according to the desired characteristic indicated by the textual data.

10. A data processing apparatus according to claim 1 , wherein the circuitry is configured to: receive user input data representing user input; determine, based on the user input data, one or more of: activation or deactivation of modification of one or more regions of a video frame of the input video data; a desired appearance of one or more modified regions of a video frame of the input video data; and a location of one or more regions of a video frame of the input video data.

11. A system comprising: a data processing apparatus according to claim 1 ; and a display apparatus configured to output the modified video frame.

12. A computer-implemented data processing method comprising: receiving input video and audio data;modifying a region of a video frame of the input video data to generate a modified video frame, the region being modified based on an audio sample of the input audio data temporally associated with the video frame; and outputting modified video data representing the modified video frame.

13. A program for controlling a computer to perform a method according to claim 12.

14. A computer-readable storage medium storing a program according to claim 13.

Citation Information

Patent Citations

  • Generation of closed captions based on various visual and non-visual elements in content

    US20230362451A1