Information processing device and information processing method

WO2025186268A8PCT designated stage Publication Date: 2025-10-02SONY GROUP CORP +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2025/055873
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-03-05
Filing Date
2025-03-04
Publication Date
2025-10-02

AI Technical Summary

Technical Problem

Existing audio quality of old movies and videos lags behind current production levels, despite advancements in visual quality enhancement techniques.

Method used

An information processing device and method utilizing a neural network to separate sound sources, recognize objects in video frames, and add spatial information to enhance audio quality, incorporating noise removal, voice and music filtering, and dynamic range improvement.

Benefits of technology

Enhances audio quality to match modern standards, providing immersive audio experiences comparable to 360-degree microphones by accurately locating and separating sound sources, even when they are out of the camera's field of view.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2025055873_02102025_PF_FP_ABST
    Figure EP2025055873_02102025_PF_FP_ABST
Patent Text Reader

Abstract

An information processing device, wherein the information processing device includes circuitry configured to: obtain video data representing a video, wherein the video data include image data representing a plurality of image frames and audio data representing a single-channel or two- channel audio recording associated with the plurality of image frames; inputting the video data into a neural network, wherein the neural network is configured to separate sound sources in the audio data into one or more active sound sources and one background sound source, wherein the neural network is further configured to recognize objects in each of the plurality of image frames and to output a type of each recognized object; and wherein the circuitry is further configured to add, for each active sound source, spatial information indicating a location of the active sound source based on a match between a type of the active sound source and the type of the recognized object.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] INFORMATION PROCESSING DEVICE AND INFORMATION

[0002] PROCESSING METHOD

[0003] TECHNICAL FIELD

[0004] The present disclosure generally pertains to an information processing device and an information processing method.

[0005] TECHNICAL BACKGROUND

[0006] Generally, video capturing is known in which a camera and one or more microphones are used to capture a sequence of images of a scene and to record sound associated with the sequence of images, for example, to produce a movie.

[0007] A visual quality improvement for old movies has been successfully implemented and used in many cases, for example, current artificial intelligence technologies have been used for quality improvement. The resolution, colorization, clarity and other visual aspects are improved to revive the old videos and bring them closer to the capabilities of current standards and capabilities of nowadays television and cinema screens. It may enable the viewer to enjoy classic movies and videos in full 4k or 8k video quality.

[0008] However, while a visual quality of old movies has been improved, an audio quality of old movies may still be behind the current production levels.

[0009] Although there exist techniques for video quality enhancement, it is generally desirable to improve the existing techniques.

[0010] SUMMARY

[0011] According to a first aspect, the disclosure provides an information processing device comprising circuitry configured to: obtain video data representing a video, wherein the video data include image data representing a plurality of image frames and audio data representing a single-channel or two- channel audio recording associated with the plurality of image frames; inputting the video data into a neural network, wherein the neural network is configured to separate sound sources in the audio data into one or more active sound sources and one background sound source, wherein the neural network is further configured to recognize objects in each of the plurality of image frames and to output a type of each recognized object; and wherein the circuitry is further configured to add, for each active sound source, spatial information indicating a location of the active sound source based on a match between a type of the active sound source and the type of the recognized object.

[0012] According to a second aspect, the disclosure provides an information processing method comprising: obtaining video data representing a video, wherein the video data include image data representing a plurality of image frames and audio data representing a single-channel or two- channel audio recording associated with the plurality of image frames; inputting the video data into a neural network, wherein the neural network is configured to separate sound sources in the audio data into one or more active sound sources and one background sound source, wherein the neural network is further configured to recognize objects in each of the plurality of image frames and to output a type of each recognized object; and adding, for each active sound source, spatial information indicating a location of the active sound source based on a match between a type of the active sound source and the type of the recognized object.

[0013] Further aspects are set forth in the dependent claims, the drawings and the following description.

[0014] BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Embodiments are explained by way of example with respect to the accompanying drawings, in which:

[0016] Fig. 1 schematically illustrates in a block diagram an embodiment of an information processing device;

[0017] Fig. 2 schematically illustrates in a flow diagram an embodiment of an information processing method; and

[0018] Fig. 3 schematically illustrates in a block diagram an embodiment of a multi-purpose computer.

[0019] DETAILED DESCRIPTION OF EMBODIMENTS

[0020] Before a detailed description of the embodiments under reference of Fig. 1 is given, general explanations are made.

[0021] As mentioned in the outset, a visual quality improvement for old movies has been successfully implemented and used in many cases, for example, current artificial intelligence technologies have been used for quality improvement. However, as further mentioned in the outset, while a visual quality of old movies has been improved, an audio quality of old movies may still be behind the current production levels.

[0022] It has thus been recognized that an enhancement of audio for old movies and other old videos may further improve video quality enhancement.

[0023] It has been recognized that artificial intelligence, as well as contextual information from the image sequence of the video, may be used for audio quality enhancement for video quality enhancement.

[0024] It has been recognized that audio quality enhancement may include separating sound sources and adding spatial information for the separated sound sources, since many old movies and videos may have limited audio channels, such as stereo in which a two-channel audio recording is used.

[0025] It has been recognized that, by utilizing contextual information from the image sequence, sound sources may be identified in relation to the camera movement and the sound may be enhanced, based on the obtained spatial information, afterwards to meet, for example, full 360 Reality Audio or Dolby Atmos standards of today.

[0026] It has further been recognized that audio quality enhancement may include audio clarity improvement by noise removal, voice and music filtering, and dynamic range improvement, for example, for each audio channel before or after separating the sound sources.

[0027] An information processing device, wherein the information processing device includes circuitry configured to: obtain video data representing a video, wherein the video data include image data representing a plurality of image frames and audio data representing a single-channel or two- channel audio recording associated with the plurality of image frames; inputting the video data into a neural network, wherein the neural network is configured to separate sound sources in the audio data into one or more active sound sources and one background sound source, wherein the neural network is further configured to recognize objects in each of the plurality of image frames and to output a type of each recognized object; and wherein the circuitry is further configured to add, for each active sound source, spatial information indicating a location of the active sound source based on a match between a type of the active sound source and the type of the recognized object.

[0028] The information processing device may be a mobile electronic device (e.g., a smartphone, a tablet computer, a laptop or the like), a computer, a server or the like. The circuitry may include one or more processors. A processor may be or may include an application processor, a central processing unit (“CPU”), a graphical processing unit (“GPU”), a digital signal processor (“DSP”), a field-programmable gate array (“FPGA”), an application specific integrated circuit (“ASIC”) etc.

[0029] The circuitry may include one or more memory components. A memory component may be or may include volatile and non-volatile memory such as static random-access memory (“SRAM”), dynamic RAM (“DRAM”), non-volatile RAM (“NVRAM”), read-only memory (“ROM”), programmable ROM (“PROM”), electrically PROM (“EPROM”), electrically erasable PROM (“EEPROM”), flash memory (e.g., NOR flash or NAND flash) etc. A memory component may be or may include one or more registers, caches, main memories, hard disk drives, solid-state drives etc.

[0030] The circuitry may include one or more data bus interfaces.

[0031] The circuitry may include one or more communication interfaces, wherein each communication interface may be configured to communicate, for example, via a local area network (LAN), a wireless local area network (WLAN), a mobile telecommunications system (GSM, UMTS, LTE, NR etc.), Bluetooth, etc.

[0032] The functionality of the circuitry may be implemented by typical electronic components configured to achieve the functionality as described herein. The functionality of the circuitry may be implemented in parts by typical electronic components and in parts by software configured to achieve the functionality as described herein. The functionality of the circuitry may be implemented by software configured to achieve the functionality as described herein.

[0033] The circuitry obtains video data representing a video, wherein the video data include image data representing a plurality of image frames and audio data representing a single-channel or two- channel audio recording associated with the plurality of image frames.

[0034] The image data thus represent a time sequence of consecutive image frames captured by a camera, wherein each image frame has time information (e g., a time stamp) for identifying its temporal position within the time sequence of consecutive image frames.

[0035] The audio data thus represent a time sequence of sound samples of an audio signal recorded with a single microphone in the case of a single-channel audio recording, wherein each sound sample has time information (e.g., a time stamp) for identifying its temporal position within the time sequence of sound samples. The audio data thus represent two separate time sequences of sound samples of two audio signals recorded with two different microphones located at two different positions in the case of a two- channel audio recording, wherein each sound sample has time information (e.g., a time stamp) for identifying its temporal position within the respective time sequence of sound samples.

[0036] The audio data are associated with the image data in the sense of a time correspondence, as generally known for videos, such that each segment of sound samples is associated with one or more image frames based on the respective time information.

[0037] The circuitry inputs the video data into a neural network.

[0038] The neural network may include or may be based on one or more fully-connected layers, one or more convolutional layers, one or more recurrent layers or the like.

[0039] As mentioned further above, it has been recognized that the audio quality enhancement may include audio clarity improvement by noise removal, voice and music filtering, and dynamic range improvement, for example, for each audio channel before or after separating the sound sources.

[0040] Given existing information on the recording quality and imperfections of old microphones and soundtracks and a possible degradation of video or audio sources over time, the circuitry may reduce the effects of noise, limited dynamic range and other imperfections of audio in the old movies. The enhancement may be comparable to current visual quality enhancement methods to enhance colors and superresolution of old videos. A combination, for example, of standard filtering and artificial intelligence-based methods may be utilized to bring old audio recordings closer to the level of current standards.

[0041] Hence, in some embodiments, the neural network is further configured to perform at least one of noise removal, voice filtering, music filtering and dynamic range enhancement.

[0042] The neural network may include a first sub-network for performing at least one of noise removal, voice filtering, music filtering and dynamic range enhancement, a second sub-network for processing the audio data and a third sub-network for processing the image data.

[0043] As mentioned further above, it has been recognized that audio quality enhancement may include separating sound sources.

[0044] Hence, the neural network is configured to separate sound sources in the audio data into one or more active sound sources and one background sound source. An active sound source corresponds to any person, animal or technical device or apparatus (e.g., an instrument, a loudspeaker, a vehicle or a machine or the like) that actively emits or causes sound, wherein the sound is represented in the audio data, wherein the sound may be voice or music or a sound that is characteristic for an animal or a sound that is characteristic for a technical device or apparatus.

[0045] The number of the one or more active sound sources the neural network is configured to separate is predetermined and may, for example, be three, four, five, etc.

[0046] The neural network is configured to generate, for each active sound source, active sound source audio data representing a separate single-channel or two-channel time sequence of sound samples and to output a type of the respective active sound source and the generated active sound source audio data.

[0047] The neural network is configured to generate background sound source audio data representing a separate single-channel or two-channel time sequence of sound samples for the one background sound source, wherein background sound source audio data corresponds to the residuum between the audio data and the generated active sound source audio data.

[0048] As mentioned further above, it has been recognized that audio quality enhancement may include adding spatial information for the separated active sound sources.

[0049] Contextual and localization information of the videos may be used to enhance the audio in terms of spatial location and sound separation of each object, for example, by tracking the changing camera position in relation to the objects in the scene producing the sound.

[0050] For illustration, a scene may include multiple people talking where sound was recorded in stereo. In the scene, where camera is facing people, the localization of sound for the viewer can be quite clear, however, if the camera starts rotating to the side, some of the actors speaking (or any other objects producing sound) may end up being out of frame on the side, back, above or below the camera and outside the field-of-view. However, their sound may still be present in the movie, but in the old recordings their positional audio information is not represented precisely.

[0051] By utilizing the analysis and tracking of the visual information and estimating the position of the camera in relation to the objects producing sound, the position of audio sources even out of field- of-view may be estimated with position changes and temporal information. In this way, some audio sources may be enhanced afterwards to provide a spatial audio experience which may, in some cases, match the current standards of 360-degree microphones used nowadays and providing the viewer with immersive audio experience. It has been recognized that visual feature tracking, camera pose change estimation, sound source separation and artificial intelligence technologies may be utilized for such enhancements.

[0052] It has further been recognized that such a technology may be applied for enhancing videos, for example, where sound sources are visible and identifiable in the video for a given scene.

[0053] Hence, the neural network is further configured to recognize objects in each of the plurality of image frames and to output a type of each recognized object, and the circuitry is further configured to add, for each active sound source, spatial information indicating a location of the active sound source based on a match between a type of the active sound source and the type of the recognized object.

[0054] The neural network is configured to recognize objects in each of the plurality of image frames and to identify a location of each recognized object in each of the plurality of image frames, wherein such neural networks are generally known.

[0055] The neural network is configured to output a type and the location of each recognized object for each of the plurality of image frames.

[0056] The circuitry is configured to determine a match between each recognized object with the active sound sources based on the respective types, which may be further based on the respective time information.

[0057] The circuitry is configured to add, for each active sound source, the location of the recognized object that has a type that matches the type of the active sound source.

[0058] For example, a recognized object may be a person and an active sound source may be a person such that a location of the person is added to the active sound source audio data corresponding to the time information.

[0059] Hence, the circuitry is configured to add, for each active sound source, spatial information indicating a location of the active sound source based on a match between a type of the active sound source and the type of the recognized obj ect.

[0060] In some embodiments, the circuitry is further configured to track the location of the active sound source.

[0061] The location of the active sound source, as identified by the matching, is thus tracked in time by the circuitry in some embodiments. The circuitry may correlate a location of the active sound source at a first time point with a location of the active sound source at a second time point later in time for tracking the location of the active sound source.

[0062] The circuitry may correlate the locations for detecting occlusions of the active sound source between the first time point and the second time point, for example, a location of the active sound source may be missing in one or more image frames between the first time point and the second time point, since the neural network could detect the matching object in the one or more image frames due to the occlusion.

[0063] The circuitry may thus be configured to interpolate, in response to the detection of an occlusion of the active sound source, the location of the active sound source between the first time point and the second time point based on the locations of the active sound source at the first time point and the second time point for tracking the location of the active sound source.

[0064] The tracking may further include the calculation of movement vectors of the active sound source.

[0065] However, the movement vector may, for example, either be due to a movement of the active sound source in the scene or due to a change in the pose of the camera.

[0066] It has thus been recognized that the pose change of the camera should be estimated in order to disentangle the movement of the active sound source and the movement (pose change) of the camera.

[0067] This may allow to estimate the location of the active sound source (or estimate the location of the active sound source more precisely) even in cases in which the active sound source is outside the field-of-view of the camera that captured the plurality of image frames.

[0068] Hence, in some embodiments, the neural network is further configured to estimate a change of a pose of a camera that captured the plurality of image frames, and wherein the circuitry is further configured to track the location of the active sound source in accordance with the estimated change of the pose of the camera.

[0069] The pose change estimation of the camera may be based on the relative movement of fixed objects in the scene, for example, the neural is configured to detect and recognize fixed objects in the plurality of image frames and to identify a location of each detected and recognized fixed object in the plurality of image frames for estimating the pose change of the camera.

[0070] Hence, in some embodiments, the circuitry is further configured to estimate the location of the active sound source outside the field-of-view of the camera. It has further been recognized that scene changes should be detected in order to reset the tracking.

[0071] Hence, in some embodiments, the neural network is further configured to detect a scene change.

[0072] The detection of a scene change may be based on a change of detected and recognized objects.

[0073] In some embodiments, the circuitry is further configured to reset the tracking when a scene change is detected.

[0074] An information processing method, wherein the information processing method includes: obtaining video data representing a video, wherein the video data include image data representing a plurality of image frames and audio data representing a single-channel or two- channel audio recording associated with the plurality of image frames; inputting the video data into a neural network, wherein the neural network is configured to separate sound sources in the audio data into one or more active sound sources and one background sound source, wherein the neural network is further configured to recognize objects in each of the plurality of image frames and to output a type of each recognized object; and adding, for each active sound source, spatial information indicating a location of the active sound source based on a match between a type of the active sound source and the type of the recognized object.

[0075] The information processing method may be performed by the information processing device as described herein.

[0076] The methods as described herein are also implemented in some embodiments as a computer program causing a computer and / or a processor to perform the method, when being carried out on the computer and / or processor. In some embodiments, also a non-transitory computer- readable recording medium is provided that stores therein a computer program product, which, when executed by a processor, such as the processor described above, causes the methods described herein to be performed.

[0077] Returning to Fig. 1, an embodiment of an information processing device 1 is discussed in the following under reference of Fig. 1, which schematically illustrates the embodiment in block diagram.

[0078] The information processing device 1 includes circuitry (not shown) which executes a neural network 2 and a matching and tracking process 3.

[0079] The neural network 2 includes a first sub-network 4, a second sub-network 5 and a third subnetwork 6. The circuitry inputs video data 7 into the neural network 2, wherein the video data 7 include audio data 8 and image data 9, wherein the audio data 8 are associated with the image data 9.

[0080] The image data 9 represent a plurality of image frames.

[0081] The audio data 8 represent a single-channel or two-channel audio recording associated with the plurality of image frames.

[0082] The first sub-network 4 processes the audio data 8 and is trained to perform at least one of noise removal, voice filtering, music filtering and dynamic range enhancement on input audio data 8.

[0083] The first sub-network 4 outputs the pre-processed audio data to the second sub-network 5.

[0084] The first sub-network 4 may, however, also be executed after the second sub-network 5.

[0085] The second sub-network 5 is trained to separate sound sources in the audio data into one or more active sound sources and one background sound source.

[0086] The second sub-network 5 outputs second audio data 11, wherein the second audio data 11 include: active sound source audio data representing a separate single-channel or two-channel time sequence of sound samples for each active sound source, a type of the respective active sound source and background sound source audio data corresponding to a residuum between the audio data and the generated active sound source audio data.

[0087] The third sub-network 6 processes the image data 9 and is trained to recognize objects in each of the plurality of image frames and outputs recognition result data 10, wherein the recognition result data 10 include: a type of each recognized object and a location of each recognized object in each of the plurality of image frames.

[0088] The matching and tracking process 3 determines a match between each recognized object with the active sound sources based on the respective types and the time information (e.g., both types match at the same time stamp or during the same time period).

[0089] Then, the matching and tracking process 3 adds, for each active sound source, spatial information indicating the location of the active sound source by adding the location of the recognized object that has a type that matches the type of the active sound source.

[0090] The matching and tracking process 3 outputs enhanced audio data 12, wherein the enhanced audio data 12 include: the active sound source audio data for each active sound source and the location of the active sound source. The matching and tracking process 3 further estimates a change of a pose of a camera that captured the plurality of image frames, as discussed herein, and tracks the location of the active sound source tracked in accordance with the estimated change of the pose of the camera.

[0091] The matching and tracking process 3 is thus able to disentangle the movement of the active sound source and the movement of the camera such that the matching and tracking process 3 is able to estimate the location of the active sound source outside the field-of-view of the camera more accurately.

[0092] Moreover, the matching and tracking process 3 detects scene changes and resets the tracking when a scene change is detected.

[0093] The first sub-network 4 may be trained as follows:

[0094] A training audio dataset may be used which includes a plurality of different single-channel (or two-channel) audio recordings, wherein each audio recording includes contributions from one or more different active sound sources and a background sound source.

[0095] Such a training audio dataset may be obtained by simulations in which various scenes are constructed. The scene constructions include placing one or more objects at different locations and attaching characteristic sounds to the various objects to simulate active sound sources. The sound wave propagation is then simulated taken into account the scene construction. Moreover, various background sounds may be added in the simulation.

[0096] The simulation may further include placing a microphone at a location and simulating the sound wave characteristics at the location of the microphone. Moreover, the microphone may be configured with different recording characteristics to simulate different types of microphones, in particular, in terms of quality of the recorded audio signal and the sampling of the recorded audio signal.

[0097] As such, the simulation may include, for each of a plurality of constructed scenes, generating a first single-channel (or two-channel) audio recording with low quality and a second singlechannel (or two-channel) audio recording with higher quality.

[0098] Hence, the first sub-network 4 uses then the low-quality audio recordings as input and the high- quality audio recordings as targets, wherein a difference between the outputs of the first subnetwork 4 and the targets is backpropagated to adjust the weights of the first sub-network 4 such that the differences may be decreased (minimized).

[0099] The second sub-network 5 may be trained as follows: The training may be based on the same training audio dataset as the training of the first subnetwork 4.

[0100] However, the training audio dataset further includes, for each of a plurality of constructed scenes, a single-channel (or two-channel) audio recording in which each sound attached to an object is recorded separately in order to allow a training of the sound source separation.

[0101] For example, in the constructed scene two people may talk simultaneously, however, for the sound source separation training, also an audio recording is simulated in which one of the people talks and another audio recording in which the other person talks.

[0102] Hence, the training audio dataset includes, for each of a plurality of constructed scenes, a singlechannel (or two-channel) audio recording in which the sound has been simulated according to the respective constructed scene - referred to as separation training input - and, for each object that actively emits sound, another audio recording in which only the respective object emitted sound - referred to as separation training targets -, and the sound has been simulated accordingly. Additionally, the training audio dataset includes, for each of the plurality of constructed scenes, the background sound audio recording.

[0103] Moreover, each of the plurality of constructed scenes may be simulated with different configurations of the microphone to account for a variety of different audio recording characteristics.

[0104] Hence, the second sub-network 4 uses then the separation training inputs as inputs and the separation training targets (including the background sound audio recording) as targets, wherein a difference between the outputs of the second sub-network 5 and the targets is backpropagated to adjust the weights of the second sub-network 5 such that the differences may be decreased (minimized).

[0105] Fig. 2 schematically illustrates in a flow diagram an embodiment of an information processing method 100, which is discussed in the following.

[0106] The information processing method may be performed by the information processing device as described herein, for example, by the information processing device 1 of Fig. 1.

[0107] At 101, video data representing a video are obtained, wherein the video data include image data representing a plurality of image frames and audio data representing a single-channel or two- channel audio recording associated with the plurality of image frames, as discussed herein.

[0108] At 102, the video data is input into a neural network, wherein the neural network is configured to separate sound sources in the audio data into one or more active sound sources and one background sound source, wherein the neural network is further configured to recognize objects in each of the plurality of image frames and to output a type of each recognized object, as discussed herein.

[0109] At 103, spatial information indicating a location of the active sound source are added, for each active sound source, based on a match between a type of the active sound source and the type of the recognized object, as discussed herein.

[0110] At 104, a change of a pose of a camera that captured the plurality of image frames is estimated, as discussed herein.

[0111] At 105, the location of the active sound source is tracked in accordance with the estimated change of the pose of the camera, as discussed herein.

[0112] At 106, the tracking is reset when a scene change is detected, as discussed herein.

[0113] Fig. 3 schematically illustrates in a block diagram an embodiment of a multi-purpose computer 130 which can be used for implementing an information processing device.

[0114] The computer 130 can be implemented such that it can basically function as any type of information processing device as described herein. The computer has components 131 to 141, which can form a circuitry, such as any one of the circuitries of the information processing device as described herein.

[0115] Embodiments which use software, firmware, programs or the like for performing the methods as described herein can be installed on computer 130, which is then configured to be suitable for the concrete embodiment.

[0116] The computer 130 has a CPU 131 (Central Processing Unit), which can execute various types of procedures and methods as described herein, for example, in accordance with programs stored in a read-only memory (ROM) 132, stored in a storage 137 and loaded into a random-access memory (RAM) 133, stored on a medium 140 which can be inserted in a respective drive 139, etc.

[0117] The CPU 131, the ROM 132 and the RAM 133 are connected with a bus 141, which in turn is connected to an input / output interface 134. The number of CPUs, memories and storages is only exemplary, and the skilled person will appreciate that the computer 130 can be adapted and configured accordingly for meeting specific requirements which arise, when it functions as an information processing device. At the input / output interface 134, several components are connected: an input 135, an output 136, the storage 137, a communication interface 138 and the drive 139, into which a medium 140 (compact disc, digital video disc, compact flash memory, or the like) can be inserted.

[0118] The input 135 can be a pointer device (mouse, graphic table, or the like), a keyboard, a microphone, a camera, a touchscreen, a time-of-fight device, etc.

[0119] The output 136 can have a display (liquid crystal display, cathode ray tube display, light emittance diode display, etc ), loudspeakers, etc.

[0120] The storage 137 can have a hard disk, a solid-state drive and the like.

[0121] The communication interface 138 can be adapted to communicate, for example, via a local area network (LAN), wireless local area network (WLAN), mobile telecommunications system (GSM, UMTS, LTE, NR etc.), Bluetooth, infrared, etc.

[0122] It should be noted that the description above only pertains to an example configuration of computer 130. Alternative configurations may be implemented with additional or other sensors, storage devices, interfaces or the like. For example, the communication interface 138 may support other radio access technologies than the mentioned UMTS, LTE and NR.

[0123] When the computer 130 functions as an information processing device, the communication interface 138 can further have a respective air interface (providing, e.g., E-UTRA protocols OFDMA (downlink) and SC-FDMA (uplink)) and network interfaces (implementing for example protocols such as Sl-AP, GTP-U, Sl-MME, X2-AP, or the like). Moreover, the computer 130 may have one or more antennas and / or an antenna array. The present disclosure is not limited to any particularities of such protocols.

[0124] It should be recognized that the embodiments describe methods with an exemplary ordering of method steps. The specific ordering of method steps is however given for illustrative purposes only and should not be construed as binding.

[0125] All units and entities described in this specification and claimed in the appended claims can, if not stated otherwise, be implemented as integrated circuit logic, for example on a chip, and functionality provided by such units and entities can, if not stated otherwise, be implemented by software.

[0126] In so far as the embodiments of the disclosure described above are implemented, at least in part, using software-controlled data processing apparatus, it will be appreciated that a computer program providing such software control and a transmission, storage or other medium by which such a computer program is provided are envisaged as aspects of the present disclosure. Note that the present technology can also be configured as described below.

[0127] (1) An information processing device, wherein the information processing device includes circuitry configured to: obtain video data representing a video, wherein the video data include image data representing a plurality of image frames and audio data representing a single-channel or two- channel audio recording associated with the plurality of image frames; inputting the video data into a neural network, wherein the neural network is configured to separate sound sources in the audio data into one or more active sound sources and one background sound source, wherein the neural network is further configured to recognize objects in each of the plurality of image frames and to output a type of each recognized object; and wherein the circuitry is further configured to add, for each active sound source, spatial information indicating a location of the active sound source based on a match between a type of the active sound source and the type of the recognized object.

[0128] (2) The information processing device of (1), wherein the circuitry is further configured to track the location of the active sound source.

[0129] (3) The information processing device of (2), wherein the neural network is further configured to estimate a change of a pose of a camera that captured the plurality of image frames, and wherein the circuitry is further configured to track the location of the active sound source in accordance with the estimated change of the pose of the camera.

[0130] (4) The information processing device of (3), wherein the circuitry is further configured to estimate the location of the active sound source outside the field-of-view of the camera.

[0131] (5) The information processing device of anyone of (2) to (4), wherein the circuitry is further configured to detect a scene change.

[0132] (6) The information processing device of (5), wherein the circuitry is further configured to reset the tracking when a scene change is detected.

[0133] (7) The information processing device of anyone of (1) to (6), wherein the neural network is further configured to perform noise removal.

[0134] (8) The information processing device of anyone of (1) to (7), wherein the neural network is further configured to perform voice filtering.

[0135] (9) The information processing device of anyone of (1) to (8), wherein the neural network is further configured to perform music filtering. (10) The information processing device of anyone of (1) to (9), wherein the neural network is further configured to perform dynamic range enhancement.

[0136] (11) An information processing method, wherein the information processing method includes: obtaining video data representing a video, wherein the video data include image data representing a plurality of image frames and audio data representing a single-channel or two- channel audio recording associated with the plurality of image frames; inputting the video data into a neural network, wherein the neural network is configured to separate sound sources in the audio data into one or more active sound sources and one background sound source, wherein the neural network is further configured to recognize objects in each of the plurality of image frames and to output a type of each recognized object; and adding, for each active sound source, spatial information indicating a location of the active sound source based on a match between a type of the active sound source and the type of the recognized object.

[0137] (12) The information processing method of (11), further including tracking the location of the active sound source.

[0138] (13) The information processing method of (12), wherein the neural network is further configured to estimate a change of a pose of a camera that captured the plurality of image frames, and further including tracking the location of the active sound source in accordance with the estimated change of the pose of the camera.

[0139] (14) The information processing method of (13), further including estimating the location of the active sound source outside the field-of-view of the camera.

[0140] (15) The information processing method of anyone of (12) to (14), further including detecting a scene change.

[0141] (16) The information processing method of (15), further including resetting the tracking when a scene change is detected.

[0142] (17) The information processing method of anyone of (11) to (16), wherein the neural network is further configured to perform noise removal.

[0143] (18) The information processing method of anyone of (11) to (17), wherein the neural network is further configured to perform voice filtering.

[0144] (19) The information processing method of anyone of (11) to (18), wherein the neural network is further configured to perform music filtering. (20) The information processing method of anyone of (11) to (19), wherein the neural network is further configured to perform dynamic range enhancement.

[0145] (21) A computer program comprising program code causing a computer to perform the information processing method according to anyone of (11) to (20), when being carried out on a computer.

[0146] (22) A non-transitory computer-readable recording medium that stores therein a computer program product, which, when executed by a processor, causes the information processing method according to anyone of (11) to (20) to be performed.

Claims

CLAIMS1. An information processing device comprising circuitry configured to: obtain video data representing a video, wherein the video data include image data representing a plurality of image frames and audio data representing a single-channel or two- channel audio recording associated with the plurality of image frames; inputting the video data into a neural network, wherein the neural network is configured to separate sound sources in the audio data into one or more active sound sources and one background sound source, wherein the neural network is further configured to recognize objects in each of the plurality of image frames and to output a type of each recognized object; and wherein the circuitry is further configured to add, for each active sound source, spatial information indicating a location of the active sound source based on a match between a type of the active sound source and the type of the recognized obj ect.

2. The information processing device of claim 1, wherein the circuitry is further configured to track the location of the active sound source.

3. The information processing device of claim 2, wherein the neural network is further configured to estimate a change of a pose of a camera that captured the plurality of image frames, and wherein the circuitry is further configured to track the location of the active sound source in accordance with the estimated change of the pose of the camera.

4. The information processing device of claim 3, wherein the circuitry is further configured to estimate the location of the active sound source outside the field-of-view of the camera.

5. The information processing device of claim 2, wherein the circuitry is further configured to detect a scene change.

6. The information processing device of claim 5, wherein the circuitry is further configured to reset the tracking when a scene change is detected.

7. The information processing device of claim 1, wherein the neural network is further configured to perform noise removal.

8. The information processing device of claim 1, wherein the neural network is further configured to perform voice filtering.

9. The information processing device of claim 1, wherein the neural network is further configured to perform music filtering.

10. The information processing device of claim 1, wherein the neural network is further configured to perform dynamic range enhancement.

11. An information processing method comprising: obtaining video data representing a video, wherein the video data include image data representing a plurality of image frames and audio data representing a single-channel or two- channel audio recording associated with the plurality of image frames; inputting the video data into a neural network, wherein the neural network is configured to separate sound sources in the audio data into one or more active sound sources and one background sound source, wherein the neural network is further configured to recognize objects in each of the plurality of image frames and to output a type of each recognized object; and adding, for each active sound source, spatial information indicating a location of the active sound source based on a match between a type of the active sound source and the type of the recognized object.

12. The information processing method of claim 11, further comprising tracking the location of the active sound source.

13. The information processing method of claim 12, wherein the neural network is further configured to estimate a change of a pose of a camera that captured the plurality of image frames, and further comprising tracking the location of the active sound source in accordance with the estimated change of the pose of the camera.

14. The information processing method of claim 13, further comprising estimating the location of the active sound source outside the field-of-view of the camera.

15. The information processing method of claim 12, further comprising detecting a scene change.

16. The information processing method of claim 15, further comprising resetting the tracking when a scene change is detected.

17. The information processing method of claim 11, wherein the neural network is further configured to perform noise removal.

18. The information processing method of claim 11, wherein the neural network is further configured to perform voice filtering.

19. The information processing method of claim 11, wherein the neural network is further configured to perform music filtering.

20. The information processing method of claim 11, wherein the neural network is further configured to perform dynamic range enhancement.