Methods and systems for generating an audio for an input video

The method and system address misalignment and customization issues in V2A frameworks by using probabilistic timestamps and user-controlled parameters to generate synchronized and customizable audio for videos, enhancing realism and user experience.

WO2026155530A1PCT designated stage Publication Date: 2026-07-23SAMSUNG ELECTRONICS CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
SAMSUNG ELECTRONICS CO LTD
Filing Date
2026-01-13
Publication Date
2026-07-23

AI Technical Summary

Technical Problem

Existing video-to-audio (V2A) frameworks struggle with accurately interpreting audio-aware visual cues, leading to temporally misaligned audio generation, limited generalization, and lack of user control over parameters like loudness and melody, resulting in diminished realism and coherence of multimedia content.

Method used

A method and system that determine temporal probabilistic audio timestamps based on object motion and scene dynamics, generate audio-aware latent embeddings, and adjust loudness and melody using predefined and user-defined parameters, employing machine learning to synchronize audio with video content.

Benefits of technology

Enables temporally synchronized, customizable audio generation that enhances realism and user experience by accurately reflecting video motion and user preferences, improving multimedia content quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2026000782_23072026_PF_FP_ABST
    Figure KR2026000782_23072026_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed is a method for generating an audio for an input video. The method includes determining one or more temporal probabilistic audio timestamps for the input video based on motion of one or more objects and scene dynamics. Further, the method includes generating one or more audio-aware latent embeddings for a plurality of frames based on the determined one or more temporal probabilistic audio timestamps. Furthermore, the method includes determining an audio loudness parameter embedding based on motion characteristics and one or more predefined parameters. Moreover, the method includes determining a melody variation factor based on one or more user defined melody parameters and a melody embedding and generating an audio embedding based on the one or more audio-aware media embeddings, the audio loudness parameter embedding, and the melody variation factor. Furthermore, generating the audio based on the audio embedding.
Need to check novelty before this filing date? Find Prior Art

Description

METHODS AND SYSTEMS FOR GENERATING AN AUDIO FOR AN INPUT VIDEO

[0001] The present disclosure relates to multimedia content generation, and more particularly relates to systems and methods forgenerating an audio for an input video.

[0002] With the advent of generative Artificial Intelligence (AI) technologies, the creation of video content from text or image inputs has witnessed significant growth. Such AI-generated videos, however, are generally devoid of any accompanying audio component, which adversely impacts their usability and overall viewer engagement. The absence of audio limits the immersive experience that users expect from multimedia content. Thus, audio generation has substantial utility in various audio-video editing workflows. For instance, noisy or distorted audio tracks in existing videos can be replaced with AI-generated audio to enhance quality. Furthermore, audio synthesis plays an important role in mobile device features such as Live Effects, Hyperlapse, and slow-motion video modes, where audio is typically absent by default. The integration of synchronized audio in such scenarios may significantly improve user experience and content appeal.

[0003] Despite these advantages, existing video-to-audio (V2A) frameworks often encounter challenges in accurately interpreting audio-aware visual cues. This limitation results in generated audio that is not temporally aligned with the corresponding video frames, thereby causing perceptual mismatches. Such discrepancies diminish the realism and coherence of multimedia output, reducing the practical applicability of such frameworks.

[0004] Further, the existing V2A frameworks predominantly rely on datasets comprising one-to-one video-audio pairs. The existing V2A frameworks restrict generalization capability of V2A frameworks due to overfitting on limited data patterns. Consequently, the existing V2A frameworks fail to adapt effectively to diverse real-world scenarios, where variations and ambiguities are inherent.

[0005] Moreover, the existing V2A frameworks primarily employ deterministic mappings rather than generating a probabilistic distribution of valid audio variations. As a result, the existing V2A frameworks lack flexibility to accommodate natural ambiguities and variations present in real-world data. Additionally, the existing V2A frameworks do not provide tunable features for user control over parameters such as loudness and melody variations in the generated audio, thereby limiting customization and overall user experience.

[0006] Hence, there is a need for a solution that overcomes the above-mentioned and other related problems.

[0007] This summary is provided to introduce a selection of concepts, in a simplified format, that are further described in the detailed description of the invention. This summary is neither intended to identify key or essential inventive concepts of the invention nor is it intended to determine the scope of the invention.

[0008] According to an embodiment of the present disclosure, disclosed herein is a method for generating an audio for an input video. The method includes determining one or more temporal probabilistic audio timestamps for the input video based on motion of one or more objects and scene dynamics in the input video. Further, the method includes generating one or more audio-aware latent embeddings for a plurality of frames from the input video based on the determined one or more temporal probabilistic audio timestamps. Furthermore, the method includes determining an audio loudness parameter embedding based on motion characteristics of the input video and one or more predefined parameters. The motion characteristics correspond to a rate of change of motion of the one or more objects in the input video. Moreover, the method includes determining a melody variation factor based on one or more user defined melody parameters and a melody embedding. The melody embedding indicates melodic characteristics of the audio to be generated. Further, the method includes generating, using the ML model, an audio embedding based on the one or more audio-aware media embeddings, the audio loudness parameter embedding, and the melody variation factor. Furthermore, generating, using the ML model, the audio for the input video based on the audio embedding.

[0009] According to an embodiment of the present disclosure, disclosed herein is a system for generating an audio for an input video. The system includes a memory and a processor. The processor is configured to determine one or more temporal probabilistic audio timestamps for the input video based on motion of one or more objects and scene dynamics in the input video. Further, the processor is configured togenerate one or more audio-aware latent embeddings for a plurality of frames from the input video based on the determined one or more temporal probabilistic audio timestamps. Furthermore, the processor is configured todetermine an audio loudness parameter embedding based on motion characteristics of the input video and one or more predefined parameters. The motion characteristics may correspond to a rate of change of motion of the one or more objects in the input video. Moreover, the processor is configured todetermine a melody variation factor based on one or more user defined melody parameters and a melody embedding. The melody embedding indicates melodic characteristics of the audio to be generated. Further, the processor is configured togenerate, using a machine learning (ML), an audio embedding based on the one or more audio-aware video embeddings, the audio loudness parameter embedding, and the melody variation factor. Furthermore, the processor is configured to generate, using the ML model, the audio for the input video based on the audio embedding.

[0010] To further clarify the advantages and features of the present invention, a more particular description of the invention will be rendered by reference to specific embodiments thereof, which are illustrated in the appended drawing. It is appreciated that these drawings depict only typical embodiments of the invention and are therefore not to be considered limiting its scope. The invention will be described and explained with additional specificity and detail in the accompanying drawings.

[0011] These and other features, aspects, and advantages of the present invention will become better understood when the following detailed description is read with reference to the accompanying drawings in which like characters represent like parts throughout the drawings, wherein:

[0012] Figure 1illustrates a block diagram depicting an environment 100 forgenerating an audio for an input video, in accordance with an embodiment of the present disclosure;

[0013] Figure 2illustrates a block diagram of the system forgenerating the audio for the input video, in accordance with an embodiment of the present disclosure.

[0014] Figure 3illustrates components of a plurality of modules of the system for generating the audio for the input video, in accordance with an embodiment of the present disclosure;

[0015] Figure 4illustrates a process flow depicting generation of the audio loudness parameter embedding, in accordance with an embodiment of the present disclosure; and

[0016] Figure 5illustrates a process flow of a method for generating the audio for the input video, in accordance with an embodiment of the present disclosure.

[0017] Further, skilled artisans will appreciate that those elements in the drawings are illustrated for simplicity and may not have necessarily been drawn to scale. For example, the flow charts illustrate the method in terms of the most prominent steps involved to help improve understanding of aspects of the present invention. Furthermore, in terms of the construction of the device, one or more components of the device may have been represented in the drawings by conventional symbols, and the drawings may show only those specific details that are pertinent to understanding the embodiments of the present invention so as not to obscure the drawings with details that will be readily apparent to those of ordinary skill in the art having the benefit of the description herein.

[0018] For the purpose of promoting an understanding of the principles of the present disclosure, reference will now be made to the various embodiments, and specific language will be used to describe the same. It will nevertheless be understood that no limitation of the scope of the present disclosure is thereby intended, such alterations and further modifications in the illustrated system, and such further applications of the principles of the present disclosure as illustrated therein, being contemplated as would normally occur to one skilled in the art to which the present disclosure relates.

[0019] It will be understood by those skilled in the art that the foregoing general description and the following detailed description are explanatory of the present disclosure and are not intended to be restrictive thereof.

[0020] Whether or not a certain feature or element was limited to being used only once, it may still be referred to as “one or more features” or “one or more elements,” “at least one feature,” or “at least one element.” Furthermore, the use of the terms “one or more” or “at least one” feature or element does not preclude there being none of that feature or element, unless otherwise specified by limiting language, including, but not limited to, “there needs to be one or more…” or “one or more elements are required.”

[0021] Reference is made herein to some “embodiments.” It should be understood that an embodiment is an example of a possible implementation of any features and / or elements of the present disclosure. Some embodiments have been described for the purpose of explaining one or more of the potential ways in which the specific features and / or elements of the proposed disclosure fulfill the requirements of uniqueness, utility, and non-obviousness.

[0022] Use of the phrases and / or terms including, but not limited to, “a first embodiment,” “a further embodiment,” “an alternate embodiment,” “one embodiment,” “an embodiment,” “multiple embodiments,” “some embodiments,” “other embodiments,” “further embodiment”, “furthermore embodiment”, “additional embodiment” or other variants thereof do not necessarily refer to the same embodiments. Unless otherwise specified, one or more particular features and / or elements described in connection with one or more embodiments may be found in one embodiment, or may be found in more than one embodiment, or may be found in all embodiments, or may be found in no embodiments. Although one or more features and / or elements may be described herein in the context of only a single embodiment, or in the context of more than one embodiment, or in the context of all embodiments, the features and / or elements may instead be provided separately or in any appropriate combination or not at all. Conversely, any features and / or elements described in the context of separate embodiments may alternatively be realized as existing together in the context of a single embodiment.

[0023] Any particular and all details set forth herein are used in the context of some embodiments and therefore should not necessarily be taken as limiting factors to the proposed disclosure.

[0024] The terms “comprises”, “comprising”, or any other variations thereof, are intended to cover a non-exclusive inclusion, such that a process or method that comprises a list of steps does not include only those steps but may include other steps not expressly listed or inherent to such process or method. Similarly, one or more devices or sub-systems or elements or structures or components preceded by “comprises... a” does not, without more constraints, preclude the existence of other devices or other sub-systems or other elements or other structures or other components or additional devices or additional sub-systems or additional elements or additional structures or additional components.

[0025] The term “couple” and the derivatives thereof refer to any direct or indirect communication between two or more elements, whether or not those elements are in physical contact with each other. The terms “transmit”, “receive”, and “communicate”, as well as the derivatives thereof, encompass both direct and indirect communication. The term “or” is an inclusive term meaning “and / or”. The phrase “associated with,” as well as derivatives thereof, refer to include, be included within, interconnect with, contain, be contained within, connect to or with, couple to or with, be communicable with, cooperate with, interleave, juxtapose, be proximate to, be bound to or with, have, have a property of, have a relationship to or with, or the like. The term “controller” refers to any device, system, or part thereof that controls at least one operation. The functionality associated with any particular controller may be centralized or distributed, whether locally or remotely. The phrase “at least one of,” when used with a list of items, means that different combinations of one or more of the listed items may be used, and only one item in the list may be needed. For example, “at least one of A, B, and C” includes any of the following combinations: A, B, C, A and B, A and C, B and C, and A and B and C, and any variations thereof. As an additional example, the expression “at least one of a, b, or c” may indicate only a, only b, only c, both a and b, both a and c, both b and c, all of a, b, and c, or variations thereof. Similarly, the term “set” means one or more. Accordingly, the set of items may be a single item or a collection of two or more items.

[0026] Moreover, multiple functions described below may be implemented or supported by one or more computer programs, each of which is formed from computer-readable program code and embodied in a computer-readable medium. The terms “application” and “program” refer to one or more computer programs, software components, sets of instructions, procedures, functions, objects, classes, instances, related data, or a portion thereof adapted for implementation in a suitable computer-readable program code. The phrase “computer-readable program code” includes any type of computer code, including source code, object code, and executable code. The phrase “computer-readable medium” includes any type of medium capable of being accessed by a computer, such as Read Only Memory (ROM), Random Access Memory (RAM), a hard disk drive, a Compact Disc (CD), a Digital Video Disc (DVD), or any other type of memory. A “non-transitory” computer-readable medium excludes wired, wireless, optical, or other communication links that transport transitory electrical or other signals. A non-transitory computer-readable medium includes media where data may be permanently stored and media where data may be stored and later overwritten, such as a rewritable optical disc or an erasable memory device.

[0027] Any particular and all details set forth herein are used in the context of some embodiments and therefore should NOT be necessarily taken as limiting factors to the attached claims. The attached claims and their legal equivalents can be realized in the context of embodiments other than the ones used as illustrative examples in the description below.

[0028] Further, skilled artisans will appreciate those elements in the drawings that are illustrated for simplicity and may not have necessarily been drawn to scale. For example, the flow charts illustrate the method in terms of the most prominent steps involved to help improve understanding of aspects of the present disclosure. Furthermore, in terms of the construction of the device, one or more components of the device may have been represented in the drawings by conventional symbols, and the drawings may show only those specific details that are pertinent to understanding the embodiments of the present disclosure so as not to obscure the drawings with details that will be readily apparent to those of ordinary skill in the art having the benefit of the description herein.

[0029] For the sake of clarity, the first digit of a reference numeral of each component of the present disclosure is indicative of the Figure number, in which the corresponding component is shown. For example, reference numerals starting with digit “1” are shown at least in Figure 1. Similarly, reference numerals starting with digit “2” are shown at least in Figure 2. Further, similar reference numerals have been used to represent similar components in the Figures.

[0030] It should be noted that the terms “evidence” and “at least one proof of evidence” have been used interchangeably throughout the description and the drawings. Further, the terms “policy” and “one or more policy configurations” have been used interchangeably throughout the description and the drawings.

[0031] Embodiments of the present disclosure will be described below in detail with reference to the accompanying drawings.

[0032] It is an object of the invention to provide a system and a method that overcome the limitations found in prior art related to multimedia content generation.

[0033] It is another object of the invention to provide a controllable audio generation system and method from an input video with high temporal synchronization.

[0034] It is yet another object of the invention to facilitate users to vary audio loudness based on motion awareness and control variation in the melody of generated audio without any additional text or audio guidance.

[0035] Figure 1illustrates a block diagram depicting an environment 100 forgenerating an audio for the input video, in accordance with an embodiment of the present disclosure.

[0036] Referring to Figure 1, the environment 100 depicts an implementation of system 106. In one implementation, the system 106 may be implemented in server 108. In an alternate implementation, the system 106 may be implemented within an electronic device 104. In a non-limiting example, the electronic device 104 may include a smartphone, tablet, laptop, desktop computer, wearable device, or any other computing device capable of facilitating communication with the system 106.

[0037] The environment 100 includes the electronic device 104 configured to generate the audio for the input video as an output 110.

[0038] In an embodiment, the system 106 may include software, hardware, a combination of software and hardware, an in-built application on the electronic device 104, or an application to be installed and operated on the electronic device 104 in communication with a network interface (not shown). The system 106 may also be accessible at the electronic device 104 via the server (a cloud-based server), and available remotely from the electronic device 104.

[0039] In the embodiment where the system 106 is located outside the user device 104, the network interface may be configured to provide network connectivity and enable communication between the system 106 and the user device 104. The network connectivity may be provided via a wireless connection or a wired connection. For example, the network connectivity may be provided via cellular technology, such as 3rd Generation (3G), 4th Generation (4G), 5th Generation (5G), pre-5G, 6th Generation (6G), Bluetooth, Local Area Network (LAN), Wi-Fi, cable, or any other wired / wireless communication technology.

[0040] In an embodiment, the system 106 may be configured to determine one or more temporal probabilistic audio timestamps for the input video based on motion of one or more objects and scene dynamics in the input video. In a non-limiting example, the motion of one or more objects may correspond to displacement or positional change of visual entities present in the input video across consecutive frames. For instance, in a video depicting a person running across a field, the motion of the person is determined by analyzing the change in pixel positions corresponding to the person’s body across successive frames, thereby generating motion flow vectors indicative of speed and direction. In a non-limiting example, the scene dynamics may correspond to contextual variations in the environment of the input video, including changes in lighting conditions, background textures, object interactions, and spatial configurations over time. In an embodiment, the scene dynamics may be derived using semantic segmentation maps, texture analysis, and temporal consistency estimators to capture environmental changes that influence audio generation. For instance, in a video depicting a car entering and moving through a tunnel, the transition from natural daylight to artificial tunnel lighting represents a change in the scene dynamics. Additional variations, such as appearance of other vehicles or alterations in road texture, further contribute to contextual changes that influence the generation of temporally synchronized audio. In a non-limiting example, the one or more temporal probabilistic audio timestamps may correspond to a probability metric indicative of an audio intensity level for each temporal instance throughout the input video. For instance, the input video depicts a drummer performing on stage. During the video, the motion of the drumsticks and the scene dynamics, such as lighting changes and audience movement, are analyzed. Based on the analysis, the system 106 may predict that a high-intensity audio event (a drumbeat) is most likely to occur at specific temporal instances where the drumstick strikes the drum surface. Each of the instances may be assigned a probability metric, such as 0.85 or 0.92, indicating the likelihood of a strong audio intensity at that point in time. The probability metric collectively forms the temporal probabilistic audio timestamps for the input video.

[0041] In an embodiment, the system 106 may be configured to generate one or more audio-aware latent embeddings for a plurality of frames from the input video based on the determined one or more temporal probabilistic audio timestamps. In a non-limiting example, the plurality of frames from the input video may correspond to two or more discrete image frames that collectively represent successive temporal instances within a pre-defined duration of the input video. Each frame may include pixel data depicting visual content at a specific point in time. The plurality of frames may be extracted at a fixed frame rate or adaptively sampled based on computational requirements. For instance, in a 10-second video recorded at 30 frames per second, the plurality of frames may include 300 individual frames, each representing a distinct visual snapshot of the scene at a corresponding temporal instance. In a non-limiting example, the one or more audio-aware latent embeddings may correspond to one or more multi-dimensional feature representations generated for the plurality of frames from the input video. The feature representations may be conditioned on audio-related characteristics.

[0042] In an embodiment, the system 106 may be configured to determine an audio loudness parameter embedding based on motion characteristics of the input video and one or more predefined parameters.

[0043] In a non-limiting example, the audio loudness parameter embedding may correspond to a multi-dimensional feature representation that encodes loudness characteristics of the audio to be generated. In a non-limiting example, the motion characteristics may correspond to a rate of change of motion of the one or more objects in the input video. The motion characteristics may relate to quantitative attributes that describe the movement of one or more objects within the input video over time. In an embodiment, the motion characteristics may include parameters such as displacement, velocity, acceleration, and direction of motion of the objects across consecutive frames. In a non-limiting example, the one or more predefined parameters may correspond to one or more control variables specified by the user 102. In an embodiment, the predefined parameters may include loudness scaling factors, amplitude thresholds, or sensitivity levels that determine how motion characteristics affect the perceived loudness of the generated audio.

[0044] In an embodiment, the system 106 may be configured to determine a melody variation factor based on one or more user defined melody parameters and a melody embedding. In a non-limiting example, the melody variation factor may correspond to a variation in a melody for the audio. In a non-limiting example, the melody embedding may correspond to at least one of a melody embedding and a non-melody embedding. In a non-limiting example, the one or more user defined melody parameters may include pitch range, tonal scale, rhythmic complexity, harmonic progression, or a Gaussian noise level for introducing controlled randomness in melody generation.

[0045] In an embodiment, the system 106 may be configured to generate, using a machine learning (ML) model, an audio embedding based on the one or more audio-aware video embeddings, the audio loudness parameter embedding, and the melody variation factor. In a non-limiting example, the audio embedding may act as an intermediate representation that captures correlations between visual features, motion dynamics, and desired acoustic properties.

[0046] In an embodiment, the system 106 may be configured to generate, using the ML model, the audio for the input video based on the audio embedding.

[0047] In an example implementation, consider the input video depicting a drummer performing on stage. The system 106 may first analyze the frames to detect motion of the one or more objects, such as drumsticks and drummer’s hands, and evaluate scene dynamics, including lighting changes and audience movement. Based on the analysis, the system 106 may determine temporal probabilistic audio timestamps, assigning probability metrics (e.g., 0.85 or 0.92) to specific temporal instances where drumstick strikes are most likely to occur. Next, the system 106 may extract the plurality of frames from the input video and generate the audio-aware latent embeddings for the plurality of frames, conditioned on the identified timestamps. The audio-aware latent embeddings may capture visual features relevant for audio synthesis. The system 106 may then compute the audio loudness parameter embedding by analyzing motion characteristics such as acceleration and velocity of the drumsticks, along with the user-defined parameters such as loudness scaling factors. For example, rapid drumstick movements correspond to higher loudness levels. Further, the system 106 may determine the melody variation factor based on the one or more user defined melody parameters, such as pitch range and rhythmic complexity, and the melody embedding. Finally, the system 106 may combine the one or more audio-aware latent embeddings, the audio loudness parameter embedding, and the melody variation factor to generate the audio embedding, which is decoded by the ML model to produce an audio waveform. The resulting audio may be temporally synchronized with the drummer’s movements and dynamically reflects changes in loudness and melody, creating a realistic and immersive experience.

[0048] In an embodiment, to determine the one or more temporal probabilistic audio timestamps for the input video, the system 106 may be configured to detect a motion of the one or more objects and the scene dynamics in the input video. In a non-limiting example, the motion may correspond to displacement or positional change of one or more visual entities present in the input video across consecutive frames. The system 106 may be configured to determine a behavioural pattern associated with the one or more objects based on the determined the motion of the one or more objects, and a contextual change in an environment of the input video based on the scene dynamics. In a non-limiting example, the behavioural patternmay correspond to a sequence or trend of motion exhibited by the one or more objects in the input video over time, which is indicative of the object’s activity or interaction within the scene. In a non-limiting example, the contextual change may correspond to variations in the environmental attributes of the input video that influence interpretation of audio generation. In an embodiment, contextual change may include alterations in lighting conditions, background textures, object interactions, or spatial configurations within the scene. The system 106 may be configured to determine the one or more temporal probabilistic audio timestamps based on the determined behavioral pattern and the contextual change.

[0049] In an embodiment, the system 106 may be configured to determine the one or more probabilistic audio timestamps based on a temporal probability of an event based on predicting a plurality of probability parameters associated with the plurality of frames. In a non-limiting example, the temporal probability of the eventmay correspond to a computed likelihood that a specific audio-related event may occur at a given temporal instance within the input video. In a non-limiting example, the plurality of probability parameters may indicate two or more statistical values associated with the prediction of audio events for the input video. In an embodiment, the plurality of probability parameters may include mean, variance, standard deviation, and confidence scores derived from a probability density function.

[0050] In an embodiment, to generate the one or more audio-aware latent embeddings, the system 106 may be configured to generate, using a video encoder, a video frame embedding for the input video. The system 106 may be configured to generate the one or more audio-aware latent embeddings based on the one or more temporal probabilistic audio timestamps and the generated video fame embeddings. In a non-limiting example, the video frame embedding may correspond to a multi-dimensional feature representation generated from a single frame of an input video using the encoder.

[0051] In an embodiment, the system 106 may be configured to train the ML model. In an embodiment, to train the ML model for estimating the temporal probabilistic audio timestamp, the system 106 may be configured to receive an input video comprising the plurality of frames encoded in Red, Green, and Blue (RGB) color format. The system 106 may be configured to predict a first set of probability density function parameters for the audio to be generated based on the input video. In a non-limiting example, the first set of probability density function parameters may indicate one or more probabilistic audio timestamps corresponding to the input video. In a non-limiting example, the first set of probability density function parameters may include at least a mean value and a variance value. In a non-limiting example, the mean value may indicate a central estimate of a temporal location of an audio event within the input video. In a non-limiting example, the variance value may indicate a measure of temporal uncertainty associated with the estimated temporal location. The system 106 may be configured to receive two or more temporally synchronised ground-truth audios corresponding to the input video. In a non-limiting example, the two or more temporally synchronised ground-truth audios may indicate audio recordings that are perfectly aligned in time with the input video. The system 106 may be configured to generate one or more timestamps for the ground-truth audios based on an analysis of intensity levels in the two or more temporally synchronized ground-truth audios that exceed a predefined threshold. In a non-limiting example, the predefined threshold may correspond to a fixed or configurable reference value used to determine whether intensity level of an audio signal exceeds a specified limit during analysis. In an embodiment, the predefined threshold may correspond to an amplitude level, energy value, or decibel range that signifies a significant audio event. In a non-limiting example, the one or more timestamps may indicate instances of elevated acoustic activity corresponding to the input video. The system 106 may be configured to obtain a second set of probability density function parameters for the ground-truth audios through fitting of a Gaussian distribution for each of the ground-truth audios based on the one or more timestamps. In a non-limiting example, the second set of probability density function parameters may indicate one or more probabilistic audio timestamps corresponding to the ground-truth audios. The system 106 may be configured to train the ML model to estimate the temporal probabilistic audio timestamp based on correlation of the first set of probability density function parameters and second set of probability density function parameters.

[0052] In an embodiment, to determine the audio loudness parameter embedding, the system 106 may be configured to extract a sound source location based on the input video. In a non-limiting example, the sound source location may be a spatial region within the input video that corresponds to origin of a sound-producing event. The system 106 may be configured to generate a bounding box based on the extracted sound source location in the input video. In a non-limiting example, the bounding box may correspond to a rectangular or polygonal region generated around the sound source location within the video frame. In an embodiment, the bounding box may be used to isolate and track motion of sound-producing object, enabling accurate computation of motion flow vectors and masking irrelevant background movements. The system 106 may be configured to determine a motion flow vector for the extracted sound source location based on the input video. The system 106 may be configured to determine a masked motion flow vector based on the motion flow vector and the bounding box, for the extracted sound source location. In a non-limiting example, the masked motion flow vector may correspond to an average rate of change of motion field in the extracted sound source location with time. The system 106 may be configured to generate a loudness embedding vector based on passing the masked motion flow vector and the one or more predefined parameters, using the ML model. The system 106 may be configured to determine the audio loudness parameter embedding based on the masked motion flow vector.

[0053] In an embodiment, prior to determining the melody variation factor, the system 106 may be configured to determine the melody embedding based on the one or more audio-aware video embedding and the audio loudness parameter embedding.

[0054] In an embodiment, to generate the audio embedding, the system 106 may be configured to receive the one or more user defined melody parameters specifying a desired level of Gaussian noise. The system 106 may be configured to generate a predicted melody embedding using the ML model based on the one or more user defined melody parameters. In a non-limiting example, the predicted melody embeddingmay correspond to a multi-dimensional feature representation generated by the ML model to encode melodic characteristics of audio to be synthesized. The system 106 may be configured to modulate the predicted melody embedding based on the one or more user defined melody parameters. The system 106 may be configured to generate the audio embedding, using the ML model, based on the modulated melody embedding and a non-melody embedding.

[0055] In an embodiment, the system 106 may be configured to decode the audio embedding using an audio decoder to reconstruct an output audio waveform. In a non-limiting example, the output audio waveform may correspond to a time-domain representation of a synthesized audio signal generated by decoding the audio embedding using the audio decoder. In an embodiment, the output audio waveform may include amplitude variations over time that correspond to the audio-related characteristics, such as loudness, pitch, and melody, synchronized with the visual content of the input video.

[0056] In an embodiment, the system 106 may be configured to train the ML model to generate the audio embedding. The system 106 may be configured to decompose an input audio signal into a melody component and a non-melody component using pre-defined decomposition methods. In a non-limiting example, the melody component may correspond to a segment of the audio signal that encodes melodic characteristics such as pitch contour, harmonic progression, and rhythmic structure. In a non-limiting example, the non-melody component may correspond to a segment of the audio signal that represents non-melodic elements, such as background noise, percussion, or ambient sounds, which do not contribute to the primary melodic pattern. The system 106 may be configured to train the ML model to generate the melody embedding based on the melody component, the non-melody component, and the input audio signal using one or more auto-encoders.

[0057] Figure 2illustrates a block diagram of the system 106 forgenerating the audio for the input video, in accordance with an embodiment of the present disclosure.

[0058] In an embodiment, the system 106 may include at least a processor 202, a memory 204, a plurality of modules 206, and a data unit 208. The processor 202, the memory 204, the plurality of modules 206, and the data unit 208 are communicably coupled with each other.

[0059] In an embodiment, the at least one processor 202 may be in communication with the memory 204. The at least one processor 202 may be a single processing unit or several units, all of which could include multiple computing units. The at least one processor 202 may be implemented as one or more microprocessors, microcomputers, microcontrollers, digital signal processors, central processing units, state machines, logic circuitries, and / or any devices that manipulate signals based on operational instructions. Among other capabilities, the at least one processor 202 may be configured to fetch and execute computer-readable instructions and data stored in the memory 204.

[0060] The processor 202 may control the overall operations of the system 106. For example, the processor 202 may cause other components of the system 106 to perform various operations by executing instructions stored in the memory 204. For example, the processor 202 may be operatively connected to the memory 204 and control the operation of the system 106. In addition, the processor 202 may control the operation of the system 106 according to the disclosure by executing one or more instructions stored in the memory 204. The processor 202 may include one or more processors.

[0061] The processor 202 may be implemented as one or more integrated circuit (or circuitry) chips and may perform various data processing operations. The processor 202 may include at least one electrical circuit and may individually or collectively distribute and process instructions (or programs, data, etc.) stored in the memory 204. The processor 202 may include a processor assembly including one or more processing circuits. The processor 202 may include any processing circuit operative to control the performance and operations of one or more components (e.g., the memory 204) of the system 106. For example, the processor 202 (e.g., an application processor (AP)) may be implemented as a system on chip (SoC) (e.g., a single chip or chipset). For example, the processor 202 may be implemented as a plurality of cores (or at least one core circuit), a plurality of chips, or a plurality of chipsets. For example, the processor 202 may include one or more processing circuits. For example, the processor 202 may include one or more processing circuits configured to individually and / or collectively perform various functions of the disclosure.

[0062] The processor 202 may include one or more of a central processing unit (CPU), a graphics processing unit (GPU), an accelerated processing unit (APU), a many integrated core (MIC), a digital signal processor (DSP), a neural processing unit (NPU), a hardware accelerator, or a machine learning accelerator. The processor 202 may control one or any combination of other components of the system 106 and perform operations related to communication or data processing. The processor 202 may execute one or more programs or instructions stored in the memory 204 of the system 106. For example, the processor 202 may perform a method according to an embodiment of the disclosure by executing one or more instructions stored in the memory 204.

[0063] In an embodiment, the memory 204 may include any non-transitory computer-readable medium known in the art including, for example, volatile memory, such as static random access memory (SRAM) and dynamic random access memory (DRAM), and / or non-volatile memory, such as read-only memory (ROM), erasable programmable ROM, flash memories, hard disks, optical disks, and magnetic tapes.

[0064] The memory 204 may store data necessary for the operation of the system 106 according to various embodiments of the disclosure. Depending on the data storage purpose, the memory 204 may be implemented as memory embedded in the system 106 (e.g., volatile memory (e.g., semi-permanent memory, such as random access memory (RAM)), non-volatile memory (e.g., permanent memory, such as read-only memory (ROM)), flash memory, a hard drive, a solid-state drive, etc.), or memory detachably attached to the system 106 (e.g., a memory card, external memory, etc.).

[0065] Instructions may be stored in the memory 204. The processor 202 may perform operations of the system 106 according to various embodiments of the disclosure by individually or collectively executing the instructions in the memory 204. In addition, programs and data for operating the system 106 may be stored in the memory 204. For example, the memory 204 may store one or more software applications, such as an operating system (or system) software application, a firmware software application, a driver software application, a plug-in (e.g., add-in, add-on, and / or applet) software application, and / or any other suitable software applications.

[0066] Meanwhile, in this disclosure, the term “memory 204” may be used to include the memory 204, ROM and RAM within the processor 202, or a memory card (e.g., a micro SD card, a memory stick) mounted on the system 106. In addition, various information within a range necessary to achieve the objectives of this disclosure may be stored in the memory 204, and the information stored in the memory 204 may be updated based on information received from an external device or input by a user.

[0067] In an embodiment, the plurality of modules 206 may be configured to generate the audio for the input video. The plurality of modules 206 may include a

[0068] temporal probabilistic audio timestamp detector 302, an audio loudness parameter modulator 304, an audio embedding generator 306, a video encoder 308, a melody variability controller 310, and an audio decoder 312. A detailed description of the plurality of modules 206 is provided with reference to Figure 3.

[0069] In some embodiments, the plurality of modules 206 may include a set of instructions that can be executed to cause the system 206 to perform any one or more of the methods disclosed. The system 206 may operate as a standalone device or may be connected, e.g., using a network, to other computer systems or peripheral devices. Further, while a single processing unit is illustrated, the term “processing unit” shall also be taken to include any collection of processing units, implemented across the system 206 that individually or jointly execute a set, or multiple sets, of instructions to perform one or more computer functions.

[0070] In an embodiment, the plurality of modules 206 may be implemented using one or more artificial intelligence (AI) units that may include a plurality of neural network layers. Examples of neural networks include, but are not limited to, Convolutional Neural Network (CNN), Deep Neural Network (DNN), Recurrent Neural Network (RNN), and Restricted Boltzmann Machine (RBM). Further, ‘learning’ may be referred to in the disclosure as a method for training a predetermined target device (for example, a robot) using a plurality of learning data to cause, allow, or control the target device to make a determination or prediction. Examples of learning techniques include, but are not limited to, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning. At least one of a plurality of CNN, DNN, RNN, RMB models and the like may be implemented to thereby achieve execution of the present subject matter’s mechanism through an AI model. A function associated with an AI unit may be performed through the non-volatile memory, the volatile memory, and the processor. The processor 202 may include one or a plurality of processors. At this time, one or a plurality of processors may be a general-purpose processor, such as a central processing unit (CPU), an application processor (AP), or the like, a graphics-only processing unit such as a graphics processing unit (GPU), a visual processing unit (VPU), and / or an AI-dedicated processor, such as a neural processing unit (NPU). One or a plurality of processors control the processing of the input data in accordance with a predefined operating rule or artificial intelligence (AI) model stored in the non-volatile memory and the volatile memory. The predefined operating rule or artificial intelligence model is provided through training or learning.

[0071] In an embodiment, the data unit 208, amongst other things, includes routines, programs, objects, components, data structures, and the like, which perform tasks or implement data types. The data unit 208 may also be implemented as signal processor(s), state machine(s), logic circuitries, and / or any other device or component that manipulates signals based on operational instructions. Further, the data unit 208 may be implemented in hardware, instructions executed by a processing unit, or by a combination thereof. The processing unit may comprise a processor, such as the at least one processor 202, a state machine, a logic array, or any other suitable device capable of processing instructions. The processing unit may be a general-purpose processor that executes instructions to cause the general-purpose processor to perform the required tasks, or the processing unit can be dedicated to performing the required functions. In another embodiment of the present disclosure, the data unit 208 may be machine-readable instructions (software) that, when executed by the processor 202, perform any of the described functionalities.

[0072] Figure 3illustrates components of the plurality of module 206 for generating the audio for the input video, in accordance with an embodiment of the present disclosure.

[0073] As shown, the plurality of modules 206 may include the temporal probabilistic audio timestamp detector 302, the audio loudness parameter modulator 304, the audio embedding generator 306, the video encoder 308, the melody variability controller 310, and the audio decoder 312, which are communicably coupled with each other.

[0074] In an embodiment, the temporal probabilistic audio timestamp detector 302 may be configured to estimate the one or more temporal probabilistic audio timestamp indicative of the probability of audio intensity at any given time instance. The temporal probabilistic audio timestamp detector 302 may achieve the estimation by analyzing the motion of the one or more objects and the scene dynamics derived from the input video frames of the input video 301-A, thereby correlating visual cues with audio characteristics. The temporal probabilistic audio timestamp detector 302 may receive at least one video frame from the input video 301-A as input and generate a latent embedding vector representing the probabilistic audio timestamp as an output.

[0075] In an embodiment, the audio loudness parameter modulator 304 may be configured to determine the audio loudness parameter embedding based on the rate of change of motion of the one or more objects in the input video 301-A. The audio loudness parameter modulator 304 may further include the one or more user-defined melody parameters 301-B to vary the loudness in accordance with user preferences. In an embodiment, the inputs to the audio loudness parameter modulator 304 may comprise the input video 301-A and the one or more user-defined melody parameters 301-B, and the output may comprise the audio loudness parameter embedding representing the loudness parameter for subsequent processing.

[0076] In an embodiment, the audio embedding generator 306 may include ML model and may be configured to generate the audio embedding by processing the one or more audio-aware media embeddings, the audio loudness parameter embedding, and the melody variation factor. The audio embedding generator 306 may receive as inputs the one or more audio-aware video embeddings, the audio loudness parameter embedding, and the melody variation factor, and may output a latent embedding vector representing the audio embedding for subsequent audio synthesis.

[0077] In an embodiment, the video encoder 308 may include the ML model and may be configured to process the input video frames to generate a compressed representation referred to as the video frame embedding. The video encoder 308 may receive the video frames from the input video 301-A as input and may output a latent embedding vector representing the video frame embedding for subsequent processing.

[0078] In an embodiment, the melody variability controller 310 may determine the melody variation factor to enable different variations in the melody of the generated audio based on the user-defined parameter. The melody variability controller 310 may receive as inputs the one or more user-defined melody parameters 301-B and the melody embedding and may output a modulated melody embedding for subsequent audio synthesis.

[0079] In an embodiment, the audio decoder 312 may include the ML model and be configured to generate the audio waveform by processing the audio embedding. The audio decoder 312 may receive the audio embedding as input and may output the audio waveform representing the synthesized sound corresponding to the input video 301-A.

[0080] Figure 4illustrates a process flow depicting the generation of the audio loudness parameter embedding, in accordance with an embodiment of the present disclosure.

[0081] In an embodiment, to determine the audio loudness parameter embedding, the system 106 may be configured to extract the sound source location based on the input video 301-A and generate the bounding box corresponding to the extracted sound source location in the input video 301-A, as shown in a sound source tracking block 402. The system 106 may determine the motion flow vector for the extracted sound source location based on the input video 301-A, as represented by a motion field calculation block 404. A motion field may be represented as M (x, y, t),where M (x, y, t) denotes the displacement of a pixel at coordinates(x, y)at timet.

[0082] The system 106 may then determine the masked motion flow vector based on the motion flow vector and the bounding box for the extracted sound source location. The masked motion flow vector may be represented as,

[0083] ,

[0084] where R(t) represents the rate of change of M (x, y, t) averaged over the sound source region.

[0085] The masked motion flow vector may correspond to an average rate of change of motion field in the extracted sound source location with respect to time, as shown in masked motion flow block 406 and rate of change computation block 408. The rate of change of motion flow may be tuned using the one or more predefined parameters, including a user-controlled scaling factor, to obtain a loudness modulation value, as shown in block 410. The loudness may be represented as,

[0086] L(t) = K*R(t) + b(t)

[0087] where K is a user-controlled scaling factor, R(t) is the rate of change of motion flow, and b(t) is the base loudness.

[0088] The tuned rate of change of motion flow value, K × R(t), may be passed through the ML model to generate the loudness embedding vector, as shown in linear layer block 412. The audio loudness parameter embedding may then be determined based on the loudness embedding vector, as represented by block 414.

[0089] Figure 5illustrates a process flow of a method 500 for generating the audio for the input video, in accordance with an embodiment of the present disclosure. The method 500 may be a computer-implemented method executed, for example, by the system 106. For the sake of brevity, constructional and operational features of the system 106 that are already explained in the description of Figures 1-4 are not explained in detail in the description of Figure 5.

[0090] At step 502, the method 500 may include determining the one or more temporal probabilistic audio timestamps for the input video based on the motion of the one or more objects and the scene dynamics in the input video. In a non-limiting example, the motion of one or more objects may correspond to the displacement or positional change of visual entities present in the input video 301-A across consecutive frames. In a non-limiting example, the scene dynamics may correspond to contextual variations in the environment of the input video, including changes in lighting conditions, background textures, object interactions, and spatial configurations over time. In a non-limiting example, the one or more temporal probabilistic audio timestamps may correspond to the probability metric indicative of the audio intensity level for each temporal instance throughout the input video.

[0091] At step 504, the method 500 may include generating the one or more audio-aware latent embeddings for the plurality of frames from the input video 301-A based on the determined one or more temporal probabilistic audio timestamps. In a non-limiting example, the plurality of frames from the input video 301-A may correspond to the two or more discrete image frames that collectively represent successive temporal instances within the pre-defined duration of the input video. In a non-limiting example, the one or more audio-aware latent embeddings may correspond to one or more multi-dimensional feature representations generated for the plurality of frames from the input video 301-A. The feature representations may be conditioned on audio-related characteristics.

[0092] At step 506, the method 500 may include determining the audio loudness parameter embedding based on the motion characteristics of the input video 301-A and the one or more predefined parameters. In a non-limiting example, the audio loudness parameter embedding may correspond to the multi-dimensional feature representation that encodes loudness characteristics of the audio to be generated. In a non-limiting example, the motion characteristics may correspond to a rate of change of motion of the one or more objects in the input video. In a non-limiting example, the one or more predefined parameters may correspond to one or more control variables specified by the user 102

[0093] At step 508, the method 500 may include determining the melody variation factor based on the one or more user defined melody parameters 301-B and the melody embedding. In a non-limiting example, the melody variation factor may correspond to the variation in the melody for the audio. In a non-limiting example, the melody embedding may correspond to at least one of a melody embedding and a non-melody embedding. In a non-limiting example, the one or more user defined melody parameters may include pitch range, tonal scale, rhythmic complexity, harmonic progression, or the Gaussian noise level for introducing controlled randomness in melody generation. In a non-limiting example, the non-melody embedding may correspond to a latent representation that captures audio characteristics other than melody, such as rhythm, timbre, loudness, and environmental sound features.

[0094] At step 510, the method 500 may include generating, using the ML model, the audio embedding based on the one or more audio-aware video embeddings, the audio loudness parameter embedding, and the melody variation factor. In a non-limiting example, the audio embedding may act as the intermediate representation that captures correlations between visual features, motion dynamics, and desired acoustic properties.

[0095] At step 512, the method 500 may include generating, using the ML model, the audio for the input video 301-A based on the audio embedding.

[0096] In an embodiment, for determining the one or more temporal probabilistic audio timestamps for the input video, the method 500 may include detecting the motion of the one or more objects and the scene dynamics in the input video. In a non-limiting example, the motion may correspond to the displacement or positional change of one or more visual entities present in the input video 301-A across consecutive frames. The method 500 may include determining the behavioural pattern associated with the one or more objects based on the determined motion of the one or more objects, and the contextual change in the environment of the input video 301-A based on the scene dynamics. In a non-limiting example, the behavioural patternmay correspond to the sequence or trend of motion exhibited by the one or more objects in the input video 301-A over time, which is indicative of the object’s activity or interaction within the scene. In a non-limiting example, the contextual change may correspond to variations in the environmental attributes of the input video that influence the interpretation of audio generation. In an embodiment, contextual change may include alterations in lighting conditions, background textures, object interactions, or spatial configurations within the scene. The method 500 may include determining the one or more temporal probabilistic audio timestamps based on the determined behavioral pattern and the contextual change.

[0097] In an embodiment, the method 500 may include determining the one or more probabilistic audio timestamps based on the temporal probability of the event based on predicting the plurality of probability parameters associated with the plurality of frames. In a non-limiting example, the temporal probability of the eventmay correspond to the computed likelihood that the specific audio-related event may occur at the given temporal instance within the input video. In a non-limiting example, the plurality of probability parameters may indicate two or more statistical values associated with the prediction of audio events for the input video 301-A. In an embodiment, the plurality of probability parameters may include mean, variance, standard deviation, and confidence scores derived from a probability density function.

[0098] In an embodiment, for generating the one or more audio-aware latent embeddings, the method 500 may include generating, using the video encoder, the video frame embedding for the input video 301-A. The method 500 may include generating the one or more audio-aware latent embeddings based on the one or more temporal probabilistic audio timestamps and the generated video frame embeddings. In a non-limiting example, the video frame embedding may correspond to a multi-dimensional feature representation generated from a single frame of the input video 301-A using the encoder.

[0099] In an embodiment, The method 500 may include training the ML model. In an embodiment, to train the ML model for estimating the temporal probabilistic audio timestamp, The method 500 may include receiving the input video 301-A comprising the plurality of frames encoded in RGB color format. The method 500 may include predicting the first set of probability density function parameters for the audio to be generated based on the input video 301-A. In a non-limiting example, the first set of probability density function parameters may indicate one or more probabilistic audio timestamps corresponding to the input video 301-A. In a non-limiting example, the first set of probability density function parameters may include at least the mean value and the variance value. In a non-limiting example, the mean value may indicate the central estimate of the temporal location of the audio event within the input video. In a non-limiting example, the variance value may indicate the measure of temporal uncertainty associated with the estimated temporal location. The method 500 may include receiving the two or more temporally synchronised ground-truth audios corresponding to the input video. In a non-limiting example, the two or more temporally synchronised ground-truth audios may indicate audio recordings that are perfectly aligned in time with the input video. The method 500 may include generating the one or more timestamps for the ground-truth audios based on the analysis of intensity levels in the two or more temporally synchronized ground-truth audios that exceed the predefined threshold. In a non-limiting example, the one or more timestamps may indicate instances of elevated acoustic activity corresponding to the input video. The method 500 may include obtaining the second set of probability density function parameters for the ground-truth audios through fitting the Gaussian distribution for each of the ground-truth audios based on the one or more timestamps. In a non-limiting example, the second set of probability density function parameters may indicate one or more probabilistic audio timestamps corresponding to the ground-truth audios. The method 500 may include training the ML model to estimate the temporal probabilistic audio timestamp based on the correlation of the first set of probability density function parameters and the second set of probability density function parameters.

[0100] In an embodiment, for determining the audio loudness parameter embedding, the method 500 may include extracting the sound source location based on the input video. In a non-limiting example, the sound source location may be the spatial region within the input video 301-A that corresponds to the origin of a sound-producing event. The method 500 may include generating the bounding box based on the extracted sound source location in the input video 301-A. In a non-limiting example, the bounding box may correspond to the rectangular or polygonal region generated around the sound source location within the video frame. In an embodiment, the bounding box may be used to isolate and track the motion of sound-producing object, enabling accurate computation of motion flow vectors and masking irrelevant background movements. The method 500 may include determining the motion flow vector for the extracted sound source location based on the input video 301-A. The method 500 may include determining the masked motion flow vector based on the motion flow vector and the bounding box for the extracted sound source location. In a non-limiting example, the masked motion flow vector may correspond to the average rate of change of the motion field in the extracted sound source location with time. The method 500 may include generating the loudness embedding vector based on passing the masked motion flow vector and the one or more predefined parameters, using the ML model. The method 500 may include determining the audio loudness parameter embedding based on the masked motion flow vector.

[0101] In an embodiment, prior to determining the melody variation factor, the method 500 may include determining the melody embedding based on the one or more audio-aware video embedding and the audio loudness parameter embedding.

[0102] In an embodiment, for generating the audio embedding, the method 500 may include receiving the one or more user defined melody parameters specifying the desired level of Gaussian noise. The method 500 may include generating the predicted melody embedding using the ML model based on the one or more user defined melody parameters. In a non-limiting example, the predicted melody embeddingmay correspond to the multi-dimensional feature representation generated by the ML model to encode melodic characteristics of audio to be synthesized. The method 500 may include modulating the predicted melody embedding based on the one or more user defined melody parameters. The method 500 may include generating the audio embedding, using the ML model, based on the modulated melody embedding and the non-melody embedding. In a non-limiting example, the non-melody embedding may correspond to a latent representation that captures audio characteristics other than melody, such as rhythm, timbre, loudness, and environmental sound features.

[0103] In an embodiment, the method 500 may include decoding the audio embedding using the audio decoder to reconstruct the output audio waveform. In a non-limiting example, the output audio waveform may correspond to a time-domain representation of synthesized audio signal generated by decoding the audio embedding using the audio decoder. In an embodiment, the output audio waveform may include amplitude variations over time that correspond to the audio-related characteristics, such as loudness, pitch, and melody, synchronized with the visual content of the input video 301-A.

[0104] In an embodiment, the method 500 may include training the ML model to generate the audio embedding. The method 500 may include decomposing the input audio signal into the melody component and the non-melody component using the pre-defined decomposition methods. In a non-limiting example, the melody component may correspond to a segment of the audio signal that encodes melodic characteristics such as pitch contour, harmonic progression, and rhythmic structure. In a non-limiting example, the non-melody component may correspond to a segment of the audio signal that represents non-melodic elements, such as background noise, percussion, or ambient sounds, which do not contribute to the primary melodic pattern. The method 500 may include training the ML model to generate the melody embedding based on the melody component, the non-melody component, and the input audio signal using one or more auto-encoders.

[0105] The present disclosure provides various advantages as mentioned below:

[0106] The present disclosure provides the system and method that enables generation of one or more audio tracks that are contextually relevant and temporally synchronized with input video content.

[0107] The present disclosure allows a tunable approach for synthesizing audio by incorporating user-controlled parameters for loudness and melody variations. Thus, the present disclosure ensures enhanced user experience by delivering audio outputs that are dynamically aligned with visual cues while offering flexibility for customization according to user preferences.

[0108] As would be apparent to a person in the art, various working modifications may be made to the method in order to implement the inventive concept as taught herein. The drawings and the foregoing description give examples of embodiments. Those skilled in the art will appreciate that one or more of the described elements may well be combined into a single functional element. Alternatively, certain elements may be split into multiple functional elements. Elements from one embodiment may be added to another embodiment. For example, orders of processes described herein may be changed and are not necessarily limited to the manner described herein.

[0109] Moreover, the actions of any signal flow diagram need not be implemented in the order shown; nor do all of the acts necessarily need to be performed. Also, those acts that are not dependent on other acts may be performed in parallel with the other acts.

Claims

1.An electronic device for generating an audio for an input video, the electronic device comprising:a memory configured to store at least one instruction;one or more processors;wherein the at least one instruction, when collectively or individually executed by the one or more processors, causes the electronic device to:determine one or more temporal probabilistic audio timestamps for the input video based on motion of one or more objects and scene dynamics in the input video;generate one or more audio-aware latent embeddings for a plurality of frames from the input video based on the determined one or more temporal probabilistic audio timestamps;determine an audio loudness parameter embedding based on motion characteristics of the input video and one or more predefined parameters, wherein the motion characteristics corresponds to a rate of change of motion of the one or more objects in the input video;determine a melody variation factor based on one or more user defined melody parameters and a melody embedding, wherein the melody embedding indicates melodic characteristics of the audio to be generated;generate, using a machine learning, an audio embedding based on the one or more audio-aware video embeddings, the audio loudness parameter embedding, and the melody variation factor; andgenerate, using the ML model, the audio for the input video based on the audio embedding.2.The electronic device as claimed in claim 1, wherein to determine the one or more temporal probabilistic audio timestamps for the input video, the at least one instruction, when collectively or individually executed by the one or more processors, causes the electronic device to:detect a motion of the one or more objects and the scene dynamics in the input video;determine a behavioural pattern associated with the one or more objects based on the determined the motion of the one or more objects indicates, and a contextual change in an environment of the input video based on the scene dynamics; anddetermine the one or more temporal probabilistic audio timestamps based on the determined behavioral pattern and the contextual change.3.The electronic device as claimed in claim 1, wherein to determine the one or more temporal probabilistic audio timestamps for the input video, the at least one instruction, when collectively or individually executed by the one or more processors, causes the electronic device to:determine the one or more probabilistic audio timestamps based on a temporal probability of an event based on predicting a plurality of probability parameters associated with the plurality of frames.4.The electronic device as claimed in claim 1, wherein to generate the one or more audio-aware latent embeddings, the at least one instruction, when collectively or individually executed by the one or more processors, causes the electronic device to:generate, using a video encoder, a video frame embedding for the input video; andgenerate the one or more audio-aware latent embeddings based on the one or more temporal probabilistic audio timestamps and the generated video fame embeddings.5.The electronic device as claimed in claim 1, wherein the one or more temporal probabilistic audio timestamps correspond to a probability metric indicative of an audio intensity level for each temporal instance throughout the input video.6.The electronic device as claimed in claim 1, wherein the melody variation factor indicates a variation in a melody for the audio, and wherein the melody embedding comprises at least one of a melody embedding and a non-melody embedding.7.The electronic device as claimed in claim 1, wherein the one or more processors are configured to train the ML model, wherein to train the ML model for estimating the temporal probabilistic audio timestamp, the at least one instruction, when collectively or individually executed by the one or more processors, causes the electronic device to:receive an input video comprising the plurality of frames encoded in Red, Green, and Blue color format;predict a first set of probability density function parameters for the audio to be generated based on the input video, wherein the first set of probability density function parameters indicates one or more probabilistic audio timestamps corresponding to the input video, wherein the first set of probability density function parameters comprising at least a mean value and a variance value, wherein the mean value indicates a central estimate of a temporal location of an audio event within the input video, and wherein the variance value indicates a measure of temporal uncertainty associated with the estimated temporal location;receive two or more temporally synchronised ground-truth audios corresponding to the input video, wherein the two or more temporally synchronised ground-truth audios indicate audio recordings that are perfectly aligned in time with the input video;generate one or more timestamps for the ground-truth audios based on an analysis of intensity levels in the two or more temporally synchronized ground-truth audios that exceed a predefined threshold, wherein the one or more timestamps indicate instances of elevated acoustic activity corresponding to the input video;obtain a second set of probability density function parameters for the ground-truth audios through fitting of a Gaussian distribution for each of the ground-truth audios based on the one or more timestamps, wherein the second set of probability density function parameters indicates one or more probabilistic audio timestamps corresponding to the ground-truth audios; andtrain the ML model to estimate the temporal probabilistic audio timestamp based on correlation of the first set of probability density function parameters and second set of probability density function parameters.8.The electronic device as claimed in claim 1, wherein to determine the audio loudness parameter embedding, the at least one instruction, when collectively or individually executed by the one or more processors, causes the electronic device to:extract a sound source location based on the input video;generate a bounding box based on the extracted sound source location in the input video;determine a motion flow vector for the extracted sound source location based on the input video;determine a masked motion flow vector based on the motion flow vector and the bounding box, for the extracted sound source location, wherein the masked motion flow vector corresponds to an average rate of change of motion field in the extracted sound source location with time;generate a loudness embedding vector based on passing the masked motion flow vector and the one or more predefined parameters, using the ML model; anddetermine the audio loudness parameter embedding based on the masked motion flow vector.9.The electronic device as claimed in claim 1, wherein prior to determining the melody variation factor, the at least one instruction, when collectively or individually executed by the one or more processors, causes the electronic device to:determine the melody embedding based on the one or more audio-aware video embedding and the audio loudness parameter embedding.10.The electronic device as claimed in claim 1, wherein to generate the audio embedding, the at least one instruction, when collectively or individually executed by the one or more processors, causes the electronic device to:receive the one or more user defined melody parameters specifying a desired level of Gaussian noise;generate a predicted melody embedding using the ML model based on the one or more user defined melody parameters;modulate the predicted melody embedding based on the one or more user defined melody parameters; andgenerate the audio embedding, using the ML model, based on the modulated melody embedding and a non-melody embedding.11.The electronic device as claimed in claim 10, the at least one instruction, when collectively or individually executed by the one or more processors, causes the electronic device to:decode the audio embedding using an audio decoder to reconstruct an output audio waveform.12.The electronic device as claimed in claim 1, wherein to train the ML model to generate the audio embedding, the at least one instruction, when collectively or individually executed by the one or more processors, causes the electronic device to:decompose an input audio signal into a melody component and a non-melody component using pre-defined decomposition methods; andtrain the ML model to generate the melody embedding based on the melody component, the non-melody component, and the input audio signal using one or more auto-encoders.13.A method for generating an audio for an input video, the method comprising:determining one or more temporal probabilistic audio timestamps for the input video based on motion of one or more objects and scene dynamics in the input video;generating one or more audio-aware latent embeddings for a plurality of frames from the input video based on the determined one or more temporal probabilistic audio timestamps;determining an audio loudness parameter embedding based on motion characteristics of the input video and one or more predefined parameters, wherein the motion characteristics correspond to a rate of change of motion of the one or more objects in the input video;determining a melody variation factor based on one or more user defined melody parameters and a melody embedding, wherein the melody embedding indicates melodic characteristics of the audio to be generated;generating, using a machine learning model, an audio embedding based on the one or more audio-aware media embeddings, the audio loudness parameter embedding, and the melody variation factor; andgenerating, using the ML model, the audio for the input video based on the audio embedding.14.The method as claimed in claim 13, for determining the one or more temporal probabilistic audio timestamps for the input video, the method comprises:detecting a motion of the one or more objects and the scene dynamics in the input video;determining a behavioural pattern associated with the one or more objects based on the determined motion of the one or more objects, and a contextual change in an environment of the input video based on the scene dynamics; anddetermining the one or more temporal probabilistic audio timestamps based on the determined behavioral pattern and the contextual change.15.The method as claimed in claim 13, for determining the one or more temporal probabilistic audio timestamps for the input video, the method comprises:determining the one or more probabilistic audio timestamps based on a temporal probability of an event based on predicting a plurality of probability parameters associated with the plurality of frames.