Method and system of foley tagging for animation
A machine learning-based system automates foley tagging in animations by predicting sound triggers and selections, addressing the challenge of manual frame-by-frame audio synchronization in games, enhancing efficiency and precision.
Patent Information
- Application Number
- GB2023016899
- Authority / Receiving Office
- GB · GB
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-11-03
- Publication Date
- 2025-05-07
AI Technical Summary
Game developers face the laborious task of manually checking animation frames to insert audio synchronously, which is exacerbated by animations generated on-the-fly in response to user input or physics simulations, making pre-checking and marking up animations difficult.
A method and system utilizing a machine learning model to associate moments in animations with sound triggers, based on supervised or unsupervised learning, using normalized visual representations and skeletal models to predict when and what sound to trigger, with optional audio activity detection for training.
Automates the process of foley tagging, ensuring accurate and efficient synchronization of audio with animations, adapting to various game states and animation types, reducing manual effort and improving synchronization precision.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
The present invention relates to a method and system of foley tagging for animation. Currently, game developers need to manually check an animation frame-by-frame to identify appropriate points at which to insert or synchronise audio (for example footsteps or exertion sounds). This is a laborious process. Furthermore, and unlike traditional animation, in a game an animation may be generated on-the-fly at runtime in response to user input, emergent events, and / or physics simulations, making it difficult for a developer to pre-check and mark up such animations in advance. The present invention sees to address of mitigate this problem. SUMMARY OF THE INVENTION Various aspects and features of the present invention are defined in the appended claims and within the text of the accompanying description. In a first aspect, a method of Foley tagging for animation is provided in accordance with claim 1. In another aspect, a Foley tagging system for animation is provided in accordance with claim 13. BRIEF DESCRIPTION OF THE DRAWINGS A more complete appreciation of the disclosure and many of the attendant advantages thereof will be readily obtained as the same becomes better understood by reference to the following detailed description when considered in connection with the accompanying drawings, wherein: Figure 1 is a schematic diagram of a Foley tagging system in accordance with embodiments of the present description. Figure 2 is a schematic diagram of frame timings with respect to animation in accordance with embodiments of the present description. Figure 3 is a flow diagram of a Foley tagging method in accordance with embodiments of the present description. DESCRIPTION OF THE EMBODIMENTS A method and system of foley tagging for animation are disclosed. In the following description, a number of specific details are presented in order to provide a thorough understanding of the embodiments of the present invention. It will be apparent, however, to a person skilled in the art that these specific details need not be employed to practice the present invention. Conversely, specific details known to the person skilled in the art are omitted for the purposes of clarity where appropriate. The problem of foley tagging involves two separate but related issues: when to trigger a sound within an animation, and what sound to trigger at that point. These issues are discussed in turn below. Timing Typically an animation sequence will depict an object (as a non-limiting example, a player or non-player character) performing a function, such as walking, running, jumping, swinging a weapon, or falling over. As noted previously, this animation sequence may be predefined, or driven algorithmically (for example at least in part by animation blending, or physics simulation). For example, different sized weapons may alter the duration of a weapon swing action, or different character types (or inventory loads) may affect the running gait and / or pace. Within a game, some foley control (i.e. decisions when to trigger audio) can optionally be determined from the game state; for example then a character is running, the character mesh intersects with the ground mesh, which can be used to trigger a footstep sound. Whilst this may be sufficient for footsteps specifically, more generally it has the disadvantage that it requires detecting in-world interactions, of which there may be a great many combinations. It also militates against foley sounds that are not triggered by such interactions, such as for example a character panting due to exertion, or a weapon swing that does not hit anything. Hence it is preferable that at least some foley control is based upon the animation of an object itself and hence not directly responsive to detecting in-world interactions - although these can still be covered where the interaction affects the animation. Accordingly, in an embodiment of the present description, a method of foley tagging for animation comprises the use of a machine learning model that has been trained to associate moments in animations with the triggering of sounds, in order to predict which point in an animation should be associated with the start of a foley sound effect. The machine learning model may use either supervised or unsupervised learning. In supervised learning, a machine learning model may be trained using labelled training data; in this case, the label would comprise at least an indication of when to trigger a sound; hence for example data representative of an animation sequence may be input and a predetermined frame / pose of that animation would be accompanied by such a label, or the label would change status (for example functioning as a flag for sound or no sound). The label or flag may thus be part of an output target for the machine learning model, so that during inference the trained model outputs a corresponding label or flag in response to appropriate frames or poses of input animation. Hence in this approach, the machine learning model learns correlations between animations and triggered sounds based upon an explicit indication of when such sounds should be triggered. In this case, the label could be manually added, or alternatively, and optionally depending upon the source material, an audio activity detector can monitor an associated soundtrack to detect the start of sounds and flag these as labels; in this way the machine learning model can be trained upon existing content (potentially both animated and live-action) to associate aspects of motion with the start of sounds. To assist the machine learning model, the inputs provided during training and subsequently inference can be optionally pre-processed as follows. In a first instance, a direct visual representation of the animation may be may be provided to the machine learning model; typically this is both computationally expensive in terms of input data and associated nodes required to handle it, and in terms of clarity of correlation between the input data and the target, and hence the time to convergence (adequate accuracy) during training. These issues can be mitigated in part by isolating the animation within a wider image, for example based on information within the game about the position of the object being animated, or based upon an indication from an animator, or selection of a relevant layer in an composite animated image, so that only a part of the scene comprising the relevant animation is used. In either case, the visual representation of the animation can be normalised in terms of scale (for example reduced in size to, say 64, 32,16, or another suitable number of pixels high) and also optionally in terms of aspect ratio (for example square, so as to produce for example 64x64, 32 x 32, or 16 x 16 input images). The animation can also be normalised in terms of brightness, and optionally in terms of colour, for example being converted to greyscale. These changes reduce incidental variabilities within the animation whilst retaining the movements that are likely to be indicative of when to trigger the sounds. Optionally, two such visual representations may be provided, one for a whole body animation, and one for a face; in this case the face is identified within the animation, and subject to a separate normalisation process at a different scale; resulting, as a non-limiting example, in a 32 x 32 normalised image of the whole body animation frame and a 32 x 32 normalised image of the face within from that same animation frame. The sizes need not be the same, but this provides scope for detecting key triggers corresponding with facial expression that may be lost in a whole-body representation of the object. Alternatively or in addition to a direct visual representation, information within a game about the object being animated may be used directly; for example a simplified representation of the object's mesh during the animation may be used, and / or a skeletal model may be used. Again optionally more detail may be retained for a face of the object. In the case of a skeletal model, if the game already utilises skeletal models then these may be used directly, but more commonly a skeletal model may be derived using known techniques, either from a visual representation of the animation, or from information within the game (e.g. the object's mesh). Typically the skeletal model will be adjusted to best fit the visible object or the mesh of the object, and after an initial match typically does so by modifying the skeletal model of the previously determined pose in the animation so that the model iterates with the animation as it progresses. Different skeletal models may be used for different types of object, such as for example bipedal humanoid, quadrupedal animal, and 1, 2, and / or 3 dimensional shape primitives (e.g. a point, a line, and a cube) for various objects such as bullets, swords, and vehicles. Hence inputs to the machine learning model may comprise a (typically pre-processed) representation of image data and / or a representation of the mesh or a skeletal model of an object, for each of successive moments within an animation of that object. It will be appreciated often sounds are triggered in response to moments of impact (whether a step or a punch) or in a predictable relationship to a characteristic action (for example a cough or sneeze, or panting). More generally, the sounds may be thought of as having a predictable relationship with points of inflection within the animation, particularly at extremities of the animation, and particularly in directions of travel (since interactions tend to occur in directions of progression). Accordingly, to facilitate a machine learning model correlating such aspects of the animation with flagged audio triggers, the input representation of the image data and / or representation of the mesh or skeletal model can optionally be supplemented with data indicative of such points inflection within the animation. Such data may include deltas of one or more various kinds, for example current image data (optionally as processed elsewhere herein) with preceding image data (e.g. corresponding to the image data input immediately prior) subtracted from it; this emphasises the points of change within the image. Similarly deltas for position data for an object mesh or skeletal model may be included. In these cases, there is a comparatively high likelihood of a correlation between animation and sound when such deltas reach zero, or change sign; a zero implies a pause in motion, whilst a sign change implies a reversal. Optionally such deltas may be calculated for the whole object at a given moment, and / or for specific parts of it (for example just the endpoints of the limbs in a skeletal model), or weightings may be given to deltas in certain parts of the animation, for example those in a predetermined portion of the image corresponding to an overall direction of travel (for example if the animation is moving to the right, the right-hand half, third, or quarter of the animation or animated object). Hence the machine learning model may receive as input generated image data (for example image data drawn by an animator, or rendered by a videogame) for each of successive frames of animation, optionally pre-processed to reduce one or more aspects of variability such as size, aspect ratio, luminance, and / or colour, and optionally relating to multiple views at optionally different scales; alternatively or in addition the machine learning model may receive as input generated image data for each of successive frames of animation comprising parametric representations of the animated object or objects, for example in the form of mesh data (optionally simplified mesh data) or skeletal model data (whether a simplified skeleton such as for example a humanoid with lower and upper jointed limbs, a torso and head, or a notional skeleton for an object such as a rock, sword, or car). Alternatively or in addition the machine learning model may receive as input one or more deltas such as data representative of the difference between successive generated images, or representative of the difference between parametric values representing the animated object or objects. Optionally input data for one form of representation may be accompanied by delta data for another form of representation; hence for example image data may be accompanied by a parametric representation of change with respect to the previous frame, or a skeletal model may be accompanied by an indication of difference between successive images. Meanwhile the machine learning model may receive as targets for output flag data corresponding towhen a sound is triggered, and optionally flag data relating to sound positioning; for example a value between 0 and 1 indicating left to right stereo positioning. Optionally, the machine learning model may also receive input data such as that described elsewhere herein for one or more preceding animation frames, so that the machine learning model can determine a correlation with the current frame within the context of one or more preceding frames; this may be of use for example when determining whether the current frame represents the last frame before a point of inflection, and hence the most appropriate with which to trigger a relevant sound. This may be of particular relevance within video games; referring to figures 2A and 2B, these show travel of a particular animation point on the y-axis, and time on the x-axis. Three animation frames capturing this travel are then illustrated by arrows evenly distributed along the time axis. Traditional animation may be depicted by figure 2A as neatly illustrating the point of maximum extension with the second frame, and in the circumstances a machine learning system can reliably learn to predict that the second frame in this sequence is the one likely to correlate with an audio trigger. By contrast referring to figure 2B, this may correspond to the frames output by a videogame, where the animation is triggered by in game events, user interactions, and / or physics interactions, and indeed the time between individual frames may vary (not shown here), and so it may not be guaranteed that any one frame neatly coincides with the point of maximum extension within the animation. Advantageously the machine learning model may generalise from such examples to predict the frame most likely to represent the last frame before a point of inflection and hence the most appropriate in which to trigger relevant sound. Optionally, the machine learning model may be trained to predict when, relative to such a frame, the point inflection would actually occur and indicate a relative timing, for example as a ratio of probabilities between the current frame and an expected next frame; hence for example in this case the machine learning model may be trained to output a flag indicating that this frame is the last frame before a point inflection, and output a second value between 0 and 1 predicting roughly where between this frame and the next frame the actual point of inflection would occur. Such an output may be used to more precisely time the triggering of a sound. Optionally such predictions may only be relevant if the frame rate falls below a predetermined lower limit such as for example 30 frames per second; 30 frames per second would correspond with roughly 33 ms per frame, and people typically start to detect acoustic delay or misalignment with differences between 15-30 ms; hence at such frame rate an output predicting whether to trigger a sound as above, or simply with the current frame or halfway between this frame and the next frame, may be sufficient to mask any misalignment between the generated image and when the user would expect an associated sound to be triggered. It will be appreciated that where animation frames are being tagged or associated with triggers for the generation of sound for use in subsequent output (in other words, not necessarily for live generation of audio in response to current rendering of animations), then optionally data relating to one or more animation frames in the future relative to the current animation frame can also be input so that the machine learning model has more complete (and non-causal) information particularly in relation to points of inflection within the animation. Hence the machine learning model may receive inputs of the type described elsewhere herein, optionally including one or more representations of one or more preceding animation frames, in addition to the current animation frame, and as applicable optionally one or more following / future animation frames. Meanwhile again the machine learning model may receive as targets for output flag data corresponding to when a sound is triggered, and optionally a further value indicating the relative timing between the current and next frame as to exactly when the sound should be triggered. Optionally, different machine learning models may be trained for different object archetypes such as humanoids, animals or particular classes of creature, vehicles, and the like; these may have different relationships between sound and movement; for example vehicles tend to have continuous sounds relating to continuous movement whereas humans and animals tend to have percussive sounds relating to movement, and the timing of these sounds may therefore differ with respect to the nature of the motion associated with animating these archetypes. Selection In some embodiments of the present description, an animation may be associated with just one sound previously selected or created for it, and so selection of that sound is by default when triggered. In other embodiments, a small set of sounds may be associated with the animation, for example for different stages of the animation, or for different applications of the animation. In this case, selection of the appropriate sound for the relevant stage and / or situation is required. In other embodiments, sounds are not a priori associated with the animation, but a library of sounds is available. In this case, choosing a sound to associate with the animation (and subsequently select as per the embodiments above) is required. Where a sound is selected by default, there is no need for further functionality of the machine learning model (or a further model). Where a small set of sounds are associated with the animation, optionally these may be respectively associated with respective key frames of the animation; consequently the sound associated with the key frame closest to the actual frame / time identified for triggering a sound can be selected as the triggered sound; this approach simplifies the process of manually associating sounds to an animation whilst ensuring that variability in the animation, for example due to frame generation rates, or modifications such as changes in terrain changing the cycle of a walk or run animation, can be automatically accommodated. In other words, the animator can indicate roughly where an animation cycle a sound would occur, and the machine learning system and the subsequent selection process then identify exactly where the sound should occur in the current instance of the animation. In principle this approach can also be applied to ragdoll physics and procedural / physics driven animation, where the animator can identify poses that should be associated with a sound effect, and the machine learning system and substance that can process can identify moments in the animation sequences corresponding to such poses and so trigger an associated sound effect. Where no sound (or an incomplete set of sounds) is a priori associated with the animation, or multiple possible sounds are associated with the animation, then choice of a sound is required. The choice may be from a full library of sounds, from a small set of sounds associated with the animation but not with a specific frame or key-frame, or from alternative equivalent sounds, whether or not associated with specific frame or key-frame. To take the example of a character walking, a full library of sounds may comprise a wide variety of audio including for example vehicle noises, birdsong, gunfire, and footsteps. Meanwhile a small set of sounds already associated with an animation may be treated as a small library. As described elsewhere herein, the machine learning model may be trained on example animations (both literal animation and / or potentially live action) that has been marked up to indicate when sounds occur. Optionally such example animations may also be marked up to indicate what sounds occur. This may be done using a consistent classification scheme. The classification scheme may for example be organised hierarchically so that if the audio relates to gunfire the label may state 'gunfire' and also a particular type of gunfire such as 'shot gun', thus providing both a general and specific classification. Similarly if the audio relates to a footstep, the label may state 'footstep', and optionally a particular type of footstep such as light / heavy or walk / run, and similarly optionally a surface characteristic such as a footstep on metal, concrete, gravel, forest floor, or the like; however, this latter surface characteristic may be of less use in relation to association with an animation, as the animation itself may be unaffected by the surface upon which it is placed; however this information (obtained from the game state or other labels in the animation) may be used for example to separately select a subset of sounds from which the machine learning model can select; hence for example if the machine learning model is not trained on surface characteristics, and so generates an output indicating footstep / heavy, the game may still have heavy and light instances of a footstep sound on the current relevant surface e.g. (a metal surface), and these may form the effective library from which to select the footstep / heavy audio indicated by the machine learning model. In this way a game (or more generally a foley system for a game or other animation) can define a subset of the audio library based upon the game state or animation tags, and the machine learning model can identify the relevant sound response to the current state of the animation within that subset of the audio library. Hence in embodiments of the present description, the machine learning model may be trained using such classification labelling as a target output. Typically a literal text output would not be efficient, and likely to produce inaccuracies, and so an alternative representation may be used; for example several output nodes each with ranges between 0 and 1 (or any suitable value depending on the operation of the machine learning model); optionally different nodes can indicate the presence of or choice between different sound classifications, and optionally some nodes can indicate different sound types (for example light / heavy, walk / run, and pistol / shot gun can all be thought of small / large examples of a particular class of sound). Meanwhile, a hierarchical classification of sounds in turn lends itself to a hierarchical representation using such nodes; hence for example a first classification may be tonal versus noise based, which would separate for example birdsong from gunfire. Within the noise based sub-classification, a second classification may be continual versus plosive, which would separate vehicle noise from gunfire. Within the plosive based sub classification, a third classification may be between gunfire and other noises such as raindrops, and so on; such a representation lends itself to a binary representation with the first classification being the most significant bit and the last classification been the least significant bit; the machine learning model can then be trained to output the relevant binary representation of the hierarchical classification / label for a sound. It will be appreciated however that any suitable representation of a classification of a sound may be used as the target output for the machine learning model. At run time, the machine learning model will then output suitable values corresponding to a classification of a sound (optionally this may double as the flag indicating the trigger point for such a sound). The system can then access a sound best matching the output classification from the current library of sounds (whether this is the full library, a level, region, or surface specific subset of the library, a character specific subset of the library, and / or an animation specific subset of the library). In this way optionally the machine learning model can indicate what sound(s) to use in an animation sequence as well as when to use it / them. In any event, the resulting output of the method is that a sound effect trigger and optionally a sound effect ID are associated with a relevant point in the animation. This in turn may be used to control output of the animation sequence and appropriate sound effects for reproduction; in the case in which the reproduction is to be output at a later date, the sound effect trigger and optionally the sound effect ID can be associated with the animation in any suitable way; for example in the form of meta data associated with the animation, or in the form of a soundtrack synchronised with the animation (in other words, implementing the triggering of the sound and recording the result in association with the animation). In the case in which the animation is being output live (for example in the case of a videogame), the trigger and optionally the sound effect ID are used to start reproduction of the sound effect for output as they occur. Training As noted elsewhere herein, training can be provided using labelled examples of animation. Also as noted herein 'animation' for the purposes of training can include live action video sequences; the machine learning model is trained to develop a correlation between movement and sound, and this can be discerned both from animation and live-action footage. Hence the term animation herein can incorporate live-action video unless explicitly indicated otherwise. The labels can be provided manually, and in some existing animation training examples may already be available; audio synchronisation markers may already exist, as may other foley data such as an identification of the relevant sound. Hence as noted elsewhere herein, the labels may only indicate timing, and / or may also indicate the sound used at that point in the animation, whether as a text label, a hierarchical representation of sound type, or any other suitable representation, such as a pointer to a relevant sound library for that animation or animation type. The labels can also be generated automatically, using audio activity detectors to detect the onset of audio during examples of animation; in this case it may be that an individual audio track within an audio mix for the animation is suitable for this purpose; that is to say the full audio mix for the animation may include music and other ambient sound, but typically one track exists for relevant foley effects, or the original recording of the live-action, in which the sounds correspond to the physical movement within the animation / live-action. Again the automatically generated labels may only indicate timing, and / or may also indicate direct properties of the sound used at that point in the animation, such as predominantly tonal or noise (and if tonal, optionally a dominant frequency), an energy envelope indicating the volume shape and duration of the sound, and / or a compressed representation of the sound such as a mel-Cepstrum of the sound (a representational format often used in voice recognition), with the number of Cepstrum bins and the time resolution of these chosen to provide a compact representation that may still be useful for determining close matches with a library of sounds for which similar mel-Cepstrum representations have been generated. The machine learning model is trained using input as discussed elsewhere herein, and as the target output labels such as those described herein. Where the machine learning model is trained to output labels that identify sounds, separately it may also still output a flag to indicate the probability that a current animation frame triggers such a sound. In addition, optionally the machine learning model may also be trained to output a separate flag, indicating the probability that a given animation frame does not trigger a sound. The or each of these flags helps to disambiguate label outputs in the event that the machine learning model outputs some noise the node(s) used to identify sounds; the output of the machine learning model may be interpreted as positively identifying a sound only when a flag indicating the sound is being triggered, or conversely when a flag does not indicate that a sound is not triggered. In either case, the machine learning system thus has separate mechanisms by which to indicate timing and sound type. Summary Turning now to figure 3, in a summary embodiment of the present description, a method of foley tagging for an animation (whether a little animation or live-action footage) comprises the following steps. In a first step, s310, providing a representation of at least part of the animation as input data to a machine learning model, as described elsewhere herein. Also as described herein, the machine learning model will have have been trained with corresponding input representations of animations and target output representations of sound effect timings, to associate moments in animations (whether on frame-wise basis or continual time basis) with the triggering of sounds so as to predict whether a point in an animation (again frame-wise or optionally between frames) should be associated with the start of a sound effect. In a second step, s320, associating the output representation of sound effect timing from the machine learning model with the corresponding point in the animation as a sound effect trigger, as described elsewhere herein; hence for example on a frame wise basis, and the output of the machine learning model is a flag, then when the machine learning model outputs a flag indicating a sound effect would start here (i.e. should be triggered), then a corresponding representation of this is associated with the current frame of the animation input to the machine learning system (or where data corresponding to multiple frames are input then the focal frame, i.e. the one for which other frames are notionally preceding or following). It will be apparent to a person skilled in the art that variations in the above method corresponding to operation of the various embodiments of the apparatus as described and claimed herein are considered within the scope of the present invention, including but not limited to that: the machine learning model has been trained using input data representative of animation sequences, and target data indicating when sound effects are triggered in the animation sequences, as described elsewhere herein; input data to the machine learning model representative of an animation comprises one or more selected from the list consisting of a normalised visual representation of the animation (e.g. to reduce variability between similar animations, and typically to reduce the dimensionality of the input), separate inputs for different views of the animation (e.g. a whole-body view and a face view, typically taken from the same source animation frame), mesh data defining a pose of an animated object; and skeletal model data defining a pose of an animated object, as described elsewhere herein; input data to the machine learning model representative of an animation comprises data representative of one or more differences between successive animation frames, as described elsewhere herein; in this instance, optionally the input data is indicative of points inflection within the animation, and / or indicative of deltas of one or more inputs to the machine learning model (whether or not these inputs are actually provided to the machine learning model; as noted elsewhere herein, optionally only deltas for some input types may be provided), as described elsewhere herein; input to the machine learning model comprises input data for a current animation frame and one or more selected from the list consisting of one or more proceeding animation frames, and one or more following animation frames, as described elsewhere herein; target output representations of the machine learning model comprise data identifying the sound effect starting at the time indicated by the output representations of sound effect timings (e.g. identifying text labels, sound library addresses, sound classes or types, or representations within a sound classification tree), as described elsewhere herein; as similarly described elsewhere herein, the data identifying the sound effect may not uniquely identify the sound effect but may identify a group of sounds that operate as alternatives to each other dependent upon other data, such as game state data or other labelling of the animation, such as for example the surface with which the animation is currently interacting; target output representations of the machine learning model comprise data representing one or more properties of the sound effect starting at the time indicated by the output representations of sound effect timings (e.g. the sound amplitude envelope, a dominant frequency - optionally as a function of time over the course of the sound, and / or a representation of the structure of the sound such as a cepstrum or mel-cepstrum representation), as described elsewhere herein; the method comprising the step of selecting from among a plurality of machine learning models, each trained on animation relating to an object of a particular archetype, a respective machine learning model to predict whether a point in a current animation should be associated with the start of a sound effect, based on the archetype of an object in the current animation, as described elsewhere herein; where each of a plurality of sounds have been associated with respective keyframes in an animation, the method comprising the steps of identifying the respective keyframe closest to the output sound effect timing, and selecting the sound associated with the identified keyframe as the one to be triggered at the output sound effect timing, as described elsewhere herein; and the method comprising the step of selecting a sound to associate with a sound effect trigger based upon data indicative of what the animated object is interacting with (e.g. a walking surface, or an object to be picked up or otherwise interacted with), as described elsewhere herein. It will be appreciated that the above methods may be carried out on hardware suitably adapted as applicable by software instruction or by the inclusion or substitution of dedicated hardware. Thus the required adaptation to parts of an equivalent device may be implemented in the form of a computer program product comprising processor implementable instructions stored on a non-transitory machine-readable medium such as a floppy disk, optical disk, hard disk, solid state disk, PROM, RAM, flash memory or any combination of these or other storage media, or realised in hardware as an ASIC (application specific integrated circuit) or an FPGA (field programmable gate array) or other configurable circuit suitable to use in adapting the conventional equivalent device. Separately, such a computer program may be transmitted via data signals on a network such as an Ethernet, a wireless network, the Internet, or any combination of these or other networks. Accordingly, and referring now also to figure 1, in a summary embodiment of the present description, a foley tagging system (10) may for example be part of an animation suite or video editing suite, in the case of traditional animation or live-action editing, or may be part of a videogame console or computer, and where configured by software instruction, the software may be integral to a videogame (for example provided as part of a development kit for incorporation into the game), or as a helper application associated with the operating system of consular computer (and for example in communication with a video game via an application program interface). The foley tagging system (10) then comprises the following. Firstly, an animation pre-processor (10) adapted to provide a representation of at least part of the animation is input data to a machine learning model. The pre-processor may for example be a graphics processing unit generating animation (for example in the case of a video games machine), or an image processing unit receiving (or generating) images from an animation, or receiving live-action footage (whether live or recorded). Secondly, a machine learning processor (30) configured to implement a machine learning model, having been trained with corresponding input representations of animations and target output representations of sound effect timings, to associate moments in animations with the triggering of sounds so as to predict whether a point in an animation should be associated with the start of a sound effect. The machine learning processor may be specifically configured to run a suitable neural network (e.g. a neural network accelerator chip or similar), or may be implemented in one or more parallel shaders within a GPU as part of a rendering pipeline to output sound cues as part of the animation process, or may be one or more CPU cores, operating under suitable software instruction as described elsewhere herein. Thirdly, a tagging processor (40) adapted to associate the output representation of sound effect timing from the machine learning model with the corresponding point in the animation as a sound effect trigger, as described elsewhere herein. This processor may optionally also associate a sound effect ID with the animation, and further optionally function to select a sound effect according to the techniques described elsewhere herein, and then associate the sound effect ID with the animation. The tagging processor may associate the sound with the animation by use of meta data, or by triggering the reproduction of the relevant sound for recording in parallel with the animation, or by triggering the reproduction of the relevant sound for life output in parallel with generated animation, for example in the case of a video game. Instances of this summary embodiment implementing the methods and techniques described herein (for example by use of suitable software instruction) are envisaged within the scope of the application, including but not limited to that: the Foley tagging system comprises a machine learning selection processor (not shown), adapted (again for example by suitable software instruction) to select from among a plurality of machine learning models, each trained on animation relating to an object of a particular archetype, a respective machine learning model to predict whether a point in a current animation should be associated with the start of a sound effect, based on the archetype of an object in the current animation, as described elsewhere herein; and the Foley tagging system comprises a sound selection processor (whether separate to or part of the tagging processor, as described previously herein) adapted (for example by suitable software instruction) to perform one or more selected from the list consisting of identifying a respective keyframe closest to the output sound effect timing and selecting a sound previously associated with the identified keyframe as the one to be triggered at the output sound effect timing; and selecting a sound to associate with a sound effect trigger based upon data indicative of what the animated object is interacting with, as described elsewhere herein. The foregoing discussion discloses and describes merely exemplary embodiments of the present invention. As will be understood by those skilled in the art, the present invention may be embodied in other specific forms without departing from the spirit or essential characteristics thereof. Accordingly, the disclosure of the present invention is intended to be illustrative, but not limiting of the scope of the 5 invention, as well as other claims. The disclosure, including any readily discernible variants of the teachings herein, defines, in part, the scope of the foregoing claim terminology such that no inventive subject matter is dedicated to the public.
Claims
1. A method of foley tagging for an animation comprises the steps of:providing a representation of at least part of the animation as input data to a machine learning model;the machine learning model, having been trained with corresponding input representations of animations and target output representations of sound effect timings, to associate moments in animations with the triggering of sounds so as to predict whether a point in an animation should be associated with the start of a sound effect; andassociating the output representation of sound effect timing from the machine learning model with the corresponding point in the animation as a sound effect trigger.
2. The method of claim 1, in which:the machine learning model has been trained using input data representative of animation sequences, and target data indicating when sound effects are triggered in the animation sequences.
3. The method of claim 1 or claim 2, in which:input data to the machine learning model representative of an animation comprises one or more selected from the list consisting of:i. a normalised visual representation of the animation;ii. separate inputs for different views of the animation;iii. mesh data defining a pose of an animated object; andiv. skeletal model data defining a pose of an animated object.
4. The method of any preceding claim, in which:input data to the machine learning model representative of an animation comprises data representative of one or more differences between successive animation frames.
5. The method of claim 4, in which:input to the machine learning model comprises one or more selected from the list consisting of:i. input data indicative of points inflection within the animation; andinput data indicative of deltas of one or more inputs to the machine learning model.
6. The method of any preceding claim, in which:input to the machine learning model comprises input data for a current animation frame and one or more selected from the list consisting of:i. one or more proceeding animation frames; andii. one or more following animation frames.
7. The method of any preceding claim, in which:target output representations of the machine learning model comprise data identifying the sound effect starting at the time indicated by the output representations of sound effect timings.
8. The method of any preceding claim, in which:target output representations of the machine learning model comprise data representing one or more properties of the sound effect starting at the time indicated by the output representations of sound effect timings.9.The method of any preceding claim, comprising the step of:selecting from among a plurality of machine learning models, each trained on animation relating to an object of a particular archetype, a respective machine learning model to predict whether a point in a current animation should be associated with the start of a sound effect, based on the archetype of an object in the current animation.
10. The method of any preceding claim, in which each of a plurality of sounds have been associated with respective keyframes in an animation, the method comprising the steps of:identifying the respective keyframe closest to the output sound effect timing; andselecting the sound associated with the identified keyframe as the one to be triggered at the output sound effect timing.
11. The method of any preceding claim, comprising the steps of:selecting a sound to associate with a sound effect trigger based upon data indicative of what the animated object is interacting with.
12. A computer program comprising computer executable instructions adapted to cause a computer system to perform the method of any one of the preceding claims.
13. A foley tagging system for an animation, comprising:an animation preprocessor adapted to provide a representation of at least part of the animation as input data to a machine learning model,a machine learning processor configured to implement a machine learning model, having been trained with corresponding input representations of animations and target output representations of sound effect timings, to associate moments in animations with the triggering of sounds so as to predict whether a point in an animation should be associated with the start of a sound effect; anda tagging processor adapted to associate the output representation of sound effect timing from the machine learning model with the corresponding point in the animation as a sound effect trigger.
14. The foley tagging system of claim 13, comprising:a machine learning selection processor, adapted to select from among a plurality of machine learning models, each trained on animation relating to an object of a particular archetype, a respective machine learning model to predict whether a point in a current animation should be associated with the start of a sound effect, based on the archetype of an object in the current animation.
15. The foley tagging system of claims 13 or 14, comprising a sound selection processor adapted to perform one or more selected from the list consisting of:i. identifying a respective keyframe closest to the output sound effect timing, and selecting a sound previously associated with the identified keyframe as the one to be triggered at the output sound effect timing; andii. selecting a sound to associate with a sound effect trigger based upon data indicative of what the animated object is interacting with.
Citation Information
Patent Citations
Method and system for determining identifiers for tagging video frames with
US20200167984A1
Mapping visual tags to sound tags using text similarity
US20200349387A1
Machine-learning Models for Tagging Video Frames
US20220254083A1