Fast grain switching for granular synthesis

The method of continuously evaluating target positions for fast grain switching in granular synthesis addresses the challenge of smooth transitions between grains, ensuring natural sound quality and dynamic adaptation in audio rendering.

WO2025149299A1PCT designated stage expired Publication Date: 2025-07-17TELEFONAKTIEBOLAGET LM ERICSSON (PUBL)
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2024/086427
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-12
Filing Date
2024-12-16
Publication Date
2025-07-17

AI Technical Summary

Technical Problem

Existing granular synthesis methods struggle with achieving fast and smooth transitions between grains, especially when using longer grains, leading to potential discontinuities and unnatural sound outputs due to delayed grain switching.

Method used

A method for granular synthesis that involves continuously evaluating an updated target position in a descriptor space to determine if a fast grain switch condition is satisfied, allowing for optimized grain transitions without introducing phase cancellations, even with longer grains.

Benefits of technology

Enables seamless and dynamic audio rendering by allowing fast grain switching, maintaining natural sound quality while adapting to real-time changes in the audio environment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2024086427_17072025_PF_FP_ABST
    Figure EP2024086427_17072025_PF_FP_ABST
Patent Text Reader

Abstract

A method for rendering audio corresponding to an audio recording divided into a plurality of grains. The method includes selecting a first grain from the plurality of grains. The method also includes rendering at least a first portion of the first grain. The method also includes, while the first grain is being rendered, obtaining information indicating an updated target position in a descriptor space and determining, based on the updated target position, whether a fast grain switch condition is satisfied. The method also includes, as a result of determining that the fast grain switch condition is satisfied, transitioning from the first grain to a second grain.
Need to check novelty before this filing date? Find Prior Art

Description

FAST GRAIN SWITCHING FOR GRANULAR SYNTHESISTECHNICAL FIELD

[0001] Disclosed are embodiments related to granular synthesis.BACKGROUND

[0002] Audio rendering is a process used for presenting audio, such as audio within an extended reality (XR) scene, such as, for example, a virtual reality (VR), an augmented reality (AR) scene, or mixed reality (MR) scene, in order to give a listener the impression that sound is coming from physical sources within the scene at a certain position. The presentation can be made through headphone speakers or other speakers. If the presentation is made via headphone speakers, the processing used is called binaural rendering and uses spatial cues of human spatial hearing that make it possible to determine from which direction sounds are coming. The cues involve inter-aural time delay (ITD), inter-aural level difference (ILD), and / or spectral difference.

[0003] Procedural audio refers to the creation of sound in real-time as a response to live input. As an example, consider the sound of a car engine in a virtual space where the sound changes based on the speed or acceleration or the car. This mechanism is commonly used in video games for better user experience. It is believed that for the use case of XR (e.g., AR or VR), there are many sounds that would benefit from being dynamically generated so that they can react to changes in the scene in real-time. For example, the sound generated when a user touches a surface or operates an engine. In reference [1], sounds of a virtual object, for example a sword, axe, or wand, are simulated based on their position and orientation. There is a base tone and an overtone. Both are modulated to change pitch, timbre, amplitude to convey speed of movement of the virtual object. The live input may come from a user via sensors, such as hand controllers or a headset, it could be control data generated in real-time by some software process such as a physics simulation or pre-defined automation data. Regardless of how the input data was generated, the audio Tenderer needs to handle incoming data and generate sound in response to this data in real-time.

[0004] There exist many different methods for procedural audio (see, e.g., reference[2]), including synthetic sound synthesis using audio processing modules, machine learning methods trained on real recordings, and concatenative synthesis methods that make use of original recordings and rearrange segments of these recordings to generate variations. Because the class of concatenative synthesis methods makes direct use of real recordings, the generated audio sounds very natural thereby enhancing user experience.

[0005] Granular synthesis is a type of concatenative synthesis where a sound recording is divided into small fragments called “grains.” (See, e.g., reference [3]). By a careful selection of the fragments (grains) at rendering time, a plausible dynamically changing sound can be generated.

[0006] A granular synthesis process includes two main steps: (1) grain extraction and(2) grain synthesis. Grain extraction refers to extraction of pertinent grains from the original longer recording. The extraction method depends on the type of sound source and the desired features to be extracted. Grain synthesis refers to the technique of selecting the appropriate order of grains; this selection of the ordering could also be based on the user input in real-time.

[0007] Many sound design tools support grain extraction and synthesis, such as, for example Soundseed grain for Audiokinectic Wwise, Alchemy for Logic Pro or AudioMotors for FMOD. The tools allow for manual or semi-automated extraction of grains by the sound designer and other simple manipulations. The designer can choose the grain length, the amplitude envelope or shape of each grain among other controls. In the case of AudioMotors, an automated grain extraction tool specialized for motor sounds is provided.

[0008] Grain extraction can be done manually by the sound designer or in a data-driven manner by identifying the relevant features of the audio for segmentation purposes. Relevant features include, for example, pitch period in the case of pitched sounds, spectral energy at a given frequency, mel-frequency cepstrum coefficients (MFCC), and local maxima of the amplitude envelope.

[0009] There are also many methods for granular synthesis. A common method is to select grains at random and perform overlap-and-add (OLA or overlap-add) operations. This method is not amenable to all types of sound sources and does not capture temporal correlation between adjacent grains.

[0010] Corpus-based concatenative synthesis (CBCS) methods are based on selecting grains from a corpus of sound segments that are sampled from a database of heterogeneous sound sources. They utilize descriptors that are associated to sound segments to organize the corpus and perform searches within the descriptor space to pick the next grain. Note that the concept of a descriptor is not limited to features of audio signal (see, e.g., reference [4]). A user can annotate grains with perceptual descriptors when a direct mapping between desired effect and feature in the audio signal is not possible.

[0011] The descriptor space is multi-dimensional with the number of dimensions being equal to the number of descriptors. Search for the appropriate grain is performed in a computationally efficient manner by utilizing weighted Euclidean distance between a target descriptor location (e.g., point or area) in the descriptor space (hereafter referred to as “target descriptor coordinate” or “target position”) and grain locations in the descriptor space.Reference [5] proposes warping functions for the distance measure to better select the set of grains and also to avoid repetitions of previously rendered grains.

[0012] For efficient search in the descriptor space, kD-tree search is used. Either k- nearest neighbors of the target descriptor coordinate or grains that are within a radius ‘r’ from the target descriptor coordinate are chosen. In reference [6], the corpus is organized as zones so that grains from different zones are not picked consequently when k-nearest neighbor search is used.

[0013] The software CATERPILLAR (see reference [7]) performs concatenative synthesis in an offline setup where a sequence of target descriptors is given. The program uses Viterbi algorithm to identify the sequence of grains to match the target descriptors. The cost function is a combination of distance from target descriptor coordinate and concatenation cost which is based on similarity of consecutive grains.

[0014] CataRT on the other hand is a real-time system and so it chooses the subsequent grain at random from a set of grains that are the k-nearest neighbors or a radius with the target descriptor coordinate being the center (see, e.g., reference [8]).

[0015] For smoother transitions in granularly synthesized sound, reference [9] uses feature descriptors like pitch, loudness, spectral centroid, fundamental frequency, periodicity, and autocorrelation coefficient at lag 1. Feature descriptors are computed for every grain andcorrelation among these feature descriptors is captured using a Gaussian Mixture Model (GMM) from which grains are sampled for synthesis.

[0016] There are also several works, such as, for example, reference

[0010] , that model or assign transition probability between adjacent grains thereby establishing Markov model to generate or synthesize sound textures. Model parameters are estimated using recordings, but synthesis of sound does not directly use recordings unlike granular synthesis.

[0017] Reference

[0011] discusses granular synthesis where the next grain is picked based on feature descriptors of the current grain for a continuity in timbre. A kD-tree search is performed to select the candidate grains closest in Euclidean distance to the current grain in the feature descriptor space.

[0018] When the corpus of sounds is sparse, reference [9] proposes corpus extension methods based on the feature descriptors. These techniques fall under the category of Feature Modulation Synthesis (FMS) (see, e.g., reference

[0015] ), that identify the appropriate transformation to apply to an audio based on the target feature descriptor. Specific methods include pitch shifting, gain adjustment and the use of filters. The methods can be quite involved with several steps and have to be performed offline, i.e. before granular synthesis or rendering

[0019] Granular synthesis typically uses overlap-add to produce a smooth output with a succession of grains without artefacts coming from sudden changes or discontinuities between the grains. The overlap-add is done in a way such that the perceived loudness is preserved within the overlap regions of consecutive grains. If not, the rendering will sound uneven, chopped up or just fluctuate in loudness.

[0020] If consecutive grains are highly correlated and mostly in-phase with each other, then a mix of the two grains will add up linearly, i.e., if two grains with the same loudness are given the gain 0.5, the combined mix will have the same loudness as the two individual grains. In this case, the two grains add constructively to each other and the gains of the two grains should be set so that total gain adds up to 1.0. In this case, the mixing follows a linear mixing rule where the mix is a linear combination of the mixed grains.

[0021] If, however, two grains, a first grain and a second grain, are largely uncorrelated, then the mix of the first and second grains will not add up linearly. This is because many of the samples of the first grain will have a different sign than the samples of the second grain with which they are being combined and when added will result in a lower amplitude . In this case, the two grains do not always add constructively to each other. In order to preserve the same perceived loudness of the mix as the two grains have, the mix should follow a different mixing rule where the gain of each mixed grain is set so that the signal power is preserved: PMIX = PI = P2. The power of a signal is proportional to the squared amplitude, which means that the gains should be set according to: (AMIX)2= (gjA-2+ (g2A2)2, where Ai and A2 are the amplitudes of the signals of the two grains to mix, and gi and g2 are the gains used when doing the mix. If the mix has the same power as signal 1 and 2 it follows that AMIX, AI and A2 are the same and therefore the gains should satisfy: (gj)2+ (g2)2= 1 For example, in the case where gl and g2 are equal, they should be set to ^1 / 2.

[0022] To achieve a smooth overlap-add result, the window function used needs to be selected according to the character of the sound to be rendered. For a sound where grains are correlated, an amplitude preserving window should be used and for a sound where grains are uncorrelated, a power preserving window should be used.

[0023] It is not only during the overlap-add processing where the mixing rule can have an impact. Also whenever doing some other form of mixing of grains, the same mixing rule would apply. For example, when the Tenderer performs interpolation between grains by doing a weighted mix of two or more grains.

[0024] Utilizing the correlation coefficient between two grains, the work described in reference

[0016] designs custom or analytical window functions for power preservation because it may be difficult to make a decision on the level of correlatedness using a threshold correlation coefficient value.SUMMARY

[0025] Certain challenges presently exist. For example, with granular synthesis, a dynamic change in sound is achieved by progressively selecting grains that best describe the wanted changes in sound output, and, when one grain is close to be completely rendered, thenext grain is selected and an overlap-add between the two grains is performed so that a seamless progression from grain to grain is achieved. In some cases, however, the grains are long and waiting to switch grains until the currently playing grain has been completed will result in a slow update where the target descriptor space coordinates may have changed greatly during the time a grain is completely rendered. This may lead to a generated sound output with large stepwise changes in character rather than a fast and smooth change. Using shorter grains would improve the dynamic response, but this is not always desirable since longer grains are sometimes needed to produce a natural sounding output.

[0026] Accordingly, in one aspect there is provided a method for rendering audio corresponding to an audio recording, wherein the audio recording is divided into a plurality of grains. The method includes selecting a first grain from the plurality of grains. The method also includes rendering at least a first portion of the first grain. The method also includes, while the first grain is being rendered, obtaining information indicating an updated target position in an N- dimensional, ND, descriptor space and determining, based on the updated target position, whether a fast grain switch condition is satisfied. The method also includes, as a result of determining that the fast grain switch condition is satisfied, transitioning from the first grain to a second grain.

[0027] In another aspect there is provided an apparatus that is configured to perform a method for rendering audio corresponding to an audio recording, wherein the audio recording is divided into a plurality of grains. The method includes selecting a first grain from the plurality of grains. The method also includes rendering at least a first portion of the first grain. The method also includes, while the first grain is being rendered, obtaining information indicating an updated target position in an ND descriptor space and determining, based on the updated target position, whether a fast grain switch condition is satisfied. The method also includes, as a result of determining that the fast grain switch condition is satisfied, transitioning from the first grain to a second grain. The apparatus may include memory and processing circuitry coupled to the memory.

[0028] In another aspect there is provided a computer program comprising instructions which when executed by processing circuitry of an apparatus causes the apparatus to perform any of the methods disclosed herein. In one embodiment, there is provided a carrier containingthe computer program wherein the carrier is one of an electronic signal, an optical signal, a radio signal, and a computer readable storage medium.

[0029] An advantage of the embodiments is that they enable fast grain transitions even if the grain currently being rendered is long. By continuously evaluating a fast-switching criterion, an optimized decision can be made when to perform a fast grain switch. The fasttransition can be performed without introducing problems with phase-cancellations with pitched sounds.BRIEF DESCRIPTION OF THE DRAWINGS

[0030] The accompanying drawings, which are incorporated herein and form part of the specification, illustrate various embodiments.

[0031] FIG. 1 illustrates a system according to an embodiment.

[0032] FIG. 2A illustrates an example two-dimensional descriptor space.

[0033] FIG. 2B illustrates an example two-dimensional descriptor space.

[0034] FIG. 3 A illustrates an example two-dimensional descriptor space.

[0035] FIG. 3B illustrates a process for determining an optimal k value for a given grain according to an embodiment.

[0036] FIG. 4A illustrates an of a descriptor trajectory of an original recording.

[0037] FIG. 4B illustrates an example of a descriptor trajectory of an original recording and a generated descriptor trajectory.

[0038] FIG. 4C illustrates an example of a descriptor trajectory of an original recording and a generated descriptor trajectory.

[0039] FIG. 5 is a flowchart illustrating a process according to an embodiment.

[0040] FIG. 6 is a flowchart illustrating a process according to an embodiment.

[0041] FIGS. 7A and 7B show a system according to some embodiments.

[0042] FIG. 8 is a block diagram of an apparatus according to some embodiments.

[0043] FIG. 9 illustrates a 2D descriptor space where the grains in GDB 104 do not evenly cover the descriptor space

[0044] FIG. 10 is a flowchart illustrating a process according to an embodiment.

[0045] FIG. 11 illustrates a grain database with several clusters of grains.

[0046] FIG. 12 is a flowchart illustrating a process according to an embodiment.

[0047] FIG. 13 is a flowchart illustrating a process according to an embodiment.

[0048] FIG. 14A illustrates the rendering of a set of grains over time.

[0049] FIG. 14B illustrates an example of fast grain switching.

[0050] FIG. 15 illustrates parameters used to determine whether fast grain switching should be triggered.

[0051] FIG. 16A illustrates an example of fast grain switching without phase compensation.

[0052] FIG. 16B illustrates an example of fast grain switching with phase compensation.

[0053] FIG. 17 is a flowchart illustrating a process according to an embodiment.DETAILED DESCRIPTION

[0054] Grain Selection

[0055] FIG. 1 illustrates a system 100, according to some embodiments, for performing granular synthesis. System 100 includes a grain extraction unit 102 which extracts grains from original audio recordings 111. That is, grain extraction unit divides the original audio recording into small fragments, called “grains.” The extracted grains are stored in a grain database (GDB) 104 (or simply “database” for short) that is accessed at rendering time by a grain scheduling unit 106, which may also be referred to as “grain selection unit” or “grain scheduler”, which is a component of a grain rendering function (GRF) 108. In some embodiments, there is one grain database per procedural audio source, each of which is available to GRF 108. Each grain stored in grain database is associated with one or more vectors of one or more descriptor values, each vector corresponding to a particular descriptor.

[0056] When creating grain database 104, an audio designer decides what aspects should be used as descriptors. In some cases, it might be features of the sound itself, such as pitch or loudness, but it could also be other aspects that relate to how the sound was generated,such as the speed of movement that generates a contact sound between two objects sliding against each other or the opening angle of a door that generates a screeching sound when opened and closed. The descriptors should be chosen so that the sound can be re-generated dynamically by the Tenderer given a target descriptor coordinate or trajectory.

[0057] As noted above, each grain stored in grain database is associated with one or more descriptor values. Accordingly, the grains of an original recording need to be annotated with the descriptor values. In the case that a descriptor is an audio feature, the descriptor value may be possible to measure directly from the audio signal itself. In other cases, the descriptor values need to be provided somehow as extra metadata of the recordings. This may be, for example, done by logging data from some sensors during the recording and providing this data in companion files. In some cases, the annotation can be done manually by creating a log of data that describes how a descriptor changes during the recording or it can be done manually for each extracted grain.

[0058] When extracting grains from the original recording(s), the descriptor values are stored as metadata for each grain. Using the descriptor values, each grain can be positioned in a multi-dimensioned descriptor space where the value of each descriptor describes a position along one axis within this space. If only one descriptor is used, the descriptor space is onedimensional (ID), but if more descriptors are used the dimensionality of the descriptor space increases.

[0059] FIG. 2A shows an example of a two-dimensional (2D) descriptor space where each circle represents a grain. As shown in FIG. 2A, each grain has a location (e.g., a point or area position) within the 2D descriptor space, this location is referred to as the grain coordinate.

[0060] In one embodiment, the descriptor metadata for a sequence of grains extracted from one recording describes a trajectory within the descriptor space, which corresponds to how the descriptors evolved during the original recording.

[0061] At rendering time, the scheduling of the grains, i.e., the selection of grains to render, is typically based on a target descriptor coordinate in the descriptor space. The target descriptor coordinate specifies what descriptor values the generated sound output should have, which means that grains close to that coordinate in the descriptor space are to be used mostprominently. A target descriptor coordinate may come from many types of sources, such as a physics engine simulating the interaction of virtual bodies, live input parameters from hand controllers or other sensors, pre-defined automation parameters.

[0062] An aspect of granular synthesis rendering is that repetition of the same grain often sounds very unrealistic and artificial. If the grains are short, less than 50ms, repeating the same grain will result in a very metallic and static sound. If the grains are longer, the repetition will be heard like a repeating pattern, which often results in a sound that is not plausible.

[0063] The selection of grains needs to avoid repetition of the same grain but at the same time select grains that are close to the target descriptor coordinate in the descriptor space. This disclosure, therefore, uses, in some embodiment, a weighted selection (e.g., a weighted random selection) procedure that will generate ever evolving sequences of grains (i.e., an ordered set of grains) that closely follow the target descriptor coordinates. An example of a sequence of grains is: [grain-7, grain-6, grain-7, grain-9, grain-11, grain-10],

[0064] In some embodiments, each grain may be assigned a predefined weight (denoted pis) as well as a set of dynamic weights that may change over time. The predefined weight can be useful in cases where a grain is an outlier that should not be used too often but can add a realistic variation to the generated sound if used every now and then. Another use case is to use predefined weights to control the frequency of grains that represent e.g., bird chirps compared to grains that represent the background sound of a forest. For example, in an embodiment where a set of one or more grain groups is defined and a first group weight (wgl) is assigned to a first grain group in the set of grain groups and grain i is a member of the first grain group, then pis may be set to wgl .

[0065] An input (e.g., a signal generated by a user interaction) that controls a procedural audio source is mapped to a target descriptor coordinate. Based on the target descriptor coordinate, one or more grains from a set of candidate grains are selected for rendering using a weighted selection (e.g., a weighted random selection or a selection where the grain with highest weight is selected). The selected grains can be rendered using standard granular synthesis methods, which include, for example, performing an overlap-add operation using two selected grains, such as crossfading a selected grain with another grain, where metadata regarding overlap percent and crossfade window can be specified by the sounddesigner beforehand. Accordingly, rendering a grain encompasses not only rending all of the grain’s samples, but also rendering a first set of the grain’s samples followed by rending a set of mixed samples that are generated by mixing a second set of the grain’s samples with samples from another grain, as well as rending a set of mixed samples that are generated by mixing a first set of the grain’s samples with samples from another grain followed by rending a second set of the grain’s sample.

[0066] The input that controls the procedural audio source can change in real-time; consequently, the target descriptor coordinate can change over time as the input changes (the target descriptor coordinate can also change over time even if the input does not change).

[0067] In one embodiment, the grain selection algorithm using weighted selection has the following steps.

[0068] Step 1 : Obtain a target descriptor coordinate (e.g., map an input, such as a user input or other input, to a target descriptor coordinate in a descriptor space).

[0069] Step 2: Determine the size of a neighborhood adaptively, e.g., calculate a k value depending on the target descriptor coordinate or calculate a radius value (r) depending on the target descriptor coordinate. Alternatively, obtain a pre-calculated value of k or radius from metadata of the grain database.

[0070] Step 3: Select a set of candidate grains from the database using the k value or radius value. For example, select the grains from the database that are the k-nearest neighbors of the target descriptor coordinate. This search can be performed using off-the-shelf computationally efficient algorithms like the kD-tree search. As another example, include in the set of candidate grains each grain having a descriptor coordinate that is within a distance of r from the target descriptor coordinate.

[0071] Step 4: Assign a final weight to each one of the grains in the set of candidate grains. The final weight assigned to a given grain may be based on:

[0072] i) the distance of the position of the grain in the descriptor space (i.e., the grain coordinate) from the target descriptor coordinate in the descriptor space,

[0073] ii) the difference in the trend of descriptors of a grain, compared to the target descriptor trajectory,

[0074] iii) the temporal history of previously used grains,

[0075] iv) the difference in time instant in the original recording between the grain and the previously used grain, if they are from the same recording, and / or

[0076] v) predefined weight of the grain in the database.

[0077] Step 5: Perform a weighted selection of grains from the set of candidate grains using the final weights assigned in step 4. For example, perform a weighted random selection, or, as another example, select the grain with the highest final weight or lowest final weight. In this manner, grains are selected based on the target descriptor coordinate and further based on the final weights assigned to the grains in the set of candidate grains.

[0078] If the target descriptor coordinate changes, the steps are repeated. Otherwise, the weighted selection of grains continues (step 4 onwards) with changes made to the weights on the basis of temporal history of previous grains.

[0079] Step 2 - Determination of value of k.

[0080] Conventionally, the value of k is a user-defined constant. It is not desirable, however, to keep the value of k fixed at all times because doing so could lead to choosing too few or too many grains which in turn could lead to under-utilization of the grains or selecting grains that are dissimilar to the target descriptor value, respectively.

[0081] Accordingly, this disclosure provides, in one embodiment, an adaptive choice of k based on the density of grains available in an area surrounding the target descriptor coordinate.

[0082] An example of why one should use different values of k for different target descriptor coordinates is illustrated in FIG. 2A and 2B. In FIG. 2A, the target descriptor coordinate is close to a cluster of 3 grains.

[0083] In FIG. 2B, however, the target descriptor coordinate is in the vicinity of more grains. In these scenarios, it is not optimal to use the same value of k. For the case of FIG. 2A, k=3 is appropriate. If k > 3, then this will lead to choosing grains that are not in the cluster and therefore lead to a discontinuity in the texture of rendered sound. In FIG. 2B, if k = 3, then too few grains are selected. A larger value here will lead to richer textures with less repetition since a variety of grains can be selected.

[0084] The value of k should be chosen adaptively based on the target descriptor coordinate in the descriptor space. In one embodiment, each grain in the database is assigned an optimum k value. Then, for a target descriptor coordinate, the value of k is set equal to the optimal k value assigned to the grain that is closest to the target descriptor coordinate.

[0085] The assignment of optimal k per grain in the database can be performed offline or during the construction of the grain database. Either all distances from grain 1 to other grains in the database are recorded or we can have a threshold on the maximum number of neighbors to stop the distance computation.

[0086] Then the distances are sorted from lowest to highest. The difference between distances for consecutive values of k will have sudden jump at a value where the distance increases drastically. This is treated as a cut-off value for k and k + 1 is assigned to the grain. The addition of 1 is to include the grain itself in the value of k.

[0087] A criterion to determine the cut-off value would be to either set an absolute threshold on the difference in sorted distances between consecutive neighbors or to use normalized percentage increases in consecutive sorted distances.

[0088] Alternatively, one could also use the concept of adaptive radius to choose the set of candidate grains. The current method in literature involves using a fixed radius with target descriptor coordinate as the center of a circle and choose all the grains within this circle of fixed radius to be included in the set of candidate grains. By the same reasoning as above, it might be beneficial to change the radius adaptively based on the position of the target descriptor coordinate. There, instead of choosing different value of k, one would use different values of radius.

[0089] FIG. 3A illustrates the computation of k for a certain grain 301 represented by the black circle. In FIG. 3A, grain 301 and its 6 corresponding neighbors are shown. In FIG. 3B, the sorted distances are shown for grain 301. Because the jump in distance values is observable for k=4, an optimum k value of 5 assigned to grain 301.

[0090] Step 3 - Weight Assignments

[0091] Each of the factors that influence the final weight assigned to a grain are described below. We denote the target descriptor as u and weight associated with grain 1 as p;.The descriptor index is denoted by j = 1,2, ... , D where D is the number of descriptors or dimensions of the descriptor space. The descriptor value of grain i at dimension j is given byand that of the target as Uj .

[0092] The criteria below that influence the probability of choosing a grain are expressed as proportional relationships since the final value of the weight is obtained after normalizing i.e., ensuring that the weights corresponding to all k grains add to 1.

[0093] i) Distance from target descriptor coordinate

[0094] A distance metric is used to define the proximity between target descriptor coordinate and other descriptor coordinates corresponding to grains. An example of the distance metric is a weighted Euclidean distance where the difference in coordinates in each dimension is weighted by the inverse of standard deviation of the corresponding descriptor values,

[0095] di= L ^^ . j ffj

[0096] The closer a grain’s descriptor coordinate to the target descriptor coordinate, the higher the weight associated. Let d, be the distance from grain 1 to the target descriptor coordinate. For example, the probability that grain 1 is chosen can be inversely proportional to the distance.

[0097] pi oc 1 / dj. Accordingly, one can set pn = 1 / dj .

[0098] In some cases, the different descriptors should not have the same amount of influence on the grain selection. For example, if one descriptor is the pitch of the sound and another is a descriptor that has less strong effect on the perceptual character of the sound, the distance in the dimension corresponding to the pitch may be given a higher weight than the distance in the dimension that corresponds to the other descriptor. This can be achieved by adding an extra variable weight, m to each dimension when calculating the distance:

[0099] di = jJnij J(uj~Xii). (Jj

[0100] ii) Difference in trend of target and original descriptor trajectories

[0101] When performing grain extraction, the descriptor coordinates are used as the main selection criterion. But the trend of the original descriptor trajectory also gives important information about the grain. For example, if an engine sound is modelled with a granular database with one descriptor that denotes the RPM of the engine, the trend of the descriptor corresponds to the acceleration or deceleration of the engine at the time instant in the recording that the grain was extracted from. A grain that was extracted from a portion of the recording when the engine was accelerating will have a pitch that is slightly lower at the start than at the end and will therefor fit best when the desired output is the sound of an accelerating engine.

[0102] In the more general case, where a multi-dimensional descriptor space is used, the descriptor trend is a vector that corresponds to the direction of the trajectory that describes how the descriptors were changing at the time of the original recording. Similarly, the trend of the target descriptor trajectory describes the direction that the target descriptor coordinate is moving in the descriptor space.

[0103] The trend describes both the direction and the rate of change. Referring again to the example of the engine, if the granular database includes grains that correspond to the same RPM but with different acceleration, the grains that correspond to a similar acceleration as that of the target descriptor trajectory should be preferred.

[0104] FIG. 4 A shows a descriptor trajectory of an original recording of an engine sound where the descriptor is the RPM of the engine. The recording is divided into 15 grains. FIG. 4B shows an example of the how grains may be selected to match a target descriptor trajectory where the descriptor trend of the grains is not considered during grain selection. As can be seen, grains with both increasing and decreasing RPM are used in combination. The resulting descriptor trajectory shows an irregular behavior which may result in a degradation in perceived quality, especially if the descriptor represents the pitch of the sound. In FIG. 4C, the grain selection also considers the descriptor trend so that only grains with decreasing RPM are used, which would result in a smoother sound.

[0105] Considering the descriptor trend during grain selection is extra important for sound sources where the character is different for different trends. For example, an engine may sound different when accelerating compared to when decelerating. Making sure to match the descriptor trend avoids problems where grains with different character are used together.

[0106] The descriptor trend of a grain can be calculated as the difference in descriptor value at the end of the grain as compared to the start of the grain divided by the duration of the grain. In the case of a multi-dimensional descriptor space the trend is a vector that describes the mean rate of change in descriptor coordinates during the grain in the original recording, for example with three descriptors the trend, to, of a grain would be a three-dimensional vector:

[0107]

[0108] When selecting grains at rendering time, the trend of the target descriptor trajectory can be calculated similarly as the difference in descriptor coordinates since they were last updated divided by the time elapsed since they were last updated, Tu.

[0109]

[0110] The difference in trend can then calculated as

[0111] tD— tG— tT

[0112] The weight assigned to a grain can then be calculated as a function of the norm of to, for example: pi2 = max (1.0, - — ■), where mr is a variable that controls the how the probability I to I decreases with increased difference in descriptor trend.

[0113] iii) Temporal history of previously used grains

[0114] A history of previously rendered grains is kept for a certain time window so as to not repeat them and decrease the probability of choosing the grain if it was already rendered. The decrease in probability is directly proportional to a function of the difference between the current time instant t and the last time instant at which the grain 1 was rendered t'1.

[0115] pi3oc f(t — tj1)

[0116] If grain 1 has not been selected in the past or in a certain time window, the value of last time instant is set to zero, tj1= 0. The function f can be linear, quadratic, logarithmic or any monotonically increasing function of the argument t — tj1with non-negative output.

[0117] iv) Difference in time instant in original recording

[0118] For some sound sources, the way that the sound evolves may not be completely described by the changes in descriptors. Sometimes sound evolves in a way that depends on what happened earlier. For example, the screeching sound of an old door may have a slightly different screeching sound every time it is opened, even if the opening of it is done at the same speed etc. In these cases, the sound of grains with similar descriptor values and descriptor trend may sound very different and combining them may result in unnatural discontinuities that were not there in the original recording. A way to avoid these discontinuities is to assign a higher probability to grains that come from the same part of the original recording as the grain that was used previously. Grains that came from the same part of the original recording are expected to be closely related and resemble each other in character and are therefore good candidates when selecting the next grain.

[0119] In order to measure how close one grain from a particular recording is to another grain from the same recording, a time difference can be calculated which corresponds to the difference in time instant in the recording that the two grains were extracted from. If the grains were close, this time difference is small. To calculate the time difference between two grains, metadata that tells from which original recording each grain was extracted and at what time instant can be used. This metadata, therefore, provides a measure of closeness between the two grains. This metadata can be specified in a compact way as two values: a recording identifier assigned to the recording of origin and a timestamp that identifies the time instant in that recording at which the grain can be found.

[0120] A weight for a grain can then be calculated based on the difference timestamps in a way that a smaller difference gives a higher probability to choose the evaluated grain. For example, the weight, pi, of grain i could be calculated as: f(ti - t0) , R; = Rob , Rj Ro

[0122] Where Ri is the recording identifier assigned to the recording from which grain i was extracted, Ro is the recording identifier assigned to the recording from which the previously rendered grain was extracted, ti and to are the respective timestamps for the two grains, and b is a design constant that sets the probability to use a grain that comes from another recording. The function f() takes the difference in time instant as input and calculates a probability for the grain.In one embodiment function f decreases linearly with an increase in difference in time instant with a slope specified with a variable a:

[0124] This has the effect that the probability reduces from 1.0 to b as the difference between the respective timestamps increases but never goes below b.

[0125] In one embodiment the recording identifier Ri can also be set to refer to segments of a recording, i.e., one recording can be divided into segments where each segment has its own index. This can be useful when one recording contain segments that are not to be seen as related by the Tenderer.

[0126] Final Weight

[0127] The final weight assigned to grain i, denoted pi, is calculated by accumulating the weights assigned to grain i from each stage, i.e., p; = PiiPi2Pi3Pi4Pis- In some embodiments, not all five weights are needed and can then be skipped by setting the corresponding weight to 1.0 or by completely exclude it from the calculation.

[0128] In one embodiment, the different weights are given different levels of influence on the final weight by modifying the individual weights with a fractional exponent, such as e.g.,

[0129] pi = p;i1 / 2pi2pi3Pi41 / 4Pi5 ,

[0130] where the weight pn from the first stage is made less influential by using a fractional exponent of Yi and the weight pi4 is made even less influential by using a fractional exponent of 1 / 4.

[0131] After the weights p;are computed for all the grains in the set of candidate grains from the steps above, they are normalized, i.e., divide each of them by the sum, so that the final weight becomes a probability value.

[0132] pi «- - i- -, where k is the size of the set of candidate grains. i=i Pi

[0133] Then grains are selected based on their final weight (e.g., by sampling from the distribution). Note that the temporal history of a grain and the descriptor trend influence the finalweight at every time instant a grain is rendered when the target descriptor coordinate does not change. Therefore, as long as the target descriptor remains constant, if at time instant t grain i is chosen, at time t+h, the final weight of that grain is influenced by temporal history and descriptor trend. The final weight of grain i becomes Pi <" Pi2Pi3 where t - tj1= h in pi3.

[0134] To avoid the repetition of the grain 1 at t + 1, we can define f(h) as

[0136] However, due to normalization operation, the probabilities of the other k — 1 grains change as well.

[0137] FIG. 5 is a flow chart illustrating a process 500, according to an embodiment, for rendering audio corresponding to an audio recording, wherein the audio recording is divided into a plurality of grains. Process 500 may begin in step s502.

[0138] Step s502 comprises obtaining a first target descriptor coordinate, wherein the first target descriptor coordinate identifies a first location in an N-dimensional descriptor space, where N > 0.

[0139] Step s504 comprises defining a first set of candidate grains based on the first target descriptor coordinate, wherein the first set of candidate grains comprises kl of the plurality of grains, where kl > 1.

[0140] Step s506 comprises assigning a final weight to each grain in the first set of candidate grains.

[0141] Step s508 comprises randomly selecting a grain from the first set of candidate grains based on the assigned final weights such that the probability that a given grain in the first set of candidate grains is selected is a function of the final weight assigned to the given grain.

[0142] Step s510 comprises rendering the selected grain.

[0143] In some embodiments, each one of the plurality of grains is associated with a grain coordinate (e.g., a set of one or more descriptor values) that identifies a location of the grain in the N-dimensional descriptor space, and the first set of candidate grains is defined based on the first target descriptor coordinate and the grains’ coordinates.

[0144] In some embodiments, defining the first set of candidate grains based on the first target descriptor coordinate and the grains’ coordinates comprises determining a nearest- neighbor set of grains consisting of kl of the plurality of grains, wherein none of the plurality of grains that are not included in the nearest-neighbor set of grains is closer to the first target descriptor coordinate than any one of the grains included in the nearest-neighbor set of grains, and the candidate set of candidate grains consists of the grains included in the nearest-neighbor set of grains.

[0145] In some embodiments, the method further comprises determining kl based on the number of grains in the plurality of grains that have a grain coordinate that is within a threshold distance of the first target descriptor coordinate.

[0146] In some embodiments, each grain included in the plurality of grains is assigned an optimal k value, and the method comprises setting kl equal to the optimal k value assigned to the grain within the plurality of grains having a grain coordinate that is closest to the target descriptor coordinate.

[0147] In some embodiments, defining the first set of candidate grains based on the first target descriptor coordinate and the grains’ coordinates comprises: determining a first radius value, rl; and including in the first set of candidate grains each one of the plurality of grains that has a grain coordinate that is within a distance of rl from the first target descriptor coordinate.

[0148] In some embodiments, the method further comprises determining rl based on the number of grains in the plurality of grains that have a grain coordinate that is within a threshold distance of the first target descriptor coordinate.

[0149] In some embodiments, assigning a final weight to each grain in the first set of candidate grains comprises: assigning a first weight to a first grain included in the set of candidate grains; determining a first final weight based on the first weight; and assigning the first final weight to the first grain.

[0150] In some embodiments, the first weight is: a function of the distance between the first target descriptor coordinate and the first grain, a function of the amount of time that has elapsed since the first grain was last rendered, a function of a trajectory associated with the first grain and a target trajectory, or a function of a measure of a closeness between the first grain andthe most recently rendered grain (e.g., a time difference indicating a difference between a timestamp for the first grain and a timestamp for the most recently rendered grain assuming both grains were extracted from the same recording or the same segment).

[0151] In some embodiments, the method further comprises after randomly selecting a grain from the first set of candidate grains based on the assigned final weights, assigning a new final weight to at least one of the grains in the first set of candidate grains or removing the rendered grain from the first set of candidate grains; after assigning a new final weight to the rendered grain or removing the rendered grain from the first set of candidate grains, randomly selecting another grain from the first set of candidate grains based on the currently assigned final weights; and rendering the selected another grain.

[0152] In some embodiments, the method further comprises after randomly selecting a grain from the first set of candidate grains, obtaining a second target descriptor coordinate; defining a second set of candidate grains based on the second target descriptor coordinate, wherein the second set of candidate grains comprises k2 of the plurality of grains, where k2 > 1; assigning a final weight to each grain in the second set of candidate grains; randomly selecting a grain from the second set of candidate grains based on the assigned final weights such that the probability that a given grain in the second set of candidate grains is selected is a function of the final weight assigned to the given grain; and rendering the grain randomly selected from the second set of candidate grains.

[0153] FIG. 6 is a flow chart illustrating a process 600, according to an embodiment, for rendering audio corresponding to an audio recording, wherein the audio recording is divided into a plurality of grains. Process 600 may begin in step s602.

[0154] Step s602 comprises obtaining a first target descriptor coordinate, wherein the first target descriptor coordinate identifies a first location in an N-dimensional descriptor space, where N > 0.

[0155] Step s604 comprises defining a first set of candidate grains based on the first target descriptor coordinate, wherein the first set of candidate grains comprises kl of the plurality of grains, where kl > 1, and the first set of candidate grains comprises a first grain and a second grain.

[0156] Step s606 comprises assigning a final weight to each grain in the first set of candidate grains, wherein assigning a final weight to each grain in the first set of candidate grains comprises assigning a first final weight to the first grain and assigning a second final weight to the second grain. The first final weight assigned to the first grain is: a function of the distance between the first target descriptor coordinate and the first grain, a function of the amount of time that has elapsed since the first grain was last rendered, a function of a trajectory associated with the first grain and a target trajectory, and / or a function of a measure of a closeness between the first grain and the most recently rendered grain.

[0157] Step s608 comprises selecting a grain from the first set of candidate grains based on the assigned final weights.

[0158] Step s610 comprises rendering the selected grain.

[0159] In some embodiments, each one of the plurality of grains is associated with a grain coordinate (e.g., a set of one or more descriptor values) that identifies a location of the grain in the N-dimensional descriptor space, and the first set of candidate grains is defined based on the first target descriptor coordinate and the grains’ coordinates.

[0160] In some embodiments, defining the first set of candidate grains based on the first target descriptor coordinate and the grains’ coordinates comprises determining a nearest- neighbor set of grains consisting of kl of the plurality of grains, wherein none of the plurality of grains that are not include in the nearest-neighbor set of grains is closer to the first target descriptor coordinate than any one of the grains included in the nearest-neighbor set of grains, and the candidate set of candidate grains consists of the grains included in the nearest-neighbor set of grains.

[0161] In some embodiments, the method further comprises determining kl based on the number of grains in the plurality of grains that have a grain coordinate that is within a threshold distance of the first target descriptor coordinate.

[0162] In some embodiments, each grain included in the plurality of grains is assigned an optimal k value, and the method comprises setting kl equal to the optimal k value assigned to the grain within the plurality of grains having a grain coordinate that is closest to the target descriptor coordinate.

[0163] In some embodiments, defining the first set of candidate grains based on the first target descriptor coordinate and the grains’ coordinates comprises: determining a first radius value, rl; and including in the first set of candidate grains each one of the plurality of grains that has a grain coordinate that is within a distance of rl from the first target descriptor coordinate.

[0164] In some embodiments, the method further comprises determining rl based on the number of grains in the plurality of grains that have a grain coordinate that is within a threshold distance of the first target descriptor coordinate.

[0165] In some embodiments, the method further comprises after selecting a grain from the first set of candidate grains based on the assigned final weights, assigning a new final weight to at least one of the grains in the first set of candidate grains or removing the rendered grain from the first set of candidate grains; after assigning a new final weight to the rendered grain or removing the rendered grain from the first set of candidate grains, selecting another grain from the first set of candidate grains based on the currently assigned final weights; and rendering the selected another grain.

[0166] In some embodiments, the method further comprises after selecting a grain from the first set of candidate grains, obtaining a second target descriptor coordinate; defining a second set of candidate grains based on the second target descriptor coordinate, wherein the second set of candidate grains comprises k2 of the plurality of grains, where k2 > 1; assigning a final weight to each grain in the second set of candidate grains; selecting a grain from the second set of candidate grains based on the assigned final weights; and rendering the grain randomly selected from the second set of candidate grains.

[0167] Grain Interpolation

[0168] As described above, at rendering time, the scheduling of the grains (i.e., the selection of grains to render) is based on a target descriptor coordinate in the descriptor space. When GDB 104 has many grains that evenly fill the whole descriptor space, the selection of grains can be done without repeating the same grain too often and without the need to use grains that are far from the target position. But in some cases, GBD 104 may have too few grains, or grains that do not evenly fill the descriptor space. In these cases, the selection of grains becomes more restricted. For some target positions in the descriptor space, there mightnot be any grains that are close, or the close grains are few so that they need to be repeated frequently.

[0169] In the case where the grain scheduler does not find a grain that is close enough to the target position, or that all grains close to the target position have already been used recently, a new grain can be created; this new grain is referred to as an “interpolated grain”. The method for creating an interpolated grain includes selecting a plurality of grains, which is called the “group of grains”, and calculating the interpolated grain based on the selected group of grains. It is important to select grains such that the group of grains surround the target position. In one embodiment, the grains in the group are close to the target position, but not all grains are on the same side of the target.

[0170] FIG. 9 illustrates a 2D descriptor space where the grains in GDB 104 do not evenly cover the descriptor space. As illustrated in FIG. 9, three clusters of grains, A, B and C, surround a target position indicated with an X. If grain 901 in cluster B is selected as the first grain of the group of grains to be used for interpolation, the next grain selected preferably is complementary to grain 901 with respect to the target position.

[0171] In one embodiment, a grain i is complementary to a grain j with respect to a target position if and only if for each grain descriptor coordinate dimension of the granular database, grains i and j are on the opposite sides of the target position. With this definition, grain i is complementary to grain 901 with respect to the target position if Gix is less than Tx and Giy is greater than Ty, where Gix is the x-coordinate of grain i, Giy is the y-coordinate of grain i, Tx is the x-coordinate of the target position, and Ty is the y-coordinate of the target position.

[0172] More generically, assuming grain j is closer to a target position (Tx, Ty) than grain i, then with respect to a 2D descriptor space, grain i is complimentary to grain j under the following conditions: if Gjx == Tx and Gjy < Ty, then grain i is complimentary to grain j if Giy > Ty, OR if Gjx == Tx and Gjy > Ty, then grain i is complimentary to grain if Giy < Ty, OR if Gjy == Ty and Gjx < Tx, then grain i is complimentary to grain if Gix > Tx, OR if Gjy == Ty and Gjx > Tx, then grain i is complimentary to grain if Gix < Tx, OR if Gjx < Tx and Gjy < Ty, then grain i is complimentary to grain if Gix >Tx and Giy > Ty, ORif Gjx < Tx and Gjy > Ty, then grain i is complimentary to grain if Gix > Tx and Giy < Ty, OR if Gjx > Tx and Gjy < Ty, then grain i is complimentary to grain if Gix < Tx and Giy > Ty, OR if Gjx > Tx and Gjy > Ty, then grain i is complimentary to grain if Gix < Tx and Giy < Ty.

[0173] Grain 902 in cluster A is an example of a grain that is complimentary to grain 901 with respect to the target position. This is indicated in FIG. 9 by fact that grain 902 is on the opposite sides of the target position, in both dimensions, compared to grain 901. Grain 902 would be ideal to add to the group of grains because grain 902 is not only complementary to grain 901, but also close to the target position. The grains from cluster C are all at the same side as grain 901 in the y-dimension and the grains in cluster B are all on the same side as grain 901 in the x-dimension; hence none of the grains in cluster B or C are complimentary to grain 901 with respect to the target position. Any grain from cluster A, however, satisfies the conditions for being complementary to grain 901.

[0174] In one embodiment, the following steps are performed:

[0175] Step 1 : Obtain a target position in the descriptor space of the database and define a set of candidate grains based on the target position using, for example, a method described above.

[0176] Step 2: Using a weighted selection, such as, for example, a weighted random selection, select a first grain from the set of candidate grains. If the selected grain is within a threshold distance of the target position, render this grain without doing any interpolation, otherwise continue with steps 3-7.

[0177] Step 3 (optional). Remove the first grain from the set of candidate grains.

[0178] Step 4. For each grain in the set of candidate grains, assign a final weight to the candidate grain, wherein the final weight is a function of whether or not the grain is complimentary to the first grain with respect to the target position. All else being equal, a candidate grain that is complimentary to the first grain will have a higher final weight than a candidate grain that is not complimentary to the first grain. If the first grain was not removed from the set of candidate grains, then assign very low final weight to the first grain (e.g., a weight of 0.00001).

[0179] In one embodiment, if there are K grains in the set of candidate grains, then for i=l to K, the final non-normalized weight assigned to grain i in the set of candidate grains can be calculated as: pi = pnpi2pi3pi4 pispicoMP, where picoMP is a weight that the depends on whether or not grain i is complimentary to the first grain. As an example PCOMP can be set as: picoMP = X, if grain i is complimentary, othwerwise picoMP is set to Y, where X > Y.Preferably X is at least an order of magnitude greater than Y. For example, in one embodiment X=1 and Y=0.0001. In some scenarios, it may be beneficial to not completely exclude the non- complementary grains in cases of a sparse grain database, where there may not be that many grains in the vicinity of the target position to choose from. Hence, for this reason, Y is typically not set equal to 0.

[0180] Step 5. Selecting at least a second grain from the set of candidate grains based on the assigned final weights.

[0181] Step 6. Set the length of the new interpolated grain. For example, if the sound of the database has a pitch, calculate the length as the weighted mean of the lengths of the selected grains, where the weight depends on the distance from each grain to the target position. If the sound in the database has no pitch, set the length according to the shortest of the selected grains.

[0182] Step 7. Generate the interpolated grain using the selected grains. For example, the interpolated grain may be a weighted mix of the selected grains, where the weights depend on the distance from each selected grain to the target position. If the sound in the database has a pitch, each selected grain can be resampled to fit the length of the interpolated grain. If the sound in the database has no pitch, only a sub section of the selected grains is used, corresponding to the length of the interpolated grain.

[0183] In the simplest case, two selected grains, GAand GBthat are closest to the target descriptor coordinate are used to generate an interpolated grain.

[0184] It is usually good to use as few grains as possible for the interpolation since mixing together many grains may result in a diffuse sound that does not resemble the original sound. Accordingly, in one embodiment, only the first selected grain and the second selected grain are used to derive the interpolated grain.

[0185] If, however, more than two grains are to be selected for interpolation especially in multi-dimensional granular databases, the condition for complimentary nature can be relaxed for a certain dimension for the choice of third grain. Another criterion could be to ensure pairwise complimentary nature for subsequently selected grains. For example, in a 3-D database, grains A and B satisfy complimentary criterion in all dimensions and grains B and C are complimentary in all the dimensions, but grains A and C may be complimentary only in the first two dimensions but not the third.

[0186] Deciding the length of the interpolated grain

[0187] In step 6, the length of the interpolated grain, which is referred to as the “target length”, is set. To do a weighted mix of grains, the grains need to have a matching length. The target grain length can be decided in different ways depending on the character of the sound described by the grain database.

[0188] For some grain databases, the length of each grain corresponds to one or a specific number of pitch periods. In this case, the target length can be calculated to be a weighted average of the lengths of the selected grains. For the case in which only two grains are used to generate the interpolated grain, the target length (L arget) can be calculated as: L arget = (wl)(Ll) + (w2)(L2), where LI and L2 are the lengths of two selected grains, and wl and w2 are weights calculated to be inversely proportional to the distance of each selected grain from the target position as: wl = d2 / (dl+d2) and w2=dl / (dl+d2).

[0189] The distances dl and d2 are the Euclidian distances in the descriptor space. In cases where only one dimension of the granular database has an influence on the pitch of the sound, and therefore the length of the grains, the weight can be calculated using a distance measure where only the distance in that dimension is considered.

[0190] For other granular databases where the length of the grains does not indicate a pitch, the target length is less critical. A straight-forward method is to select a target length that matches the length of the shortest of the grains that were selected.

[0191] Creating Temporary Grains for the Weighted Mix

[0192] Before a weighted mix can be calculated to form the interpolated grain, temporary versions of the selected grains with the selected target grain length are created.

[0193] In the case of a granular database representing a pitched sound, the selected grains can be resampled using a resampling method that allows arbitrary resampling ratios, such as for example a linear resampling, sine interpolation or Lanczos resampling or similar well-known techniques.

[0194] If the granular database does not represent a pitched sound, sub-sections of the selected grains that matches the target grain length can be used. Either the first part of each selected grain is used, or a subsection within each grain with a randomly chosen starting point, where the randomly selected starting point is limited so that the remaining samples of the grain corresponds at least to the target grain length.

[0195] Generating the interpolated grain using a weighted mix

[0196] In one embodiment, the interpolated grain is a weighted mix of the selected grains (assuming the selected grains have the same length) or a weighted mix of one selected grain and one temporary grain derived from the other selected grain (assuming the selected grain’s length is equal to the target length) or a weighted mix of temporary grains (assuming none of the grains have a length equal to the target length). Accordingly, the interpolated grain is equal to: (wl)(gl) + w2(g2), or (wl)(tgl) + w2(g2), or (wl)(gl) + w2(tg2), or (wl)(tgl) + w2(tg2), where tgl is a first temporary grain derived from the first grain, and tg2 is a second temporary grain derived from the second grain. In one embodiment, the weights can be calculated to be inversely proportional to the distance of the selected grains to the target position in a similar way as was done when selecting the target grain length.

[0197] For example, if two grains gl and g2 are selected and the distances of the selected from the target position are dl and d2, respectively, then, in one embodiment, the weights used in performing the weighted mix are calculated as: wl = (d2 / (dl+d2)) and w2 = (dl / (dl+d2)). In the case where consecutive grains are expected to be mostly uncorrelated2, the weight can be calculated according to a constant power mixing rule rather than a linear mixing rule. In this case the weights can be calculated as wl = sqrt(d2 / (dl+d2)) and w2 = sqrt(dl / (dl+d2)). Other ways of creating the weights (also known as gains) are described in the section below regarding determining an optimal mixing window coefficient.

[0198] Caching of interpolated grains

[0199] After an interpolated grain has been generated, the interpolated grain may be stored for later use. This can decrease the complexity of calculating new interpolated grains later. In this case, the interpolated grain should be assigned an index and / or descriptor coordinates, as the other grains in the database, so that the Tenderer can avoid repetitions of the same interpolated grain. While caching interpolated grains may decrease the complexity of the Tenderer, caching increases the memory consumption, so it may be beneficial to limit the number of cached grains. Cached interpolated grains may also be written into the granular database, if desired, making the proposed interpolation technique another method of corpus extension.

[0200] FIG. 10 is a flow chart illustrating a process 1000, according to an embodiment, for rendering audio corresponding to an audio recording, wherein the audio recording is divided into a plurality of grains. Process 1000 may begin in step sl002. Step sl002 comprises obtaining a target descriptor coordinate, wherein the target descriptor coordinate identifies a target position in an N-dimensional descriptor space, where N > 0. Step si 004 comprises based on the target descriptor coordinate, selecting a group of grains from the grain database, wherein the selected group of grains comprises a first grain and a second grain. Step si 006 comprises determining a first weight, wl, for the first grain. Step si 008 comprises determining a second weight, w2, for the second grain. Step slOlO comprises producing an interpolated grain using the first weight, the second weight, the first grain, and the second grain. Step sl012 comprises rendering the interpolated grain.

[0201] In some embodiments, the second grain has a position in the N-dimensional space, and the first grain is in a complementary position within the N-dimensional space with respect to the position of the second grain in the N-dimensional space.

[0202] In some embodiments, wherein the interpolated grain is equal to: (wl)(gl) + w2(g2), or (wl)(tgl) + w2(g2), or (wl)(gl) + w2(tg2), or (wl)(tgl) + w2(tg2), where gl is the first grain, g2 is the second grain, tgl is a first temporary grain derived from the first grain, and tg2 is a second temporary grain derived from the second grain.

[0203] In some embodiments, the first grain has a length and the second grain has a length, and producing the interpolated grain comprises: determining a target grain length using a first length value specifying the length of the first grain and a second length value specifying thelength of the second grain, wherein the length of the first grain is not equal to the target grain length; deriving, from the first grain, a first temporary grain having a length equal to the target grain length; and using the first temporary grain, the second grain, the first weight, and the second weight to produce the interpolated grain.

[0204] In some embodiments, the first grain has a length and the second grain has a length, and producing the interpolated grain comprises: determining a target grain length using a first length value specifying the length of the first grain and a second length value specifying the length of the second grain, wherein the length of the first grain is not equal to the target grain length and the length of the second grain is not equal to the target grain length; deriving, from the first grain, a first temporary grain having a length equal to the target grain length; deriving, from the second grain, a second temporary grain having a length equal to the target grain length; and using the first temporary grain, the second temporary grain, the first weight, and the second weight to produce the interpolated grain.

[0205] In some embodiments, deriving the first temporary grain from the first grain comprises: resampling the first grain to produce the first temporary grain, or selecting a subsection of the first grain, the first temporary grain is the selected sub-section.

[0206] In some embodiments, determining the target grain length using the first length value and the second length value comprises determining the target grain length using the first length value, the second length value, the first weight, and the second weight.

[0207] In some embodiments, the target grain length is equal to (wl)(Ll) + (w2)(L2), where LI is the first length value, and L2 is the second length value.

[0208] In some embodiments, the first grain has a position in the N-dimensional space, and the second grain has a position in the N-dimensional space, the first weight is based on (e.g., inversely proportional to) the distance from the position of the first grain in the N-dimensional space to the target position in the N-dimensional space, and the second weight is based on (e.g., inversely proportional to) the distance from the position of the second grain in the N-dimensional space to the target position in the N-dimensional space.

[0209] In some embodiments, the first weight is equal to: d2 / (d 1 + d2), the second weight is equal to: dl / (dl + d2), dl is the distance from the position of the first grain in the N-dimensional space to the target position in the N-dimensional space, and d2 is the distance from the position of the second grain in the N-dimensional space to the target position in the N- dimensional space.

[0210] In some embodiments, selecting a group of grains from the grain database comprises: defining a first set of candidate grains based on the target position; selecting the first grain from the first set of candidate grains; after selecting the first grain from the first set of candidate grains, removing the first grain from the first set of grains, thereby forming a second set of grains; for each grain included in the second set of grains, assigning a final weight to the grain; selecting a grain from the second set of candidate grains based on the assigned final weights, wherein the grain selected from the second set of grains is the second grain.

[0211] In some embodiments, selecting a group of grains from the grain database comprises: defining a set of candidate grains based on the target position; selecting the first grain from the set of candidate grains; for each grain included in the set of candidate grains, assigning a final weight to the grain; after assigning the final weights to each grain included in the set of candidate grains, selecting a grain from the set of candidate grains based on the assigned final weights, wherein the grain selected from the second set of grains is the second grain.

[0212] In some embodiments, assigning a final weight to the second grain comprises: determining whether the second grain is complimentary to the first grain with respect to the target position; assigning a complimentary weight, PCOMP, to the second grain, wherein the value of PCOMP is X if the second grain is complimentary to the first grain with respect to the target position, otherwise the value of PCOMP is Y, wherein X is greater than Y; and calculating the final weight for the second grain using the complimentary weight assigned to the second grain. In some embodiments, X is at least ten times greater than Y.

[0213] Determination of an Optimal Mixing Window Coefficient.

[0214] When performing granular synthesis, overlap-add techniques can be used to provide a smooth transition from one grain to the next in the sequence of grains. As previously described, it may be beneficial to adapt the choice of overlap-add window to the character of the signals in the grains to be mixed. If the signals are highly correlated, a linear mixing window should be used. If, however, the signals are uncorrelated, a power preserving mixing windowshould be used. In many cases the optimal mixing window is somewhere in-between the linear and the power preserving window because it may be that only parts of the signals are correlated.

[0215] To generate an optimized mixing window (denoted Wo), a mixing window coefficient (denoted “m”) (also known as mixing window weight) can be used, which gives a continuous control over the mixing window, from linear to power preserving. In one embodiment, the mixing window coefficient (m) is a value between 0.0 and 1.0, and Wo = mWP+ (l-m)WL, where WL is the linear overlap window and Wp is the power preserving overlap window. Accordingly, in this embodiment, the optimized mixing window is a weighted average of the power preserving overlap window and the linear overlap window.

[0216] When m=l, only the power preserving overlap window will be used. When m=0, only the linear overlap window will be used. For values in-between 0 and 1, a mix of the two windows will be used. So, by tuning the mixing window coefficient correctly, an optimized mixing window can be found.

[0217] In one embodiment, WL for a pair of grains subject to the overlap-add processing (the pair of grains to mixed) is a vector of X scalar values, i.e., WL= [WI[0], WI[1], ..., wi[X-l], where wi[x] = x / (X-l) and X = min (LI, L2) * p, where LI is the length of the first grain of the pair and L2 is the length of the second grain of the pair and p is an overlap percentage. In one embodiment, Wp is also a vector of X scalar values, i.e., Wp= [wP[0], wP[l], ..., wP[X-l], where wP[x] = (wi[x])1 / 2.

[0218] If the mixing window coefficient is m for both the first grain of the pair and the second grain of the pair, then the weights gl and g2 for the two grains at sample index n are, gl[n] = m(wi[n])1 / 2+ (l-m)wi[n] and g2[n] = m(l-wi[n])1 / 2+ (l-m)(l-wi[n]).

[0219] The X overlap-add samples, s[n] for n=0 to X-l, would then be a linear combination of X samples from the two grains with weights gl and g2, i.e., s[n] = gl[n]grainl[Spl+n] + g2[n]grain2[Sp2+n], for n=0 to X-l, where grainlf] is the set of samples that comprise the first grain, grain2[] is the set of samples that comprise the second grain, and Spl and Sp2 are starting positions for the first grain and second grain respectively.

[0220] When using a granular database where the grains in the database have a position in a descriptor space, typically there may be clusters of grains that are highly correlated with each other, and other clusters of grains that are not correlated. This may for example happen when a granular database describes a sound which has a strong pitch in parts of the descriptor space but is more noisy in character in other parts. By specifying a mixing window coefficient for the different parts of the descriptor space, an overlap window can be selected that works optimally for the grains in that part.

[0221] FIG. 11 illustrates a grain database with several clusters of grains. By specifying a mixing window coefficient in several position in the descriptor space, the mixing window that is used can be tuned for each cluster of grains.

[0222] In one embodiment, a set of mixing window coefficients is specified, where each coefficient is associated with a position in the descriptor space, i.e., each coefficient is associated with a set of coordinates that specifies a position in the descriptor space. The coefficients can be stored as metadata in the grain database. This makes it possible to pre-calculate optimized mixing window coefficients by, for example, performing correlation checks between grains or other methods to find optimized mixing windows. The set of coefficients could also be manually created, and the coefficients tuned by the sound designer that creates the granular database.

[0223] Having a list of mixing window coefficients along with a descriptor space position is a compact, efficient, and scalable way to specify the mixing windows for a database. In some cases, it may be enough to have only one mixing window specified and, in some cases, a large number of mixing windows specified in different positions may be beneficial.

[0224] An example of a data structure (table) storing mixing window coefficients for a 3- dimensional granular database is shown below:

[0225] Here each entry (row of the table) has a field storing coordinates identifying a position in a 3D descriptor space and an associated field storing a mixing window coefficient specified for that position. Other data structures, such as for example, lists, can be used to store the coefficients and corresponding coordinates.

[0226] At rendering time, an optimal (or final) mixing window coefficient can then be found for every grain based on the position in the descriptor space of the grain. In the simplest case, the coefficient in the list that is closest to the grain in the descriptor space is selected. Another embodiment may use a weighted sum of coefficients, where the weight of each coefficient is based on the distance between the coefficient and the grain, i.e., the optimal mixing window coefficient is: m = r|imi, where T], is a function of the distance between the ith coefficient in the list and the grain.

[0227] In order to efficiently retrieve the closest coefficients for a particular grain, the list of coefficients can be ordered in, for example, a K-D tree structure.

[0228] When a mix, such as, for example, an overlap-add, of two grains is to be performed, a coefficient for each of the two grains can be determined as described above and then the optimal mixing window coefficient would then be based on these two determined coefficients. In one embodiment, the optimal mixing window coefficient is the maximum of the two determined mixing window coefficients, i.e., m = max(ml, m2), where ml is the coefficient determined for the first grain and m2 is the coefficient determined for the second grain. The logic behind this operation is that higher value of mixing window coefficient means that the grains to be mixed are potentially less correlated.

[0229] For example, if one grain has a mixing window coefficient of 0.32 and the other grain has a mixing window of 0.93 then m= 0.93 will be used. Even though grain with mixingwindow coefficient of 0.32 is from a cluster of correlated grains it will probably not be so correlated with a grain from a cluster with low correlation. So one can assume that the grains are mostly uncorrelated if they have different mixing window values. Accordingly, it is advantageous to use the largest mixing window coefficient.

[0230] In one embodiment, the final mixing window coefficient is not based on the positions of the grains that are to be mixed, but instead on the current target position in the descriptor space. In this case, for example, some number of mixing window coefficients that are closest to the current target position are used as a basis for determining the optimized mixing window coefficient. This could be done by picking the closest one, or doing some form of weighted mix of a set of mixing window coefficients that are specified in points close to the target position.

[0231] The optimal mixing window coefficient may be used whenever a mix of two or more grains are to be performed. In the case of overlap-add, the mixing happens within an overlap region.

[0232] The optimal mixing window coefficient (m) may also be used when mixing grains together, e.g., to form an interpolated grain, as described above. As an example, the samples, Sint[], of an interpolated grain may be:Sintfn] = gl *grainl[n] + g2*grain2[n], for n=0 to L-l, whereL is the length of the grains (the grains in this embodiment have equal length), grainlf] is the set of samples that comprise the first grain, grain2[] is the set of samples that comprise the second grain, gl = m (wl)1 / 2+ (l-m)wl; and g2 = m (w2)1 / 2+ (l-m)w2.

[0233] In one embodiment, wl = (d2 / (dl+d2) and w2 = (dl / (dl+d2), where dl is the Euclidian distances between the first grain and the target position and d2 is the Euclidian distances between the second grain and the target position. But in cases where only one dimension of the granular database has an influence on the pitch of the sound, and therefore thelength of the grains, the weights wl and w2 can be calculated using a distance measure where only the distance in that dimension is considered.

[0234] FIG. 12 is a flow chart illustrating a process 1200, according to an embodiment, for rendering audio corresponding to an audio recording, wherein the audio recording is divided into a plurality of grains. Process 1200 may begin in step sl202.

[0235] Step sl202 comprises selecting a first grain from the plurality of grains, wherein the first grain is associated with a first position in an N-dimensional, ND, descriptor space.

[0236] Step sl204 comprises selecting a first set of one or more mixing window coefficients based on the first position in the ND descriptor space.

[0237] Step sl206 comprises selecting a second grain from the plurality of grains, wherein the second grain is associated with a second position in the N-dimensional descriptor space.

[0238] Step sl208 comprises selecting a second set of one or more mixing window coefficients based on the second position in the ND descriptor space.

[0239] Step sl210 comprises obtaining a final mixing window coefficient, m, using the first set of mixing window coefficients and the second set of mixing window coefficients.

[0240] Step sl212 producing mixed samples, S, by mixing at least a portion of the first grain with at least a portion of the second grain using the final mixing window coefficient.

[0241] In some embodiments, the first grain comprises a first set of samples, the second grain comprises a second set of samples, and producing the mixed samples using the final mixing window coefficient comprises producing a first mixed sample, s[0], by calculating s = gl x grainl_sample + g2 x grain2_sample, where gl is function of the final mixing window coefficient, grainl sample is one of the samples from the first set of samples, g2 is function of the final mixing window coefficient, and grain2_sample is one of the samples from the second set of samples.

[0242] In some embodiments, gl is a further function of a first weight, wl, associated with a first mixing window and a second weight, w2, associated with a second mixing window, and g2 is a further function of a third weight, w2, associated with the first mixing window and a fourth weight, w4, associated with a second mixing window.

[0243] In some embodiments, gl = m x wl + ( l -m) xw2, and g2 = m x w3 + ( l -m) x w4.

[0244] In some embodiments, wl = (w2)1 / 2, and w3 = (w4)1 / 2.

[0245] In some embodiments, w4 = 1 - w2.

[0246] In some embodiments, w2 = n / (L-l), n > 0 and n < (L-l), L = min(Ll,L2) x p, LI is the length of the first grain, L2 is the length of the second grain, and p is a predetermined overlap percentage.

[0247] In some embodiments, w2= dl / (dl+d2), w4= d2 / (dl+d2), dl is a distance from the first position in the ND descriptor space to a target position in the descriptor space, and d2 is a distance from the second position in the ND descriptor space to the target position in the descriptor space.

[0248] In some embodiments, the first set of mixing window coefficients consists of a first mixing window coefficient, ml, obtaining a final mixing window coefficient comprises setting m equal to max(ml,m2), and m2 is a mixing window coefficient included in the second set of mixing window coefficients or is based on the mixing window coefficients included in the second set of mixing window coefficients.

[0249] In some embodiments, obtaining the final mixing window coefficient comprises: assigning a weight to each mixing window coefficient included in the first set of mixing window coefficients, using the weights assigned to each mixing window coefficient included in the first set of mixing window coefficients, obtaining a weighted average of the mixing window coefficients included in the first set of mixing window coefficients, and setting m equal to max(ml,m2), where ml is the weighted average of the mixing window coefficients included in the first set of mixing window coefficients, and m2 is mixing window coefficient included in the second set of mixing window coefficients or is based on the mixing window coefficients included in the second set of mixing window coefficients.

[0250] FIG. 13 is a flow chart illustrating a process 1300, according to an embodiment, for rendering audio corresponding to an audio recording, wherein the audio recording is divided into a plurality of grains. Process 1300 may begin in step sl302.

[0251] Step sl302 comprises selecting a first grain from the plurality of grains.

[0252] Step sl304 comprises selecting a second grain from the plurality of grains.

[0253] Step sl306 comprises selecting i) a set of two or more mixing window coefficients based on a target position in an N-dimensional, ND, descriptor space or ii) a single mixing window coefficient based on the target position in the ND descriptor space.

[0254] Step sl308 comprises producing mixed samples, S, by mixing at least a portion of the first grain with at least a portion of the second grain using i) the single mixing window coefficient or ii) a derived mixing window coefficient derived using the set of two or more mixing window coefficients.

[0255] In some embodiments, the first grain comprises a first set of samples, the second grain comprises a second set of samples, and producing the mixed samples using the final mixing window coefficient comprises producing a first mixed sample, s[0], by calculating s = gl x grainl_sample + g2 x grain2_sample, where gl is function of the final mixing window coefficient, grainl sample is one of the samples from the first set of samples, g2 is function of the final mixing window coefficient, and grain2_sample is one of the samples from the second set of samples.

[0256] In some embodiments, gl is a further function of a first weight, wl, associated with a first mixing window and a second weight, w2, associated with a second mixing window, and g2 is a further function of a third weight, w2, associated with the first mixing window and a fourth weight, w4, associated with a second mixing window.

[0257] In some embodiments, gl = m x wl + ( l -m) xw2, and g2 = m x w3 + ( l -m) x w4, where m is the single mixing window coefficient or the derived mixing window coefficient.

[0258] In some embodiments, wl = (w2)1 / 2, and w3 = (w4)1 / 2.

[0259] In some embodiments, w4 = 1 - w2.

[0260] In some embodiments, w2 = n / (L-l), n > 0 and n < (L-l), L = min(Ll,L2) x p, LI is the length of the first grain, L2 is the length of the second grain, and p is a predetermined overlap percentage.

[0261] In some embodiments, w2= dl / (dl+d2), w4= d2 / (dl+d2), dl is a distance from the first position in the ND descriptor space to a target position in the descriptor space, and d2 isa distance from the second position in the ND descriptor space to the target position in the descriptor space.

[0262] Fast Grain Switching

[0263] Described above are methods for scheduling grains from a database of grains where each grain has specified position in a descriptor space. A target position in the descriptor space controls what grain should be selected next. As the target position is updated in real-time, grains are selected that are close to the target position. This makes it possible to control how the sound evolves dynamically. A simplified example of this feature is illustrated in FIG. 14A.

[0264] FIG. 14A illustrates how consecutive grains can be selected following a changing target position. In the case shown in FIG. 14 A, the granular database has only one dimension, descriptor 1. On the x-axis, a timeline shows when grains are selected and how long each grain is used to produce an audio output, i.e., how long each grain is rendered. The dotted line in FIG.14A shows how the target position varies over time. In figure 14A, each entire grain is rendered before transitioning to the next grain.

[0265] Typically, however, samples from the currently selected grain are played until the play position reaches a predetermined sample position, which, typically, is a sample position close to the end of the grain. When the play position reaches the predetermined sample position, the next grain is selected and a cross-fade (e.g., an overlap-add) between the two grains is initiated.

[0266] In some cases, the grains of a granular database need to be rather long to retain the natural sound of the original sound source. In this case, the selection of grains may not be frequent enough to allow a fast and smooth transition. For example, if the target position is moving quickly through the descriptor space, the output of the granular synthesis may have large jumps due to the fact that the target position having moved a substantial distance in the descriptor space in the time one grain was played. In this case, the switching of grains is too slow to follow the dynamic changes of the target position.

[0267] At the same time, using a forced, higher, rate of switching grains may produce a less natural output. If the granular database was designed with long grains, the intention of thesound designer was that the grains should be long and using a forced grain switching rate would not allow the full grains to be played as intended.

[0268] Accordingly, this disclosure proposes to frequently evaluate the amount by which the target position has moved since the last time a grain was selected. A criterion is then evaluated to determine if a fast grain switch is to be initiated. In this way, longer grains will play out completely as long as the target position does not move too much. If the target position does move fast, a fast grain switch can be triggered, and a fast dynamic behavior can be had. This feature is illustrated in FIG. 14B.

[0269] As shown in FIG. 14B, a new grain can be selected and transitioned to before the end of the currently used grain has been reached. In this way, a smoother change in sound can be had that follows the target position more closely without big jumps. In FIG. 14B, the overlap regions are not shown to keep the illustration simple.

[0270] Accordingly, in one embodiment there is a method that includes: obtaining an updated target position in the descriptor space, determining a distance using the target position, for example, determining a distance the target position has moved since a prior point in time, and, based on the determined distance, determining whether a fast grain switch condition is satisfied.

[0271] In one embodiment, as a result of determining that the fast grain switch condition is satisfied, the Tenderer immediately transitions to a new grain. For example, the new grain is selected based on the updated target position and an overlap-add is initiated to transition to the new grain. As another example, the new grain is selected based on the updated target position, the Tenderer ceases playing the current grain, and the Tenderer immediately begins playing the new grain.

[0272] Criterion for triggering a fast grain switch

[0273] In one embodiment, determining whether the fast grain switch condition is satisfied includes determining the amount by which the target position has moved since the last grain was selected. For example, in one embodiment, the distance that the target position has moved since the last grain was selected is determined and if this determined distance is greater than a threshold, then a fast switch is triggered, i.e., the fast grain switch condition is satisfied.

[0274] In another embodiment, determining whether the fast grain switch condition is satisfied includes determining how well the currently used grain(s) match(es) the updated position. For example, in one embodiment, in the case where a single grain is currently used for rendering, the updated target position is compared with the position of this single grain. If the target position is more than a threshold distance away from the grain, then the fast grain switch is triggered. Alternatively, a fast switch is triggered if this distance has increased more than a threshold since the current grain was selected.

[0275] As another example, in the case where more than one grain is currently used for rendering, e.g., if the Tenderer uses a cross-fade of two or more grains, the updated target position can be compared to the positions of all the current grains. For instance, the distance from the updated target position to the closest of the used grains is calculated and if that distance is greater than a threshold, a fast grain switch is triggered. Alternatively, a fast switch is triggered if this distance has increased more than a threshold since the current grain was selected.

[0276] FIG. 15 illustrates a scenario where two grains (grain 1501 and grain 1502) are currently used to produce the audio output, i.e., both grains are being rendered in a mixed fashion.

[0277] The distance (d2) from an updated target position 1512 from a line 1520 between the two used grains is calculated, and, in one embodiment, if this distance (d2) is greater than a predetermined value (a.k.a., threshold) a fast grain switch is triggered.

[0278] In another embodiment, in order to avoid constant triggering of fast grain switching in a sparse database where no grains are found close to the target position, the change of this distance since the current grains were selected can be compared to a threshold, i.e., the value (d2 - dl) is compared to the threshold, where dl is the distance a target position 1511 that was used when selecting grains 1501 and 1502. In other words, how much the distance from the target position to a line between the currently used grains has increased since the currently used grains were selected. If the increase in distance is more than a threshold, a fast grain switch can be triggered. If the increase in distance is less than the threshold, the currently selected grains are likely still valid for the updated target position.

[0279] In another embodiment, d2 is compared to dl, and, if the distance has increased more than a threshold (T), then a fast grain switch is triggered, i.e., if d2 > dl + T, then the a fastgrain switch is triggered. The value of T can be configured for each granular database. It can, for example, be written as metadata into the granular database. In one embodiment, the threshold is specified using a list of points in the descriptor space where each point has its own specific threshold value. Such a threshold can be calculated as a function of inter-grain distances in the database.

[0280] If the database is organized in clusters, then a list of thresholds is more appropriate. For each cluster, its centroid can be chosen as a point where the threshold is specified. The value of the threshold could then be a function of the mean of distances of grains from the centroid. Thresholds can also be defined based on inter-cluster distances.

[0281] Fast grain switching with phase compensation for pitched sounds

[0282] In the case that a granular database describes a sound that has a prominent pitch, where the length of each grain is proportional to the pitch cycle, a switch of grains could cause problems with phase cancellation if a grain is not played completely before transitioning to the next using an overlap-add. If the play position of the currently playing grain is at 50% of the grain length and an overlap-add is initiated with the next grain, the two grains may be 180 degrees out-of-phase during the overlap-add period, as illustrated in FIG. 16A, and this could lead to severe cancellations where the two signals more or less cancel each other out.

[0283] As shown in FIG. 16A, grain n is the currently rendered grain. At some time instant, t, into the grain a fast grain switch is triggered. In this case the next grain, Grain n+1, that was selected to transition to, is out-of-phase with Grain n in the overlap region. This will cause severe waveform distortion during the overlapping and a smooth transition will not be possible.

[0284] By keeping track of the current play position and calculating a corresponding starting position in the next grain, the phase can be preserved during the grain switch and phase cancellations are avoided, as illustrated in FIG. 16B. More specifically, FIG. 16B show that phase compensation is used so that grain n+1 is in phase with grain n during the overlap. The phase compensation is done by skipping the first part of Grain n+1 that corresponds to how much of Grain n was used before the fast grain switch was triggered.

[0285] For example, in one embodiment, in response to detecting that the fast grain condition is satisfied, the current play position, sp, of the first grain is stored. For instance, if, forexample, at the time the condition is determined to be satisfied, sample i in the grain currently being rendered has been played, but the sample i+1 has not, then set sp equal to i+1. Then obtain the following set, s[], of mixed samples as follows: s[n] = gl[n]grain_n[sp+n] + g2[n]grain_n+l[sp+n], for n=0 to X-l, where grain nf] is the set of samples that comprise the grain n, grain_n+l[] is the set of samples that comprise grain n+1, and X is the length of the overlap area.

[0286] FIG. 17 is a flowchart illustrating a process 1700 according to an embodiment. Process 1700 may begin with step si 702. Step si 702 comprises selecting a first grain from the plurality of grains. Step sl704 comprises rendering at least a first portion of the first grain. Step sl706 comprises, while the first grain is being rendered, obtaining information indicating an updated target position in the ND descriptor space and determining, based on the updated target position, whether a fast grain switch condition is satisfied. Step sl708 comprises, as a result of determining that the fast grain switch condition is satisfied, transitioning from the first grain to a second grain.

[0287] In some embodiments, transitioning from the first grain to the second grain comprises: ceasing the rendering of the first grain; and rendering at least a portion of the second grain.

[0288] In some embodiments, transitioning from the first grain to the second grain comprises: producing mixed samples by mixing at least a portion of the first grain with at least a first portion of the second grain; and rendering the mixed samples.

[0289] In some embodiments, the method further comprises: after rendering the mixed samples, rendering at least a second portion of the second grain.

[0290] In some embodiments, the method further comprises: prior to selecting the first grain, obtaining information indicating a first target position in the ND descriptor space, wherein the first grain is selected from the plurality of grains based on the first target position; and determining a distance between the updated target position and the first target position, wherein determining whether the fast grain switch condition is satisfied comprises comparing the distance to a predetermined value.

[0291] In some embodiments, the method further comprises determining a distance between a position of the first grain in the ND descriptor space and the updated target position, and determining whether the fast grain switch condition is satisfied comprises comparing the distance to a predetermined value.

[0292] In some embodiments, rendering at least a portion of the first grain comprises producing mixed samples using samples from the first grain and samples from a third grain, the first grain has a position in the ND descriptor space, the third grain has a position in the ND descriptor space, and the determination as to whether the fast grain switch condition is satisfied is further based on the position of the first grain in the ND descriptor space and the position of the third grain in the ND descriptor space.

[0293] In some embodiments, determining whether the fast grain switch condition is satisfied comprises: determining a first distance between the updated target position and the position of the first grain; determining a second distance between the updated target position and the position of the third grain; comparing the first and second distances; based on the comparison, determining that the first distance is less than the second distance; comparing the first distance to a predefined value; and determining whether the first grain switch condition is satisfied based on the comparison.

[0294] In some embodiments, determining whether the fast grain switch condition is satisfied comprises: determining a first distance between the updated target position and a straight line passing through the position of the first grain and the position of the third grain; comparing the first distance to a predefined value; and determining whether the fast grain switch condition is satisfied based on the comparison.

[0295] In some embodiments, the method further comprises, prior to selecting the first grain, obtaining information indicating a first target position in the ND descriptor space, wherein the first grain is selected from the plurality of grains based on the first target position, and determining whether the fast grain switch condition is satisfied comprises: determining a first distance between the first target position and a straight line passing through the position of the first grain and the position of the third grain; determining a second distance between the updated target position and the straight line; determining a difference between the first and seconddistances; comparing the difference to a predefined value; and determining whether the fast grain switch condition is satisfied based on the comparison.

[0296] Example Use Case

[0297] FIG. 7A illustrates an XR system 700, according to one embodiments, in which the embodiments disclosed herein may be applied. As shown in FIG. 7A, XR system 700 comprises an XR headset 720 (e.g., XR goggles, XR glasses, XR head mounted display (HMD), etc.) that is configured to be worn by a user and that is operable to display to the user an XR scene, such as, for example, a VR scene in which the user is virtually immersed, speakers 734 and 735 for producing sound for the user, and an input device 750 for receiving input from the user. In this example, input device 750 is in the form of a joystick.

[0298] As shown in FIG. 7B, XR headset 720 may comprise an orientation sensing unit 721, a position sensing unit 722, and an XR rendering device (XRRD) 724. In this embodiments, XRRD 724 includes an audio Tenderer (AR) 799 that includes GDB 104 and GRF 108.

[0299] Orientation sensing unit 721 is configured to detect a change in the orientation of the user and provides information regarding the detected change to XR rendering device 724. In some embodiments, XR rendering device 724 determines the absolute orientation (in relation to some coordinate system) given the detected change in orientation detected by orientation sensing unit 721. In some embodiments, orientation sensing unit 721 may comprise one or more accelerometers and / or one or more gyroscopes.

[0300] In addition to receiving input from sensing units 721 and 722, XR rendering device 724 may also receive input from input device 750 and may also obtain XR scene configuration information (e.g., the grain metadata). Based on these inputs and the XR scene configuration, XR rendering device 724 renders an XR scene in real-time for the user. That is, in real-time, XR rendering device produces XR content, including, for example, video data that is provided to a display driver 726 so that display driver 726 will display on a display screen 727 images included in the XR scene and audio data that is provided to speaker driver 728 so that speaker driver 728 will play audio for the using speakers 734 and 735. The audio data or portion thereof may be generated by GRF 108 using grains selected from GDB 104, as described above. While XR rendering device 724 is shown as being within XR headset 720 in this embodiment,in other embodiments XR rendering device 724, or one or more components thereof, such as, for example, GDB 104 and GRF 108, are located remotely from XR headset 720, in which case XR headset 720 and XR rendering device 724 have communication means (transmitter, receiver) for enabling XR rendering device 724 to transmit XR content to XR headset 720 (e.g., XR rendering device or components thereof may be implemented in the “cloud”).

[0301] FIG. 8 is a block diagram of an XR rendering device 724, according to some embodiments, for performing the methods disclosed herein. As shown in FIG. 8, XR rendering device 724 may comprise: processing circuitry (PC) 802, which includes one or more processors (P) 855 such as, for example, one or more general purpose microprocessors and / or one or more other processors, such as an application specific integrated circuit (ASIC), field-programmable gate arrays (FPGAs), and the like, which processors may be co-located in a single housing or in a single data center or may be geographically distributed (e.g., XR rendering device 724 may be a distributed computing apparatus comprising two or more computers or a monolithic computing apparatus consisting of a single computer); at least one network interface 848 (e.g., a physical interface or air interface) comprising a transmitter (Tx) 845 and a receiver (Rx) 847 for enabling XR rendering device 724 to transmit data to and receive data from other nodes connected to a network 110 (e.g., an Internet Protocol (IP) network) to which network interface 848 is connected (physically or wirelessly) (e.g., network interface 848 may be coupled to an antenna arrangement comprising one or more antennas for enabling XR rendering device 724 to wirelessly transmit / receive data); and a storage unit (a.k.a., “data storage system”) 808, which may include one or more non-volatile storage devices and / or one or more volatile storage devices. As illustrated in FIG. 8, GDB 104 may be stored in storage unit 808. In embodiments where PC 802 includes a programmable processor, a computer readable storage medium (CRSM) 842 may be provided. CRSM 842 may store a computer program (CP) 843 comprising computer readable instructions (CRI) 844. CRSM 842 may be a non-transitory computer readable medium, such as, magnetic media (e.g., a hard disk), optical media, memory devices (e.g., random access memory, flash memory), and the like. In some embodiments, the CRI 844 of computer program 843 is configured such that when executed by PC 802, the CRI causes XR rendering device 724 to perform steps described herein (e.g., steps described herein with reference to the flow charts). In other embodiments, XR rendering device 724 may be configured to perform steps described herein without the need for code. That is, for example, PC 802 mayconsist merely of one or more ASICs. Hence, the features of the embodiments described herein may be implemented in hardware and / or software.

[0302] Summary of Various Embodiments

[0303] Al. A method for rendering audio corresponding to an audio recording, wherein the audio recording is divided into a plurality of grains, the method comprising: selecting a first grain from the plurality of grains; rendering at least a first portion of the first grain; while the first grain is being rendered, obtaining information indicating an updated target position in an N- dimensional, ND, descriptor space and determining, based on the updated target position, whether a fast grain switch condition is satisfied; and as a result of determining that the fast grain switch condition is satisfied, transitioning from the first grain to a second grain.

[0304] A2. The method of embodiment Al, wherein transitioning from the first grain to the second grain comprises: ceasing the rendering of the first grain; and rendering at least a portion of the second grain.

[0305] A3. The method of embodiment Al, wherein transitioning from the first grain to the second grain comprises: producing mixed samples by mixing at least a portion of the first grain with at least a first portion of the second grain; and rendering the mixed samples.

[0306] A4. The method of embodiment A3, wherein the method further comprises: after rendering the mixed samples, rendering at least a second portion of the second grain.

[0307] A5. The method of any one of embodiments A1-A4, wherein the method further comprises: prior to selecting the first grain, obtaining information indicating a first target position in the ND descriptor space, wherein the first grain is selected from the plurality of grains based on the first target position; and determining a distance between the updated target position and the first target position, wherein determining whether the fast grain switch condition is satisfied comprises comparing the distance to a predetermined value.

[0308] A6. The method of any one of embodiments A1-A4, wherein the method further comprises determining a distance between a position of the first grain in the ND descriptor space and the updated target position, and determining whether the fast grain switch condition is satisfied comprises comparing the distance to a predetermined value.

[0309] A7. The method of any one of embodiments A1-A4,, wherein rendering at least a portion of the first grain comprises producing mixed samples using samples from the first grain and samples from a third grain, the first grain has a position in the ND descriptor space, the third grain has a position in the ND descriptor space, and the determination as to whether the fast grain switch condition is satisfied is further based on the position of the first grain in the ND descriptor space and the position of the third grain in the ND descriptor space.

[0310] A8. The method of embodiment A7, wherein determining whether the fast grain switch condition is satisfied comprises: determining a first distance between the updated target position and the position of the first grain; determining a second distance between the updated target position and the position of the third grain; comparing the first and second distances; based on the comparison, determining that the first distance is less than the second distance; comparing the first distance to a predefined value; and determining whether the first grain switch condition is satisfied based on the comparison.

[0311] A8-1. The method of embodiment 7, wherein determining whether the fast grain switch condition is satisfied comprises: determining a first distance between the updated target position and the position of the first grain; determining a second distance between the updated target position and the position of the third grain; comparing the first distance to a predefined value if the first distance is less than the second distance, otherwise comprising the second distance to the predefined value; and determining whether the first grain switch condition is satisfied based on the comparison.

[0312] A9. The method of embodiment A7, wherein determining whether the fast grain switch condition is satisfied comprises: determining a first distance between the updated target position and a straight line passing through the position of the first grain and the position of the third grain; comparing the first distance to a predefined value; and determining whether the fast grain switch condition is satisfied based on the comparison.

[0313] A10. The method of embodiment A7, wherein the method further comprises, prior to selecting the first grain, obtaining information indicating a first target position in the ND descriptor space, wherein the first grain is selected from the plurality of grains based on the first target position, and determining whether the fast grain switch condition is satisfied comprises: determining a first distance between the first target position and a straight line passing throughthe position of the first grain and the position of the third grain; determining a second distance between the updated target position and the straight line; determining a difference between the first and second distances; comparing the difference to a predefined value; and determining whether the fast grain switch condition is satisfied based on the comparison.

[0314] Al l. The method of embodiment Al, wherein transitioning from the first grain to the second grain comprises: determining a starting position in the second grain based on a current play position within the first grain; and rendering at least a portion of the second grain starting at the starting position.

[0315] A12. The method of any one of embodiment A1-A4 or Al l, wherein the method further comprises: determining a first distance between a position of the first grain in the ND descriptor space and a first target position, determining a second distance between a position of the first grain in the ND descriptor space and the updated target position, and determining a difference between the first distance and the second distance; and determining whether the fast grain switch condition is satisfied comprises comparing the difference to a predetermined value.

[0316] A13. The method of embodiment Al, wherein grain_n[] is an ordered set of samples that comprise the first grain, grain_n[i] is the sample of the first grain that was being rendered at the time it was determined that the fast grain switch condition is satisfied, transitioning from the first grain to the second grain comprises obtaining a set of X mixed samples, s[], wherein s[n] = gl[n]grain_n[sp+n] + g2[n]grain_n+l[sp+n], for n=0 to X-l, where gl[] is a first set of weight; g2[] is a second set of weights; sp is a function of i, grain_n+l[] is the set of samples that comprise the second grain n+1, and X is a determined overlap value.

[0317] Bl. A computer program comprising instructions which when executed by processing circuitry of a rendering device causes the rendering device to perform the method of any one of embodiments Al -Al 3.

[0318] B2. A carrier containing the computer program of embodiment Bl, wherein the carrier is one of an electronic signal, an optical signal, a radio signal, and a computer readable storage medium.

[0319] Cl. A rendering device for rendering audio corresponding to an audio recording, wherein the audio recording is divided into a plurality of grains, the rendering devicecomprising: memory; and processing circuitry coupled to the memory, wherein the rendering device is configured to perform a method that comprises: selecting a first grain from the plurality of grains; rendering at least a first portion of the first grain; while the first grain is being rendered, obtaining information indicating an updated target position in an N-dimensional, ND, descriptor space and determining, based on the updated target position, whether a fast grain switch condition is satisfied; and as a result of determining that the fast grain switch condition is satisfied, transitioning from the first grain to a second grain.

[0320] C2. The rendering device of embodiment Cl, wherein the rendering device is further configured to perform the method of any one of claims A2-A13.

[0321] While various embodiments are described herein, it should be understood that they have been presented by way of example only, and not limitation. Thus, the breadth and scope of this disclosure should not be limited by any of the above-described exemplary embodiments. Moreover, any combination of the above-described elements in all possible variations thereof is encompassed by the disclosure unless otherwise indicated herein or otherwise clearly contradicted by context.

[0322] Additionally, while the processes described above and illustrated in the drawings are shown as a sequence of steps, this was done solely for the sake of illustration. Accordingly, it is contemplated that some steps may be added, some steps may be omitted, the order of the steps may be re-arranged, and some steps may be performed in parallel. Further, as used herein “a” means “at least one” or “one or more.”

[0323] References

[0324] [1] US20180068487A1 : Systems and methods for simulating sounds of a virtual object using procedural audio (Disney Enterprises Inc.).

[0325] [2] Famell, Andy, "An introduction to procedural audio and its application in computer games,” Audio mostly conference Vol. 23. 2007.

[0326] [3] D. Gabor, “Theory of communication. Part 1 : The analysis of information,”Journal of the Institution of Electrical Engineers-Part III: Radio and Communication Engineering, vol. 93, no. 26, pp. 429-441, 1946.

[0327] [4] D. Schwarz, "Corpus-Based Concatenative Synthesis," in IEEE SignalProcessing Magazine, vol. 24, no. 2, pp. 92-104, March 2007, doi: 10.1109 / MSP.2007.323274.

[0328] [5] D. Schwarz "Distance mapping for corpus-based concatenative synthesis."Sound and Music Computing (SMC) 2011.

[0329] [6] Aaron Einbond and Diemo Schwarz. "Spatializing timbre with corpus-based concatenative synthesis." International Computer Music Conference Proceedings. Vol. 2010. International Computer Music Association, 2010.

[0330] [7] D. Schwarz. "A system for data-driven concatenative sound synthesis." DigitalAudio Effects (DAFx). 2000.

[0331] [8] D. Schwarz et al. "Real-time corpus-based concatenative synthesis with catart." 9th International Conference on Digital Audio Effects (DAFx). 2006.

[0332] [9] Diemo Schwarz and Norbert Schnell. "Descriptor-based sound texture sampling." Sound and music computing (SMC) 2010.

[0333]

[0010] Stefan Kersten and Hendrik Purwins. “Sound texture synthesis with HiddenMarkov Tree models in the wavelet domain”. In Proceedings of the International Conference on Sound and Music Computing (SMC), Barcelona, Spain, July 2010.

[0334]

[0011] Diemo Schwarz and Sean O'Leary. "Smooth granular sound texture synthesis by control of timbral similarity." Sound and Music Computing (SMC) 2015.

[0335]

[0012] Zechen Zhang, Nikunj Raghuvanshi, John Snyder, and Steve Marschner.2019. Acoustic texture rendering for extended sources in complex scenes. ACM Trans. Graph. 38, 6, Article 222 (December 2019), 9 pages, https: / / doi.org / 10.1145 / 3355089.3356566.

[0336]

[0013] Barrass, Stephen, and Matt Adcock. "Interactive granular synthesis of haptic contact sounds." Audio Engineering Society conference: 22nd international conference: virtual, synthetic, and entertainment audio. Audio Engineering Society, 2002.

[0337]

[0014] US20190094975A1 : Haptic Effect Conversion System Using GranularSynthesis: Immersion Corp.

[0338]

[0015] T. Park, J. Biguenet, Z. Li, C. Richardson, and T. Scharr, “Feature modulation synthesis (FMS),” in Proc. ICMC, (Copenhagen, Denmark), 2007.

[0339]

[0016] Fink, Marco, Martin Holters, and Udo Zolzer. "Signal-matched power- complementary cross-fading and dry -wet mixing." Proceedings of the 19th International Conference on Digital Audio Effects (DAFx-16). 2016.

Claims

CLAIMS1. A method (1700) for rendering audio corresponding to an audio recording (111), wherein the audio recording is divided into a plurality of grains, the method comprising: selecting (si 702) a first grain from the plurality of grains; rendering (sl704) at least a first portion of the first grain; while the first grain is being rendered, obtaining (sl706) information indicating an updated target position in an N-dimensional, ND, descriptor space and determining, based on the updated target position, whether a fast grain switch condition is satisfied; and as a result of determining that the fast grain switch condition is satisfied, transitioning (sl708) from the first grain to a second grain.

2. The method of claim 1, wherein transitioning from the first grain to the second grain comprises: ceasing the rendering of the first grain; and rendering at least a portion of the second grain.

3. The method of claim 1, wherein transitioning from the first grain to the second grain comprises: producing mixed samples by mixing at least a portion of the first grain with at least a first portion of the second grain; and rendering the mixed samples.

4. The method of claim 3, wherein the method further comprises: after rendering the mixed samples, rendering at least a second portion of the second grain.

5. The method of claim 1, wherein transitioning from the first grain to the second grain comprises: determining a starting position in the second grain based on a current play position within the first grain; and rendering at least a portion of the second grain starting at the starting position.

6. The method of any one of claims 1-5, wherein the method further comprises: prior to selecting the first grain, obtaining information indicating a first target position in the ND descriptor space, wherein the first grain is selected from the plurality of grains based on the first target position; and determining a distance between the updated target position and the first target position, wherein determining whether the fast grain switch condition is satisfied comprises comparing the distance to a predetermined value.

7. The method of any one of claims 1-5, wherein the method further comprises determining a distance between a position of the first grain in the ND descriptor space and the updated target position, and determining whether the fast grain switch condition is satisfied comprises comparing the distance to a predetermined value.

8. The method of any one of claims 1-5, wherein the method further comprises: determining a first distance between a position of the first grain in the ND descriptor space and a first target position; determining a second distance between a position of the first grain in the ND descriptor space and the updated target position; and determining a difference between the first distance and the second distance, and determining whether the fast grain switch condition is satisfied comprises comparing the difference to a predetermined value.

9. The method of any one of claims 1-5, wherein rendering at least a portion of the first grain comprises producing mixed samples using samples from the first grain and samples from a third grain, the first grain has a position in the ND descriptor space, the third grain has a position in the ND descriptor space, andthe determination as to whether the fast grain switch condition is satisfied is further based on the position of the first grain in the ND descriptor space and the position of the third grain in the ND descriptor space.

10. The method of claim 9, wherein determining whether the fast grain switch condition is satisfied comprises: determining a first distance between the updated target position and the position of the first grain; determining a second distance between the updated target position and the position of the third grain; comparing the first distance to a predefined value if the first distance is less than the second distance, otherwise comprising the second distance to the predefined value; and determining whether the first grain switch condition is satisfied based on the comparison.

11. The method of claim 9, wherein determining whether the fast grain switch condition is satisfied comprises: determining a first distance between the updated target position and a straight line passing through the position of the first grain and the position of the third grain; comparing the first distance to a predefined value; and determining whether the fast grain switch condition is satisfied based on the comparison.

12. The method of claim 9, wherein the method further comprises, prior to selecting the first grain, obtaining information indicating a first target position in the ND descriptor space, wherein the first grain is selected from the plurality of grains based on the first target position, and determining whether the fast grain switch condition is satisfied comprises: determining a first distance between the first target position and a straight line passing through the position of the first grain and the position of the third grain; determining a second distance between the updated target position and the straight line; determining a difference between the first and second distances;comparing the difference to a predefined value; and determining whether the fast grain switch condition is satisfied based on the comparison.

13. The method of claim 1, wherein grain nf] is an ordered set of samples that comprise the first grain, grain nfi] is the sample of the first grain that was being rendered at the time it was determined that the fast grain switch condition is satisfied, transitioning from the first grain to the second grain comprises obtaining a set of X mixed samples, s[], wherein s[n] = gl[n]grain_n[sp+n] + g2[n]grain_n+l[sp+n], for n=0 to X-l, where gl[] is a first set of weight; g2[] is a second set of weights; sp is a function of i, grain_n+l[] is the set of samples that comprise the second grain n+1, andX is a determined overlap value.

14. A computer program (843) comprising instructions (844) which when executed by processing circuitry of a rendering device (724) causes the rendering device to perform the method of any one of claims 1-13.

15. A carrier containing the computer program of claim 14, wherein the carrier is one of an electronic signal, an optical signal, a radio signal, and a computer readable storage medium (842).

16. A rendering device (724) for rendering audio corresponding to an audio recording (111), wherein the audio recording is divided into a plurality of grains, the rendering device comprising: memory (842); and processing circuitry (802) coupled to the memory, wherein the rendering device is configured to perform a method (1700) that comprises:selecting (si 702) a first grain from the plurality of grains; rendering (sl704) at least a first portion of the first grain; while the first grain is being rendered, obtaining (sl706) information indicating an updated target position in an N-dimensional, ND, descriptor space and determining, based on the updated target position, whether a fast grain switch condition is satisfied; and as a result of determining that the fast grain switch condition is satisfied, transitioning (sl708) from the first grain to a second grain.

17. The rendering device of claim 16, wherein the rendering device is further configured to perform the method of any one of claims 2-13.

Citation Information

Patent Citations

  • Systems and methods for simulating sounds of a virtual object using procedural audio

    US20180068487A1

  • Haptic effect conversion system using granular synthesis

    US20190094975A1