Grain selection for granular synthesis

The weighted grain selection method in granular synthesis addresses adaptive grain selection challenges, ensuring natural audio output by considering proximity, trends, and temporal history, suitable for devices with limited resources.

WO2025149312A1PCT designated stage expired Publication Date: 2025-07-17TELEFONAKTIEBOLAGET LM ERICSSON (PUBL)
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2024/086640
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-12
Filing Date
2024-12-16
Publication Date
2025-07-17

AI Technical Summary

Technical Problem

Existing granular synthesis systems face challenges in selecting grains adaptively and efficiently, leading to spurious outputs and unnatural sounding audio due to fixed neighborhood selection and random grain choices, which do not account for the dynamic nature of user inputs and virtual environments.

Method used

A method for selecting grains based on weighted selection, where each grain or grain group is assigned weights, considering proximity, descriptor trends, temporal history, and grain group probabilities, allowing for adaptive neighborhood size and real-time updates.

Benefits of technology

This approach enables natural-sounding audio progressions with low computational complexity, supporting devices with limited resources and ensuring responsive, reactive sound generation without delays.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2024086640_17072025_PF_FP_ABST
    Figure EP2024086640_17072025_PF_FP_ABST
Patent Text Reader

Abstract

A method for rendering audio corresponding to an audio recording; the audio recording is divided into a plurality of grains. The method includes selecting a grain from the plurality of grains and rendering the audio using the selected grain. Selecting a grain from the plurality of grains includes: i) a) selecting a grain group from a set of grain groups, wherein each grain group included in the set of grain groups is assigned a weight and the selection of the grain group from the set of grain groups is based on the assigned weights and b) selecting a grain from the selected grain group, or ii) a) assigning a final weight to each grain in a set of candidate grains and b) selecting a grain from the set of candidate grains based on the assigned final weights.
Need to check novelty before this filing date? Find Prior Art

Description

GRAIN SELECTION FOR GRANULAR SYNTHESISTECHNICAL FIELD

[0001] Disclosed are embodiments related to granular synthesis.BACKGROUND

[0002] Audio rendering is a process used for presenting audio, such as audio within an extended reality (XR) scene, such as, for example, a virtual reality (VR), an augmented reality (AR) scene, or mixed reality (MR) scene, in order to give a listener the impression that sound is coming from physical sources within the scene at a certain position. The presentation can be made through headphone speakers or other speakers. If the presentation is made via headphone speakers, the processing used is called binaural rendering and uses spatial cues of human spatial hearing that make it possible to determine from which direction sounds are coming. The cues involve inter-aural time delay (ITD), inter-aural level difference (ILD), and / or spectral difference.

[0003] Procedural audio refers to the creation of sound in real-time as a response to live input. As an example, consider the sound of a car engine in a virtual space where the sound changes based on the speed or acceleration or the car. This mechanism is commonly used in video games for better user experience. It is believed that for the use case of XR (e.g., AR or VR), there are many sounds that would benefit from being dynamically generated so that they can react to changes in the scene in real-time. For example, the sound generated when a user touches a surface or operates an engine. In reference [1], sounds of a virtual object, for example a sword, axe, or wand, are simulated based on their position and orientation. There is a base tone and an overtone. Both are modulated to change pitch, timbre, amplitude to convey speed of movement of the virtual object. The live input may come from a user via sensors, such as hand controllers or a headset, it could be control data generated in real-time by some software process such as a physics simulation or pre-defined automation data. Regardless of how the input data was generated, the audio Tenderer needs to handle incoming data and generate sound in response to this data in real-time.

[0004] There exist many different methods for procedural audio (see, e.g., reference[2]), including synthetic sound synthesis using audio processing modules, machine learning methods trained on real recordings, and concatenative synthesis methods that make use of original recordings and rearrange segments of these recordings to generate variations. Because the class of concatenative synthesis methods makes direct use of real recordings, the generated audio sounds very natural thereby enhancing user experience.

[0005] Granular synthesis is a type of concatenative synthesis where a sound recording is divided into small fragments called “grains.” (See, e.g., reference [3]). By a careful selection of the fragments (grains) at rendering time, a plausible dynamically changing sound can be generated.

[0006] A granular synthesis process includes two main steps: (1) grain extraction and(2) grain synthesis. Grain extraction refers to extraction of pertinent grains from the original longer recording. The extraction method depends on the type of sound source and the desired features to be extracted. Grain synthesis refers to the technique of selecting the appropriate order of grains; this selection of the ordering could also be based on the user input in real-time.

[0007] Many sound design tools support grain extraction and synthesis, such as, for example Soundseed grain for Audiokinectic Wwise, Alchemy for Logic Pro or AudioMotors for FMOD. The tools allow for manual or semi-automated extraction of grains by the sound designer and other simple manipulations. The designer can choose the grain length, the amplitude envelope or shape of each grain among other controls. In the case of AudioMotors, an automated grain extraction tool specialized for motor sounds is provided.

[0008] Grain extraction can be done manually by the sound designer or in a data-driven manner by identifying the relevant features of the audio for segmentation purposes. Relevant features include, for example, pitch period in the case of pitched sounds, spectral energy at a given frequency, mel-frequency cepstrum coefficients (MFCC), and local maxima of the amplitude envelope.

[0009] There are also many methods for granular synthesis. A common method is to select grains at random and perform overlap and add (OLA) operation. This method is not amenable to all types of sound sources and does not capture temporal correlation between adjacent grains.

[0010] Corpus-based concatenative synthesis (CBCS) methods are based on selecting grains from a corpus of sound segments that are sampled from a database of heterogeneous sound sources. They utilize descriptors that are associated to sound segments to organize the corpus and perform searches within the descriptor space to pick the next grain. Note that the concept of a descriptor is not limited to features of audio signal (see, e.g., reference [4]). A user can annotate grains with perceptual descriptors when a direct mapping between desired effect and feature in the audio signal is not possible.

[0011] The descriptor space is multi-dimensional with the number of dimensions being equal to the number of descriptors. Search for the appropriate grain is performed in a computationally efficient manner by utilizing weighted Euclidean distance between a target descriptor location (e.g., point or area) in the descriptor space (hereafter referred to as “target descriptor coordinate”) and grain locations in the descriptor space. Reference [5] proposes warping functions for the distance measure to better select the set of grains and also to avoid repetitions of previously rendered grains.

[0012] For efficient search in the descriptor space, kD-tree search is used. Either k- nearest neighbors of the target descriptor coordinate or grains that are within a radius ‘r’ from the target descriptor coordinate are chosen. In reference [6], the corpus is organized as zones so that grains from different zones are not picked consequently when k-nearest neighbor search is used.

[0013] The software CATERPILLAR (see reference [7]) performs concatenative synthesis in an offline setup where a sequence of target descriptors is given. The program uses Viterbi algorithm to identify the sequence of grains to match the target descriptors. The cost function is a combination of distance from target descriptor coordinate and concatenation cost which is based on similarity of consecutive grains.

[0014] CataRT on the other hand is a real-time system and so it chooses the subsequent grain at random from a set of grains that are the k-nearest neighbors or a radius with the target descriptor coordinate being the center (see, e.g., reference [8]).

[0015] For smoother transitions in granularly synthesized sound, reference [9] uses feature descriptors like pitch, loudness, spectral centroid, fundamental frequency, periodicity, and autocorrelation coefficient at lag 1. Feature descriptors are computed for every grain andcorrelation among these feature descriptors is captured using a Gaussian Mixture Model (GMM) from which grains are sampled for synthesis.

[0016] There are also several works, such as, for example, reference

[0010] , that model or assign transition probability between adjacent grains thereby establishing Markov model to generate or synthesize sound textures. Model parameters are estimated using recordings, but synthesis of sound does not directly use recordings unlike granular synthesis.

[0017] Reference

[0011] discusses granular synthesis where the next grain is picked based on feature descriptors of the current grain for a continuity in timbre. A kD-tree search is performed to select the candidate grains closest in Euclidean distance to the current grain in the feature descriptor space.SUMMARY

[0018] Certain challenges presently exist. For instance, in existing systems, the method of selecting grains, for example, choosing the k closest neighbors of a target descriptor coordinate in the descriptor space, is fixed and not adaptive, and this can lead to spurious outputs when the neighborhood is not compact and the granular database consists of clusters that do not have similar sizes. While in some systems (e.g., CataRT) a user can adaptively choose the neighborhood size, it is still limited to 3-D beyond which visualization is not possible and therefore the user cannot adjust it. Also, the grains in the neighborhood of the target descriptor coordinate are selected uniformly at random without making a distinction between most and least likely grains to be played. Accordingly, the conventional systems may not produce natural sounding output, especially when a user input does not stay long enough at a target descriptor coordinate. In such a case, it is desirable to render the most representative grain closest to the target descriptor coordinate than choose at random. Additionally, while some systems (e.g., CATERPILLAR) use a Viterbi algorithm to find the next grain that minimizes a cost function describing the discontinuity introduced by the concatenation of the next grain, this optimization serves to minimize the discontinuity but does not avoid unnatural sounding results due to repeating the same grain, or a pattern of grains. Furthermore, in many cases, it would be beneficial to be able control the probability of both individual and groups of grains and to be able to update these probabilities in real-time.

[0019] Accordingly, in one aspect there is provided a method for rendering audio corresponding to an audio recording, wherein the audio recording is divided into a plurality of grains. The method includes selecting a grain from the plurality of grains. The method also includes rendering the audio using the selected grain. Selecting a grain from the plurality of grains includes: i) a) selecting a grain group from a set of grain groups comprising a first grain group and a second grain group, wherein each grain group included in the set of grain groups is assigned a weight and the selection of the grain group from the set of grain groups is based on the assigned weights and b) selecting a grain from the selected grain group, or ii) a) assigning a final weight to each grain in a set of candidate grains comprising a first grain and a second grain, wherein assigning a final weight to each grain in the set of candidate grains comprises assigning a first final weight to the first grain and assigning a second final weight to the second grain and b) selecting a grain from the set of candidate grains based on the assigned final weights, wherein the first final weight assigned to the first grain is a function of a weight associated with a grain group to which the first grain belongs.

[0020] In another aspect there is provided an apparatus that is configured to perform a method for rendering audio corresponding to an audio recording, wherein the audio recording is divided into a plurality of grains. The method includes selecting a grain from the plurality of grains. The method also includes rendering the audio using the selected grain. Selecting a grain from the plurality of grains includes: i) a) selecting a grain group from a set of grain groups comprising a first grain group and a second grain group, wherein each grain group included in the set of grain groups is assigned a weight and the selection of the grain group from the set of grain groups is based on the assigned weights and b) selecting a grain from the selected grain group, or ii) a) assigning a final weight to each grain in a set of candidate grains comprising a first grain and a second grain, wherein assigning a final weight to each grain in the set of candidate grains comprises assigning a first final weight to the first grain and assigning a second final weight to the second grain and b) selecting a grain from the set of candidate grains based on the assigned final weights, wherein the first final weight assigned to the first grain is a function of a weight associated with a grain group to which the first grain belongs. The apparatus may include memory and processing circuitry coupled to the memory.

[0021] In another aspect there is provided a computer program comprising instructions which when executed by processing circuitry of an apparatus causes the apparatus to performany of the methods disclosed herein. In one embodiment, there is provided a carrier containing the computer program wherein the carrier is one of an electronic signal, an optical signal, a radio signal, and a computer readable storage medium.

[0022] An advantage of the embodiments in which one or more grains are selected based on a target descriptor coordinate is that they facilitate granular synthesis that generates natural progressions of grains. By using a weighted selection (i.e., a selection process that uses weights assigned to grains, or grain groups, to make the grain selection), a combination of exact control and natural variation of the sound can be achieved. The complexity of the solution is low, which enables the rendering of many granular sound sources also on devices with low computational power. Low real-time complexity also enables the solution to produce reactive and responsive sounds without delays.

[0023] An advantage of the embodiments in which grain groups are defined and each grain group is assigned a weight is that the burden of complexity is taken away from the database author and transferred instead to the rendering algorithm. The embodiments allow the database author to test many different weights by rendering and changing the pre-assigned weights. This involves less overhead than changing a model and re-estimating parameters. These embodiment provide an efficient way to update the weights of grains that is easy to use and requires little bitrate when updates are sent from a device that is not co-located with the renderer.BRIEF DESCRIPTION OF THE DRAWINGS

[0024] The accompanying drawings, which are incorporated herein and form part of the specification, illustrate various embodiments.

[0025] FIG. 1 illustrates a system according to an embodiment.

[0026] FIG. 2A illustrates an example two-dimensional descriptor space.

[0027] FIG. 2B illustrates an example two-dimensional descriptor space.

[0028] FIG. 3 A illustrates an example two-dimensional descriptor space.

[0029] FIG. 3B illustrates a process for determining an optimal k value for a given grain according to an embodiment.

[0030] FIG. 4A illustrates an example of a descriptor trajectory of an original recording.

[0031] FIG. 4B illustrates an example of a descriptor trajectory of an original recording and a generated descriptor trajectory.

[0032] FIG. 4C illustrates an example set of a descriptor trajectory of an original recording and a generated descriptor trajectory.

[0033] FIG. 5 is a flowchart illustrating a process according to an embodiment.

[0034] FIG. 6 is a flowchart illustrating a process according to an embodiment.

[0035] FIGS. 7A and 7B show a system according to some embodiments.

[0036] FIG. 8 is a block diagram of an apparatus according to some embodiments.

[0037] FIG. 9 is a flowchart illustrating a process according to an embodiment.DETAILED DESCRIPTION

[0038] FIG. 1 illustrates a system 100, according to some embodiments, for performing granular synthesis. System 100 includes a grain extraction unit 102 which extracts grains from original audio recordings 111. That is, grain extraction unit divides the original audio recordings into small fragments, called “grains.” The extracted grains are stored in a grain database (GDB) 104 (or simply “database” for short) that is accessed at rendering time by a grain scheduling unit 106, which may also be referred to as “grain selection unit” or “grain scheduler”, which is a component of a grain rendering function (GRF) 108. In some embodiments, there is one grain database per procedural audio source, each of which is available to GRF 108. Each grain stored in grain database is associated with one or more vectors of one or more descriptor values, each vector corresponding to a particular descriptor.

[0039] When creating grain database 104, an audio designer decides what aspects should be used as descriptors. In some cases, it might be features of the sound itself, such as pitch or loudness, but it could also be other aspects that relate to how the sound was generated, such as the speed of movement that generates a contact sound between two objects sliding against each other or the opening angle of a door that generates a screeching sound when opened and closed. The descriptors should be chosen so that the sound can be re-generated dynamically by the Tenderer given a target descriptor coordinate or trajectory.

[0040] As noted above, each grain stored in grain database is associated with one or more descriptor values. Accordingly, the grains of an original recording need to be annotated with the descriptor values. In the case that a descriptor is an audio feature, the descriptor value may be possible to measure directly from the audio signal itself. In other cases, the descriptor values need to be provided somehow as extra metadata of the recordings. This may be, for example, done by logging data from some sensors during the recording and providing this data in companion files. In some cases, the annotation can be done manually by creating a log of data that describes how a descriptor changes during the recording or it can be done manually for each extracted grain.

[0041] When extracting grains from the original recording(s), the descriptor values are stored as metadata for each grain. Using the descriptor values, each grain can be positioned in a multi-dimensioned descriptor space where the value of each descriptor describes a position along one axis within this space. If only one descriptor is used, the descriptor space is onedimensional (ID), but if more descriptors are used the dimensionality of the descriptor space increases.

[0042] FIG. 2A shows an example of a two-dimensional (2D) descriptor space where each circle represents a grain. As shown in FIG. 2A, each grain has a location (e.g., a point or area position) within the 2D descriptor space, this location is referred to as the grain coordinate.

[0043] In one embodiment, the descriptor metadata for a sequence of grains extracted from one recording describes a trajectory within the descriptor space, which corresponds to how the descriptors evolved during the original recording.

[0044] At rendering time, the scheduling of the grains, i.e., the selection of grains to render, is typically based on a target descriptor coordinate in the descriptor space. The target descriptor coordinate specifies what descriptor values the generated sound output should have, which means that grains close to that coordinate in the descriptor space are to be used most prominently. A target descriptor coordinate may come from many types of sources, such as a physics engine simulating the interaction of virtual bodies, live input parameters from hand controllers or other sensors, pre-defined automation parameters.

[0045] An aspect of granular synthesis rendering is that repetition of the same grain often sounds very unrealistic and artificial. If the grains are short, less than 50ms, repeating the same grain will result in a very metallic and static sound. If the grains are longer, the repetition will be heard like a repeating pattern, which often results in a sound that is not plausible.

[0046] The selection of grains needs to avoid repetition of the same grain but at the same time select grains that are close to the target descriptor coordinate in the descriptor space. This disclosure, therefore, uses, in some embodiment, a weighted selection (e.g., a weighted random selection) procedure that will generate ever evolving sequences of grains (i.e., an ordered set of grains) that closely follow the target descriptor coordinates. An example of a sequence of grains is: [grain-7, grain-6, grain-7, grain-9, grain-11, grain- 10],

[0047] In some embodiments, each grain may be assigned a predefined weight (denoted Pis) as well as a set of dynamic weights that may change over time. The predefined weight can be useful in cases where a grain is an outlier that should not be used too often but can add a realistic variation to the generated sound if used every now and then. Another use case is to use predefined weights to control the frequency of grains that represent e.g., bird chirps compared to grains that represent the background sound of a forest.

[0048] An input (e.g., a signal generated by a user interaction) that controls a procedural audio source is mapped to a target descriptor coordinate. Based on the target descriptor coordinate, one or more grains from a set of candidate grains are selected for rendering using a weighted selection (e.g., a weighted random selection or a selection where the grain with highest weight is selected). The selected grains can be rendered using standard granular synthesis methods where metadata regarding overlap percent and crossfade window can be specified by the sound designer beforehand.

[0049] The input that controls the procedural audio source can change in real-time; consequently, the target descriptor coordinate can change over time as the input changes (the target descriptor coordinate can also change over time even if the input does not change).

[0050] In one embodiment, the grain selection algorithm using weighted selection has the following steps.

[0051] Step 1 : Obtain a target descriptor coordinate (e.g., map an input, such as a user input or other input, to a target descriptor coordinate in a descriptor space).

[0052] Step 2: Determine the size of a neighborhood adaptively, e.g., calculate a k value depending on the target descriptor coordinate or calculate a radius value (r) depending on the target descriptor coordinate. Alternatively, obtain a pre-calculated value of k or radius from metadata of the grain database.

[0053] Step 3: Select a set of candidate grains from the database using the k value or radius value. For example, select the grains from the database that are the k-nearest neighbors of the target descriptor coordinate. This search can be performed using off-the-shelf computationally efficient algorithms like the kD-tree search. As another example, include in the set of candidate grains each grain having a grain coordinate that is within a distance of r from the target descriptor coordinate.

[0054] Step 4: Assign a final weight to each one of the grains in the set of candidate grains. The final weight assigned to a given grain may be based on:

[0055] i) the distance of the position of the grain in the descriptor space (i.e., the grain coordinate) from the target descriptor coordinate in the descriptor space,

[0056] ii) the difference in the trend of descriptors of a grain, compared to the target descriptor trajectory,

[0057] iii) the temporal history of previously used grains,

[0058] iv) the difference in time instant in the original recording between the grain and the previously used grain, if they are from the same recording, and / or

[0059] v) predefined weight of the grain in the database.

[0060] Step 5: Perform a weighted selection of grains from the set of candidate grains using the final weights assigned in step 4. For example, perform a weighted random selection, or, as another example, select the grain with the highest final weight or lowest final weight. In this manner, grains are selected based on the target descriptor coordinate and further based on the final weights assigned to the grains in the set of candidate grains.

[0061] If the target descriptor coordinate changes, the steps are repeated. Otherwise, the weighted selection of grains continues (step 4 onwards) with changes made to the weights on the basis of temporal history of previous grains.

[0062] Step 2 - Determination of value of k.

[0063] Conventionally, the value of k is a user-defined constant. It is not desirable, however, to keep the value of k fixed at all times because doing so could lead to choosing too few or too many grains which in turn could lead to under-utilization of the grains or selecting grains that are dissimilar to the target descriptor value, respectively.

[0064] Accordingly, this disclosure provides, in one embodiment, an adaptive choice of k based on the density of grains available in an area surrounding the target descriptor coordinate.

[0065] An example of why one should use different values of k for different target descriptor coordinates is illustrated in FIG. 2A and 2B. In FIG. 2 A, the target descriptor coordinate is close to a cluster of 3 grains.

[0066] In FIG. 2B, however, the target descriptor coordinate is in the vicinity of more grains. In these scenarios, it is not optimal to use the same value of k. For the case of FIG. 2A, k=3 is appropriate. If k > 3, then this will lead to choosing grains that are not in the cluster and therefore lead to a discontinuity in the texture of rendered sound. In FIG. 2B, if k = 3, then too few grains are selected. A larger value here will lead to richer textures with less repetition since a variety of grains can be selected.

[0067] The value of k should be chosen adaptively based on the target descriptor coordinate in the descriptor space. In one embodiment, each grain in the database is assigned an optimum k value. Then, for a target descriptor coordinate, the value of k is set equal to the optimal k value assigned to the grain that is closest to the target descriptor coordinate.

[0068] The assignment of optimal k per grain in the database can be performed offline or during the construction of the grain database. Either all distances from grain i to other grains in the database are recorded or we can have a threshold on the maximum number of neighbors to stop the distance computation.

[0069] Then the distances are sorted from lowest to highest. The difference between distances for consecutive values of k will have sudden jump at a value where the distanceincreases drastically. This is treated as a cut-off value for k and k + 1 is assigned to the grain.The addition of 1 is to include the grain itself in the value of k.

[0070] A criterion to determine the cut-off value would be to either set an absolute threshold on the difference in sorted distances between consecutive neighbors or to use normalized percentage increases in consecutive sorted distances.

[0071] Alternatively, one could also use the concept of adaptive radius to choose the set of candidate grains. The current method in literature involves using a fixed radius with target descriptor coordinate as the center of a circle and choose all the grains within this circle of fixed radius to be included in the set of candidate grains. By the same reasoning as above, it might be beneficial to change the radius adaptively based on the position of the target descriptor coordinate. There, instead of choosing different value of k, on would use different values of radius.

[0072] FIG. 3 A illustrates the computation of k for a certain grain 301 represented by the black circle. In FIG. 3A, grain 301 and its 6 corresponding neighbors are shown. In FIG. 3B, the sorted distances are shown for grain 301. Because the jump in distance values is observable for k=4, an optimum k value of 5 assigned to grain 301.

[0073] Step 3 - Weight Assignments

[0074] Each of the factors that influence the final weight assigned to a grain are described below. We denote the target descriptor as u and weight associated with grain i as pt. The descriptor index is denoted by j = 1,2,where D is the number of descriptors or dimensions of the descriptor space. The descriptor value of grain i at dimension j is given by xtj and that of the target as Uj.

[0075] The criteria below that influence the probability of choosing a grain are expressed as proportional relationships since the final value of the weight is obtained after normalizing i.e., ensuring that the weights corresponding to all k grains add to 1.

[0076] i) Distance from target descriptor coordinate

[0077] A distance metric is used to define the proximity between target descriptor coordinate and other descriptor coordinates corresponding to grains. An example of the distancemetric is a weighted Euclidean distance where the difference in coordinates in each dimension is weighted by the inverse of standard deviation of the corresponding descriptor values,

[0079] The closer a grain’s descriptor coordinate to the target descriptor coordinate, the higher the weight associated. Let dtbe the distance from grain i to the target descriptor coordinate. For example, the probability that grain i is chosen can be inversely proportional to the distance.

[0080] pi oc 1 / dj. Accordingly, one can set pii = 1 / dj.

[0081] In some cases, the different descriptors should not have the same amount of influence on the grain selection. For example, if one descriptor is the pitch of the sound and another is a descriptor that has less strong effect on the perceptual character of the sound, the distance in the dimension corresponding to the pitch may be given a higher weight than the distance in the dimension that corresponds to the other descriptor. This can be achieved by adding an extra variable weight, my, to each dimension when calculating the distance:

[0083] ii) Difference in trend of target and original descriptor trajectories

[0084] When performing grain extraction, the descriptor coordinates are used as the main selection criterion. But the trend of the original descriptor trajectory also gives important information about the grain. For example, if an engine sound is modelled with a granular database with one descriptor that denotes the RPM of the engine, the trend of the descriptor corresponds to the acceleration or deceleration of the engine at the time instant in the recording that the grain was extracted from. A grain that was extracted from a portion of the recording when the engine was accelerating will have a pitch that is slightly lower at the start than at the end and will therefor fit best when the desired output is the sound of an accelerating engine.

[0085] In the more general case, where a multi-dimensional descriptor space is used, the descriptor trend is a vector that corresponds to the direction of the trajectory that describes how the descriptors were changing at the time of the original recording. Similarly, the trend of thetarget descriptor trajectory describes the direction that the target descriptor coordinate is moving in the descriptor space.

[0086] The trend describes both the direction and the rate of change. Referring again to the example of the engine, if the granular database includes grains that correspond to the same RPM but with different acceleration, the grains that correspond to a similar acceleration as that of the target descriptor trajectory should be preferred.

[0087] FIG. 4 A shows a descriptor trajectory of an original recording of an engine sound where the descriptor is the RPM of the engine. The recording is divided into 15 grains. FIG. 4B shows an example of the how grains may be selected to match a target descriptor trajectory where the descriptor trend of the grains is not considered during grain selection. As can be seen, grains with both increasing and decreasing RPM are used in combination. The resulting descriptor trajectory shows an irregular behavior which may result in a degradation in perceived quality, especially if the descriptor represents the pitch of the sound. In FIG. 4C, the grain selection also considers the descriptor trend so that only grains with decreasing RPM are used, which would result in a smoother sound.

[0088] Considering the descriptor trend during grain selection is extra important for sound sources where the character is different for different trends. For example, an engine may sound different when accelerating compared to when decelerating. Making sure to match the descriptor trend avoids problems where grains with different character are used together.

[0089] The descriptor trend of a grain can be calculated as the difference in descriptor value at the end of the grain as compared to the start of the grain divided by the duration of the grain. In the case of a multi-dimensional descriptor space the trend is a vector that describes the mean rate of change in descriptor coordinates during the grain in the original recording, for example with three descriptors the trend, tr,, of a grain would be a three-dimensional vector:100901 t ~ (U1END~U1START U2END ~U2START U3END ~U3STARTI \ Duration ' Duration ' Duration )

[0091] When selecting grains at rendering time, the trend of the target descriptor trajectory can be calculated similarly as the difference in descriptor coordinates since they were last updated divided by the time elapsed since they were last updated, Tu.

[0093] The difference in trend can then calculated as

[0094] -D ~ -G ~ ty

[0095] The weight assigned to a grain can then be calculated as a function of the norm of to, for example: pi2 = max (1.0, -— ), where mr is a variable that controls the how the probability I tD I decreases with increased difference in descriptor trend.

[0096] iii) Temporal history of previously used grains

[0097] A history of previously rendered grains is kept for a certain time window so as to not repeat them and decrease the probability of choosing the grain if it was already rendered. The decrease in probability is directly proportional to a function of the difference between the current time instant t and the last time instant at which the grain i was rendered.

[0099] If grain i has not been selected in the past or in a certain time window, the value of last time instant is set to zero, t-1= 0. The function f can be linear, quadratic, logarithmic or any monotonically increasing function of the argument t — t-1with non-negative output.

[0100] iv) Difference in time instant in original recording

[0101] For some sound sources, the way that the sound evolves may not be completely described by the changes in descriptors. Sometimes sound evolves in a way that depends on what happened earlier. For example, the screeching sound of an old door may have a slightly different screeching sound every time it is opened, even if the opening of it is done at the same speed etc. In these cases, the sound of grains with similar descriptor values and descriptor trend may sound very different and combining them may result in unnatural discontinuities that were not there in the original recording. A way to avoid these discontinuities is to assign a higher probability to grains that come from the same part of the original recording as the grain that was used previously. Grains that came from the same part of the original recording are expected to be closely related and resemble each other in character and are therefore good candidates when selecting the next grain.

[0102] In order to measure how close one grain from a particular recording is to another grain from the same recording, a time difference can be calculated which corresponds to the difference in time instant in the recording that the two grains were extracted from. If the grains were close, this time difference is small. To calculate the time difference between two grains, metadata that tells from which original recording each grain was extracted and at what time instant can be used. This metadata, therefore, provides a measure of closeness between the two grains. This metadata can be specified in a compact way as two values: a recording identifier assigned to the recording of origin and a timestamp that identifies the time instant in that recording at which the grain can be found.

[0103] A weight for a grain can then be calculated based on the difference timestamps in a way that a smaller difference gives a higher probability to choose the evaluated grain. For example, the weight, pt, of grain i could be calculated as:

[0104]

[0105] Where Rt is the recording identifier assigned to the recording from which grain i was extracted, Ro is the recording identifier assigned to the recording from which the previously rendered grain was extracted, ti and to are the respective timestamps for the two grains, and b is a design constant that sets the probability to use a grain that comes from another recording. The function f() takes the difference in time instant as input and calculates a probability for the grain. In one embodiment function f decreases linearly with an increase in difference in time instant with a slope specified with a variable a.

[0107] This has the effect that the probability reduces from 1.0 to b as the difference between the respective timestamps increases but never goes below b.

[0108] In one embodiment the recording identifier R, can also be set to refer to segments of a recording, i.e., one recording can be divided into segments where each segment has its own index. This can be useful when one recording contain segments that are not to be seen as related by the Tenderer.

[0109] Final Weight

[0110] The final weight assigned to grain i, denoted pi, is calculated by accumulating the weights assigned to grain i from each stage, i.e., pi = PtiPi2Pi3Pi4Pi5- In some embodiments, not all five weights are needed and can then be skipped by setting the corresponding weight to 1.0 or by completely exclude it from the calculation.

[0111] In one embodiment, the different weights are given different levels of influence on the final weight by modifying the individual weights with a fractional exponent, such as e.g.,

[0112] pi = pil1 / 2pi2Pi3Pi41 / 4Pi5,

[0113] where the weight pn from the first stage is made less influential by using a fractional exponent of % and the weight pt4 is made even less influential by using a fractional exponent of 1 / 4.

[0114] After the weights pi are computed for all the grains in the set of candidate grains from the steps above, they are normalized, i.e., divide each of them by the sum, so that the final weight becomes a probability value. p

[0115] Pi «- — , where k is the size of the set of candidate grains.Sj=i Pi

[0116] Then grains are selected based on their final weight (e.g., by sampling from the distribution). Note that the temporal history of a grain and the descriptor trend influence the final weight at every time instant a grain is rendered when the target descriptor coordinate does not change. Therefore, as long as the target descriptor remains constant, if at time instant t grain i is chosen, at time t+h, the final weight of that grain is influenced by temporal history and descriptor trend. The final weight of grain i becomes

[0117] To avoid the repetition of the grain i at t + 1, we can define / ( / i) as[Gons]

[0119] However, due to normalization operation, the probabilities of the other k — 1 grains change as well.

[0120] FIG. 5 is a flow chart illustrating a process 500, according to an embodiment, for rendering audio corresponding to an audio recording, wherein the audio recording is divided into a plurality of grains. Process 500 may begin in step s502.

[0121] Step s502 comprises obtaining a first target descriptor coordinate, wherein the first target descriptor coordinate identifies a first location in an N-dimensional descriptor space, where N > 0.

[0122] Step s504 comprises defining a first set of candidate grains based on the first target descriptor coordinate, wherein the first set of candidate grains comprises kl of the plurality of grains, where kl > 1.

[0123] Step s506 comprises assigning a final weight to each grain in the first set of candidate grains.

[0124] Step s508 comprises randomly selecting a grain from the first set of candidate grains based on the assigned final weights such that the probability that a given grain in the first set of candidate grains is selected is a function of the final weight assigned to the given grain.

[0125] Step s510 comprises rendering the selected grain.

[0126] FIG. 6 is a flow chart illustrating a process 600, according to an embodiment, for rendering audio corresponding to an audio recording, wherein the audio recording is divided into a plurality of grains. Process 600 may begin in step s602.

[0127] Step s602 comprises obtaining a first target descriptor coordinate, wherein the first target descriptor coordinate identifies a first location in an N-dimensional descriptor space, where N > 0.

[0128] Step s604 comprises defining a first set of candidate grains based on the first target descriptor coordinate, wherein the first set of candidate grains comprises kl of the plurality of grains, where kl > 1, and the first set of candidate grains comprises a first grain and a second grain.

[0129] Step s606 comprises assigning a final weight to each grain in the first set of candidate grains, wherein assigning a final weight to each grain in the first set of candidate grains comprises assigning a first final weight to the first grain and assigning a second final weight to the second grain. The first final weight assigned to the first grain is: a function of thedistance between the first target descriptor coordinate and the first grain, a function of the amount of time that has elapsed since the first grain was last rendered, a function of a trajectory associated with the first grain and a target trajectory, and / or a function of a measure of a closeness between the first grain and the most recently rendered grain.

[0130] Step s608 comprises selecting a grain from the first set of candidate grains based on the assigned final weights.

[0131] Step s610 comprises rendering the selected grain.

[0132] Grain Grouping Embodiments

[0133] An alternative way of authoring a granular database would be to define grain groups, i.e., define at least a first group of grains. An example is that of a recording of forest sounds wherein one group of grains includes grains containing bird sounds, another group of grains includes grains containing ambient forest sounds, such as rustling of winds, and another group of grains includes grains containing water sounds (e.g., the sound of a babbling brook). Each group can be assigned a weight.

[0134] For some sound sources, where it is easy to define descriptors that represent some aspect of the sound, e.g., the RPM of an engine, it is often beneficial to select grains based on a target position in the descriptor space. For other sound sources that are more stochastic in nature, it may be better to base the grain selection on grain groups instead. External input in the case of grain selection without a descriptor space could be an update to the weight assigned to a grain group. However, it is also possible to combine the two aspects, by using a descriptor space with a target position and at the same time define grain groups.

[0135] Granular rendering with the combination of descriptor space and grain groups

[0136] As described above, the weight assigned to grain i (pi) is function of a set of weights, such as weights, pn, pi2, pi3, pi4, and pis described above. In one embodiment, when grain i is a member of a grain group, pis is set equal to the weight assigned to the grain group to which grain i belongs (this weight is denoted PK;). Accordingly, in one embodiment the final non-normalized weight of a grain in the weighted random selection can be calculated as: pz =PiiPi2Pi3Pi4PiG.

[0137] In one embodiment, for each defined grain group, metadata for the grain group is stored in the grain database. The metadata includes a grain group identifier (GID) and a weight assigned to the grain group. In one embodiment the grain group metadata is stored as a list of GIDs with an associated weight, wherein the weight associated with a GID is the weight assigned to the grain group identified by the GID. Table 1 below illustrates such an example list:TABLE 1

[0138] In this example, four grain groups are defined, with GIDs 0-3, each with its initial weight. The weight of any group can be reduced by setting the weight to a low number. Setting the initial weight to a higher number will increase the probability that grains belonging to that group are selected at rendering time. The initial weight may also be set to numbers above 1.0. The list of pre-defined grain groups can be stored as global metadata for the granular database.

[0139] For each grain that is assigned to a grain group, the metadata for the grain may include the GID of the group to which the grain is assigned. If a grain is not assigned to any group then this is equivalent to a grain being assigned to a grain group with the weight set to 1.0.

[0140] With this embodiment, it is easy to increase or decrease the probability that grain i will be selected with respect to, for example, grains not assigned to any group. For example, assigning grain i to a grain group J and assigning a weight of less than 1.0 to group J, the probability of the grains belonging to this group J will be lower than the grains not belonging to any group. Similarly, if the weight of the group is set to above 1.0, the probability of the grains belonging to this group J will be higher than the grains not belonging to any group.

[0141] Granular rendering for databases with only grain groups

[0142] For granular databases that do not have the concept of a descriptor space but have grain groups defined, granular rendering may include selecting a grain group from the database using a weighted selection (e.g., using a weighted random selection algorithm based on the assigned weights for the groups or a selection where the grain group with highest weight is selected) and then selecting a grain from the selected grain group. In some embodiments, selecting a grain from the selected grain group may be done using a weighted selection procedure because, in some embodiments, each grain i in the selected grain group may be assigned a different weight, such as, for example a weight equal to pi. For example, in some embodiments, pi = pi3pi4pio. More generically, for granular databases that do not have the concept of a descriptor space, pi can be function of weights that are independent of the concept of a descriptor space, such as, for example, pi3 and pi4.

[0143] Using the same example as in the previous section with 4 probability groups, probability group 1 would be sampled or chosen more often compared to probability group 2. Then, once the probability group is chosen, any grain from that group can be chosen at random for rendering. In another embodiment, every grain group can also have a separate probability mass function (PMF) which means that different grains within a group have different assigned weights and in such an embodiment a grain from the grain group can be selected using a weighted random selection algorithm.

[0144] Dynamic Update of Grain Group Weights

[0145] At rendering time, a grain group’s weight may be updated dynamically. This means that the initial weight assigned to the grain group is updated and that all grains assigned to that group may be affected by this updated weight. By changing the weight assigned to one or more grain groups, the rendering can be changed to reflect a change in user input or other changes in the virtual scene. That is, a user input or change in the virtual scene may cause GRF 108 to change the weight assigned to one or more grain groups. For example, based on user location in the virtual scene, the probability of hearing birds can be higher than the probability of hearing a car; accordingly GRF 108 may increase the weight assigned to the grain group to which grains containing bird sounds belong and at the same time decrease the weight assigned to the grain group to which grains containing cars sounds belong. In another example where the sound of an engine is being rendered, GRF 108 may increase the weight of a grain grouprepresenting the sound an engine makes as the engine is running out of gas in response to detecting that a virtual engine in the virtual scene is running out of gas. Hence, by controlling the probability of this group of grains, the grains in the group may be made more frequent as the level of gas in the tank is getting critically low.

[0146] The updates of the weights of the grain groups may be done by direct functional calls to GRF 108, by a higher-layer application that monitors the state of the virtual scene and receives user input, where the GID of the grain group is specified along with a new updated weight, for example: UpdateGrainGroupWeight(GID, new weight).

[0147] Another way to update the probabilities of the grain groups is to include the updates into a bitstream update package that is decoded by a decoder and then fed to the renderer. The proposed method then offer a bitrate efficient way to update the probability of many grains, since one update is just a group index (GID) and the new probability, and this update could potentially change the probability of thousands of grains. In the bitstream update package only the group index to update and the new probability need to be added.

[0148] Conditional Weights

[0149] Each grain can be assigned one or more conditional weights with respect to one or more grain groups. For example, if a grain belongs to two grain groups, then the grain may be assigned two conditional weights, a first conditional weight associated with the first grain group and a second conditional weight associated with the second grain group. Table 2 below illustrates a metadata table for a grain i, which metadata table includes a record for each grain group to which grain i belongs, where the record includes a first field storing the GID of a grain group to which grain i belongs and a second field storing a conditional weight for grain i.TABLE 2 - Metadata Record for Grain i

[0150] In some embodiments, a grain may belong only to one group, in which case the table will have a single record.

[0151] When defining a grain group, an author could specify that, if a grain group is selected (or all candidate grains belong to the same grain group), then GRF 108 should treat each grain included the grain group as having the same conditional weight, in which case no further grain group metadata is needed to be added to the metadata record for the grain group. This is referred to as a “uniform” assignment of weights to the grains in a grain group. Likewise, the author could specify that, if the grain group is selected by GRF 108 (or all candidate grains belong to the same grain group), then GRF 108 assigns a conditional weight to each grain in the grain group such that it is possible that one or more grains in the group have a unique weight assigned to it; in this case additionally metadata (e.g. a flag) should be added to the metadata record for the grain group to indicate that each grain in the group should be assigned a group weight. This is referred to as a “non-uniform assignment”. Accordingly, in some embodiments, if a grain belongs to more than one group, the final weight assigned to the grain depends on the selected grain group.

[0152] The non-uniform assignment is beneficial when the quality of grains belonging to a certain grain group can differ. Higher quality grains in a grain group can be assigned higher per-grain weight values. Similarly, another criterion could be the length of the grains wherein longer grains are assigned higher per-grain weight values compared to shorter grains.

[0153] A conditional weight may be used in the case where a set of candidate grains are first selected by GRF 108 and each candidate grain belongs to the same group. The candidate grains may be selected from the database based on a target position and using a KD-tree search, as described above, where some of the grains closest to the target position in the descriptor space are chosen for further evaluation. As an example, assuming that all of the candidate grains belong to group J and that grain i is one of the candidate grains, then the final weight assigned to grain i may be calculated as: pz = piipi2pi3pi4 pispicj, where picj is grain i’s conditional weight with respect to group J. As a concrete example when group is the group with GID=2, then, as shown in Table 2 above, picj is equal to 1.3. In some embodiments, pis = PiG (i.e. the weight assigned to the grain group with GID=2). In this embodiment, the conditional weight is only active along with the weight value for the grain group (pio), i.e., picjcannot be used as an individual grain multiplicative factor in the absence of PiGor when grains from a different probability group are among the candidate grains. This is because picj a conditional probability value.

[0154] In another embodiment, the final weight assigned to grain i may be calculated as: pz = piipi2Pi3Pi4Picj, In other words, in this embodiment, pis is set equal to picj rather than being equal to piG.

[0155] In the case that the grains of a granular database are not specified with a descriptor space position, the step of selecting a group of candidate grains may instead be done by first selecting a grain group. Then the grains belonging to the selected grain group are the candidate grains and the conditional per-grain weights of those grains will be applied. As an example, assuming that the grain group with GID=1 is selected, then the final weight assigned to grain i, which is a member of the grain group with GID=1, may be calculated as: pz = PiiPi2Pi3Pi4Pici, where as shown in table 2 above pici = 0.7.

[0156] As noted above, the selection of a grain group can be done using a weighted random selection algorithm where the probabilities in the random selection are based on the weights of the different groups. For example, assuming three grain groups assigned weights wl, w2, and w3, respectively, then the probability that the first group is selected is wl / (wl+w2+w3), the probability that the second group is selected is w2 / (wl+w2+w3), and the probability that the third group is selected is w3 / (wl+w2+w3).

[0157] FIG. 9 is a flow chart illustrating a process 900, according to an embodiment, for rendering audio corresponding to an audio recording, wherein the audio recording is divided into a plurality of grains. Process 900 may begin in step s902. Step s902 comprises selecting a grain from the plurality of grains. Step s904 comprises rendering the audio using the selected grain.

[0158] In one embodiment, step s902 comprises selecting a grain group from a set of grain groups comprising a first grain group and a second grain group, wherein each grain group included in the set of grain groups is assigned a weight and the selection of the grain group from the set of grain groups is based on the assigned weights and b) selecting a grain from the selected gram group.

[0159] In another embodiment, step s902 comprises assigning a final weight to each grain in a set of candidate grains comprising a first grain and a second grain, wherein assigning a final weight to each grain in the set of candidate grains comprises assigning a first final weight to the first grain and assigning a second final weight to the second grain and b) selecting a grain from the set of candidate grains based on the assigned final weights, wherein the first final weight assigned to the first grain is a function of a weight associated with a grain group to which the first grain belongs.

[0160] Example Use Case

[0161] FIG. 7A illustrates an XR system 700, according to one embodiments, in which the embodiments disclosed herein may be applied. As shown in FIG. 7A, XR system 700 comprises an XR headset 720 (e.g., XR goggles, XR glasses, XR head mounted display (HMD), etc.) that is configured to be worn by a user and that is operable to display to the user an XR scene, such as, for example, a VR scene in which the user is virtually immersed, speakers 734 and 735 for producing sound for the user, and an input device 750 for receiving input from the user. In this example, input device 750 is in the form of a joystick.

[0162] As shown in FIG. 7B, XR headset 720 may comprise an orientation sensing unit 721, a position sensing unit 722, and an XR rendering device (XRRD) 724. In this embodiments, XRRD 724 includes an audio Tenderer (AR) 799 that includes GDB 104 and GRF 108.

[0163] Orientation sensing unit 721 is configured to detect a change in the orientation of the user and provides information regarding the detected change to XR rendering device 724. In some embodiments, XR rendering device 724 determines the absolute orientation (in relation to some coordinate system) given the detected change in orientation detected by orientation sensing unit 721. In some embodiments, orientation sensing unit 721 may comprise one or more accelerometers and / or one or more gyroscopes.

[0164] In addition to receiving input from sensing units 721 and 722, XR rendering device 724 may also receive input from input device 750 and may also obtain XR scene configuration information (e.g., the grain metadata). Based on these inputs and the XR scene configuration, XR rendering device 724 renders an XR scene in real-time for the user. That is, in real-time, XR rendering device produces XR content, including, for example, video data that isprovided to a display driver 726 so that display driver 726 will display on a display screen 727 images included in the XR scene and audio data that is provided to speaker driver 728 so that speaker driver 728 will play audio for the using speakers 734 and 735. The audio data or portion thereof may be generated by GRF108 using grains selected from GDB 104, as described above. While XR rendering device 724 is shown as being within XR headset 720 in this embodiment, in other embodiments XR rendering device 724, or one or more components thereof, such as, for example, GDB 104 and GRF 108, are located remotely from XR headset 720, in which case XR headset 720 and XR rendering device 724 have communication means (transmitter, receiver) for enabling XR rendering device 724 to transmit XR content to XR headset 720 (e.g., XR rendering device or components thereof may be implemented in the “cloud”).

[0165] FIG. 8 is a block diagram of an XR rendering device 724, according to some embodiments, for performing the methods disclosed herein. As shown in FIG. 8, XR rendering device 724 may comprise: processing circuitry (PC) 802, which includes one or more processors (P) 855 such as, for example, one or more general purpose microprocessors and / or one or more other processors, such as an application specific integrated circuit (ASIC), field-programmable gate arrays (FPGAs), and the like, which processors may be co-located in a single housing or in a single data center or may be geographically distributed (e.g., XR rendering device 724 may be a distributed computing apparatus comprising two or more computers or a monolithic computing apparatus consisting of a single computer); at least one network interface 848 (e.g., a physical interface or air interface) comprising a transmitter (Tx) 845 and a receiver (Rx) 847 for enabling XR rendering device 724 to transmit data to and receive data from other nodes connected to a network 110 (e.g., an Internet Protocol (IP) network) to which network interface 848 is connected (physically or wirelessly) (e.g., network interface 848 may be coupled to an antenna arrangement comprising one or more antennas for enabling XR rendering device 724 to wirelessly transmit / receive data); and a storage unit (a.k.a., “data storage system”) 808, which may include one or more non-volatile storage devices and / or one or more volatile storage devices. As illustrated in FIG. 8, GDB 104 may be stored in storage unit 808. In embodiments where PC 802 includes a programmable processor, a computer readable storage medium (CRSM) 842 may be provided. CRSM 842 may store a computer program (CP) 843 comprising computer readable instructions (CRI) 844. CRSM 842 may be a non-transitory computer readable medium, such as, magnetic media (e.g., a hard disk), optical media, memory devices(e.g., random access memory, flash memory), and the like. In some embodiments, the CRI 844 of computer program 843 is configured such that when executed by PC 802, the CRI causes XR rendering device 724 to perform steps described herein (e.g., steps described herein with reference to the flow charts). In other embodiments, XR rendering device 724 may be configured to perform steps described herein without the need for code. That is, for example, PC 802 may consist merely of one or more ASICs. Hence, the features of the embodiments described herein may be implemented in hardware and / or software.

[0166] Summary of Various Embodiments

[0167] Al . A method for rendering audio corresponding to an audio recording, wherein the audio recording is divided into a plurality of grains, the method comprising: selecting a grain from the plurality of grains; and rendering the audio using the selected grain, wherein selecting a grain from the plurality of grains comprises: i) a) selecting a grain group from a set of grain groups comprising a first grain group and a second grain group, wherein each grain group included in the set of grain groups is assigned a weight and the selection of the grain group from the set of grain groups is based on the assigned weights and b) selecting a grain from the selected grain group, or ii) a) assigning a final weight to each grain in a set of candidate grains comprising a first grain and a second grain, wherein assigning a final weight to each grain in the set of candidate grains comprises assigning a first final weight to the first grain and assigning a second final weight to the second grain and b) selecting a grain from the set of candidate grains based on the assigned final weights, wherein the first final weight assigned to the first grain is a function of a weight associated with a grain group to which the first grain belongs.

[0168] A2. The method of embodiment Al, wherein selecting a grain from the plurality of grains comprises: selecting a grain group from a set of grain groups comprising a first grain group and a second grain group, wherein each grain group included in the set of grain groups is assigned a weight and the selection of the grain group from the set of grain groups is based on the assigned weights; and selecting a grain from the selected grain group.

[0169] A3. The method of embodiment A2, wherein selecting a grain group from the set of grain groups comprises using a weighted random selection algorithm and the assigned weights to select a grain group from the set of grain groups.

[0170] A4. The method of embodiment A2 or A3, wherein selecting a grain from the selected grain group comprises: assigning a final weight to each grain in the selected grain group; and selecting a grain from the selected grain group based on the assigned final weights.

[0171] A5. The method of embodiment A4, wherein assigning a final weight to each grain in the selected grain group comprises assigning a first final weight to a first grain included in the selected grain group, and the first final weight is a function of a conditional weight associated with both the first grain and the selected grain group.

[0172] A6. The method of embodiment A5, wherein the method further comprises: prior to assigning the first final weight to the first grain, obtaining metadata for the first grain, and the metadata for the first grain comprises a record comprising a grain group identifier, GID, assigned to the selected grain group and the conditional weight.

[0173] A7. The method of embodiment A5 or A6, wherein the first final weight is also a function of: a distance between a first target position and a position of the first grain, an amount of time that has elapsed since the first grain was last rendered, a trajectory associated with the first grain and a target trajectory, and / or a measure of a closeness between the first grain and a previously rendered grain.

[0174] A8. The method of any one of embodiments A2-A7, wherein the method further comprises obtaining first metadata for the first grain group second metadata for the second grain group, the first metadata comprises a first grain group identifier, GID, assigned to the first grain group and a first weight assigned to the first grain group, and the second metadata comprises a second GID assigned to the second grain group and a second weight assigned to the second grain group.

[0175] A9. The method of embodiment Al, wherein selecting a grain from the plurality of grains comprises: assigning a final weight to each grain in the set of candidate grains; and selecting a grain from the set of candidate grains based on the assigned final weights, wherein the first final weight assigned to the first grain is a function of a weight associated with a grain group to which the first grain belongs.

[0176] A10. The method of embodiment A9, wherein selecting a grain from the plurality of grains further comprises, prior to assigning a final weight to each grain in the set of candidategrains, obtaining information identifying a first target position in an N-dimensional descriptor space, where N > 0, and the first set of candidate grains is defined based on the first target position.

[0177] Al 1. The method of embodiment A9 or Al 0, wherein selecting a grain from the plurality of grains further comprises, prior to assigning the first final weight to the first grain: determining the grain group to which the first grain belongs; and obtaining metadata for the grain group to which the first grain belongs, wherein the metadata for the grain group to which the first grain belongs comprises a weight assigned to the grain group, and the weight associated with the grain group to which the first grain belongs is the weight assigned to the grain group to which the first grain belongs.

[0178] Al 2. The method of embodiment A9 or A10, wherein selecting a grain from the plurality of grains further comprises, prior to assigning the first final weight to the first grain: obtaining metadata for the first grain; and obtaining from the metadata for the first grain the weight associated with the grain group to which the first grain belongs.

[0179] Al 3. The method of embodiment Al 2, wherein selecting a grain from the plurality of grains further comprises, prior to assigning the first final weight to the first grain, determining that each grain included in the first set of candidate grains is a member of the grain group to which the first grain belongs, and the step of obtaining the weight associated with the grain group to which the first grain belongs is performed as a result of determining that each grain included in the first set of candidate grains is a member of the grain group to which the first grain belongs.

[0180] Al 4. The method of any one of embodiments A9-A13, wherein the first final weight is also a function of: a distance between a first target position and a position of the first grain, an amount of time that has elapsed since the first grain was last rendered, a trajectory associated with the first grain and a target trajectory, and / or a measure of a closeness between the first grain and a previously rendered grain.

[0181] Bl. A computer program comprising instructions which when executed by processing circuitry of a rendering device causes the rendering device to perform the method of any one of embodiments Al -Al 4.

[0182] B2. A carrier containing the computer program of embodiment Bl, wherein the carrier is one of an electronic signal, an optical signal, a radio signal, and a computer readable storage medium.

[0183] Cl. A rendering device for rendering audio corresponding to an audio recording, wherein the audio recording is divided into a plurality of grains, the rendering device comprising: memory; and processing circuitry coupled to the memory, wherein the rendering device is configured to perform a method that comprises: selecting a grain from the plurality of grains; and rendering the audio using the selected grain, wherein selecting a grain from the plurality of grains comprises: i) a) selecting a grain group from a set of grain groups comprising a first grain group and a second grain group, wherein each grain group included in the set of grain groups is assigned a weight and the selection of the grain group from the set of grain groups is based on the assigned weights and b) selecting a grain from the selected grain group, or ii) a) assigning a final weight to each grain in a set of candidate grains comprising a first grain and a second grain, wherein assigning a final weight to each grain in the set of candidate grains comprises assigning a first final weight to the first grain and assigning a second final weight to the second grain and b) selecting a grain from the set of candidate grains based on the assigned final weights, wherein the first final weight assigned to the first grain is a function of a weight associated with a grain group to which the first grain belongs.

[0184] C2. The rendering device of embodiment Cl, wherein the rendering device is further configured to perform the method of any one of claims A2-A14.

[0185] While various embodiments are described herein, it should be understood that they have been presented by way of example only, and not limitation. Thus, the breadth and scope of this disclosure should not be limited by any of the above-described exemplary embodiments. Moreover, any combination of the above-described elements in all possible variations thereof is encompassed by the disclosure unless otherwise indicated herein or otherwise clearly contradicted by context.

[0186] Additionally, while the processes described above and illustrated in the drawings are shown as a sequence of steps, this was done solely for the sake of illustration. Accordingly, it is contemplated that some steps may be added, some steps may be omitted, the order of the stepsmay be re-arranged, and some steps may be performed in parallel. Further, as used herein “a” means “at least one” or “one or more.”

[0187] References

[0188] [1] US20180068487A1 : Systems and methods for simulating sounds of a virtual object using procedural audio (Disney Enterprises Inc.).

[0189] [2] Farnell, Andy, "An introduction to procedural audio and its application in computer games,” Audio mostly conference Vol. 23. 2007.

[0190] [3] D. Gabor, “Theory of communication. Part 1 : The analysis of information,”Journal of the Institution of Electrical Engineers-Part III: Radio and Communication Engineering, vol. 93, no. 26, pp. 429-441, 1946.

[0191] [4] D. Schwarz, "Corpus-Based Concatenative Synthesis," in IEEE SignalProcessing Magazine, vol. 24, no. 2, pp. 92-104, March 2007, doi: 10.1109 / MSP.2007.323274.

[0192] [5] D. Schwarz "Distance mapping for corpus-based concatenative synthesis."Sound and Music Computing (SMC) 2011.

[0193] [6] Aaron Einbond and Diemo Schwarz. "Spatializing timbre with corpus-based concatenative synthesis." International Computer Music Conference Proceedings. Vol. 2010. International Computer Music Association, 2010.

[0194] [7] D. Schwarz. "A system for data-driven concatenative sound synthesis." DigitalAudio Effects (DAFx). 2000.

[0195] [8] D. Schwarz et al. "Real-time corpus-based concatenative synthesis with catart." 9th International Conference on Digital Audio Effects (DAFx). 2006.

[0196] [9] Diemo Schwarz and Norbert Schnell. "Descriptor-based sound texture sampling." Sound and music computing (SMC) 2010.

[0197]

[0010] Stefan Kersten and Hendrik Purwins. “Sound texture synthesis with HiddenMarkov Tree models in the wavelet domain”. In Proceedings of the International Conference on Sound and Music Computing (SMC), Barcelona, Spain, July 2010.

[0198]

[0011] Diemo Schwarz and Sean O'Leary. "Smooth granular sound texture synthesis by control of timbral similarity." Sound and Music Computing (SMC) 2015.

[0199]

[0012] Zechen Zhang, Nikunj Raghuvanshi, John Snyder, and Steve Marschner.2019. Acoustic texture rendering for extended sources in complex scenes. ACM Trans. Graph.38, 6, Article 222 (December 2019), 9 pages, https: / / doi.org / 10.1145 / 3355089.3356566.

[0200]

[0013] Barrass, Stephen, and Matt Adcock. "Interactive granular synthesis of haptic contact sounds." Audio Engineering Society conference: 22nd international conference: virtual, synthetic, and entertainment audio. Audio Engineering Society, 2002.

[0201]

[0014] US20190094975A1: Haptic Effect Conversion System Using GranularSynthesis: Immersion Corp.

Claims

CLAIMS1. A method (900) for rendering audio corresponding to an audio recording (111), wherein the audio recording is divided into a plurality of grains, the method comprising: selecting (s902) a grain from the plurality of grains; and rendering (s904) the audio using the selected grain, wherein selecting a grain from the plurality of grains comprises: i) a) selecting a grain group from a set of grain groups comprising a first grain group and a second grain group, wherein each grain group included in the set of grain groups is assigned a weight and the selection of the grain group from the set of grain groups is based on the assigned weights and b) selecting a grain from the selected grain group, or ii) a) assigning a final weight to each grain in a set of candidate grains comprising a first grain and a second grain, wherein assigning a final weight to each grain in the set of candidate grains comprises assigning a first final weight to the first grain and assigning a second final weight to the second grain and b) selecting a grain from the set of candidate grains based on the assigned final weights, wherein the first final weight assigned to the first grain is a function of a weight associated with a grain group to which the first grain belongs.

2. The method of claim 1 , wherein selecting a grain from the plurality of grains comprises: selecting a grain group from a set of grain groups comprising a first grain group and a second grain group, wherein each grain group included in the set of grain groups is assigned a weight and the selection of the grain group from the set of grain groups is based on the assigned weights; and selecting a grain from the selected grain group.

3. The method of claim 2, wherein selecting a grain group from the set of grain groups comprises using a weighted random selection algorithm and the assigned weights to select a grain group from the set of grain groups.

4. The method of claim 2 or 3, wherein selecting a grain from the selected grain group comprises: assigning a final weight to each grain in the selected grain group; and selecting a grain from the selected grain group based on the assigned final weights.

5. The method of claim 4, wherein assigning a final weight to each grain in the selected grain group comprises assigning a first final weight to a first grain included in the selected grain group, and the first final weight is a function of a conditional weight associated with both the first grain and the selected grain group.

6. The method of claim 5, wherein the method further comprises: prior to assigning the first final weight to the first grain, obtaining metadata for the first grain, and the metadata for the first grain comprises a record comprising a grain group identifier, GID, assigned to the selected grain group and the conditional weight.

7. The method of claim 5 or 6, wherein the first final weight is also a function of: a distance between a first target position and a position of the first grain, an amount of time that has elapsed since the first grain was last rendered, a trajectory associated with the first grain and a target trajectory, and / or a measure of a closeness between the first grain and a previously rendered grain.

8. The method of any one of claims 2-7, wherein the method further comprises obtaining first metadata for the first grain group second metadata for the second grain group,the first metadata comprises a first grain group identifier, GID, assigned to the first grain group and a first weight assigned to the first grain group, and the second metadata comprises a second GID assigned to the second grain group and a second weight assigned to the second grain group.

9. The method of claim 1 , wherein selecting a grain from the plurality of grains comprises: assigning a final weight to each grain in the set of candidate grains; and selecting a grain from the set of candidate grains based on the assigned final weights, wherein the first final weight assigned to the first grain is a function of a weight associated with a grain group to which the first grain belongs.

10. The method of claim 9, wherein selecting a grain from the plurality of grains further comprises, prior to assigning a final weight to each grain in the set of candidate grains, obtaining information identifying a first target position in an N-dimensional descriptor space, where N > 0, and the first set of candidate grains is defined based on the first target position.

11. The method of claim 9 or 10, wherein selecting a grain from the plurality of grains further comprises, prior to assigning the first final weight to the first grain: determining the grain group to which the first grain belongs; and obtaining metadata for the grain group to which the first grain belongs, wherein the metadata for the grain group to which the first grain belongs comprises a weight assigned to the grain group, and the weight associated with the grain group to which the first grain belongs is the weight assigned to the grain group to which the first grain belongs.

12. The method of claim 9 or 10, wherein selecting a grain from the plurality of grains further comprises, prior to assigning the first final weight to the first grain: obtaining metadata for the first grain; andobtaining from the metadata for the first grain the weight associated with the grain group to which the first grain belongs.

13. The method of claim 12, wherein selecting a grain from the plurality of grains further comprises, prior to assigning the first final weight to the first grain, determining that each grain included in the first set of candidate grains is a member of the grain group to which the first grain belongs, and the step of obtaining the weight associated with the grain group to which the first grain belongs is performed as a result of determining that each grain included in the first set of candidate grains is a member of the grain group to which the first grain belongs.

14. The method of any one of claims 9-13, wherein the first final weight is also a function of: a distance between a first target position and a position of the first grain, an amount of time that has elapsed since the first grain was last rendered, a trajectory associated with the first grain and a target trajectory, and / or a measure of a closeness between the first grain and a previously rendered grain.

15. A computer program (843) comprising instructions (844) which when executed by processing circuitry (502) of a rendering device (724) causes the rendering device to perform the method of any one of claims 1-14.

16. A carrier containing the computer program of claim 15, wherein the carrier is one of an electronic signal, an optical signal, a radio signal, and a computer readable storage medium (842).

17. A rendering device (724) for rendering audio corresponding to an audio recording (111), wherein the audio recording is divided into a plurality of grains, the rendering device comprising: memory (842); andprocessing circuitry (802) coupled to the memory, wherein the rendering device is configured to perform a method that comprises: selecting (s902) a grain from the plurality of grains; and rendering (s904) the audio using the selected grain, wherein selecting a grain from the plurality of grains comprises: i) a) selecting a grain group from a set of grain groups comprising a first grain group and a second grain group, wherein each grain group included in the set of grain groups is assigned a weight and the selection of the grain group from the set of grain groups is based on the assigned weights and b) selecting a grain from the selected grain group, or ii) a) assigning a final weight to each grain in a set of candidate grains comprising a first grain and a second grain, wherein assigning a final weight to each grain in the set of candidate grains comprises assigning a first final weight to the first grain and assigning a second final weight to the second grain and b) selecting a grain from the set of candidate grains based on the assigned final weights, wherein the first final weight assigned to the first grain is a function of a weight associated with a grain group to which the first grain belongs.

18. The rendering device of claim 17, wherein the rendering device is further configured to perform the method of any one of claims 2-14.

Citation Information

Patent Citations

  • Systems and methods for simulating sounds of a virtual object using procedural audio

    US20180068487A1

  • Haptic effect conversion system using granular synthesis

    US20190094975A1