Particle mixing for particle synthesis

CN122826618APending Publication Date: 2026-09-25TELEFONAKTIEBOLAGET LM ERICSSON (PUBL)
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202480088599.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-01-12
Filing Date
2024-12-17
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

如果不是这样的话,则渲染听起来将是不均匀的、断断续续的或者只是在响度上波动

Benefits of technology

[0031]本文公开的实施例的优点在于,它们提供了一种确定最佳混合窗口系数的有效方式。可以选择为其指定混合窗口系数的点数以及这些点的位置,使得它适合给定颗粒数据库的内容。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122826618A_ABST
    Figure CN122826618A_ABST
Patent Text Reader

Abstract

A method for rendering audio corresponding to an audio recording, the audio recording being divided into a plurality of grains. The method comprises selecting i) a first grain associated with a first position in a descriptor space, and ii) a second grain associated with a second position in the descriptor space, from the plurality of grains. The method further comprises selecting a first set of one or more mixing window coefficients based on the first position, selecting a second set of one or more mixing window coefficients based on the second position, and obtaining final mixing window coefficients using the first and second sets of mixing window coefficients. The method further comprises producing a mixed sample by mixing at least a portion of the first grain with at least a portion of the second grain using the final mixing window coefficients.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The disclosed examples relate to particle synthesis. Background Technology

[0002] Audio rendering is the process of presenting audio, such as within extended reality (XR) scenes (e.g., virtual reality (VR), augmented reality (AR), or mixed reality (MR) scenes), to give the listener the impression that the sound originates from a physical source at a specific location within the scene. Presentation can be done via headphone speakers or other loudspeakers. If presentation is done via headphone speakers, the processing used is called binaural rendering and utilizes spatial cues of human spatial hearing, which make it possible to determine the direction from which the sound is coming. These cues involve interaural time delay (ITD), interaural level difference (ILD), and / or spectral difference.

[0003] Procedural audio refers to the creation of sound in real time as a response to live input. As an example, consider the sound of a car engine in virtual space, where the sound changes based on speed or acceleration or the car itself. This mechanism is commonly used in video games to achieve a better user experience. It is believed that for XR (e.g., AR or VR) use cases, there are many that would benefit from dynamically generated sounds, allowing them to react in real time to changes in the scene. For example, the sound generated when a user touches a surface or operates an engine. In reference [1], the sound of virtual objects (e.g., swords, axes, or wands) is simulated based on their position and orientation. There are fundamental and overtones. Both are modulated to change pitch, timbre, and amplitude to convey the speed of movement of the virtual object. Live input can come from the user via sensors (such as handheld controllers or headsets), and it can be control data or predefined automated data generated in real time by some software process such as physics simulation. Regardless of how the input data is generated, the audio renderer needs to process the incoming data and generate sound in real time in response to that data.

[0004] There are many different approaches for procedural audio (see, for example, reference [2]), including synthesized sound synthesis using audio processing modules, machine learning methods trained on real recordings, and splicing synthesis methods that utilize original recordings and rearrange segments of those recordings to generate variations. Because splicing synthesis methods use real recordings directly, the generated audio sounds very natural, thus enhancing the user experience.

[0005] Particle synthesis is a type of splicing synthesis in which sound recordings are divided into small segments called “particles”. (See, for example, reference [3]). By carefully selecting segments (particles) at render time, it is possible to generate seemingly realistic, dynamically changing sounds.

[0006] The particle synthesis process comprises two main steps: (1) particle extraction and (2) particle synthesis. Particle extraction refers to the extraction of relevant particles from a long, raw recording. The extraction method depends on the type of sound source and the desired features to be extracted. Particle synthesis refers to the technique of selecting an appropriate order of particles; this order selection can also be based on real-time user input.

[0007] Many sound design tools support grain extraction and synthesis, such as Soundseed Grain for Audiokinectic Wwise, Alchemy for Logic Pro, and AudioMotors for FMOD. These tools allow sound designers to perform manual or semi-automatic grain extraction, along with other simple operations. Among other controls, designers can select grain length, amplitude envelope of each grain, or shape. In the case of AudioMotors, an automated grain extraction tool specifically designed for motor sounds is provided.

[0008] Particle extraction can be done manually by the sound designer or in a data-driven manner by identifying relevant features of the audio for segmentation purposes. Relevant features include, for example, pitch period in the case of pitched sounds, spectral energy at a given frequency, Mel-frequency cepstral coefficients (MFCC), and local maxima of the amplitude envelope.

[0009] There are also many methods for particle synthesis. One common method is to randomly select particles and perform an overlap-add (OLA or overlap-add) operation. This method is not suitable for all types of sound sources and does not capture the temporal correlation between adjacent particles.

[0010] Corpus-based concatenation synthesis (CBCS) methods select particles from a corpus of sound segments sampled from a database of heterogeneous sound sources. They organize the corpus using descriptors associated with the sound segments and perform a search within the descriptor space to pick up the next particle. Note that the concept of descriptors is not limited to features of the audio signal (see, for example, reference [4]). When a direct mapping between the desired effect and features in the audio signal is not possible, the user can annotate the particles with perceptual descriptors.

[0011] The descriptor space is multidimensional, where the number of dimensions equals the number of descriptors. The search for appropriate particles is performed in a computationally efficient manner by utilizing the weighted Euclidean distance between the target descriptor location (e.g., a point or region) in the descriptor space (hereinafter referred to as "target descriptor coordinates" or "target location") and the particle location in the descriptor space. Reference [5] proposes a warping function for the distance measurement to better select the set of particles and also avoid duplication of previously rendered particles.

[0012] To perform an efficient search in the descriptor space, a kD-tree search is used. The k nearest neighbors of the target descriptor coordinates are selected, or particles within a radius 'r' of the target descriptor coordinates are selected. In reference [6], the corpus is organized into regions such that particles from different regions are not subsequently picked up when using a k nearest neighbor search.

[0013] The software CATERPILLAR (see reference [7]) performs splicing synthesis in an offline setting given a sequence of target descriptors. The program uses the Viterbi algorithm to identify the particle sequence to match the target descriptor. The cost function is a combination of the distance from the target descriptor coordinates and the splicing cost based on the similarity of consecutive particles.

[0014] On the other hand, CataRT is a real-time system, and therefore it randomly selects subsequent particles from a set of particles that are either the k nearest neighbors or the radii centered on the target descriptor coordinates (see, for example, reference [8]).

[0015] To achieve smoother transitions in the particle-synthesized sound, reference [9] uses feature descriptors such as pitch, loudness, spectral centroid, fundamental frequency, periodicity, and autocorrelation coefficient at hysteresis 1. Feature descriptors are computed for each particle, and Gaussian mixture models (GMMs) are used to capture the correlation among these feature descriptors, from which particles are sampled for synthesis.

[0016] Several works, such as reference

[10] , model or assign transition probabilities between adjacent particles to build Markov models for generating or synthesizing sound textures. Records are used to estimate model parameters, but unlike particle synthesis, sound synthesis does not directly use records.

[0017] Reference

[11] discusses particle synthesis, in which the next particle is picked based on the feature descriptor of the current particle to achieve timbre continuity. A kD-tree search is performed to select the candidate particle that is closest to the current particle in the feature descriptor space in terms of Euclidean distance.

[0018] When the audio corpus is sparse, reference [9] proposes a corpus expansion method based on feature descriptors. These techniques fall under the category of Feature Modulation Synthesis (FMS) (see, for example, reference

[15] ), which identifies the appropriate transformations to be applied to the audio based on target feature descriptors. Specific methods involve pitch shifting, gain adjustment, and the use of filters. The methods can be quite complex and have several steps, and must be performed offline, i.e., before particle synthesis or rendering.

[0019] Particle synthesis typically uses overlay addition to produce a smooth output with a continuous stream of particles, without any artificial artifacts from abrupt changes or discontinuities between particles. Overlay addition is done in such a way that perceived loudness is preserved within the overlapping areas of consecutive particles. Otherwise, the rendering will sound uneven, discontinuous, or simply fluctuate in loudness.

[0020] If consecutive particles are highly correlated and mostly in phase with each other, the mixture of two particles will be linearly additive; that is, if two particles with the same loudness are given a gain of 0.5, the combined mixture will have the same loudness as the two individual particles. In this case, the two particles are additive, and the gains of the two particles should be set so that the total gain adds up to 1.0. In this case, the mixture follows the linear mixing rule, where the mixture is a linear combination of the mixed particles.

[0021] However, if the first and second particles are largely uncorrelated, their mixture will not be linearly additive. This is because many samples from the first particle will have different signs than those from the second particle with which they are being combined, resulting in lower amplitudes when added. In this case, the two particles are not always constructively additive. To preserve the same perceived loudness as the mixture, the mixture should follow different mixing rules, where the gain of each mixed particle is set such that the signal power is preserved: P MIX = P1 = P2. The power of the signal is proportional to the square of the amplitude, which means the gain should be set according to the following: (A MIX ) 2 = (g1A1) 2 + (g2A2) 2 Where A1 and A2 are the amplitudes of the signals of the two particles to be mixed, and g1 and g2 are the gains used during mixing. If the mixture has the same power as signals 1 and 2, then A1 and g2 are the amplitudes of the signals of the two particles to be mixed. MIX A1 and A2 are identical, and therefore the gain should satisfy: (g1) 2 + (g2) 2 =1. For example, in cases where g1 and g2 are equal, they should be set to 1. .

[0022] To achieve a smooth, overlapping sum, the window function used needs to be selected based on the characteristics of the sound being rendered. For sounds with graininess, an amplitude-preserving window should be used, while for sounds without graininess, a power-preserving window should be used.

[0023] The blending rules don't only have an effect during overlap-addition processing. They apply whenever any other form of particle blending is performed. For example, when the renderer performs interpolation between particles by weighted blending of two or more particles.

[0024] Using the correlation coefficient between two particles, the work described in reference

[16] designed a custom or analytical window function for power preservation, since it may be difficult to use threshold correlation coefficient values ​​to determine the level of correlation. Summary of the Invention

[0025] There are some challenges. For example, in a multidimensional particle database, there may be particle clusters that are highly correlated with each other and thus benefit from using gain-preserving blending rules. At the same time, there may be other particle clusters that are largely uncorrelated with each other and thus benefit from using power-preserving blending rules. Therefore, specifying a blending rule for the entire database may be disadvantageous. On the other hand, since the database can hold thousands of particles, it is inefficient to store additional metadata for each particle that specifies the optimized choice of blending rule for that particle. While the approach to designing custom window functions disclosed in reference

[16] can be used at render time based on the selected particles to be blended, the process may be slow because it requires calculating correlation coefficients. Storing these correlation coefficients is also inefficient, as they must be calculated pairwise for all possible particle pairs to be rendered.

[0026] Therefore, in one aspect, a method is provided for rendering audio corresponding to an audio recording, wherein the audio recording is divided into a plurality of particles. In one embodiment, the method includes selecting a first particle from the plurality of particles, wherein the first particle is associated with a first position in an N-dimensional ND descriptor space. The method further includes selecting a first set of one or more blending window coefficients based on the first position in the ND descriptor space. The method further includes selecting a second particle from the plurality of particles, wherein the second particle is associated with a second position in the N-dimensional descriptor space. The method further includes selecting a second set of one or more blending window coefficients based on the second position in the ND descriptor space. The method further includes obtaining a final blending window coefficient m using the first set and the second set of blending window coefficients. The method further includes generating a blended sample S by blending at least a portion of the first particle with at least a portion of the second particle using the final blending window coefficients.

[0027] In another embodiment, the method includes selecting a first particle from a plurality of particles. The method also includes selecting a second particle from the plurality of particles. The method further includes i) selecting a set of two or more mixing window coefficients based on a target location in the ND descriptor space, or ii) selecting a single mixing window coefficient based on a target location in the ND descriptor space. The method also includes generating a mixed sample S by mixing at least a portion of the first particle with at least a portion of the second particle using either i) the single mixing window coefficient or ii) a derived mixing window coefficient derived from the set of two or more mixing window coefficients.

[0028] In another aspect, an apparatus is provided configured to perform a method for rendering audio corresponding to an audio recording, wherein the audio recording is divided into a plurality of particles. The method includes selecting a first particle from the plurality of particles, wherein the first particle is associated with a first position in an N-dimensional descriptor space. The method further includes selecting a first set of one or more blending window coefficients based on the first position in the N-dimensional descriptor space. The method further includes selecting a second particle from the plurality of particles, wherein the second particle is associated with a second position in an N-dimensional descriptor space. The method further includes selecting a second set of one or more blending window coefficients based on the second position in the N-dimensional descriptor space. The method further includes obtaining a final blending window coefficient m using the first set and the second set of blending window coefficients. The method further includes generating a blended sample S by blending at least a portion of the first particle with at least a portion of the second particle using the final blending window coefficients. The apparatus may include a memory and processing circuitry coupled to the memory.

[0029] In another aspect, an apparatus is provided configured to perform a method for rendering audio corresponding to an audio recording, wherein the audio recording is divided into multiple particles. The method includes selecting a first particle from the multiple particles. The method also includes selecting a second particle from the multiple particles. The method further includes i) selecting a set of two or more blending window coefficients based on a target location in the ND descriptor space, or ii) selecting a single blending window coefficient based on a target location in the ND descriptor space. The method also includes generating a blended sample S by blending at least a portion of the first particle with at least a portion of the second particle using either i) the single blending window coefficient or ii) a derived blending window coefficient derived from the set of two or more blending window coefficients.

[0030] In another aspect, a computer program is provided, including instructions that, when executed by processing circuitry of a device, cause the device to perform any of the methods disclosed herein. In one embodiment, a carrier containing the computer program is provided, wherein the carrier is one of electronic signals, optical signals, radio signals, and computer-readable storage media.

[0031] The advantage of the embodiments disclosed herein is that they provide an efficient way to determine the optimal mixing window coefficients. The number of points for the mixing window coefficients and the positions of these points can be optionally specified to suit the content of a given granular database. Attached Figure Description

[0032] The accompanying drawings, which are incorporated herein and form part of the specification, illustrate various embodiments.

[0033] Figure 1 The system according to an embodiment is illustrated.

[0034] Figure 2A The illustration shows an example two-dimensional descriptor space.

[0035] Figure 2B The illustration shows an example two-dimensional descriptor space.

[0036] Figure 3A The illustration shows an example two-dimensional descriptor space.

[0037] Figure 3B The illustration shows the process for determining the optimal k value for a given particle according to an embodiment.

[0038] Figure 4A The diagram illustrates the descriptor trajectory of the original record.

[0039] Figure 4B The illustration shows examples of the original record's descriptor trajectory and the generated descriptor trajectory.

[0040] Figure 4C The illustration shows examples of the original record's descriptor trajectory and the generated descriptor trajectory.

[0041] Figure 5 This is a flowchart illustrating a process according to an embodiment.

[0042] Figure 6 This is a flowchart illustrating a process according to an embodiment.

[0043] Figure 7A and Figure 7B A system according to some embodiments is shown.

[0044] Figure 8 This is a block diagram of a device according to some embodiments.

[0045] Figure 9 The illustration shows a 2D descriptor space where the particles in GDB 104 do not uniformly cover the descriptor space.

[0046] Figure 10 This is a flowchart illustrating a process according to an embodiment.

[0047] Figure 11 The diagram illustrates a particle database with several particle clusters.

[0048] Figure 12 This is a flowchart illustrating a process according to an embodiment.

[0049] Figure 13 This is a flowchart illustrating a process according to an embodiment.

[0050] Figure 14A The illustration shows the collection of particles rendered over time.

[0051] Figure 14B An example of fast particle switching is illustrated.

[0052] Figure 15 The diagram illustrates the parameters used to determine whether a fast granular switching should be triggered.

[0053] Figure 16A The illustration shows an example of fast particle switching without phase compensation.

[0054] Figure 16B An example of fast particle switching with phase compensation is illustrated.

[0055] Figure 17 This is a flowchart illustrating a process according to an embodiment. Detailed Implementation

[0056] Particle selection Figure 1 The illustration shows a system 100 for performing particle synthesis according to some embodiments. System 100 includes a particle extraction unit 102 that extracts particles from a raw audio recording 111. That is, the particle extraction unit divides the raw audio recording into small segments, referred to as "particles". The extracted particles are stored in a particle database (GDB) 104 (or simply "database"), which is accessed during rendering by a particle scheduling unit 106, also referred to as a "particle selection unit" or "particle scheduler," which is a component of the particle rendering function (GRF) 108. In some embodiments, there is one particle database for each procedural audio source, where each is available to the GRF 108. Each particle stored in the particle database is associated with one or more vectors of one or more descriptor values, each vector corresponding to a specific descriptor.

[0057] When creating the particle database 104, the audio designer decides which aspects should be used as descriptors. In some cases, it might be a characteristic of the sound itself, such as pitch or loudness, but it could also be other aspects related to how the sound is generated, such as the speed at which contact sound is generated between two sliding objects, or the opening angle of a door that creaks as it opens and closes. Descriptors should be chosen so that the renderer dynamically regenerates the sound given the target descriptor coordinates or trajectory.

[0058] As mentioned above, each particle stored in the particle database is associated with one or more descriptor values. Therefore, the particles in the original recording need to be annotated with descriptor values. Where the descriptor is an audio feature, it may be possible to measure the descriptor value directly from the audio signal itself. In other cases, the descriptor value needs to be provided as additional metadata for the recording in some way. For example, this can be done by recording data from some sensors during recording and providing this data in a companion file. In some cases, annotation can be done manually by creating a data log describing how the descriptor changes during recording, or annotation can be done manually for each extracted particle.

[0059] When granules are extracted from the original (one or more) records, descriptor values ​​are stored as metadata for each granule. Using descriptor values, each granule can be positioned in a multidimensional descriptor space, where the value of each descriptor describes its position along an axis within that space. If only one descriptor is used, the descriptor space is one-dimensional (1D), but the dimensionality of the descriptor space increases if more descriptors are used.

[0060] Figure 2A An example of a two-dimensional (2D) descriptor space is shown, where each circle represents a particle. (e.g.) Figure 2A As shown, each particle has a location (e.g., a point or region location) in the 2D descriptor space, which is called particle coordinates.

[0061] In one embodiment, the descriptor metadata of a granular sequence extracted from a record describes a trajectory within the descriptor space, corresponding to how the descriptor evolved during the original record.

[0062] During rendering, particle scheduling (i.e., the selection of particles to render) is typically based on target descriptor coordinates in descriptor space. Target descriptor coordinates specify what descriptor values ​​the generated sound output should have, meaning that particles in descriptor space closest to those coordinates will be used most prominently. Target descriptor coordinates can come from many types of sources, such as physics engines simulating the interactions of virtual objects, live input parameters from hand controllers or other sensors, or predefined automation parameters.

[0063] One aspect of particle compositing rendering is that repeating the same particles often sounds very unrealistic and artificial. If the particles are short, less than 50ms, repeating the same particles will result in a very metallic and static sound. If the particles are longer, the repetition will sound like a repetitive pattern, which usually results in an unrealistic sound.

[0064] The selection of particles needs to avoid repeating the same particles, but at the same time select particles in the descriptor space that are close to the target descriptor coordinates. Therefore, in some embodiments, this disclosure uses a weighted selection (e.g., weighted random selection) process, which generates a sequence of particles (i.e., an ordered set of particles) that closely follows the evolving target descriptor coordinates. An example of a particle sequence is: [particle-7, particle-6, particle-7, particle-9, particle-11, particle-10].

[0065] In some embodiments, each particle can be assigned a predefined weight (denoted as p). i5 This includes a set of dynamic weights that may change over time. Predefined weights are useful when particles are outliers, which should not be used too frequently, but can add realistic variation to the generated sound if used occasionally. Another use case is to use predefined weights to control the frequency of particles representing, for example, birdsong relative to particles representing forest background sounds. For example, in an embodiment where a set of one or more particle groups is defined and a first set of weights (wg1) is assigned to a first particle group in that set of particle groups, and particle i is a member of the first particle group, then p... i5 Set it to wg1.

[0066] The input to the programmable audio source (e.g., a signal generated by user interaction) is mapped to target descriptor coordinates. Based on the target descriptor coordinates, one or more particles are selected from a set of candidate particles for rendering using weighted selection (e.g., weighted random selection or selection of the particle with the highest weight). The selected particles can be rendered using standard particle synthesis methods, which involve, for example, performing an overlap-addition operation with two selected particles, such as crossfading a selected particle with another particle, where metadata about the overlap percentage and crossfade window can be pre-specified by the sound designer. Therefore, rendering particles encompasses not only all samples of the rendered particles, but also a first set of samples of the rendered particles, followed by a blended sample set generated by mixing a second set of samples of the particles with samples from another particle, and then rendering a second set of samples of the rendered particles.

[0067] The input to the programmed audio source can change in real time; therefore, the target descriptor coordinates can change over time as the input changes (even if the input does not change, the target descriptor coordinates can still change over time).

[0068] In one embodiment, the granular selection algorithm using weighted selection has the following steps.

[0069] Step 1: Obtain the target descriptor coordinates (e.g., map the input (e.g., user input or other input) to the target descriptor coordinates in the descriptor space).

[0070] Step 2: Adaptively determine the size of the neighborhood, for example, by calculating the k value based on the target descriptor coordinates, or by calculating the radius value (r) based on the target descriptor coordinates. Alternatively, obtain pre-calculated values ​​for k or radius from the metadata of the particle database.

[0071] Step 3: Select a set of candidate particles from the database using either the k-value or the radius value. For example, select the particles from the database that are the k nearest neighbors of the target descriptor coordinates. This search can be performed using readily available computationally efficient algorithms such as kD-tree search. As another example, the candidate particle set includes each particle with descriptor coordinates within a distance r from the target descriptor coordinates.

[0072] Step 4: Assign the final weights to each particle in the candidate particle set. The final weights assigned to a given particle can be based on: i) The distance between the particle's position in the descriptor space (i.e., particle coordinates) and the target descriptor coordinates in the descriptor space. ii) Differences in particle descriptor trends compared to target descriptor trajector trajectories. iii) The time history of previously used particles, iv) The temporal difference between particles in the original record and previously used particles (if they come from the same record), and / or v) Predefined weights of particles in the database.

[0073] Step 5: Using the final weights assigned in Step 4, perform a weighted selection of particles from the candidate particle set. For example, perform a weighted random selection, or, as another example, select the particle with the highest or lowest final weight. In this way, particles are selected based on the target descriptor coordinates and further based on the final weights assigned to the particles in the candidate particle set.

[0074] If the target descriptor coordinates change, repeat these steps. Otherwise, continue with the weighted selection of particles (starting from step 4) based on changes to the weights made from the time history of previous particles.

[0075] Step 2 - Determining the value of k.

[0076] Traditionally, the k value is a user-defined constant. However, keeping the k value fixed all the time is undesirable, as doing so may result in selecting too few or too many granules, which in turn may lead to underutilization of granules or selection of granules that are not similar to the target descriptor value, respectively.

[0077] Therefore, in one embodiment, this disclosure provides an adaptive selection of k based on the density of available particles in the region surrounding the target descriptor coordinates.

[0078] exist Figure 2A and 2B The diagram illustrates an example of why different k values ​​should be used for different target descriptor coordinates. Figure 2A In the middle, the target descriptor coordinates are close to the cluster of 3 particles.

[0079] However, in Figure 2B In these scenarios, the target descriptor coordinates are closer to more particles. Using the same k value is not optimal in these situations. For Figure 2A In this case, k=3 is appropriate. If k>3, this will result in selecting particles that are not in the cluster, and thus cause discontinuities in the rendered sound texture. Figure 2B In this case, if k = 3, too few particles are selected. Larger values ​​here will result in richer textures with less repetition, as a wider variety of particles can be chosen.

[0080] The k-value should be adaptively selected based on the target descriptor coordinates in the descriptor space. In one embodiment, each particle in the database is assigned an optimal k-value. Then, for the target descriptor coordinates, the k-value is set to be equal to the optimal k-value assigned to the particle closest to the target descriptor coordinates.

[0081] The assignment of the optimal k for each particle in the database can be performed offline or during the construction of the particle database. All distances from particle i to other particles in the database are recorded, or we can have a threshold for the maximum number of neighbors to stop distance calculation.

[0082] The distances are then sorted from lowest to highest. The difference between consecutive k values ​​will show a sudden jump at values ​​where the distance increases sharply. This is considered the cutoff value for k, and k+1 is assigned to the particle. The increment is to include the particle itself in the k value.

[0083] The criterion for determining the cutoff value would be to set an absolute threshold for the difference in sorted distances between consecutive neighbors, or to increase the normalized percentage in the consecutive sorted distances.

[0084] Alternatively, one could use the concept of an adaptive radius to select the candidate particle set. Current methods in the literature involve using a fixed radius, centered on the target descriptor coordinates, and selecting all particles within this fixed-radius circle to be included in the candidate particle set. Following the same reasoning as above, it might be beneficial to adaptively change the radius based on the position of the target descriptor coordinates. There, instead of choosing different k values, different radius values ​​would be used.

[0085] Figure 3A The diagram illustrates the k calculation for a specific particle 301, represented by a black circle. Figure 3A The image shows particle 301 and its six neighbors. Figure 3B The diagram shows the sorting distance used for particle 301. Because a jump in distance values ​​can be observed for k=4, the optimal k value of 5 is assigned to particle 301.

[0086] Step 3 - Weight Assignment The following describes each factor that influences the final weight assigned to a particle. We denote the target descriptor as u, and the weight associated with particle i as p. i The descriptor index is represented by j = 1, 2, ..., D, where D is the number of descriptors or the dimension of the descriptor space. The descriptor value of particle i at dimension j is represented by x. ij The target's descriptor value is given as u. j .

[0087] The following criterion affecting the probability of selecting a particle is expressed as a proportional relationship, because the final value of the weights is obtained after normalization (i.e., ensuring that the weights corresponding to all k particles add up to 1).

[0088] i) Distance from the target descriptor coordinates Distance metrics are used to define the proximity between the coordinates of a target descriptor and the coordinates of other descriptors corresponding to a particle. An example of a distance metric is the weighted Euclidean distance, where the differences in coordinates across each dimension are weighted by the inverse of the standard deviation of the corresponding descriptor values. .

[0089] The closer the particle's descriptor coordinates are to the target descriptor coordinates, the higher the association weight. Let d i Let be the distance from particle i to the target descriptor coordinates. For example, the probability of selecting particle i can be inversely proportional to the distance.

[0090] Therefore, it is possible to set... .

[0091] In some cases, different descriptors should not have the same amount of influence on particle selection. For example, if one descriptor is the pitch of a sound, and another descriptor has a weaker influence on the perceived characteristics of the sound, then the distance in the dimension corresponding to pitch can be assigned a higher weight than the distance in the dimension corresponding to the other descriptor. This can be achieved by adding an additional variable weight m to each dimension when calculating the distance. j To achieve: .

[0092] ii) Trend differences between the target descriptor trajectory and the original descriptor trajectory When performing particle extraction, descriptor coordinates are used as the primary selection criterion. However, the trend of the original descriptor trajectory also provides important information about the particles. For example, if engine sound is modeled using a particle database with a descriptor representing the engine's RPM, the trend of the descriptor corresponds to the engine's acceleration or deceleration at the time when particles are extracted from that record. Particles extracted from a portion of the record when the engine is accelerating will have a slightly lower pitch at the beginning than at the end, and will therefore be best suited when the desired output is the sound of an accelerating engine.

[0093] In a more general case, where a multidimensional descriptor space is used, the descriptor trend is a vector corresponding to the trajectory direction describing how the descriptor changes during the original record. Similarly, the trend of the target descriptor trajectory describes the direction in which the target descriptor coordinates move in the descriptor space.

[0094] The trend describes both direction and rate of change. Referring again to the engine example, if the particle database contains particles corresponding to the same RPM but with different accelerations, particles corresponding to accelerations similar to those of the target descriptor trajectory should be preferred.

[0095] Figure 4A The descriptor trace of the raw recording of the engine sound is shown, where the descriptor is the engine's RPM. The recording is divided into 15 particles. Figure 4B An example of how particles can be selected to match a target descriptor trajectory is shown, where the descriptor trend of the particles is not considered during particle selection. As can be seen, particles with increasing and decreasing RPM are used in combination. The resulting descriptor trajectory exhibits irregular behavior, which may lead to a decrease in perceived quality, especially if the descriptor represents the pitch of a sound. Figure 4C In addition, particle selection also takes into account descriptor trends, so that only particles with reduced RPM are used, which will produce a smoother sound.

[0096] Considering descriptor trends during particle selection is especially important for sound sources where different trend characteristics differ. For example, an engine may sound different when accelerating versus decelerating. Ensuring matching descriptor trends avoids problems caused by using particles with different characteristics together.

[0097] The descriptor trend of a particle can be calculated as the difference in descriptor values ​​at the end of the particle compared to the start of the particle, divided by the duration of the particle. In the case of a multidimensional descriptor space, the trend is a vector describing the average rate of change of descriptor coordinates during the particle's duration in the original record; for example, in the case of three descriptors, the particle trend t... G It will be a three-dimensional vector: .

[0098] When graininess is selected during rendering, the trend of the target descriptor trajectory can be similarly calculated as the difference in descriptor coordinates since the last update divided by the time T elapsed since the last update. U .

[0099] Then the difference in trend can be calculated. Then the weight assigned to the particle can be calculated as t. D Functions of the norm, for example: , where m T It is a variable that controls how the probability decreases as the trend difference of the descriptor increases.

[0100] iii) Time history of previously used particles For a specific time window, the history of previously rendered particles is preserved to avoid duplication, and the probability of selecting a particle is reduced if it has already been rendered. This reduction in probability is directly proportional to the difference between the current time t and the time of the last rendered particle i. A function of the difference between them.

[0101] If particle i was not selected in the past or within a specific time window, the value at the last time point is set to zero. The function f can be linear, quadratic, logarithmic, or have parameters with non-negative output. Any monotonically increasing function.

[0102] iv) Differences in time points in the original records For some sound sources, the way a sound evolves may not be fully described by changes in descriptors. Sometimes, the way a sound evolves depends on what happened before. For example, the creaking sound of an old door might be slightly different each time it opens, even if it opens at the same speed. In these cases, particles with similar descriptor values ​​and trends can sound very different, and combining them can lead to unnatural discontinuities not present in the original recording. A way to avoid these discontinuities is to assign higher probabilities to particles from the same part of the original recording as previously used particles. Particles from the same part of the original recording are expected to be closely related and similar in characteristics, and are therefore good candidates for selecting the next particle.

[0103] To measure how close one particle from a specific record is to another particle from the same record, a time difference can be calculated, which corresponds to the difference in time points between the records from which the two particles were extracted. If the particles are close, this time difference is small. To calculate the time difference between two particles, metadata describing the original record from which each particle was extracted and at what time point can be used. This metadata thus provides a measure of the closeness between the two particles. This metadata can be specified in a compact way as two values: the record identifier assigned to the original record, and a timestamp indicating the time point at which the particle can be found in that record.

[0104] The weights of particles can then be calculated by giving a higher probability of selecting the evaluated particle based on the difference in timestamps, with smaller differences. For example, the weight p of particle i. i It can be calculated as follows: Where R i R0 is the record identifier assigned to the record from which particle i is extracted, and R0 is the record identifier assigned to the record from which previously rendered particles are extracted. i t0 and t0 are the respective timestamps of the two particles, and b is a design constant that sets the probability of using a particle from the other record. The function f() takes the difference in time points as input and calculates the probability of the particle. In one embodiment, the function f decreases linearly with the difference in time points, where the slope is specified by the variable a: The effect of this is that as the difference between their respective timestamps increases, the probability decreases from 1.0 to b, but never falls below b.

[0105] In one embodiment, the record identifier R iIt can also be set to reference segments of a record; that is, a record can be divided into multiple segments, each with its own index. This can be useful when a record contains segments that the renderer does not consider relevant.

[0106] Final weight The final weight assigned to particle i (denoted as p) i ) is calculated by accumulating the weights assigned to particle i from each stage, i.e. In some embodiments, all five weights are not required and can therefore be skipped by setting the corresponding weight to 1.0 or by excluding it entirely from the calculation.

[0107] In one embodiment, by modifying individual weights using a fractional index, different weights are given different degrees of influence on the final weight, for example: , Here, by using a fractional exponent of 1 / 2, the weight p from the first stage is... i1 The impact is relatively small, and the weight p is made easier by using a 1 / 4 fractional exponent. i4 The impact is even smaller.

[0108] The weight p of all particles in the candidate particle set is calculated from the above steps. i Then, they are normalized by dividing each of them by the sum, so that the final weights become probability values.

[0109] , where k is the size of the candidate particle set.

[0110] Then, particles are selected based on their final weights (e.g., by sampling from a distribution). Note that while the target descriptor coordinates remain constant, the particle's temporal history and descriptor trend affect the final weights at each rendering time of the particle. Therefore, as long as the target descriptor remains constant, if particle i is selected at time t, then at time t+h, the particle's final weight is influenced by its temporal history and descriptor trend. The final weight of particle i becomes: in .

[0111] To avoid repeating particle i at t+1, we can define f(h) as: However, due to the normalization operation, the probabilities of the other k-1 particles will also change.

[0112] Figure 5This is a flowchart illustrating a process 500 for rendering audio corresponding to an audio recording according to an embodiment, wherein the audio recording is divided into multiple particles. Process 500 may begin at step s502.

[0113] Step s502 includes obtaining the coordinates of a first target descriptor, wherein the coordinates of the first target descriptor identify a first position in an N-dimensional descriptor space, where N>0.

[0114] Step s504 includes defining a first set of candidate particles based on the coordinates of the first target descriptor, wherein the first set of candidate particles includes k1 of the plurality of particles, where k1>1.

[0115] Step s506 includes assigning a final weight to each particle in the first set of candidate particles.

[0116] Step s508 includes randomly selecting a particle from a first set of candidate particles based on the assigned final weight, such that the probability of selecting a given particle from the first set of candidate particles is a function of the final weight assigned to that given particle.

[0117] Step s510 includes rendering the selected particles.

[0118] In some embodiments, each of the plurality of particles is associated with particle coordinates (e.g., a set of one or more descriptor values) that identify the particle’s position in an N-dimensional descriptor space, and a first set of candidate particles is defined based on the first target descriptor coordinates and the particle coordinates.

[0119] In some embodiments, defining a first set of candidate particles based on the coordinates of a first target descriptor and the coordinates of a particle includes: determining a set of nearest neighbor particles consisting of k1 particles from the plurality of particles, wherein none of the plurality of particles not included in the set of nearest neighbor particles is closer to the first target descriptor coordinates than any of the particles included in the set of nearest neighbor particles, and the candidate set of candidate particles consists of particles included in the set of nearest neighbor particles.

[0120] In some embodiments, the method further includes determining k1 based on the number of particles among the plurality of particles having particle coordinates within a threshold distance of the first target descriptor coordinates.

[0121] In some embodiments, each of the plurality of particles is assigned an optimal k value, and the method includes setting k1 to be equal to the optimal k value assigned to the particle among the plurality of particles that has the closest coordinates to the target descriptor coordinates.

[0122] In some embodiments, defining a first set of candidate particles based on the coordinates of a first target descriptor and the coordinates of the particles includes: determining a first radius value r1; and including each of the plurality of particles having particle coordinates within a distance of the first target descriptor coordinate r1 in the first set of candidate particles.

[0123] In some embodiments, the method further includes determining r1 based on the number of particles among the plurality of particles having particle coordinates within a threshold distance of the first target descriptor coordinates.

[0124] In some embodiments, assigning a final weight to each particle in a first set of candidate particles includes: assigning a first weight to a first particle included in the set of candidate particles; determining a first final weight based on the first weight; and assigning the first final weight to the first particle.

[0125] In some embodiments, the first weight is: a function of the distance between the coordinates of the first target descriptor and the first particle, a function of the amount of time elapsed since the last rendering of the first particle, a function of the trajectory associated with the first particle and the target trajectory, or a function of the proximity measure between the first particle and the most recently rendered particle (e.g., a time difference indicating the difference between the timestamp of the first particle and the timestamp of the most recently rendered particle, assuming that both particles were extracted from the same record or the same segment).

[0126] In some embodiments, the method further includes: after randomly selecting a particle from a first set of candidate particles based on an assigned final weight, assigning a new final weight to at least one particle in the first set of candidate particles, or removing a rendered particle from the first set of candidate particles; after assigning a new final weight to the rendered particle or removing a rendered particle from the first set of candidate particles, randomly selecting another particle from the first set of candidate particles based on the currently assigned final weight; and rendering the selected other particle.

[0127] In some embodiments, the method further includes: obtaining second target descriptor coordinates after randomly selecting particles from a first set of candidate particles; defining a second set of candidate particles based on the second target descriptor coordinates, wherein the second set of candidate particles includes k2 of the plurality of particles, where k2>1; assigning a final weight to each particle in the second set of candidate particles; randomly selecting particles from the second set of candidate particles based on the assigned final weights, such that the probability of selecting a given particle in the second set of candidate particles is a function of the final weight assigned to that given particle; and rendering the particle randomly selected from the second set of candidate particles.

[0128] Figure 6This is a flowchart illustrating a process 600 for rendering audio corresponding to an audio recording according to an embodiment, wherein the audio recording is divided into multiple particles. Process 600 may begin at step s602.

[0129] Step s602 includes obtaining the coordinates of a first target descriptor, wherein the coordinates of the first target descriptor identify a first position in an N-dimensional descriptor space, where N>0.

[0130] Step s604 includes defining a first set of candidate particles based on the coordinates of a first target descriptor, wherein the first set of candidate particles includes k1 of the plurality of particles, where k1>1, and the first set of candidate particles includes a first particle and a second particle.

[0131] Step s606 includes assigning a final weight to each particle in a first set of candidate particles, wherein assigning a final weight to each particle in the first set of candidate particles includes assigning a first final weight to a first particle and assigning a second final weight to a second particle. The first final weight assigned to the first particle is: a function of the distance between the coordinates of a first target descriptor and the first particle, a function of the amount of time elapsed since the last rendering of the first particle, a function of the trajectory associated with the first particle and the target trajectory, and / or a function of a proximity metric between the first particle and the most recently rendered particle.

[0132] Step s608 includes selecting particles from the first set of candidate particles based on the final weights assigned.

[0133] Step s610 includes rendering the selected particles.

[0134] In some embodiments, each of the plurality of particles is associated with particle coordinates (e.g., a set of one or more descriptor values) that identify the particle’s position in an N-dimensional descriptor space, and a first set of candidate particles is defined based on the first target descriptor coordinates and the particle coordinates.

[0135] In some embodiments, defining a first set of candidate particles based on the coordinates of a first target descriptor and the coordinates of a particle includes: determining a set of nearest neighbor particles consisting of k1 particles from the plurality of particles, wherein none of the plurality of particles not included in the set of nearest neighbor particles is closer to the first target descriptor coordinates than any particle included in the set of nearest neighbor particles, and the candidate set of candidate particles consists of particles included in the set of nearest neighbor particles.

[0136] In some embodiments, the method further includes determining k1 based on the number of particles among the plurality of particles having particle coordinates within a threshold distance of the first target descriptor coordinates.

[0137] In some embodiments, each of the plurality of particles is assigned an optimal k value, and the method includes setting k1 to be equal to the optimal k value assigned to the particle among the plurality of particles that has the closest particle coordinates to the target descriptor coordinates.

[0138] In some embodiments, defining a first set of candidate particles based on the coordinates of a first target descriptor and the coordinates of the particles includes: determining a first radius value r1; and including each of the plurality of particles having particle coordinates within a distance of the first target descriptor coordinate r1 in the first set of candidate particles.

[0139] In some embodiments, the method further includes determining r1 based on the number of particles among the plurality of particles having particle coordinates within a threshold distance of the first target descriptor coordinates.

[0140] In some embodiments, the method further includes: after selecting a particle from a first set of candidate particles based on an assigned final weight, assigning a new final weight to at least one particle in the first set of candidate particles, or removing a rendered particle from the first set of candidate particles; after assigning a new final weight to a rendered particle or removing a rendered particle from the first set of candidate particles, selecting another particle from the first set of candidate particles based on the currently assigned final weight; and rendering the selected other particle.

[0141] In some embodiments, the method further includes: after selecting a particle from a first set of candidate particles, obtaining second target descriptor coordinates; defining a second set of candidate particles based on the second target descriptor coordinates, wherein the second set of candidate particles includes k2 of the plurality of particles, where k2>1; assigning a final weight to each particle in the second set of candidate particles; selecting a particle from the second set of candidate particles based on the assigned final weight; and rendering a particle randomly selected from the second set of candidate particles.

[0142] Particle interpolation As mentioned above, during rendering, particle scheduling (i.e., the selection of particles to render) is based on the target descriptor coordinates in the descriptor space. When GDB 104 has many particles that uniformly fill the entire descriptor space, particle selection can be completed without frequently repeating the same particles and without needing to use particles far from the target location. However, in some cases, GDB 104 may have too few particles, or particles that do not uniformly fill the descriptor space. In these cases, particle selection becomes more restricted. For some target locations in the descriptor space, there may be no nearby particles, or very few nearby particles, making it necessary to repeat them frequently.

[0143] A new particle can be created in a case where the particle scheduler does not find any particle sufficiently close to the target position, or all particles close to the target position have been recently used; this new particle is referred to as an "interpolated particle". The method for creating an interpolated particle comprises selecting a plurality of particles (this is referred to as a "particle group") and calculating the interpolated particle based on the selected particle group. It is important that particles are selected such that the particle group surrounds the target position. In one embodiment, the particles in the group are close to the target position, but not all particles are on the same side of the target.

[0144] Figure 9 illustrates a 2D descriptor space, in which particles in GDB 104 do not uniformly cover the descriptor space. As is illustrated in Figure 9 , three particle clusters A, B and C surround the target position indicated by X. If particle 901 in cluster B is selected as a first particle of a particle group for interpolation, the next selected particle is preferably complementary to particle 901 with respect to the target position.

[0145] In one embodiment, particle i is complementary to particle j with respect to the target position if and only if for each particle descriptor coordinate dimension of the particle database, particles i and j are located on opposite sides of the target position. With this definition, particle i is complementary to particle 901 with respect to the target position if Gix is less than Tx and Giy is greater than Ty, where Gix is the x-coordinate of particle i, Giy is the y-coordinate of particle i, Tx is the x-coordinate of the target position, and Ty is the y-coordinate of the target position.

[0146] More generally, assuming particle j is closer to the target position (Tx, Ty) than particle i, with respect to a 2D descriptor space, particle i is complementary to particle j under the following conditions: if Gjx == Tx and Gjy < Ty, then particle i is complementary to particle j if Giy > Ty, or if Gjx == Tx and Gjy > Ty, then particle i is complementary to particle j if Giy < Ty, or if Gjy == Ty and Gjx < Tx, then particle i is complementary to particle j if Gix > Tx, or if Gjy == Ty and Gjx > Tx, then particle i is complementary to particle j if Gix < Tx, or if Gjx < Tx and Gjy < Ty, then particle i is complementary to particle j if Gix ≥ Tx and Giy ≥ Ty, or if Gjx < Tx and Gjy > Ty, then particle i is complementary to particle j if Gix ≥ Tx and Giy ≤ Ty, or If Gjx>Tx and Gjy<Ty, then if Gix ≤ Tx and Giy ≥ Ty, particle i is complementary to particle j, or If Gjx>Tx and Gjy>Ty, then if Gix ≤ Tx and Giy ≤ Ty, particle i is complementary to particle j.

[0147] Particle 902 in cluster A is an example of a particle that is complementary to particle 901 relative to the target position. In Figure 9 , this is indicated by the fact that compared with particle 901, particle 902 is on the opposite side of the target position in both dimensions. Adding particle 902 to the particle group is desirable, because particle 902 is not only complementary to particle 901, but also close to the target position. All particles from cluster C are on the same side as particle 901 in the y-dimension, and all particles in cluster B are on the same side as particle 901 in the x-dimension; therefore, relative to the target position, none of the particles in cluster B or C are complementary to particle 901. However, any particle from cluster A satisfies the condition of being complementary to particle 901.

[0148] In one embodiment, the following steps are performed: Step 1: Obtain a target position in the descriptor space of a database, and define a candidate particle set based on the target position, for example, using the method described above.

[0149] Step 2: Use weighted selection, such as weighted random selection, to select a first particle from the candidate particle set. If the selected particle is within a threshold distance of the target position, render the particle without any interpolation, otherwise proceed to steps 3-7.

[0150] Step 3 (optional): Remove the first particle from the candidate particle set.

[0151] Step 4: For each particle in the candidate particle set, assign a final weight to the candidate particle, wherein the final weight is a function of whether the particle is complementary to the first particle relative to the target position. Other things being equal, a candidate particle complementary to the first particle will have a higher final weight than a candidate particle not complementary to the first particle. If the first particle is not removed from the candidate particle set, assign a very low final weight to the first particle (e.g., a weight of 0.00001).

[0152] In one embodiment, if there are K particles in the candidate particle set, then for i=1 to K, the final non-normalized weight assigned to particle i in the candidate particle set can be calculated as: p i = p i1 p i2 p i3 p i4 p i5p iCOMP , where p iCOMP The weight depends on whether particle i is complementary to the first particle. As an example, p COMP It can be set as follows: if particle i is complementary, then p iCOMP = X, otherwise p iCOMP Let Y be the value of X, where X > Y. Preferably, X is at least an order of magnitude larger than Y. For example, in one embodiment, X = 1 and Y = 0.0001. In some scenarios, in the case of sparse particle databases, it may be beneficial to use non-complementary particles, where there may not be many particles available near the target location. Therefore, for this reason, Y is generally not set to 0.

[0153] Step 5. Based on the final weight assigned, select at least the second particle from the candidate particle set.

[0154] Step 6. Set the length of the new interpolation particle. For example, if the sounds in the database have pitch, the length is calculated as a weighted average of the lengths of the selected particles, where the weights depend on the distance of each particle to the target location. If the sounds in the database do not have pitch, the length is set based on the shortest of the selected particles.

[0155] Step 7. Generate interpolation particles using the selected particles. For example, the interpolation particles can be a weighted mixture of the selected particles, where the weights depend on the distance from each selected particle to the target location. If the sounds in the database have pitch, each selected particle can be resampled to fit the length of the interpolation particles. If the sounds in the database do not have pitch, only the sub-parts of the selected particles corresponding to the length of the interpolation particles are used.

[0156] In the simplest case, the two selected particles G that are closest to the target descriptor coordinates will be used. A and G B To generate interpolation particles.

[0157] Generally, it's good to use as many particles as possible for interpolation, as mixing many particles together can result in a diffuse sound that is dissimilar to the original sound. Therefore, in one embodiment, only the first and second selected particles are used to derive the interpolation particles.

[0158] However, if more than two particles are selected for interpolation (especially in multidimensional particle databases), the condition for complementarity can be relaxed for a certain dimension when selecting a third particle. Another criterion could be to ensure that subsequently selected particles are paired and complementary. For example, in a 3-D database, particles A and B satisfy the complementarity criterion in all dimensions, and particles B and C are also complementary in all dimensions, but particles A and C might only be complementary in the first two dimensions and not in the third dimension.

[0159] Determine the length of the interpolation particles In step 6, the length of the interpolated particles (referred to as the "target length") is set. For weighted mixing of particles, the particles need to have a matching length. The target particle length can be determined in different ways based on the sound characteristics described by the particle database.

[0160] For some particle databases, the length of each particle corresponds to one or a specific number of pitch cycles. In this case, the target length can be calculated as a weighted average of the lengths of the selected particles. For cases where only two particles are used to generate the interpolated particles, the target length (L) 目标 It can be calculated as: L 目标 = (w1)(L1) + (w2)(L2), where L1 and L2 are the lengths of the two selected particles, and w1 and w2 are weights calculated inversely proportional to the distance of each selected particle to the target position, with the formulas: w1 = d2 / (d1+d2) and w2 = d1 / (d1+d2).

[0161] Distances d1 and d2 are Euclidean distances in the descriptor space. In a particle database where only one dimension affects sound pitch and consequently particle length, weights can be calculated using a distance metric that considers only that dimension.

[0162] For other particle databases where particle length does not indicate pitch, the target length is less critical. A straightforward approach is to select a target length that matches the shortest of the selected particles.

[0163] Create temporary particles for weighted blending Before the weighted mixture can be computed to form the interpolated particles, a temporary version of the selected particles with the selected target particle length is created.

[0164] When the particle database represents the pitch of a sound, the selected particles can be resampled using a resampling method that allows for arbitrary resampling ratios (such as linear resampling, Sinc interpolation, or Lanczos resampling, or similar well-known techniques).

[0165] If the particle database does not represent pitched sounds, a sub-part of the selected particles that matches the target particle length can be used. This can be done by using the first part of each selected particle, or by using a sub-part within each particle with a randomly selected starting point, where the randomly selected starting point is constrained such that the remaining samples of that particle correspond at least to the target particle length.

[0166] Interpolation particles are generated using weighted mixing. In one embodiment, the interpolating particle is a weighted mixture of selected particles (assuming the selected particles have the same length), or a weighted mixture of a selected particle and a temporary particle derived from another selected particle (assuming the length of the selected particle is equal to the target length), or a weighted mixture of temporary particles (assuming no particle has a length equal to the target length). Accordingly, the interpolating particle is equal to: (w1)(g1) + w2(g2), or (w1)(tg1) + w2(g2), or (w1)(g1) + w2(tg2), or (w1)(tg1) + w2(tg2), where tg1 is a first temporary particle derived from a first particle, and tg2 is a second temporary particle derived from a second particle. In one embodiment, the weights can be calculated in a manner similar to that used when selecting the target particle length, as inversely proportional to the distance from the selected particle to the target location.

[0167] For example, if two particles g1 and g2 are selected, and the distances from the selected particles to the target location are d1 and d2, respectively, then in one embodiment, the weights used to perform weighted mixing are calculated as: w1 = (d2 / (d1+d2)) and w2 = (d1 / (d1+d2)). Where it is expected that consecutive particles are largely uncorrelated, the weights can be calculated according to a constant power mixing rule instead of a linear mixing rule. In this case, the weights can be calculated as: w1 = sqrt(d2 / (d1+d2)) and w2 = sqrt(d1 / (d1+d2)). Other methods for creating weights (also known as gains) will be described below in the section on determining the optimal mixing window coefficients.

[0168] Cache interpolation granules After generating interpolation particles, these particles can be stored for later use. This reduces the complexity of calculating new interpolation particles later. In this case, the interpolation particle should be assigned the same index and / or descriptor coordinates as other particles in the database, so that the renderer can avoid duplicating the same interpolation particles. While caching interpolation particles can reduce renderer complexity, caching increases memory consumption, so limiting the number of cached particles may be beneficial. If needed, cached interpolation particles can also be written to a particle database, making the proposed interpolation technique another method for corpus extension.

[0169] Figure 10 This is a flowchart illustrating a process 1000 for rendering audio corresponding to an audio recording according to an embodiment, wherein the audio recording is divided into multiple particles. Process 1000 may begin at step s1002. Step s1002 includes obtaining target descriptor coordinates, wherein the target descriptor coordinates identify a target location in an N-dimensional descriptor space (where N>0). Step s1004 includes selecting a set of particles from a particle database based on the target descriptor coordinates, wherein the selected set of particles includes a first particle and a second particle. Step s1006 includes determining a first weight w1 for the first particle. Step s1008 includes determining a second weight w2 for the second particle. Step s1010 includes using the first weight, the second weight, the first particle, and the second particle to generate an interpolated particle. Step s1012 includes rendering the interpolated particle.

[0170] In some embodiments, the second particle has a position in N-dimensional space, and the first particle is located at a complementary position in N-dimensional space relative to the position of the second particle in N-dimensional space.

[0171] In some embodiments, the interpolation particle is equal to: (w1)(g1) + w2(g2), or (w1)(tg1) + w2(g2), or (w1)(g1) + w2(tg2), or (w1)(tg1) + w2(tg2), where g1 is a first particle, g2 is a second particle, tg1 is a first temporary particle derived from the first particle, and tg2 is a second temporary particle derived from the second particle.

[0172] In some embodiments, the first particle has a length and the second particle has a length, and generating the interpolated particle includes: determining a target particle length using a first length value specifying the length of the first particle and a second length value specifying the length of the second particle, wherein the length of the first particle is not equal to the target particle length; deriving a first temporary particle with a length equal to the target particle length from the first particle; and generating the interpolated particle using the first temporary particle, the second particle, the first weight, and the second weight.

[0173] In some embodiments, the first particle has a length and the second particle has a length, and generating the interpolated particle includes: determining a target particle length using a first length value specifying the length of the first particle and a second length value specifying the length of the second particle, wherein the length of the first particle is not equal to the target particle length and the length of the second particle is not equal to the target particle length; deriving a first temporary particle with a length equal to the target particle length from the first particle; deriving a second temporary particle with a length equal to the target particle length from the second particle; and generating the interpolated particle using the first temporary particle, the second temporary particle, the first weight, and the second weight.

[0174] In some embodiments, deriving a first temporary particle from a first particle includes: resampling the first particle to generate a first temporary particle, or selecting a sub-part of the first particle, wherein the first temporary particle is the selected sub-part.

[0175] In some embodiments, determining the target particle length using a first length value and a second length value includes: determining the target particle length using a first length value, a second length value, a first weight, and a second weight.

[0176] In some embodiments, the target particle length is equal to (w1)(L1) + (w2)(L2), where L1 is a first length value and L2 is a second length value.

[0177] In some embodiments, a first particle has a position in N-dimensional space, and a second particle has a position in N-dimensional space, a first weight is based on (e.g., inversely proportional to) the distance from the position of the first particle in N-dimensional space to the target position in N-dimensional space, and a second weight is based on (e.g., inversely proportional to) the distance from the position of the second particle in N-dimensional space to the target position in N-dimensional space.

[0178] In some embodiments, the first weight is equal to d2 / (d1 + d2), the second weight is equal to d1 / (d1 + d2), d1 is the distance from the position of the first particle in N-dimensional space to the target position, and d2 is the distance from the position of the second particle in N-dimensional space to the target position.

[0179] In some embodiments, selecting a set of particles from a particle database includes: defining a first set of candidate particles based on a target location; selecting a first particle from the first set of candidate particles; after selecting the first particle from the first set of candidate particles, removing the first particle from the first set of candidate particles to form a second set of particles; assigning a final weight to each particle included in the second set of candidate particles; and selecting a particle from the second set of candidate particles based on the assigned final weight, wherein the particle selected from the second set of particles is a second particle.

[0180] In some embodiments, selecting a set of particles from a particle database includes: defining a set of candidate particles based on a target location; selecting a first particle from the set of candidate particles; assigning a final weight to each particle contained in the set of candidate particles; and, after assigning the final weight to each particle contained in the set of candidate particles, selecting a particle from the set of candidate particles based on the assigned final weight, wherein the particle selected from a second set of particles is a second particle.

[0181] In some embodiments, assigning a final weight to the second particle includes: determining whether the second particle is complementary to the first particle relative to the target position; and assigning a complementary weight p to the second particle. COMP Where the second particle is complementary to the first particle relative to the target position, then p COMP The value of p is X, otherwise p COMP The value of X is Y, where X is greater than Y; and the final weight of the second particle is calculated using the complementary weights assigned to the second particle. In some embodiments, X is at least ten times Y.

[0182] Determine the optimal mixing window coefficient When performing granular synthesis, the overlap-add technique can be used to provide a smooth transition from one particle to the next in a particle sequence. As mentioned earlier, it can be beneficial to adapt the choice of the overlap-add window to the characteristics of the particle signals to be mixed. If the signals are highly correlated, a linear mixing window should be used. However, if the signals are uncorrelated, a power-preserving mixing window should be used. In many cases, the optimal mixing window lies somewhere between the linear and power-preserving windows, as only a portion of the signal may be correlated.

[0183] To generate an optimized blending window (denoted as W) OThe mixing window coefficient (denoted as m) (also known as the mixing window weight) can be used to provide continuous control over the mixing window (from linear to power hold). In one embodiment, the mixing window coefficient (m) is a value between 0.0 and 1.0, and: W O = m W P + (1-m) W L W L It is a linear overlapping window, while W P It is a power preservation overlapping window. Therefore, in this embodiment, the optimized hybrid window is a weighted average of the power preservation overlapping window and the linear overlapping window.

[0184] When m=1, only the power-preserving overlap window will be used. When m=0, only the linear overlap window will be used. For values ​​between 0 and 1, a mixture of these two windows will be used. Therefore, by properly adjusting the mixing window coefficients, an optimal mixing window can be found.

[0185] In one embodiment, the W of a pair of particles (the particle pair to be mixed) undergoing overlapping addition processing L W is a vector of X scalar values. L = [w l [0], w l [1], ..., w l [X-1], where w l [x] = x / (X-1) and X = min(L1, L2) × p, where L1 is the length of the first particle in the pair, L2 is the length of the second particle in the pair, and p is the overlap percentage. In one embodiment, W P It is also a vector of X scalar values, i.e.: W P = [w p [0], w p [1],..., w p [X-1], where w p [x]= (w l [x]) 1 / 2 .

[0186] If the mixing window coefficient is m for both the first and second particles of the pair, then at sample index n, the weights g1 and g2 of the two particles are respectively: g1[n] = m(w l [n]) 1 / 2 + (1-m)w l [n] and g2[n] = m(1-w l [n]) 1 / 2 + (1-m)(1-w l [n]).

[0187] Next, the X overlapping samples s[n] from n=0 to X-1 will be a linear combination of X samples from two particles with weights g1 and g2, that is, s[n] = g1[n]grain1[Sp1+n] + g2[n]grain2[Sp2+n] (within n=0 to X-1), where grain1[] is the sample set that makes up the first particle, grain2[] is the sample set that makes up the second particle, and Sp1 and Sp2 are the starting positions of the first particle and the second particle, respectively.

[0188] When using a particle database where particles have positions in the descriptor space, it's common to encounter highly correlated particle clusters alongside other uncorrelated clusters. This can happen, for example, when the particle database describes a sound with strong pitch in some parts of the descriptor space but more noisy characteristics in others. By specifying blending window coefficients for different parts of the descriptor space, an overlap window that works best in that part can be selected.

[0189] Figure 11 The diagram illustrates a particle database with multiple particle clusters. The mixing window used can be adjusted for each particle cluster by specifying the mixing window coefficients at several locations in the descriptor space.

[0190] In one embodiment, a set of blending window coefficients is specified, where each coefficient is associated with a location in the descriptor space, i.e., each coefficient is associated with a set of coordinates for the specified location in the descriptor space. These coefficients can be stored as metadata in a granular database. This makes it possible to pre-compute optimized blending window coefficients (e.g., by performing correlation checks or other methods to find the optimal blending window). This set of coefficients can also be created manually and tuned by the sound designer who created the granular database.

[0191] Combining the list of blending window coefficients with the spatial location of the descriptors is a compact, efficient, and scalable way to specify blending windows for a database. In some cases, specifying only one blending window may be sufficient, while in others, specifying a large number of blending windows in different locations may be beneficial.

[0192] An example of a data structure (table) for storing the mixed window coefficients of a 3D particle database is shown below: Specify the coordinates of a point in the 3D descriptor space. coefficient [0.0, 0.0, 0.0] 0.00 [0.3, 0.0, 0.0] 0.32 [0.3, 0.5, 0.0] 0.45 [0.7, 0.5, 0.0] 0.67 [0.8, 0.7, 0.0] 0.86 [1.0, 1.0, 1.0] 0.23 [0.7, 0.5, 1.0] 0.37 [0.8, 0.7, 1.0] 0.66 [1.0, 1.0, 1.0] 0.73 [0.7, 0.5, 0.5] 0.67 [0.8, 0.7, 0.5] 0.86 [1.0, 1.0, 0.5] 0.93 Here, each entry (a row in the table) has a field that stores the coordinates of a location in the 3D descriptor space, and an associated field that stores the blending window coefficients specified for that location. Other data structures (such as lists) can be used to store the coefficients and their corresponding coordinates.

[0193] During rendering, optimal (or final) blending window coefficients can be found for each particle based on its position in descriptor space. In the simplest case, the coefficient closest to the particle in the descriptor space is selected from the list. Another embodiment can use a weighted sum of coefficients, where the weight of each coefficient is based on the distance between that coefficient and the particle; that is, the optimal blending window coefficients are: ,in It is a function of the distance between the i-th coefficient in the list and the particle.

[0194] To efficiently retrieve the nearest coefficient for a specific particle, the coefficient list can be sorted, for example, by a KD-tree structure.

[0195] When performing a blending of two particles (e.g., overlapping addition), the coefficients for each of the two particles can be determined as described above, and then the optimal blending window coefficient is based on these two determined coefficients. In one embodiment, the optimal blending window coefficient is the maximum of the two determined blending window coefficients, i.e., m = max(m1, m2), where m1 is the coefficient determined for the first particle and m2 is the coefficient determined for the second particle. The logic behind this operation is that the higher the value of the blending window coefficient, the less potentially relevant the particles to be blended are.

[0196] For example, if one particle has a mixing window coefficient of 0.32 and another particle has a mixing window coefficient of 0.93, then m = 0.93 would be used. Even if the particle with a mixing window coefficient of 0.32 comes from a correlated particle cluster, its correlation with particles from a cluster with low correlation may not be as high. Therefore, if particles have different mixing window values, it can be assumed that the particles are largely uncorrelated. Thus, using the largest mixing window coefficient is advantageous.

[0197] In one embodiment, the final blending window coefficients are not based on the positions of the particles to be blended, but rather on the current target position in the descriptor space. In this case, for example, a certain number of blending window coefficients closest to the current target position are used as the basis for determining the optimized blending window coefficients. This can be done by picking the closest coefficient, or by performing some form of weighted blending of the set of blending window coefficients specified among points near the target position.

[0198] When mixing two or more particles, an optimal mixing window coefficient can be used. In the case of overlapping addition, mixing occurs within the overlapping region.

[0199] When mixing particles together, an optimal mixing window coefficient (m) can also be used, for example, to form interpolated particles as described above. As an example, the sample s of interpolated particles... int [] can be: s int [n]= g1 × grain1[n] + g2 × grain2[n] Within the range of n=0 to L-1, where L is the length of the particle (in this embodiment, the particles have equal lengths). grain1[] is the set of samples that make up the first grain. grain2[] is the set of samples that make up the second grain. g1 = m (w1) 1 / 2 + (1-m)w1; and g2 = m (w2) 1 / 2 + (1-m)w2.

[0200] In one embodiment, w1 = (d2 / (d1+d2)) and w2 = (d1 / (d1+d2)), where d1 is the Euclidean distance between the first particle and the target location, and d2 is the Euclidean distance between the second particle and the target location. However, in cases where only one dimension of the particle database affects the pitch of the sound and therefore the length of the particles, a distance metric that considers only that dimension can be used to calculate the weights w1 and w2.

[0201] Figure 12 This is a flowchart illustrating a process 1200 for rendering audio corresponding to an audio recording according to an embodiment, wherein the audio recording is divided into multiple particles. Process 1200 may begin at step s1202.

[0202] Step s1202 includes selecting a first particle from the plurality of particles, wherein the first particle is associated with a first position in the N-dimensional ND descriptor space.

[0203] Step s1204 includes selecting a first set of one or more blending window coefficients based on the first position in the ND descriptor space.

[0204] Step s1206 includes selecting a second particle from the plurality of particles, wherein the second particle is associated with a second position in the N-dimensional descriptor space.

[0205] Step s1208 includes selecting a second set of one or more blending window coefficients based on the second position in the ND descriptor space.

[0206] Step s1210 includes using a first set of mixed window coefficients and a second set of mixed window coefficients to obtain the final mixed window coefficients m.

[0207] Step s1212 includes generating a mixed sample S by mixing at least a portion of the first particle with at least a portion of the second particle using a final mixing window coefficient.

[0208] In some embodiments, the first particle includes a first set of samples, the second particle includes a second set of samples, and generating a mixed sample using the final mixing window coefficient includes generating a first mixed sample s[0] by calculating s = g1 × grain1_sample + g2 × grain2_sample, where g1 is a function of the final mixing window coefficient, grain1_sample is one of the samples from the first set of samples, g2 is a function of the final mixing window coefficient, and grain2_sample is one of the samples from the second set of samples.

[0209] In some embodiments, g1 is a further function of a first weight w1 associated with a first blending window and a second weight w2 associated with a second blending window, and g2 is a further function of a third weight w2 associated with the first blending window and a fourth weight w4 associated with the second blending window.

[0210] In some embodiments, g1 = m × w1 + (1-m) × w2, and g2 = m × w3 + (1-m) × w4.

[0211] In some embodiments, w1 = (w2) 1 / 2 And w3 = (w4) 1 / 2 .

[0212] In some embodiments, w4 = 1 - w2.

[0213] In some embodiments, w2 = n / (L-1), n ​​≥ 0 and n ≤ (L-1), L = min(L1,L2) × p, L1 is the length of the first particle, L2 is the length of the second particle, and p is a predetermined overlap percentage.

[0214] In some embodiments, w2 = d1 / (d1+d2), w4 = d2 / (d1+d2), where d1 is the distance from the first position in the ND descriptor space to the target position in the descriptor space, and d2 is the distance from the second position in the ND descriptor space to the target position in the descriptor space.

[0215] In some embodiments, the first set of mixed window coefficients consists of a first mixed window coefficient m1, and obtaining the final mixed window coefficients includes setting m to be equal to max(m1,m2), where m2 is a mixed window coefficient contained in a second set of mixed window coefficients, or is based on a mixed window coefficient contained in a second set of mixed window coefficients.

[0216] In some embodiments, obtaining the final blended window coefficients includes: assigning weights to each blended window coefficient contained in a first set of blended window coefficients, obtaining a weighted average of the blended window coefficients contained in the first set of blended window coefficients using the weights assigned to each blended window coefficient contained in the first set of blended window coefficients, and setting m to be equal to max(m1,m2), where m1 is the weighted average of the blended window coefficients contained in the first set of blended window coefficients, and m2 is a blended window coefficient contained in a second set of blended window coefficients, or is based on the blended window coefficients contained in the second set of blended window coefficients.

[0217] Figure 13 This is a flowchart illustrating a process 1300 for rendering audio corresponding to an audio recording according to an embodiment, wherein the audio recording is divided into multiple particles. Process 1300 may begin at step s1302.

[0218] Step s1302 includes selecting a first particle from the plurality of particles.

[0219] Step s1304 includes selecting a second particle from the plurality of particles.

[0220] Step s1306 includes i) selecting a set of two or more blending window coefficients based on the target location in the N-dimensional ND descriptor space, or ii) selecting a single blending window coefficient based on the target location in the ND descriptor space.

[0221] Step s1308 includes mixing at least a portion of the first particle with at least a portion of the second particle by using i) a single mixing window coefficient, or ii) a derived mixing window coefficient derived from the set of two or more mixing window coefficients, thereby producing a mixed sample S.

[0222] In some embodiments, the first grain includes a first set of samples, the second grain includes a second set of samples, and generating a mixed sample using the final mixing window coefficients includes generating a first mixed sample s[0] by calculating s = g1 × grain1_sample + g2 × grain2_sample, where g1 is a function of the final mixing window coefficients, grain1_sample is one of the samples from the first set of samples, g2 is a function of the final mixing window coefficients, and grain2_sample is one of the samples from the second set of samples.

[0223] In some embodiments, g1 is a further function of a first weight w1 associated with a first blending window and a second weight w2 associated with a second blending window, and g2 is a further function of a third weight w2 associated with the first blending window and a fourth weight w4 associated with the second blending window.

[0224] In some embodiments, g1 = m × w1 + (1-m) × w2, and g2 = m × w3 + (1-m) × w4, where m is a single mixed window coefficient or a derived mixed window coefficient.

[0225] In some embodiments, w1 = (w2) 1 / 2 And w3 = (w4) 1 / 2 .

[0226] In some embodiments, w4 = 1 - w2.

[0227] In some embodiments, w2 = n / (L-1), n ​​≥ 0 and n ≤ (L-1), L = min(L1,L2) × p, L1 is the length of the first particle, L2 is the length of the second particle, and p is a predetermined overlap percentage.

[0228] In some embodiments, w2 = d1 / (d1+d2), w4 = d2 / (d1+d2), where d1 is the distance from the first position in the ND descriptor space to the target position in the descriptor space, and d2 is the distance from the second position in the ND descriptor space to the target position in the descriptor space.

[0229] Fast particle switching The above describes a method for scheduling particles from a particle database, where each particle has a designated position in a descriptor space. The target position in the descriptor space controls which particle should be selected next. As the target position is updated in real time, particles close to the target position are selected. This allows for dynamic control over how the sound evolves. Figure 14A The diagram below illustrates a simplified example of this feature.

[0230] Figure 14A The diagram illustrates how to select successive particles as the target position changes. Figure 14A In the case shown, the particle database has only one dimension (descriptor 1). On the x-axis, the time axis shows when particles are selected and how long each particle is used to generate audio output (i.e., how long each particle is rendered). Figure 14A The dashed lines in the diagram illustrate how the target's position changes over time. Figure 14A In this process, each complete particle is rendered before transitioning to the next particle.

[0231] However, typically, samples from the currently selected particle are played until a predetermined sample position is reached (which is usually near the end of the particle). When the predetermined sample position is reached, the next particle is selected, and crossfading (e.g., overlap addition) is initiated between the two particles.

[0232] In some cases, the particles in the particle database need to be quite long to preserve the natural sound of the original sound source. In such cases, particle selection may not be frequent enough to allow for fast and smooth transitions. For example, if the target position moves rapidly in the descriptor space, the output of the particle synthesis may exhibit large jumps due to the fact that the target position has already moved a substantial distance in the descriptor space within the time it takes to play one particle. In this situation, particle switching is too slow to keep up with the dynamic changes in the target position.

[0233] At the same time, using a forced, higher particle switching rate may produce unnatural output. If the particle database is designed for long particles, the sound designer intends the particles to be longer, and using a forced particle switching rate will not allow the complete particles to play as expected.

[0234] Therefore, this disclosure proposes to frequently evaluate the amount of time the target position has moved since the last particle selection. A criterion is then evaluated to determine whether a fast particle switch should be initiated. In this way, longer particles will play out completely as long as the target position hasn't moved too much. If the target position has indeed moved very quickly, a fast particle switch can be triggered, and it can have fast dynamic behavior. This feature is... Figure 14B The diagram in the middle is shown.

[0235] like Figure 14B As shown, a new particle can be selected and transitioned to before reaching the end of the currently used particle. This allows for smoother changes in sound and a closer following of the target position without large jumps. Figure 14B To keep the illustration simple, overlapping areas are not shown.

[0236] Therefore, in one embodiment there is a method that includes: obtaining an updated target location in a descriptor space, using the target location to determine a distance (e.g., determining the distance the target location has moved since a previous time point), and based on the determined distance, determining whether a fast granular switching condition is met.

[0237] In one embodiment, the renderer immediately transitions to the new grain because the fast grain switching conditions are determined to be met. For example, a new grain is selected based on the updated target location, and an overlap transition is initiated to the new grain. As another example, a new grain is selected based on the updated target location, the renderer stops playing the current grain, and the renderer immediately begins playing the new grain.

[0238] Criteria for triggering fast granular switching In one embodiment, determining whether the fast particle switching condition is met includes determining the amount by which the target location has moved since the last particle selection. For example, in one embodiment, the distance the target location has moved since the last particle selection is determined, and if the determined distance is greater than a threshold, a fast switch is triggered, i.e., the fast particle switching condition is met.

[0239] In another embodiment, determining whether a fast particle switching condition is met includes determining the degree of match between the currently used particle(s) and the updated location. For example, in one embodiment where a single particle is currently being used for rendering, the updated target location is compared to the location of that single particle. If the target location is more than a threshold distance away from the particle, a fast particle switching is triggered. Alternatively, a fast switching is triggered if that distance has increased more than a threshold since the current particle was selected.

[0240] As another example, in situations where more than one particle is currently being used for rendering (e.g., if the renderer uses crossfade of two or more particles), the updated target location can be compared to the locations of all current particles. For instance, the distance from the updated target location to the nearest particle being used can be calculated, and if that distance is greater than a threshold, a fast particle switching is triggered. Alternatively, a fast switching is triggered if that distance has increased by more than a threshold since the current particle was selected.

[0241] Figure 15 The illustration shows a scenario where two particles (particle 1501 and particle 1502) are currently used to generate audio output, meaning that these two particles are being rendered in a blended manner.

[0242] The distance (d2) from the updated target position 1512 to the line 1520 between the two used particles is calculated, and in one embodiment, if the distance (d2) is greater than a predetermined value (also known as a threshold), a fast particle switching is triggered.

[0243] In another embodiment, to avoid repeatedly triggering fast particle switching in a sparse database where no particles are found near the target location, the change in distance since the current particle was selected can be compared to a threshold. Specifically, the value (d2 - d1) is compared to a threshold, where d1 is the distance to target location 1511 used when particles 1501 and 1502 were selected. In other words, how much the distance between the target location and the currently used particle has increased since the current particle was selected. If the distance increase exceeds the threshold, a fast particle switching can be triggered. If the distance increase is less than the threshold, the currently selected particle may still be valid for the updated target location.

[0244] In another embodiment, d2 is compared to d1, and a fast granular switch is triggered if the distance increases by more than a threshold (T); that is, if d2 > d1 + T, a fast granular switch is triggered. The value of T can be configured for each granular database. For example, it can be written as metadata to the granular database. In one embodiment, a list of points in a descriptor space is used to specify the threshold, where each point has its own specific threshold. Such a threshold can be calculated as a function of the inter-granular distances in the database.

[0245] If the database is organized into clusters, a list of thresholds would be more appropriate. For each cluster, its centroid can be chosen as the point where a specified threshold is applied. The threshold value can then be a function of the mean distance from the particle to the centroid. Alternatively, the threshold can be defined based on the inter-cluster distance.

[0246] Fast particle switching with phase compensation for pitch-sensitive sounds In particle databases describing sounds with prominent pitches (where the length of each particle is proportional to its pitch period), particle switching can lead to phase cancellation issues if the current particle hasn't been fully played before transitioning to the next particle using overlap addition. If the currently playing particle is at 50% of its length and overlap addition with the next particle is initiated, the two particles might be 180 degrees out of phase during the overlap addition period (e.g., ...). Figure 16A (as illustrated in the diagram), and this can lead to severe cancellation, where the two signals more or less cancel each other out.

[0247] like Figure 16A As shown, particle n is the particle currently being rendered. At some point t when entering this particle, a fast particle switch is triggered. In this case, the next particle to transition to (particle n+1) is out of phase with particle n in the overlapping region. This will cause severe waveform distortion during the overlap, and a smooth transition will be impossible.

[0248] By tracking the current playback position and calculating the corresponding starting position in the next particle, phase can be maintained and phase cancellation can be avoided during particle switching, such as... Figure 16B As shown. More specifically, Figure 16B Phase compensation is shown to ensure that particle n+1 is in phase with particle n during overlap. Phase compensation is accomplished by skipping the first portion of particle n+1, which corresponds to the amount of particle n used before triggering a fast particle switch.

[0249] For example, in one embodiment, in response to detecting that the fast particle condition is met, the current playback position sp of the first particle is stored. For example, if, for example, when the condition is determined to be met, sample i in the currently rendering particle has been played, but sample i+1 has not yet been played, then sp is set to equal i+1. The following set s[] of the mixed samples is then obtained as follows: s[n] = g1[n]grain_n[sp+n] + g2[n]grain_n+1[sp+n], in the range n = 0 to X-1, where grain_n[] is the set of samples that make up particle n, grain_n+1[] is the set of samples that make up particle n+1, and X is the length of the overlapping region.

[0250] Figure 17 This is a flowchart illustrating process 1700 according to an embodiment. Process 1700 may begin at step s1702. Step s1702 includes selecting a first particle from a plurality of particles. Step s1704 includes rendering at least a first portion of the first particle. Step s1706 includes, while rendering the first particle, obtaining information indicating an updated target location in the ND descriptor space, and determining whether a fast particle switching condition is met based on the updated target location. Step s1708 includes, since the fast particle switching condition is determined to be met, transitioning from the first particle to a second particle.

[0251] In some embodiments, transitioning from the first particle to the second particle includes: stopping rendering the first particle; and rendering at least a portion of the second particle.

[0252] In some embodiments, transitioning from a first particle to a second particle includes: generating a hybrid sample by mixing at least a portion of the first particle with at least a first portion of the second particle; and rendering the hybrid sample.

[0253] In some embodiments, the method further includes rendering at least a second portion of the second particle after rendering the blended sample.

[0254] In some embodiments, the method further includes: obtaining information indicating a first target location in the ND descriptor space before selecting the first particle, wherein the first particle is selected from a plurality of particles based on the first target location; and determining the distance between the updated target location and the first target location, wherein determining whether a fast particle switching condition is met includes comparing the distance with a predetermined value.

[0255] In some embodiments, the method further includes determining the distance between the location of the first particle in the ND descriptor space and the updated target location, and determining whether a fast particle switching condition is met includes comparing the distance with a predetermined value.

[0256] In some embodiments, rendering at least a portion of the first particle includes generating a mixed sample using samples from the first particle and samples from the third particle, the first particle having a position in the ND descriptor space, the third particle having a position in the ND descriptor space, and the determination of whether the fast particle switching condition is met is also based on the positions of the first particle and the third particle in the ND descriptor space.

[0257] In some embodiments, determining whether the fast particle switching condition is met includes: determining a first distance between the updated target location and the location of the first particle; determining a second distance between the updated target location and the location of the third particle; comparing the first distance and the second distance; determining, based on the comparison, that the first distance is less than the second distance; comparing the first distance with a predefined value; and determining, based on the comparison, whether the fast particle switching condition is met.

[0258] In some embodiments, determining whether the fast particle switching condition is met includes: determining a first distance between the updated target location and a straight line passing through the locations of the first particle and the third particle; comparing the first distance with a predefined value; and determining whether the fast particle switching condition is met based on the comparison.

[0259] In some embodiments, the method further includes obtaining information indicating a first target location in the ND descriptor space before selecting the first particle, wherein selecting the first particle from a plurality of particles based on the first target location and determining whether a fast particle switching condition is met includes: determining a first distance between the first target location and a straight line passing through the location of the first particle and the location of the third particle; determining a second distance between the updated target location and the straight line; determining the difference between the first distance and the second distance; comparing the difference with a predefined value; and determining whether a fast particle switching condition is met based on the comparison.

[0260] Example use cases Figure 7AThe illustration shows an XR system 700 according to one embodiment, and the embodiments disclosed herein can be applied to this system. Figure 7A As shown, the XR system 700 includes an XR headset 720 (e.g., XR goggles, XR glasses, XR head-mounted display (HMD), etc.) configured to be worn by a user and operable to display an XR scene (e.g., a VR scene in which the user is virtually immersed), speakers 734 and 735 for generating sound for the user, and an input device 750 for receiving input from the user. In this example, the input device 750 takes the form of a joystick.

[0261] like Figure 7B As shown, the XR headset 720 may include an orientation sensing unit 721, a position sensing unit 722, and an XR rendering device (XRRD) 724. In this embodiment, the XRRD 724 includes an audio renderer (AR) 799, which includes a GDB 104 and a GRF 108.

[0262] The orientation sensing unit 721 is configured to detect changes in the user's orientation and provide information about the detected changes to the XR rendering apparatus 724. In some embodiments, the XR rendering apparatus 724 determines an absolute orientation (relative to a coordinate system) given an orientation change detected by the orientation sensing unit 721. In some embodiments, the orientation sensing unit 721 may include one or more accelerometers and / or one or more gyroscopes.

[0263] In addition to receiving input from sensing units 721 and 722, the XR rendering apparatus 724 can also receive input from input device 750 and obtain XR scene configuration information (e.g., particle metadata). Based on these inputs and the XR scene configuration, the XR rendering apparatus 724 renders the XR scene for the user in real time. That is, the XR rendering apparatus generates XR content in real time, including, for example, video data provided to display driver 726 so that display driver 726 will display images contained in the XR scene on display screen 727, and audio data provided to speaker driver 728 so that speaker driver 728 will play audio for the user using speakers 734 and 735. The audio data, or a portion thereof, can be generated by GRF 108 using particles selected from GDB 104, as described above. Although the XR rendering device 724 is shown within the XR headset 720 in this embodiment, in other embodiments, the XR rendering device 724 or one or more of its components (such as GDB 104 and GRF 108) are located remotely from the XR headset 720. In this case, the XR headset 720 and the XR rendering device 724 have communication components (transmitters, receivers) to enable the XR rendering device 724 to transmit XR content to the XR headset 720 (e.g., the XR rendering device or its components may be implemented in the "cloud").

[0264] Figure 8 This is a block diagram of an XR rendering apparatus 724 for performing the methods disclosed herein, according to some embodiments. Figure 8 As shown, the XR rendering device 724 may include: a processing circuitry (PC) 802, which includes one or more processors (P) 855, such as one or more general-purpose microprocessors and / or one or more other processors, such as application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), etc., which may coexist in a single housing or a single data center, or may be geographically distributed (e.g., the XR rendering device 724 may be a distributed computing device containing two or more computers, or a monolithic computing device consisting of only a single computer); at least one network interface 848 (e.g., a physical interface or an over-the-air interface). The XR rendering apparatus 724 includes a transmitter (Tx) 845 and a receiver (Rx) 847, which enable the XR rendering apparatus 724 to transmit and receive data from other nodes connected to the network 110 (e.g., an Internet Protocol (IP) network), the network interface 848 being (physical or wireless) connected to the network 110 (e.g., the network interface 848 may be coupled to an antenna device including one or more antennas to enable the XR rendering apparatus 724 to wirelessly transmit / receive data); and a storage unit (also known as a "data storage system") 808, which may include one or more non-volatile storage devices and / or one or more volatile storage devices. Figure 8 As shown, GDB 104 can be stored in storage unit 808. In embodiments where PC 802 includes a programmable processor, a computer-readable storage medium (CRSM) 842 may be provided. CRSM 842 may store a computer program (CP) 843 containing computer-readable instructions (CRI) 844. CRSM 842 may be a non-transitory computer-readable medium, such as a magnetic medium (e.g., a hard disk), an optical medium, a memory device (e.g., random access memory, flash memory), etc. In some embodiments, the CRI 844 of computer program 843 is configured such that when executed by PC 802, the CRI causes XR rendering apparatus 724 to perform the steps described herein (e.g., the steps described with reference to the flowchart). In other embodiments, XR rendering apparatus 724 may be configured to perform the steps described herein without requiring code. That is, for example, PC 802 may consist of only one or more ASICs. Therefore, the features of the embodiments described herein can be implemented in hardware and / or software.

[0265] Overview of various embodiments A1. A method for rendering audio corresponding to an audio recording, wherein the audio recording is divided into a plurality of particles, the method comprising: selecting a first particle from the plurality of particles, wherein the first particle is associated with a first position in an N-dimensional ND descriptor space; selecting a first set of one or more mixing window coefficients based on the first position in the ND descriptor space; selecting a second particle from the plurality of particles, wherein the second particle is associated with a second position in the N-dimensional descriptor space; selecting a second set of one or more mixing window coefficients based on the second position in the ND descriptor space; obtaining a final mixing window coefficient m using the first set of mixing window coefficients and the second set of mixing window coefficients; and generating a mixed sample S by mixing at least a portion of the first particle with at least a portion of the second particle using the final mixing window coefficients.

[0266] A2. The method according to embodiment A1, wherein the first particle includes a first set of samples, the second particle includes a second set of samples, and generating a mixed sample using a final mixing window coefficient includes generating a first mixed sample s[0] by calculating s=g1×grain1_sample+g2×grain2_sample, where g1 is a function of the final mixing window coefficient, grain1_sample is one of the samples from the first set of samples, g2 is a function of the final mixing window coefficient, and grain2_sample is one of the samples from the second set of samples.

[0267] A3. The method according to embodiment A2, wherein g1 is a further function of a first weight w1 associated with a first mixing window and a second weight w2 associated with a second mixing window, and g2 is a further function of a third weight w3 associated with the first mixing window and a fourth weight w4 associated with the second mixing window.

[0268] A4. The method according to embodiment A3, wherein g1 = m × w1 + (1-m) × w2, and g2 = m × w3 + (1-m) × w4.

[0269] A5. The method according to embodiment A4, wherein w1 = (w2) 1 / 2 And w3 = (w4) 1 / 2 .

[0270] A6. The method according to embodiment A5, wherein w4 = 1 - w2.

[0271] A7. The method according to any one of embodiments A3-A6, wherein w2=n / (L-1), n≥0 and n≤(L-1), L=min(L1,L2)×p, L1 is the length of the first particle, L2 is the length of the second particle, and p is a predetermined overlap percentage.

[0272] A8. The method according to any one of embodiments A3-A6, wherein w2=d1 / (d1+d2), w4=d2 / (d1+d2), d1 is the distance from the first position in the ND descriptor space to the target position in the descriptor space, and d2 is the distance from the second position in the ND descriptor space to the target position in the descriptor space.

[0273] A9. The method according to any one of embodiments A1-A8, wherein the first set of mixed window coefficients consists of a first mixed window coefficient m1, and obtaining the final mixed window coefficients includes setting m to be equal to max(m1,m2), and m2 is a mixed window coefficient contained in a second set of mixed window coefficients, or is based on a mixed window coefficient contained in a second set of mixed window coefficients.

[0274] A10. The method according to any one of embodiments A1-A8, wherein obtaining the final mixed window coefficients comprises: assigning weights to each mixed window coefficient contained in a first set of mixed window coefficients, obtaining a weighted average of the mixed window coefficients contained in the first set of mixed window coefficients using the weights assigned to each mixed window coefficient contained in the first set of mixed window coefficients, and setting m to be equal to max(m1,m2), where m1 is the weighted average of the mixed window coefficients contained in the first set of mixed window coefficients, and m2 is a mixed window coefficient contained in a second set of mixed window coefficients, or based on the mixed window coefficients contained in the second set of mixed window coefficients.

[0275] B1. A method for rendering audio corresponding to an audio recording, wherein the audio recording is divided into a plurality of particles, the method comprising: selecting a first particle from the plurality of particles; selecting a second particle from the plurality of particles; i) selecting a set of two or more blending window coefficients based on a target location in an N-dimensional ND descriptor space, or ii) selecting a single blending window coefficient based on the target location in the ND descriptor space; and blending at least a portion of the first particle with at least a portion of the second particle using the single blending window coefficient in i) or a derived blending window coefficient derived using the set of the two or more blending window coefficients in ii) to generate a blended sample S.

[0276] B2. The method according to embodiment B1, wherein the first particle includes a first set of samples, the second particle includes a second set of samples, and generating a mixed sample using a final mixing window coefficient includes generating a first mixed sample s[0] by calculating s=g1×grain1_sample+g2×grain2_sample, where g1 is a function of the final mixing window coefficient, grain1_sample is one of the samples from the first set of samples, g2 is a function of the final mixing window coefficient, and grain2_sample is one of the samples from the second set of samples.

[0277] B3. The method according to embodiment B2, wherein g1 is a further function of a first weight w1 associated with a first mixing window and a second weight w2 associated with a second mixing window, and g2 is a further function of a third weight w3 associated with the first mixing window and a fourth weight w4 associated with the second mixing window.

[0278] B4. The method according to embodiment B3, wherein g1 = m × w1 + (1-m) × w2 and g2 = m × w3 + (1-m) × w4, where m is the single mixed window coefficient or the derived mixed window coefficient.

[0279] B5. The method according to embodiment B4, wherein w1 = (w2) 1 / 2 And w3 = (w4) 1 / 2 .

[0280] B6. The method according to embodiment B5, wherein w4 = 1 - w2.

[0281] B7. The method according to any one of embodiments B3-B6, wherein w2=n / (L-1), n≥0 and n≤(L-1), L=min(L1,L2)×p, L1 is the length of the first particle, L2 is the length of the second particle, and p is a predetermined overlap percentage.

[0282] B8. The method according to any one of embodiments B3-B6, wherein w2=d1 / (d1+d2), w4=d2 / (d1+d2), d1 is the distance from the first position in the ND descriptor space to the target position in the descriptor space, and d2 is the distance from the second position in the ND descriptor space to the target position in the descriptor space.

[0283] C1. A computer program containing instructions that, when executed by a processing circuit of a rendering apparatus, cause the rendering apparatus to perform the method described in any one of embodiments A1-A10 or B1-B8.

[0284] C2. A carrier comprising the computer program described in embodiment C1, wherein the carrier is one of an electronic signal, an optical signal, a radio signal, and a computer-readable storage medium.

[0285] D1. A rendering apparatus for rendering audio corresponding to an audio recording, wherein the audio recording is divided into a plurality of particles, the rendering apparatus comprising: a memory; and processing circuitry coupled to the memory, wherein the rendering apparatus is configured to perform a method comprising the steps of: selecting a first particle from the plurality of particles, wherein the first particle is associated with a first position in an N-dimensional ND descriptor space; selecting a first set of one or more mixing window coefficients based on the first position in the ND descriptor space; selecting a second particle from the plurality of particles, wherein the second particle is associated with a second position in the N-dimensional descriptor space; selecting a second set of one or more mixing window coefficients based on the second position in the ND descriptor space; obtaining a final mixing window coefficient m using the first set of mixing window coefficients and the second set of mixing window coefficients; and generating a mixed sample S by mixing at least a portion of the first particle with at least a portion of the second particle using the final mixing window coefficients.

[0286] D2. The rendering apparatus according to embodiment D1, wherein the rendering apparatus is further configured to perform the method of any one of claims A2-A10.

[0287] E1. A rendering apparatus for rendering audio corresponding to an audio recording, wherein the audio recording is divided into a plurality of particles, the rendering apparatus comprising: a memory; and processing circuitry coupled to the memory, wherein the rendering apparatus is configured to perform a method comprising the steps of: selecting a first particle from the plurality of particles; selecting a second particle from the plurality of particles; i) selecting a set of two or more blending window coefficients based on a target location in an N-dimensional ND descriptor space, or ii) selecting a single blending window coefficient based on the target location in the ND descriptor space; and blending at least a portion of the first particle with at least a portion of the second particle using the single blending window coefficient of i) or a derived blending window coefficient derived using the set of two or more blending window coefficients, thereby generating a blended sample S.

[0288] E2. The rendering apparatus according to embodiment D1, wherein the rendering apparatus is further configured to perform the method of any one of claims B2-B8.

[0289] While various embodiments have been described herein, it should be understood that they are presented by way of example only and not as limitations. Therefore, the breadth and scope of this disclosure should not be limited to any of the exemplary embodiments described above. Furthermore, unless otherwise indicated herein or there is a clear contradiction in the context, this disclosure covers any combination of the foregoing elements in all possible variations.

[0290] Furthermore, although the processes illustrated above and in the accompanying figures are shown as a series of steps, this is for illustrative purposes only. Therefore, it is expected that some steps may be added, some steps may be omitted, the order of steps may be rearranged, and some steps may be performed in parallel. Further, as used herein, "a" means "at least one" or "one or more".

[0291] References [1]US20180068487A1: Systems and methods for simulating sounds of avirtual object using procedural audio (Disney Enterprises Inc.). [2]Farnell, Andy, "An introduction to procedural audio and itsapplication in computer games,”Audio mostly conference Vol. 23. 2007. [3]D. Gabor, “Theory of communication. Part 1: The analysis ofinformation,”Journal of the Institution of Electrical Engineers-Part III:Radio and CommunicationEngineering, vol. 93, no. 26, pp. 429–441, 1946. [4]D. Schwarz, "Corpus-Based Concatenative Synthesis," in IEEE SignalProcessing Magazine, vol. 24, no. 2, pp. 92-104, March 2007, doi:10.1109 / MSP.2007.323274. [5]D. Schwarz "Distance mapping for corpus-based concatenativesynthesis." Sound and Music Computing (SMC) 2011. [6]Aaron Einbond and Diemo Schwarz. "Spatializing timbre with corpus-based concatenative synthesis." International Computer MusicConferenceProceedings. Vol. 2010. International Computer Music Association,2010. [7]D. Schwarz. "A system for data-driven concatenative soundsynthesis." Digital Audio Effects (DAFx). 2000. [8]D. Schwarz et al. "Real-time corpus-based concatenative synthesiswith catart." 9th International Conference on Digital Audio Effects (DAFx).2006. [9]Diemo Schwarz and Norbert Schnell. "Descriptor-based sound texturesampling." Sound and music computing (SMC) 2010.

[10] Stefan Kersten and Hendrik Purwins. “Sound texture synthesis withHidden Markov Tree models in the wavelet domain”. In Proceedings of theInternationalConference on Sound and Music Computing (SMC), Barcelona, Spain,July 2010.

[11] Diemo Schwarz and Sean O'Leary. "Smooth granular sound texturesynthesis by control of timbral similarity." Sound and Music Computing (SMC)2015.

[12] Zechen Zhang, Nikunj Raghuvanshi, John Snyder, and SteveMarschner. 2019. Acoustic texture rendering for extended sources in complexscenes. ACM Trans.Graph. 38, 6, Article 222 (December 2019), 9 pages.https: / / doi.org / 10.1145 / 3355089.3356566.

[13] Barrass, Stephen, and Matt Adcock. "Interactive granularsynthesis of haptic contact sounds." Audio Engineering Society conference:22ndinternational conference: virtual, synthetic, and entertainment audio.Audio Engineering Society, 2002.

[14] US20190094975A1: Haptic Effect Conversion System Using GranularSynthesis: Immersion Corp.

[15] T. Park, J. Biguenet, Z. Li, C. Richardson, and T. Scharr,“Feature modulation synthesis (FMS),”in Proc. ICMC, (Copenhagen, Denmark),2007.

[16] Fink,Marco, Martin Holters, and Udo Zölzer. "Signal-matchedpower-complementary cross-fading and dry-wet mixing." Proceedings of the 19thInternationalConference on Digital Audio Effects (DAFx-16). 2016。

Claims

1. A method (1200) for rendering audio corresponding to an audio record (111), wherein the audio record is divided into multiple particles, the method comprising: Select (s1202) a first particle from the plurality of particles, wherein the first particle is associated with a first position in the N-dimensional ND descriptor space; Based on the first position in the ND descriptor space, select (s1204) a first set of one or more mixed window coefficients; Select (s1206) a second particle from the plurality of particles, wherein the second particle is associated with a second position in the ND descriptor space; Based on the second position in the ND descriptor space, select (s1208) a second set of one or more mixed window coefficients; Using the first set of mixed window coefficients and the second set of mixed window coefficients, the final mixed window coefficient m (s1210) is obtained; as well as A (s1212) mixed sample S is generated by mixing at least a portion of the first particle with at least a portion of the second particle using the final mixing window coefficient.

2. The method of claim 1, wherein The first particle comprises a first set of samples. The second particle comprises a second set of samples, and Generating the mixed sample using the final mixing window coefficients includes generating a first mixed sample s[0] by calculating s = g1 × grain1_sample + g2 × grain2_sample, where g1 is a function of the final mixing window coefficients. grain1_sample is one of the samples from the first set of samples. g2 is a function of the final mixing window coefficients, and grain2_sample is one of the samples from the second set of samples.

3. The method of claim 2, wherein g1 is a further function of the first weight w1 associated with the first blending window and the second weight w2 associated with the second blending window, and g2 is a further function of the third weight w3 associated with the first blending window and the fourth weight w4 associated with the second blending window.

4. The method of claim 3, wherein g1 = m × w1 + (1-m) × w2, and g2 = m × w3 + (1-m) × w4.

5. The method of claim 4, wherein w1 = (w2) 1 / 2 ,as well as w3 = (w4) 1 / 2 。 6. The method of claim 5, wherein w4 = 1 - w2.

7. The method according to any one of claims 3-6, wherein w2 = n / (L-1), n ≥ 0 and n ≤ (L-1), L = min(L1,L2) × p, L1 is the length of the first particle. L2 is the length of the second particle, and p is the predetermined overlap percentage.

8. The method according to any one of claims 3-6, wherein w2 = d1 / (d1+d2), w4 = d2 / (d1+d2), d1 is the distance from the first position in the ND descriptor space to the target position in the descriptor space, and d2 is the distance from the second position in the ND descriptor space to the target position in the descriptor space.

9. The method according to any one of claims 1-8, wherein The first set of mixed window coefficients consists of the first mixed window coefficient m1. Obtaining the final mixing window coefficients involves setting m to be equal to max(m1, m2), and m2 is a set of mixed window coefficients included in the second set of mixed window coefficients, or is based on the mixed window coefficients included in the second set of mixed window coefficients.

10. The method according to any one of claims 1-8, wherein, Obtaining the final mixing window coefficients includes: Assign weights to each of the first set of mixed window coefficients contained in the mixed window coefficients. Using the weights assigned to each mixed window coefficient in the first set of mixed window coefficients, obtain the weighted average of the mixed window coefficients in the first set of mixed window coefficients, and Set m to be equal to max(m1, m2), where m1 is the weighted average of the mixed window coefficients contained in the first set of mixed window coefficients, and m2 is a set of mixed window coefficients included in the second set of mixed window coefficients, or is based on the mixed window coefficients included in the second set of mixed window coefficients.

11. A method (1300) for rendering audio corresponding to an audio record (111), wherein the audio record is divided into multiple particles, the method comprising: Select (s1302) a first particle from the plurality of particles; Select a second particle (s1304) from the plurality of particles; i) Selecting a set of two or more blending window coefficients based on the target location in the N-dimensional ND descriptor space (s1306), or ii) Selecting a single blending window coefficient based on the target location in the ND descriptor space (s1306). as well as A (s1308) mixed sample S is generated by mixing at least a portion of the first particle with at least a portion of the second particle using either i) the single mixing window coefficient or ii) a derived mixing window coefficient derived from the set of two or more mixing window coefficients.

12. The method of claim 11, wherein The first particle comprises a first set of samples. The second particle comprises a second set of samples, and Generating the mixed sample using the final mixing window coefficients includes generating a first mixed sample s[0] by calculating s = g1 × grain1_sample + g2 × grain2_sample, where g1 is a function of the final mixing window coefficients. grain1_sample is one of the samples from the first set of samples. g2 is a function of the final mixing window coefficients, and grain2_sample is one of the samples from the second set of samples.

13. The method of claim 12, wherein g1 is a further function of the first weight w1 associated with the first blending window and the second weight w2 associated with the second blending window, and g2 is a further function of the third weight w3 associated with the first blending window and the fourth weight w4 associated with the second blending window.

14. The method of claim 13, wherein g1 = m × w1 + (1-m) × w2, and g2 = m × w3 + (1-m) × w4, where m is either the single mixed window coefficient or the derived mixed window coefficient.

15. The method of claim 14, wherein w1 = (w2) 1 / 2 ,as well as w3 = (w4) 1 / 2 。 16. The method of claim 15, wherein w4 = 1 - w2.

17. The method of any one of claims 13-16, wherein w2 = n / (L-1), n ≥ 0 and n ≤ (L-1), L = min(L1,L2) × p, L1 is the length of the first particle. L2 is the length of the second particle, and p is the predetermined overlap percentage.

18. The method of any one of claims 13-16, wherein w2 = d1 / (d1+d2), w4 = d2 / (d1+d2), d1 is the distance from the first position in the ND descriptor space to the target position in the descriptor space, and d2 is the distance from the second position in the ND descriptor space to the target position in the descriptor space.

19. A computer program (843) comprising instructions (844) which, when executed by a processing circuit (802) of a rendering apparatus (724), cause the rendering apparatus to perform the method as described in any one of claims 1-18.

20. A carrier containing the computer program as described in claim 29, wherein, The carrier is one of electrical signals, optical signals, radio signals, and computer-readable storage media (842).

21. A rendering apparatus (724) for rendering audio corresponding to an audio record (111), wherein the audio record is divided into multiple particles, the rendering apparatus comprising: Memory (842); as well as A processing circuit (802) coupled to the memory, wherein the rendering device is configured to execute a method (1200) comprising: Select (s1202) a first particle from the plurality of particles, wherein the first particle is associated with a first position in the N-dimensional ND descriptor space; Based on the first position in the ND descriptor space, select (s1204) a first set of one or more mixed window coefficients; Select (s1206) a second particle from the plurality of particles, wherein the second particle is associated with a second position in the N-dimensional descriptor space; Based on the second position in the ND descriptor space, select (s1208) a second set of one or more mixed window coefficients; Using the first set and the second set of mixed window coefficients, the final mixed window coefficient m (s1210) is obtained; and A (s1212) mixed sample S is generated by mixing at least a portion of the first particle with at least a portion of the second particle using the final mixing window coefficient.

22. The rendering apparatus of claim 21, wherein, The rendering apparatus is further configured to perform the method as described in any one of claims 2-10.

23. A rendering apparatus (724) for rendering audio corresponding to an audio record (111), wherein the audio record is divided into multiple particles, the rendering apparatus comprising: Memory (842); as well as A processing circuit (802) coupled to the memory, wherein the rendering device is configured to execute a method (1300) comprising: Select (s1302) a first particle from the plurality of particles; Select a second particle (s1304) from the plurality of particles; i) selecting a set of two or more blending window coefficients based on the target location in the N-dimensional ND descriptor space (s1306), or ii) selecting a single blending window coefficient based on the target location in the ND descriptor space (s1306); and A (s1308) mixed sample S is generated by mixing at least a portion of the first particle with at least a portion of the second particle using either i) the single mixing window coefficient or ii) a derived mixing window coefficient derived from the set of two or more mixing window coefficients.

24. The rendering apparatus of claim 23, wherein, The rendering apparatus is further configured to perform the method as described in any one of claims 12-18.

Citation Information

Patent Citations

  • Systems and methods for simulating sounds of a virtual object using procedural audio

    US20180068487A1

  • Haptic effect conversion system using granular synthesis

    US20190094975A1