Method for producing soundscapes
By adjusting audio content properties to align with therapeutic frequency distributions, the method enhances user engagement and therapeutic efficacy in sound therapy, allowing personalized and enjoyable soundscapes.
Patent Information
- Authority / Receiving Office
- GB · GB
- Patent Type
- Applications
- Current Assignee / Owner
- KP ACOUSTICS
- Filing Date
- 2024-09-24
- Publication Date
- 2026-04-29
AI Technical Summary
Existing sound therapy methods often rely on pure noise, which may not engage users effectively, and there is a need for audio content with therapeutic properties that users can enjoy and distinguish individually.
A method to adjust the properties of multiple audio content items, such as sound power levels, to align their frequency distribution with a desired distribution, using regression analysis and gain adjustments, allowing users to select meaningful sounds that approximate therapeutic frequency distributions like white or pink noise.
Enhances the therapeutic effects of sound therapy by engaging users with personalized, enjoyable soundscapes that approximate desired frequency distributions, improving treatment adherence and emotional connection.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
Field of the disclosure The present disclosure relates to provision of audio, and more particularly to processing of a plurality of audio content items into a soundscape with a frequency distribution approximating a desired frequency distribution. Background to the Disclosure Sound or noise therapy is a technique involving subjecting a user to audio content which has certain pre-determined properties (such as frequency properties) which are intended to have a therapeutic or otherwise beneficial effect on the user. In many applications, the audio content is selected for its therapeutic / beneficial properties, and may consist of noise. It is desired to produce audio content which has the desired properties but which is formed of distinct and meaningful sounds, which may for example be selected by a user. Summary of the Disclosure According to an aspect of the present disclosure, there is described: a method comprising adjusting at least one property of at least one of a plurality of audio content items such that a difference between a frequency distribution of a mix of the plurality of audio content items and a desired frequency distribution is reduced. Reducing the difference between the mix and a desired frequency distribution allows the mix to be used for the same purposes as are known for the desired frequency distributions, such as for therapeutic purposes. Advantageously, the provision of a mix of audio content items, rather than content consisting of pure noise, may enhance the therapeutic / beneficial effects, for example by encouraging users to listen to a mix of sounds that they enjoy. The resultant adjusted mix of the plurality of audio content items is preferably configured such that a user may distinguish individual content items. In examples, adjusting at least one property of at least one of a plurality of audio content items comprises adjusting a sound power level of at least one of the plurality of audio content items. In examples, the at least one property is a sound power level. In examples, the desired frequency distribution is defined by an equation relating frequency and power. In examples, the desired frequency distribution is linear in a logarithmic space. In examples, the desired frequency distribution corresponds to an exponential relationship between frequency and power. In examples, the method comprises adjusting at least one property iteratively. In examples, adjusting at least one property of at least one of a plurality of audio content items comprises applying a regression analysis; preferably a non-linear regression analysis; more preferably a non-linear least squares analysis, optionally comprising use of the Levenberg-Marquardt algorithm. In examples, adjusting at least one property of at least one of a plurality of audio content items comprises adjusting a gain of at least one audio content item of the plurality of audio content items. Preferably, the method further comprises adjusting a gain of each audio content item of the plurality of audio content items. More preferably, the method comprises adjusting a gain of each audio content item of the plurality of audio content items so as to minimise the sum of square residuals associated with the plurality of audio content items. In some examples, the method comprises dividing each of the plurality of audio content items into frequency bands, preferably one-third octave bands, and determining the sound pressure level of each of the plurality of audio content items in each frequency band. In examples, determining the sound pressure level is done by determining the power in a frequency band based on the signal with no gain applied. Preferably, the method comprises determining the sound pressure level for each frequency band by calculating a sum of: the sound pressure level of each of the plurality of audio content items in each frequency band; and the gain associated with each of the audio content items. Since gain is multiplicative while sound pressure level is logarithmic, this sum corresponds to the power in each frequency band. The gain in this context may be a dummy variable that is adjusted to determine adjusted gains. More preferably, the method comprises determining the standard deviation in respect of the sound pressure level for each frequency band; and adjusting the gain of each audio content item in order to minimise the standard deviation. In examples, the method comprises further adjusting the gain of at least one of the plurality of audio content items, optionally such that a difference between a frequency distribution of a mix of the plurality of audio content items and a further desired frequency distribution is reduced. Adjusting the gain in two stages allows the gains of the plurality of audio content items to be normalised (e.g. with reference to white noise), before being further adjusted to correspond to a further desired frequency distribution having an improved therapeutic effect. In this regard, , first gain adjustments may be computed relative to white noise, and then the gains may be adjusted to minimise the difference with a pink noise distribution. That is, the desired frequency distribution may be white noise (i.e. having equal power throughout its bandwidth), and the further desired frequency distribution may be a different frequency distribution, preferably wherein the further desired frequency distribution may be pink noise (i.e. 1 / f noise). Adjusting the gain in two stages may also simplify storage and retrieval of pre-calculated gain values. For example, gains could be stored in a database to minimise the difference between a sum of the audio content items and a first frequency distribution (the desired frequency distribution), and then these gains could be adjusted to minimise the difference between the final output and any of a plurality of frequency distributions (the further desired frequency distribution). In examples, further adjusting the gain of at least one of the plurality of audio content items comprises increasing the gain of at least one of the content items; and reducing the gain of at least one of the content items. Preferably, further adjusting the gain of at least one of the plurality of audio content items comprises increasing the gain of at least one low frequency content item; and reducing the gain of at least one high frequency content item. In examples, the method further comprises calculating a sound pressure level of the mix of the audio content items, and additionally adjusting the gain for each of the plurality of audio content items to adjust the sound pressure level. Preferably, additionally adjusting the gain for each of the plurality of audio content items to adjust the sound pressure level is in dependence on calibration information. More preferably, additionally adjusting the gain comprises adjusting the gain of each of the plurality of audio content items equally. Preferably, additionally adjusting the gain is done in dependence on information regarding a desired audio output method. In examples, a desired audio output method may be a speaker array. In examples, a desired audio output method may be headphones. In examples, a desired audio output method may be five speaker surround sound, or any equivalent sound system layout as known in the art (such as 5.1 surround sound). In examples, the method comprises retrieving at least one pre-calculated adjustment to at least one property from a database. In examples, the at least one property may be a gain of the audio. In examples, the plurality of audio content items comprises at least one content item selected by a user from a superset of audio content items. In examples, the method further comprises selecting, by a user, at least one content item of the plurality of audio content items from a superset of audio content items. In examples, the superset of audio content items comprises a plurality of categories of audio content items, and wherein selecting, by a user, the plurality of audio content items comprises selecting an audio content item from each category of audio content items. Preferably, the method further comprises categorising the plurality of audio content items into the plurality of categories based on their frequency characteristics. In examples, the plurality of categories of audio content items comprise the following categories: low frequency audio content items; medium frequency audio content items; and high frequency audio content items. In examples, at least one of the plurality of categories of audio content items is defined as follows: audio content items in the low frequency category have at least 70% of their power spectrum in the range 20 - 800 Hz; audio content items in the medium frequency category have at least 70% of their power spectrum in the range 800 Hz - 2 kHz; and audio content items in the high frequency category have at least 70% of their power spectrum in the range 2 kHz - 20 kHz. In examples, the method comprises calculating, for all combinations of audio content items from each of the categories, adjustments to at least one property of a plurality of audio content items such that a difference between a frequency distribution of a mix of the plurality of audio content items and a desired frequency distribution is reduced; and recording the calculated adjustments in a database. For example, each combination of a low frequency, a medium frequency, and a high frequency, may be analysed to calculate the adjustments to at least one property that would reduce / minimise the difference between a mix of the audio content items and a desired frequency distribution. In examples, the plurality of audio content items comprises at least one high frequency audio content item; at least one medium frequency audio content item; and at least one low frequency audio content item. Preferably, the plurality of audio content items consists of one high frequency audio content item, one medium frequency audio content item, and one low frequency audio content item. In examples, the difference between the frequency distribution of the mix of the plurality of audio content items and the desired frequency distribution is minimised. In examples, the plurality of audio content items comprises at least one audio content item provided by the user. For example, the user may record an audio content item, and then the method may categorise the audio content item into low, medium or high frequency, and may then allow a user to select further audio content items from the remaining frequency categories. The method may then comprise calculating adjusted gains for the combination of the at least one audio content item provided by the user with the further audio content items. In examples, the desired frequency distribution is 1 / f noise. In examples, the desired frequency distribution has equal power throughout its bandwidth. That is, the desired frequency distribution may be described as ‘flat’. Preferably, a final frequency distribution (or a further desired frequency distribution) is 1 / f noise. In examples, the desired frequency distribution has therapeutic properties. Preferably, the desired frequency distribution has therapeutic properties for individuals with attention deficit hyperactive disorder. In examples, the method comprises playing the mix of audio content items, after the at least one property has been adjusted, to a user. In examples, the method comprises converting the mix of audio content items into a spatial reproduction format, preferably an Ambisonics format. According to a second aspect of the present disclosure, there is described: a device comprising computer-executable instructions for executing the method of any example of the first aspect. According to a third aspect of the present disclosure, there is described: a non-transitory computer-readable storage medium comprising instructions that, when executed, perform the method of any example of the first aspect. According to a fourth aspect of the present disclosure, there is described: a method for storing suitable gains for combinations of audio content items, the method comprising: for each combination of audio content items, calculating at least one gain adjustment to reduce a difference between a frequency distribution of a mix of the combination of audio content items and a desired frequency distribution; and storing the gain adjustments in a database. According to a fifth aspect of the present disclosure, there is described: a database comprising a plurality of audio content items, and a plurality of sets of parameters relating to a respective combination of a subset of the plurality of audio content items, wherein each set of parameters minimises the difference between the relevant combination of a subset of the plurality of audio content items and a desired frequency distribution; preferably wherein the sets of parameters are sets of gain adjustments. The audio content items may include sound information which is interpretable by the user (and so do not consist of pure noise). For example, the audio content items may relate to various natural or manmade sounds (being obtained via recording, or being generated by some other process). Preferably, the audio content items relate to generally continuous sounds. The mix of audio content items may be referred to as a ‘soundscape’ (where this term in this contest refers to an acoustic environment being formed of generally continuous sounds). As used herein, unless otherwise specified a frequency distribution refers only to a non-zero frequency distributions, i.e. including a finite quantity of sound power distributed across the spectrum, such that the spectrum does not correspond to silence. As used herein, unless otherwise specified the term “mix” refers to a combination of a plurality of content items. As used herein the term ‘sound pressure level’ refers to a measure of the pressure variation in a medium caused by a sound wave relative to a reference pressure of 20 micropascals. The term ‘sound power level’ is used herein to refer to a total acoustic energy emitted by a sound source in decibels. The term “gain” is used to refer to an amount by which an audio signal is amplified, as a ratio relative to the input signal. It is measured in decibels. Increasing the gain of a signal will increase the sound power level when the signal is played. Any feature in one aspect of the disclosure may be applied to other aspects of the invention, in any appropriate combination. In particular, method aspects may be applied to apparatus aspects, and vice versa. Furthermore, features implemented in hardware may be implemented in software, and vice versa. Any reference to software and hardware features herein should be construed accordingly. Any apparatus feature as described herein may also be provided as a method feature, and vice versa. As used herein, means plus function features may be expressed alternatively in terms of their corresponding structure, such as a suitably programmed processor and associated memory. It should also be appreciated that particular combinations of the various features described and defined in any aspects of the disclosure can be implemented and / or supplied and / or used independently. The disclosure also provides a computer program and a computer program product comprising software code adapted, when executed on a data processing apparatus, to perform any of the methods described herein, including any or all of their component steps. The disclosure also provides a computer program and a computer program product comprising software code which, when executed on a data processing apparatus, comprises any of the apparatus features described herein. The disclosure also provides a computer program and a computer program product having an operating system which supports a computer program for carrying out any of the methods described herein and / or for embodying any of the apparatus features described herein. The disclosure also provides a computer readable medium having stored thereon the computer program as aforesaid. The disclosure also provides a signal carrying the computer program as aforesaid, and a method of transmitting such a signal. The disclosure extends to methods and / or apparatus substantially as herein described with reference to the accompanying drawings. The disclosure will now be described, by way of example, with reference to the accompanying drawings. Description of the Drawings Figure 1 shows an example method for producing a soundscape according to an embodiment; Figure 2 shows an example method for calculating gain adjustments for use in producing a soundscape; Figure 3 shows example frequency spectra for a plurality of audio content items and a mix of audio content items; Figure 4 shows example frequency spectra for a plurality of audio content items; Figure 5 shows a graphical user interface for selecting a plurality of audio content items; Figure 6 shows an example audio setup; Figure 7 shows an example method for calculating gain adjustments; Figure 8 shows an application enabling participants to personalize their soundscape; and Figure 9 shows an example method for calculating gain adjustments. Description of the preferred embodiments Referring to Figure 1, there is shown a high-level explanation of a method 100 for producing a soundscape (i.e. a mix of audio content items) using a computer-implemented system. At step 110, the system receives an indication of a plurality of audio content items. In examples, this may be an indication of audio content items held in an audio content library or database. For example, the indication may be an index of each audio content item in the database. In other examples, the indication of a plurality of audio content items may instead comprise transmitting the audio content items directly, so that instead of providing a reference or index of the audio content items, the audio content items themselves are provided. In further examples, the indication may comprise providing an indication of a first audio content item and transmitting a second audio content item directly. For example, one audio content item might be recorded by a user and the indication of this audio content item might be a transmission of the recording, while the indication of a second audio content item of the plurality of audio content items might be an index of the second audio content item in a database. At step 120, the system adjusts at least one property of at least one of the plurality of audio content items to reduce a difference between a mix of the audio content items and a desired frequency distribution. For example, the system may adjust the gain of at least one of the audio content items. The adjustment reduces a difference between a mix of the audio content items and a desired frequency distribution. For example, when the property is a gain of the audio content items, a different gain may be chosen for each audio content item of the plurality of audio content items in order to reduce the difference between a mix of the audio content items (each with their respective gain) and a desired frequency distribution. When the property is a frequency of the audio content items, a frequency of at least one of the audio content items may be increased or decreased so that the combination of audio content items, once the frequency has been adjusted, is closer to a desired frequency distribution than the combination of audio content items would be without the adjustment. At step 130, the mix of the audio content items is played back to a user. This mix comprises the adjustments of step 120, and so is closerto the desired frequency distribution than the unadjusted combination of audio content items would be. The mix of audio content items may be played to the user through the same system that performs the method. Alternatively, the mix of audio content items may be transmitted to an audio device, such as a set of speakers or headphones, in order to play back the mix of audio content items to a user. In a particular example, the gains of all of the plurality of audio content items are adjusted. Three audio content items may be used, being respectively a low frequency audio content item, a medium frequency content item, and a high frequency content item (it will be appreciated that the frequencies of all of the content items are selected so as to be audible to humans). In examples, the audio content items primarily consist of narrower-band noise centered around specific frequencies, rather than covering the entire spectrum. In examples, the audio content items are categorized into three groups based on their central frequency and narrow-band characteristics: Low-frequency band samples: 20Hz - 800Hz Mid-frequency band samples: 800Hz - 2kHz High-frequency band samples: 2kHz - 20kHz For example, low-frequency band samples may have 70% of their power within the 20Hz - 800Hz range. Figure 2 shows a method 200 for calculating gain adjustments for use in producing a soundscape, for example producing a soundscape according to the method shown in Figure 1. Determining gain adjustments that may be applied to each audio content item in order to reduce a difference between the frequency distribution of a mix of the audio content items and a desired frequency distribution can be considered to be a minimisation problem. “Minimisation” here refers to finding the optimum gain adjustments for the audio content items within a set of constraints. This problem can be considered to be of the form: mjn [[f(x)]]2 = min (AW2 + AW2 + AW2) X X wherein f^x), f3(x) and f3(x) are the residuals for each of the audio content items (i.e. the low frequency, medium frequency, and high frequency audio content items), and the min function aims to minimise the sum of the squared residuals across all values of x. The input objective function (which may be referred to as a loss function) may be a function that uses proposed gains for the unknown gain variables to optimise the overall frequency response of the combined low, medium and high frequency audio content items, such as by iteratively adjusting the gains in the input objective function until the min function returns a minimum value. Thus, the solution to the problem can be considered to be of the form: Opt.Resp. = G|0W(LFrSample)+Gmid(MFrSample)+Ghigh(HFrSample) wherein Glow, Gmid and Ghigh are the optimal gains for each of the low frequency audio content item, the medium frequency audio content item, and the high frequency audio content item respectively, and Opt.Resp. is the combined frequency distribution with optimal gains. The method comprises, at step 210, receiving a plurality of audio content items. For example, a low frequency audio content item, a medium frequency audio content item, and a high frequency audio content item may all be received. At step 220, each audio content item is subdivided into a plurality of segments in frequency space. For example, each audio content item may be subdivided into 1 / 3 octave segments, each corresponding to a respective frequency band. In examples, each audio content item may be subdivided into 31 segments. An octave is the interval between a first frequency and a second frequency, wherein the second frequency is double the first frequency. One third of an octave 1 therefore corresponds to a frequency ratio of 2^ , i.e. 1 / 3 octave segments each span from a 1 frequency f± to a frequency f± * 2T Where the audible range is considered to be 20 Hz - 20 kHz, this corresponds to 31 frequency segments. The use of 1 / 3 octave segments may provide an acceptable trade-off between granularity of adjustment and processing efficiency. Once subdivided, a function may be used to analyse the energy of each audio sample in the one-third octave frequency bands. At step 230, a matrix is formed, comprising the sound pressure level of each audio content item in each segment. For example, the columns of the matrix may each correspond to a respective audio content item, and the rows may each correspond to a respective frequency band, so that one row would contain the sound pressure level in a given frequency band for each of the audio content items. Such a matrix may take the following form: / LowSPL1 MidSPL1 HighSP^ \ A = = = = \LowSPL31 MidSPL31 HighSPL31J At step 240, a regression function is used to determine gain adjustments that reduce a difference between the combined frequency distribution, and a desired frequency distribution. For example, the desired frequency distribution may correspond to white noise or a “flat” distribution, and the gains may be adjusted to minimise the difference between the combined frequency distribution and white noise. In an example, the regression function is non-linear, in order to account for the fact that the frequencies are in log space. In an example, the regression function is a nonlinear least-squares function based on the Levenberg-Marquardt approach. In an example, step 240 comprises calculating the total sound pressure level all of the 1 / 3 octave segments by adding the sound pressure levels of each segment with the relevant gains. For example, the total sound pressure level for the first frequency band would be as follows: SPLtotfrl = LowSPL1 + Glow + MidSPL1 + Gmid + HighSP^ + Ghigh In an example, step 240 further comprises calculating the standard deviation of all SPLtot^^ : / SPLtotfrl \ Std = I j \SPLtotfr31J The standard deviation reflects the deviation of the combined distribution from a flat distribution corresponding to white noise. Minimising the standard deviation therefore minimises a difference between the combined distribution and white noise. In examples, the standard deviation is minimised by iterating the calculations with different values of a gain variable for each audio content item, and determining the values corresponding to optimal gain adjustments. In summary, the optimization process iteratively runs a loss function to optimize the gains, Gtow, Gmid and Ghighm order to minimize the standard deviation. The optimized gains , Glow, Gmid and Ghigh f°r each combination of sounds from the low, mid, and high frequency band folders may then be saved. In an example, the calculated gain adjustments are then stored in a database or lookup table. Alternatively and / or additionally, the optimum gain adjustments are transmitted to a program to be applied to the audio content items for playback. In some examples, the calculated gain adjustments are then further adjusted to produce second gain adjustments. For example, the gains may be adjusted to approximate a frequency distribution other than white noise, such as pink noise (i.e. 1 / f noise - that is, noise with a power spectral density of the form S(f) oc where f is frequency, and 0 <a <2). There is a -3dB per octave band decay in the frequency response towards the high frequencies of pink noise. In other words, pink noise comprises diminished high-frequency components. In this example, the gain for the low frequency audio content item Gtow may therefore be amplified by +6 dB, while the gain for the high frequency audio content item Ghigh may be reduced by -6 dB, in orderto adjust the combined distribution from approximating white noise to one approximating pink noise. The resulting second gain adjustments may then be stored and / or applied to the audio content items directly for playback. In some examples, adjusting the calculated gain adjustments to produce second gain adjustments may be done at a later time and / or by a different device. For example, a first device may calculate the calculated gain adjustments, and store these in a database, and a second device may retrieve the calculated gain adjustments and generate the second gain adjustments. This may be based on a user request, such as a user indicating a frequency distribution which the mix of audio content items should approximate. Figure 3 shows example frequency distributions for a plurality of audio content items. Line 302 shows the frequency distribution of a low frequency audio content item. This peaks towards the left of the graph, as the majority of the power in this audio content item is at low frequencies, but has some power even at high frequencies. Line 304 shows the frequency distribution of a medium frequency audio content item. This has the majority of its power in the middle of the graph, at medium frequencies. Line 306 shows the frequency distribution of a high frequency audio content item. This has the majority of its power at high frequencies. Line 308 shows the sum of the three audio content items represented by lines 302, 304 and 306. This sum shows a peak at lower frequencies, generally decreasing with frequency. Line 310 shows the frequency distribution of a sum of the three audio content items, wherein the gains of the audio content items have been adjusted in order to minimise the difference between the frequency distribution and white noise. As shown, the line 310 is closer to flat than the line 308, and so more closely approximates a white noise distribution. Figure 4 shows exemplary frequency spectra of three possible audio content items, which have been adjusted so that the overall distribution approximates pink noise. The combined distribution may then be played to a listener. The described method may be particularly beneficial in noise therapy for the treatment of attention deficit hyperactivity disorder (ADHD), in particular in children. Research has shown that individuals with ADHD may benefit from noise therapy treatment, in particular treatment based on pink noise. The present disclosure may improve such treatment by allowing users to be treated with meaningful sound (ratherthan noise), where the user may select their preferred sounds. This may make users more likely to listen to the produced soundscape. It is also anticipated that the use of meaningful sounds may allow users to attach emotional significance to the sounds that are used, which may improve treatment and / or further encourage users to listen to the soundscape more regularly. Likewise, the ability of the user to select various different options may improve a user’s engagement with treatment. Personalization of the therapeutic intervention is an important factor in its effectiveness. In examples, the sound content chosen for the audio content items has characteristics similar to pink noise. Recordings of natural soundscapes may be used, featuring broadband noise with random spectral textures and minimal distinct tonal or rhythmic elements. Figure 5 shows an example user interface 500 for an application. The user interface 500 comprises buttons in columns. The first column comprises low frequency sound recordings and is labelled by label 502 as “Low”. Buttons in the first column correspond to a “bonfire” recording 512, a “waves” recording 514, and a “raindrops” recording 516. The second column comprises sound recordings of medium frequencies, and is labelled by label 504 as “Mid”. Buttons in the second column correspond to a “stream” recording 522, a “rain on a tin roof’ recording 524, and an “autumn wind” recording 526. The third column comprises sound recordings of high frequencies, and is labelled by label 506 as “High”. Buttons in the third column correspond to a “rain on trees” recording 532, a “wind on trees” recording 534, and a “fireplace” recording 536. It will be appreciated that other audio content items could instead be provided (though the audio content items should desirably consist of generally continuous sound, in orderto improve results). Buttons 512, 522 and 532 are shown selected in the image. This is illustrated by a stronger border around these buttons. In other examples, selection may be shown in other ways, such as by changing the button text colour, size and / or boldness, and / or by changing a shadow of the button, a background colour of the button, and / or a reflectiveness of the button. Other methods of illustrating selection of a button will occur to the skilled reader. Stop button 542 allows a user to stop all audio from playing. Save button 544 allows the user to save the soundscape created. In use, a user is able to select one recording from each column. The chosen recordings are then combined using a different respective gain for each recording. The gains are chosen in order to reduce the difference between a frequency distribution of a combination of the recordings and a desired frequency distribution, as previously described. For example, the desired frequency distribution may be white noise, and the gain may be set to be lower on the “Bonfire” clip and higher on the “Rain on trees” clip in order to approximate this distribution. In embodiments, audio plays automatically once a recording has been selected. In such embodiments, the user may press the Stop button 542 in order to suspend playback. When a user is happy with their selected audio, the user may press the Save button 544. In embodiments, this saves the selection of the audio content items, and the associated gains. In other embodiments, this saves a combined audio file comprising the mix of audio content items at their associated gains. The gains may be pre-calculated, and retrieved from a database by the device on selection of the sound recordings. Alternatively, the gains may be calculated by the device on receiving the user selection. In still further embodiments, the gains may be partially pre-calculated, and may be adjusted on receiving the user selection. In examples, the user may be able to select a plurality of sound recordings in each frequency category. In examples, the user may be able to provide a selection of sound recordings from only a subset of the frequency categories. For example, the user may be able to provide a selection of sound recordings that does not include a selection of a medium frequency sound recording. In other examples, the user may be limited to selecting one sound recording from each category. In examples, the number of categories may be larger or smaller than shown. For example, four categories may be provided, sorting the sound recordings into categories of very low frequency, medium low frequency, medium high frequency, or very high frequency. The categorisation of sound recordings may be done manually. Alternatively, the categorisation of sound recordings may be automated. Other example audio content items may be recordings of birdsong, canteen or cafe background noise, fire, traffic, or jungle noises. The resulting soundscape may then be played back to the user. In examples, this may be using binaural audio, such as for playback through headphones. In other examples, this may be for playback using Ambisonics. Figure 6 shows an example audio calibration system designed to maintain a consistent sound pressure level (SPL) at the listener’s ears when reproducing audio through both speakers and headphones, specifically focusing on dual mono and 3rd order Ambisonics audio reproduction methods. Ambisonics is a full-sphere surround sound format, allowing a user to be immersed in a 3d soundscape. Dual mono refers to an audio reproduction method where the same mono audio signal is sent to both the left and right channels of a stereo system. The system comprises a plurality of speakers 602 distributed across the space. A dummy head 604 is positioned at the centre of the space. A sound level meter 606 is positioned at the dummy head 604, preferably on top of the dummy head, for measuring the sound pressure level generated by the plurality of speakers 602. Preferably, the dummy head 604 comprises ear canals, so that a sound pressure level within the ear can be determined by microphones 605 within the ears based on the artificial ear acoustics. Headphones 608 are positioned around the dummy head 604. The calibration method comprises two stages. In a first stage, the microphones 605 are calibrated. This is done by playing a sound over the speakers 602 that is registered by the sound level meter 606 as corresponding to an 80 dBA sound power level at the dummy head. The sound may be pink noise, i.e. 1 / f noise. The headphones 608 may be removed for this stage to prevent the headphones from obstructing the microphones. The input level at microphones 605 corresponding to a sound power level of 80 dBA can thereby be determined. Sound level meters are bulky and so may not be practically workable with artificial ear acoustics, such as an artificial ear canal. Calibrating a microphone with a sound power level in this way allows a microphone to be used in a realistic configuration while still obtaining desired information on a sound power level. In a second stage, the microphones 605 are used to compare Ambisonics audio decoded as 3rd order Ambisonics audio, played over the plurality of speakers, with Ambisonics audio decoded played over the headphones, along with audio recorded as dual mono audio played over the speakers, and audio recorded as dual mono played over the headphones. The level of the output audio is adjusted until the input level at microphones 605 corresponds to the reading observed at a sound power level of 80 dBA in the first calibration stage. This allows the output level needed for an 80 dBA sound power level to be received by a user to be observed. In experiments, this corresponded to a -13 dBFS input level when using the plurality of speakers with 3rd order Ambisonics audio, a -17 dBFS output level when using the headphones to play 3rd order Ambisonics audio, a -13 dBFS input level when using the plurality of speakers to play dual mono audio, and a -15 dBFS output level when using the headphones to play dual mono audio. Preferably, when a mix of audio content items is listened to by a user, the user receives a sound power level of 80 dBA. The calibration indicates that, to achieve this, the output level of a set of headphones should be set at -17 dBFS when using Ambisonics audio, and at -15 dBFS when using dual mono audio. In an example embodiment, the system may comprise 50 speakers. In another example embodiment, the system may comprise 26 speakers. The speakers may be arranged in a Levedevgrid configuration. In examples, the received audio signal may be modified by a head-related transfer function in order to more accurately approximate the sound as perceived by a listener. In examples, a generalized head-related transfer function may be used, which may be an average for the listener’s demographic, such as an average adult head-related transfer function, or an average child head-related transfer function. Figure 7 shows an example method for calculating gain adjustments, similar to the method discussed in relation to Figure 2. At step 712, a low frequency natural sound is received. This is an audio content item comprising a recording of a natural sound, having low characteristic frequencies. At step 714, a mid frequency natural sound is received. This is an audio content item comprising a recording of a natural sound, having medium characteristic frequencies. At step 716, a high frequency natural sound is received. This is an audio content item comprising a recording of a natural sound, having high characteristic frequencies. At step 722, the low frequency natural sound is fed through a 1 / 3 octave filter bank. This calculates the sound pressure level of the audio content in each 1 / 3 octave segment of the audio content item. At step 724, the mid frequency natural sound is fed through a 1 / 3 octave filter bank. This calculates the sound pressure level of the audio content in each 1 / 3 octave segment of the audio content item. At step 726, the high frequency natural sound is fed through a 1 / 3 octave filter bank. This calculates the sound pressure level of the audio content in each 1 / 3 octave segment of the audio content item. At step 730, the outputs of each of steps 722, 724 and 726 are combined into a matrix, wherein the rows each comprise the sound pressure level in a respective 1 / 3 octave segment, and wherein the columns each comprise the sound pressure level for a respective audio content item. At step 740, the matrix is fed into a least square non-linear function, which minimises a loss function comprising the standard deviations of the sums of each row in the matrix. At step 750, correction gains are output, corresponding to a respective gain adjustment for each audio content item of the low frequency natural sound, the mid frequency natural sound, and the high frequency natural sound. Figure 8 shows an application enabling participants to personalize their soundscape. The objective of this application is to promote individualism in personal preferences and enhance participant engagement in the process. The application comprises a GUI 810. This may be a GUI as shown in Figure 5. The GUI allows a user to select a plurality of audio content items from a superset of audio content items, provided in low frequency sound bank 822, mid frequency sound bank 824, and high frequency sound bank 826. The GUI then translates this selection into corresponding indices. Where a low frequency, medium frequency, and high frequency audio content item are each selected, the resulting indices may be indicated as idxjow, idx_mid and idx_high respectively. The application uses these indices to find corresponding gain adjustments from a lookup table 830. The lookup table may be produced in Matlab. The lookup table may comprise gain adjustments calculated according to the methods of any of Figures 2, 7 and 9. The gain adjustments are then sent to playback engine binaural audio 840. This receives the gain adjustments and the audio content items, and combines the audio content items, each with their respective gain adjustment, in the application Supercollider during real-time execution of the application. Supercollider is a platform for audio synthesis. The audio is then transmitted to headphones 850. Figure 9 shows an example method 900 for calculating gain adjustments, similar to the method discussed in relation to Figures 2 and 7. At step 910, a plurality of audio content items are received. In examples, these may be a low frequency audio content item, a medium frequency audio content item, and a high frequency audio content item. At step 920, gain adjustments are calculated so that a mix of the audio content items, with the applied gain adjustments, has the flattest frequency response possible with each audio content item at non-zero gain. In examples, a MATLAB algorithm calculates the best fit of a selected triad of audio content items that produces the flattest possible frequency response within a specified frequency range. In examples, the specified frequency range may correspond to the range of frequencies audible to humans. In examples, the algorithm may comprise elements of the method described in relation to Figure 2. At step 930, the gain adjustments for the low and high frequency audio content items are adjusted to achieve the high-frequency decay characteristic of pink noise. Alternatives and modifications It will be understood that the present invention has been described above purely by way of example, and modifications of detail can be made within the scope of the invention. For example, while in the examples, three categories of audio content item are used, in other examples more or fewer categories of audio content items may be used, such as two categories, four categories, or a single category. In some examples, audio content items are categorised based on other features of the audio content. For example, the audio content items may be categorised based on the sound which they reflect, such as waves, cafes, waterfalls, woodland sounds, and other types of sound content. In examples, the audio content items may comprise computer-generated sounds. In examples, the user may be able to provide audio content items. For example, the user may be able to record an environment and send it to the system. In some such examples, the system may then analyse the sound recording to determine whether it is suitable for use in producing a soundscape. For example, the system may determine whether there is a change overtime in the frequency spectrum of the recording that exceeds a threshold, whether there is an impulse exceeding a threshold power level, and whether the recording comprises tonal characteristics exceeding a threshold level. If the system determines that one or more conditions are not met, the system may reject the recording as unsuitable. For example, the system may reject the recording as unsuitable if the change overtime in the frequency spectrum exceeds a threshold. In examples, the system may directly play back the audio to a user. In other examples, the system may stream or otherwise transfer the audio data to another device. In still further examples, the 5 system may save the audio data for later retrieval. In examples, as an alternative or an addition to adjusting the gains, the system may adjust other properties. For example, the system may adjust a frequency spectrum of at least one of the audio content items, such as by shifting the entire frequency spectrum to lower or higher frequencies. In some embodiments, after calculating or retrieving the gains that reduce the difference between 10 the audio content items and a desired frequency distribution, the user may manually adjust the gains to suit their preferences. In such embodiments, the gains may be displayed to the user in a text box, allowing the user to type in new gains, may be displayed as a slider, or by any other suitable means. Reference numerals appearing in the claims are by way of illustration only and shall have no 15 limiting effect on the scope of the claims.
Claims
1. A method comprising:adjusting at least one property of at least one of a plurality of audio content items such that a difference between a frequency distribution of a mix of the plurality of audio content items and a desired frequency distribution is reduced.
2. The method of claim 1, comprising adjusting at least one property iteratively.
3. The method of claim 1 or 2, wherein adjusting at least one property of at least one of a plurality of audio content items comprises applying a regression analysis; preferably a non-linear regression analysis; more preferably a non-linear least squares analysis.
4. The method of any preceding claim, wherein adjusting at least one property of at least one of a plurality of audio content items comprises adjusting a gain of at least one audio content item of the plurality of audio content items.
5. The method of claim 4, comprising adjusting a gain of each audio content item of the plurality of audio content items.
6. The method of claim 5, comprising adjusting a gain of each audio content item of the plurality of audio content items so as to minimise the sum of square residuals associated with the plurality of audio content items.
7. The method of any of claims 4 to 6, comprising dividing each of the plurality of audio content items into frequency bands, preferably one-third octave bands, and determining the sound pressure level of each of the plurality of audio content items in each frequency band, preferably comprising determining the sound pressure level for each frequency band by calculating a sum of: the sound pressure level of each of the plurality of audio content items in each frequency band; and the gain associated with each of the audio content items.
8. The method of claim 7, comprising determining the standard deviation in respect of the sound pressure level for each frequency band; and adjusting the gain of each audio content item in order to minimise the standard deviation.
9. The method of any of claims 4 to 8, comprising further adjusting the gain of at least one of the plurality of audio content items, preferably wherein further adjusting the gain of at least one of the plurality of audio content items comprises increasing the gain of at least one of the content items; and reducing the gain of at least one of the content items.
10. The method of claim 9, wherein further adjusting the gain of at least one of the plurality of audio content items comprises increasing the gain of at least one low frequency content item; and reducing the gain of at least one high frequency content item.
11. The method of any of claims 4 to 10, further comprising calculating a sound pressure level of the mix of the audio content items, and additionally adjusting the gain for each of the plurality of audio content items to adjust the sound pressure level, preferably wherein additionally adjusting the gain for each of the plurality of audio content items to adjust the sound pressure level is in dependence on calibration information.
12. The method of any preceding claim, wherein the method comprises retrieving at least one pre-calculated adjustment to at least one property from a database.
13. The method of any preceding claim, further comprising selecting, by a user, at least one content item of the plurality of audio content items from a superset of audio content items.
14. The method of claim 13, wherein the superset of audio content items comprises a plurality of categories of audio content items, and wherein selecting, by a user, the plurality of audio content items comprises selecting an audio content item from each category of audio content items, preferably further comprising categorising the plurality of audio content items into the plurality of categories based on their frequency characteristics.
15. The method of claim 14, wherein the plurality of categories of audio content items comprise the following categories:low frequency audio content items;medium frequency audio content items; and high frequency audio content items.
16. The method of claim 15, wherein at least one of the plurality of categories of audio content items is defined as follows:audio content items in the low frequency category have at least 70% of their power spectrum in the range 20 - 800 Hz;audio content items in the medium frequency category have at least 70% of their power spectrum in the range 800 Hz - 2 kHz; andaudio content items in the high frequency category have at least 70% of their power spectrum in the range 2 kHz - 20 kHz.
17. The method of any of claims 14 to 16, comprising calculating, for all combinations of audio content items in each of the categories, adjustments to at least one property of a plurality of audio content items such that a difference between a frequency distribution of a mix of the plurality of audio content items and a desired frequency distribution is reduced; and recording the calculated adjustments in a database.
18. The method of any preceding claim, wherein the plurality of audio content items comprises at least one high frequency audio content item; at least one medium frequency audio content item; and at least one low frequency audio content item.
19. The method of any preceding claim, wherein the difference between the frequency distribution of the mix of the plurality of audio content items and the desired frequency distribution is minimised.
20. The method of any preceding claim, wherein the plurality of audio content items comprises at least one audio content item provided by the user.
21. The method of any preceding claim, wherein the desired frequency distribution is 1 / f noise.
22. The method of any of claims 1 to 21, wherein the desired frequency distribution has equal power throughout its bandwidth.
23. The method of any preceding claim, wherein the desired frequency distribution has therapeutic properties.
24. The method of any preceding claim, comprising playing the mix of audio content items, after the at least one property has been adjusted, to a user, preferably further comprising converting the mix of audio content items into a spatial reproduction format, more preferably an Ambisonics format.
25. A device comprising computer-executable instructions for executing the method of any of claims 1-24.
26. A non-transitory computer-readable storage medium comprising instructions that, when executed, perform the method of any of claims 1-24.T +44(0)30 0300 2000A
Citation Information
Patent Citations
White noise output system and white noise output method
KR101865449B1
Method and apparatus for environmental setting and information for environmental setting
US20100204540A1
Sound masking in open-plan spaces using natural sounds
US20180122353A1
Audio Conflict Resolution
US20210208841A1
Extraction and classification of audio events in gaming systems
US20230064627A1