Device and method for selecting audio effects for processing an audio signal
An audio effects selection device calculates distances between audio files and environments to automate the selection of suitable audio effects, addressing the challenge of diverse audio content and enhancing the listening experience.
Patent Information
- Authority / Receiving Office
- FR · FR
- Patent Type
- Patents
- Current Assignee / Owner
- ARKAMYS
- Filing Date
- 2024-01-09
- Publication Date
- 2026-05-22
AI Technical Summary
The selection of suitable audio effects for audio content is difficult due to the diversity of audio content, making it unlikely that users will perform audio processing on a case-by-case basis, especially when using simple devices like smartphones.
An audio effects selection device that stores datasets with reference values and weighting coefficients for different audio environments, calculates distances between audio files and environments, and selects audio effects based on the smallest distance to apply automated audio processing.
Automates the selection of appropriate audio effects, enhancing the listening experience without requiring users to manually choose effects for each audio file.
Smart Images

Figure 00000020_0000 
Figure 00000020_0001 
Figure 00000021_0000
Abstract
Description
Title of the invention: Device and method for selecting audio effects for processing an audio signal
[0001] The field of the invention relates to audio processing, and more specifically to the addition of audio effects to an audio signal.
[0002] Audio effects refer to various techniques used in audio processing that aim to alter or enhance the characteristics of an audio signal. Such techniques are usually implemented with electronic devices—themselves called "audio effects" or simply "effects"—that allow the audio signal to be manipulated, and more specifically, its dynamics, timing, or frequency response to be modified. Audio effects are used in a variety of contexts, including video games, film post-production, and, of course, music.
[0003] Dynamic effects include, among others, the de-esser, which allows certain sibilant frequencies to be limited on a recorded voice, the compressor, which reduces dynamic range by attenuating the gain when the sound level exceeds a certain threshold, or the Y expander, which increases dynamic range by attenuating the gain when the sound level is below a certain threshold.
[0004] Time effects include, among others, delay, which reproduces a sound signal with a delay in order to simulate an echo, or reverberation (sometimes abbreviated "reverb"), which creates a persistence of sound similar to that of a concert hall and thus gives an impression of depth.
[0005] Equalization (commonly denoted as "EQ") consists of modifying the frequency response of an audio signal, that is, attenuating or boosting different frequency bands. Equalization allows one to act on the harmonic content of an audio signal, that is, the set of frequency peaks around the fundamental frequency of a sound, to remove unwanted high frequencies or accentuate the bass, for example. It is possible, through equalization, to restore tonal balance by preventing one frequency component from dominating over the others.
[0006] Audio effects enrich the listening experience of music lovers or podcast listeners. Such audio processing is all the more relevant today given the recent development of digital distribution platforms, such as Spotify (registered trademark), Deezer (registered trademark) or Napster (registered trademark), which now make it easy to access audio files for download or streaming.
[0007] However, the selection of suitable audio effects is made difficult by the diversity of audio content and it is rather unlikely that audio processing will be carried out on a case-by-case basis by the user, especially when the latter uses a simple smartphone to access a digital distribution platform.
[0008] There is therefore a need to automate audio processing without requiring users to do more than choose the audio content they wish to listen to.
[0009] The present invention improves the situation.
[0010] In this respect, the invention relates to an audio effects selection device for processing an audio signal comprising: - a memory arranged to store datasets, each associated with a respective audio environment, each dataset comprising a vector of reference values corresponding to a respective audio component, a vector of weighting coefficients corresponding to a respective audio component, and audio effects, and - a calculator arranged to receive a vector of values associated with an audio file, each value corresponding to a respective audio component, to calculate a respective distance between the audio file and each audio ambience, which distance is calculated based on respective differences between a value associated with the audio file and the reference value associated with the considered audio ambience corresponding to the same audio component, each difference being weighted by the weighting coefficient associated with the considered audio ambience corresponding to the same audio component, and to select, as audio effects to be applied to an audio signal taken from the audio file, the audio effects from the dataset associated with the audio ambience for which the distance is the smallest.
[0011] In one or more embodiments, the calculator is arranged to calculate the distance between the audio file and an audio environment as follows: 1) -
[0012] where: O i is an index corresponding to an audio environment, O7 is the value vector associated with the audio file, OJ is the number of audio components, O wij is the reference value corresponding to the j-th audio component of the audio ambience i, O is the weighting coefficient corresponding to the j-th audio component of the audio ambience i, and O0 is the value corresponding to the j-th audio component of the value vector T.
[0013] In one or more embodiments, the sum of the weighting coefficients the value of an audio ambiance is equal to 1.
[0014] In one or more embodiments, the reference value vector and the weighting coefficient vector of a given audio ambiance are constructed from a plurality of value vectors associated respectively with audio files labeled with the given audio ambiance, each reference value corresponding to an audio component depending on an average of the values associated with the audio files corresponding to each of the audio components, each weighting coefficient corresponding to an audio component depending on the inverse of the average deviation of the values associated with the audio files corresponding to each of the audio components.
[0015] Alternatively, the reference value vector and the weighting coefficient vector of a given audio ambiance are constructed by supervised learning from training data comprising a plurality of value vectors associated respectively with audio files labeled with the given audio ambiance.
[0016] In one or more embodiments, the audio components are chosen from: danceability, energy, loudness, speech, acoustics, instrumentality, livability, valence, tempo and duration.
[0017] The invention also relates to a method for selecting audio effects for processing an audio signal implemented by computer means and comprising the following operations: - to store datasets, each associated with a respective audio environment, each dataset comprising a vector of reference values corresponding to a respective audio component, a vector of weighting coefficients corresponding to a respective audio component, and audio effects, - receive a vector of values associated with an audio file, each value corresponding to a respective audio component, - calculate a respective distance between the audio file and each audio ambience, which distance is calculated based on the respective differences between a value associated with the audio file and the reference value associated with the audio ambience considered, corresponding to the same audio component, each difference being weighted by the weighting coefficient associated with the audio ambience considered, corresponding to the same audio component, and - select, as audio effects to apply to an audio signal taken from the audio file, the audio effects from the dataset associated with the audio environment for which the distance is the smallest.
[0018] The invention also relates to an audio processing system comprising: - a device as described above arranged to select audio effects to be applied to an audio signal taken from an audio file, and - a processing unit arranged to apply the selected audio effects to the audio signal.
[0019] The invention further relates to an audio processing method implemented by computer means and comprising the following operations: - to select audio effects to apply to an audio signal extracted from an audio file by implementing the audio effects selection process described previously, - Apply the selected audio effects to the audio signal.
[0020] Finally, the invention relates to a computer program comprising instructions whose execution, by at least one processor, results in the implementation of the audio effect selection method and / or audio processing method described above.
[0021] Other features, details and advantages will become apparent upon reading the detailed description below, and upon analysis of the accompanying drawings on which:
[0022] [Fig.1] illustrates a sound diffusion system;
[0023] [Fig.2] illustrates an audio processing system according to the invention;
[0024] [Fig. 3] illustrates a method for selecting audio effects according to the invention; and
[0025] [Fig.4] illustrates an audio processing method according to the invention.
[0026] Fig. 1 illustrates an environment 100 in which a sound diffusion system 110 is installed.
[0027] The sound diffusion system 110 is arranged to broadcast audio content, for example music or a podcast, to one or more individuals present in the environment 100.
[0028] In the example of [Fig.1], the environment 100 is a car and the sound diffusion system 110 allows audio content to be broadcast to the passengers present in the passenger compartment.
[0029] Environment 100 may be a vehicle other than a car. It should also be noted that environment 100 is not necessarily a vehicle and may, for example, be a home audio system, while the sound diffusion system 110 may include headphones. Generally speaking, the sound diffusion system 110 may refer to any type of audio system comprising at least one loudspeaker, and environment 100 may refer to any environment in which such an audio system may be installed.
[0030] The sound diffusion system 110 includes at least one loudspeaker 112 and an audio processing system 114.
[0031] The loudspeaker 112 is arranged to produce an audio signal from an electrical signal. More specifically here, the loudspeaker 112 is arranged to convert an audio signal processed by the audio processing system 114 into an audio signal.
[0032] In the example of [Fig. 1], the sound diffusion system 110 comprises four speakers 112. However, the sound diffusion system 110 may also include only a single speaker 112.
[0033] The audio processing system 114 is arranged to perform audio processing, that is, to add audio effects to an audio signal. Such audio processing makes it possible to modify the dynamics, timing, and / or frequency response of the audio signal. To do this, the audio processing system 114 can use various audio effects such as dynamic effects (de-esser, compressor, expander, etc.), time-based effects (delay, reverb, etc.), and / or filters (equalization, etc.).
[0034] Furthermore, the audio processing system 114 is arranged to transmit the processed audio signal to each loudspeaker 112 for the purpose of broadcasting a sound signal in the environment 100.
[0035] To achieve this, the audio processing system 114 can communicate with each speaker 112 using wired technology. Alternatively, the audio processing system 114 can communicate with each speaker 112 using wireless short-range communication technology. Examples of wireless technologies include near-field communication (NFC), ZigBee (registered trademark), Bluetooth (registered trademark), and Wi-Fi.
[0036] Fig. 2 schematically illustrates the audio processing system 114.
[0037] The audio processing system 114 includes an audio source 200, an audio effects selection device 210 and a processing unit 220.
[0038] The audio source 200 is configured to store audio files. Each audio file contains the audio data necessary for the production of an audio signal. In addition, each audio file is accompanied by metadata. The combination of the audio data and metadata can be referred to as an audio track.
[0039] Typically, the audio source 200 has the necessary means to connect to a digital distribution platform and access the catalog offered by that platform in order to retrieve audio files and their associated metadata. For example, the audio source 200 is capable of connecting to a digital distribution platform via a wide area network (WAN), such as the Internet. Communication between the audio source 200 and the digital distribution platform can be based on a client-server or peer-to-peer (P2P) architecture.
[0040] In particular, the audio source 200 can connect to the digital distribution platform by means of a player created using a software development kit (also known by the English acronym SDK for "software development kit"), which can take the form of an application programming interface (also known by the English acronym API for "application programming interface" made available by the digital distribution platform.
[0041] The audio source 200 is configured to retrieve audio files and associated metadata on request from a user, who can consult the digital distribution platform catalogue via a website, software or application on a user terminal such as a smartphone.
[0042] The digital distribution platform is for example Spotify (registered trademark), Amazon Music (registered trademark), Apple Music (registered trademark), Tidal (registered trademark), Napster (registered trademark), YouTube Music or Deezer (registered trademark).
[0043] The audio source 200 is arranged to communicate with the device 210 and with the processing unit 220. More particularly, the audio source 200 is arranged to transmit the audio data of an audio file to the processing unit 220 and to transmit the metadata associated with the audio file to the device 210.
[0044] In the context of the invention, the metadata associated with an audio file comprises at least one vector of values. Each value corresponds to a respective audio component. In a paradigm in which the set of audio content is considered a vector space, the family of audio components can be viewed as a basis of such a vector space.
[0045] By way of example, Spotify (registered trademark) provides a value for each of the following audio features: danceability, energy, loudness, speechiness, acousticness, instrumentalness, liveness, valence, tempo and duration.
[0046] In the following description, the value vector is denoted ' and the value corresponding to the j-th audio component is denoted 0-. Finally, it is assumed that a number J of audio components are used. Therefore:
[0047] ...
[0048] The device 210 is arranged to select audio effects from the metadata associated with an audio file, and more precisely from the value vector
[0049] Such audio effects are intended to be added to the audio signal produced from the audio data of the audio file to which are associated the metadata including the vector of values f received by the device 210.
[0050] As illustrated in [Fig. 2], the device 210 comprises a memory 212 and a cal- 214 cylinder.
[0051] Memory 212 is arranged to store data sets each associated with a respective audio ambiance.
[0052] In the context of the invention, an audio environment is characterized by a vector of reference values, each reference value corresponding to a respective audio component. According to the paradigm described above, the vector of reference values of an audio environment corresponds to the representation of this audio environment in the vector space of audio content.
[0053] The Applicant has defined the following audio environments for the purpose of configuring device 210: dynamic, acoustic, cinema, surround, and natural. Of course, other audio environments can be defined.
[0054] To determine the reference value vector of each audio environment, it is possible to use the metadata associated with audio files available on one or more digital distribution platforms.
[0055] Consider a given audio ambiance, for example, a cinema audio ambiance. The Applicant identified, in the Spotify (registered trademark) catalogue, the following four pieces of music corresponding to this audio ambiance: H. Zimmer, JN Howard. (2012). A Dark Knight (hereinafter ^); H. Jackman. (2011). First Class (hereinafter O); L. Williams. (1977). Star Wars (Main Title) (hereinafter / 3); and D. Arnold. (2006). The Name's Bond... James Bond (hereinafter £4).
[0056] The same audio components are used to determine the reference value vector for each of the audio ambiences. In the case described here, the values retrieved for each of the music pieces are those corresponding to the following audio components: danceability, energy, speech, acoustics, instrumentality, livability, and valence.
[0057] The following value vectors are extracted:
[0058] 11 = [ 0.277; 0.182; 0.034; 0.482; 0.782; 0.125; 0.0357]
[0059] t2 = [0.534; 0.487; 0.0302; 0.74; 0.885; 0.109, 0.237]
[0060] = [0.245; 0.321; 0.0379; 0.865; 0.855; 0.156; 0.142]
[0061] = [ 0.264; 0.57; 0.0338; 0.21; 0.706; 0.0783; 0.105]
[0062] It should be noted that Spotify (registered trademark) provides normalized values, which explains why all the values of the vectors t2, and ^4 are between 0 and 1.
[0063] This information is summarized in the table below: Music: A Dark Knight First Class, Star Wars (Main Title), The Name 's Bond... James Bond. Composition: Audio Health, Danceability: 0.277, 0.534, 0.245, 0.264; Energy: 0.182, 0.487, 0.321, 0.57; Lyrics: 0.034, 0.0302, 0.0379, 0.0338; Acoustics: 0.482, 0.74, 0.865, 0.21; Instrumentality: 0.782, 0.885, 0.855, 0.706; Liveability: 0.125, 0.109, 0.156, 0.0783; Valence: 0.0357, 0.237, 0.142, 0.105
[0064] The reference value vector of the cinema audio ambiance can be constructed by calculating an average of the vectors 1], h and ^4. More particularly, the reference value corresponding to an audio component is calculated as a function of the average of the values of the vectors 1[, U, ^3 and / 4 corresponding to that audio component.
[0065] The following reference value vector is obtained for the cinema audio environment (rounded to the nearest 105):
[0066] tref = [0.33000; 0.39000; 0.03398; 0.57425; 0.80700,0.11708; 0.12993]
[0067] In the example given here, the mean used is the arithmetic mean. However, it is possible to use a geometric mean or a harmonic mean.
[0068] Such a vector of reference values can be refined by using more music tagged with cinema audio ambiance.
[0069] By proceeding in this manner for the other audio environments mentioned above, the Applicant was thus able to construct the following table in which each column contains the reference values of the corresponding audio environment: Dynamic Audio Environments, Acoustics, Natural Surround Sound Audio Components Danceability 0.68480 0.45725 0.33000 0.77400 0.63500 Energy 0.75960 0.11743 0.39000 0.59950 0.81420 Speech 0.07426 0.05710 0.03398 0.19680 0.12084 Acoustics 0.03357 0.90750 0.57425 0.21620 0.08786 Instrumentality 0.10801 0.67450 0.80700 0.00156 0.03888 Liveability 0.19814 0.11280 0.11708 0.91750 0.16792 Valence 0.54080 0.22935 0.12993 0.67400 0.62540
[0070] It is possible to proceed in a manner other than calculating the average. For example, the reference value vector of a given audio ambiance can be constructed by supervised learning from training data comprising a plurality of value vectors associated with audio files labeled with that given audio ambiance.
[0071] We can thus define a matrix W = { where / is the number of atmospheres audio and wi.j is the reference value corresponding to the j-th audio component of the audio ambience i. Each row of the matrix W thus corresponds to the vector of reference values of an audio ambience.
[0072] In addition to a vector of reference values, each dataset associated with an audio ambiance also includes a vector of weighting coefficients, each weighting coefficient corresponding to a respective audio component.
[0073] The weighting coefficients allow the respective contributions of the audio components to be ranked for a given audio environment.
[0074] Here again, to determine the weighting coefficient vector of each audio ambiance, it is possible to use the metadata associated with audio files available on one or more digital distribution platforms.
[0075] We again consider the cinema audio ambiance and the four music pieces mentioned above for which the vectors of respective values 1], h and ^4 have been extracted.
[0076] The vector of weighting coefficients for the cinema audio ambiance can be constructed by calculating the inverse of the mean absolute difference of the vectors t], O, ^3 and ^4. More specifically, the weighting coefficient corresponding to an audio component is calculated as a function of the inverse of the mean difference of the values of the vectors h and £4 corresponding to that audio component.
[0077] The following weighting coefficient vector is obtained for the cinema audio environment (rounded to the nearest 105):
[0078]
[0079]
[0080]
[0081]
[0082]
[0083]
[0084]
[0085] = [9,80392:7,22022 ; 506,32911 ; 4,38116 :15,87302 ; 42,68943 ; 16,78556] Again, such a vector of weighting coefficients can be refined by using more music tagged with cinema audio ambiance. By proceeding in this manner for the other audio environments mentioned above, the Applicant was thus able to construct the following table in which each column contains the weighting coefficients of the corresponding audio environment: Audio Ambiences Dynamic Acoustics Natural Surround Cinema Audio Components Danceability 8.51499 13.74570 9.80392 9.80392 8.03859 Energy 15.94388 17.21911 7.22022 9.04977 12.3517 8 Speech 58.16659 30.34901 506.3291 1 6.65779 10.7610 2 Acoustics 29.95734 20.61856 4.38116 6.63130 8.11330 Instrumentality 5.92436 2.96516 15.87302 643.05373 18.6405 5 Livableity 13.60766 54.94505 42.68943 2000.00000 9.79777 Valence 6.13648 12.87830 16.78556 20.00000 4.60236 It is possible to proceed in a different way than by calculating the inverse of the mean deviation. For example, the weighting coefficient vector of a given audio ambiance can be constructed by supervised learning from training data comprising a plurality of value vectors associated with audio files tagged with that given audio ambiance. We can thus define a matrix, where 7 is the number of atmospheres audio and Pj j is the weighting coefficient corresponding to the j-th audio component of the audio ambience i. Each row of the matrix B thus corresponds to the vector of weighting coefficients of an audio ambience. Once calculated, the weighting coefficients of an audio environment can be normalized as follows: where: . is the normalized weighting coefficient corresponding to the jth audio component of the audio ambience i.
[0086] It is also possible, for a given audio environment, to disregard one or more audio components by assigning a zero value to the respective weighting coefficient of this or these audio components.
[0087] Finally, in addition to a vector of reference values and a vector of weighting coefficients, each dataset associated with an audio ambiance also includes audio effects.
[0088] In other words, for a number / of audio ambiences, a number / of audio effect sets AE^, ■ ■ ■, AE, are stored in memory 212.
[0089] As explained previously, such audio effects can be dynamic effects, time-domain effects, and / or filters. The audio effects AE^, ■ ■ ■, AE1 can be stored in memory 212 as data or coefficients suitable for conversion by the processing unit 220 into a specific audio processing operation. Alternatively, the audio effects AE^, ■ ■ ■, AE1 can be stored in memory 212 as an identifier from which the processing unit 220 can retrieve the corresponding audio effects in a specific storage medium.
[0090] Memory 212 can refer to any data storage medium designed to receive and retain digital data, for example, a hard drive, a solid-state drive (SSD), or more generally, any computer hardware that allows data storage on flash memory. Memory 212 can also be random access memory (RAM) or a magneto-optical disk. A combination of several storage media can also be considered. In this latter case, each dataset associated with an audio environment can be divided into several subsets distributed across different storage media.
[0091] Typically, the reference value vector and the weighting coefficient vector of a dataset are stored in a first storage medium, while the audio effects of this dataset are stored in a second storage medium. Memory 212 then designates the combination of the first and second storage media.
[0092] Furthermore, the memory 212 can also be arranged to store instructions whose execution by the computer 214 results in the operation of the device 210.
[0093] The computer 214 is configured to receive a vector of values 1 and to identify, using the data sets stored in memory 212 and each associated with a respective audio ambiance, the audio ambiance corresponding to the audio signal. Identifying the appropriate audio ambiance allows the computer 214 to select the suitable audio effects, i.e., the audio effects from the data set associated with the identified audio ambiance.
[0094] The operation of the calculator 214 will be described in more detail below, with reference to [Fig.3].
[0095] The computer 214 can be implemented in any known form, for example, as a microprocessor, a programmable logic device (PLD), a dedicated chip such as a field-programmable gate array (FPGA) or system-on-chip (SoC), a computing resource grid, a microcontroller, or any other proprietary form with sufficient computing power for selecting audio effects. One or more of these elements can also be implemented as specialized electronic circuits such as application-specific integrated circuits (ASICs). A combination of processors and electronic circuits is also possible.
[0096] Finally, the processing unit 220 is arranged to receive, on the one hand, the audio data from an audio file and, on the other hand, the audio effects selected by the device 210 from the value vector { associated with the audio file, and to produce an audio signal presenting the selected audio effects.
[0097] Equivalently, the audio data of an audio file received by the processing unit 220 can be considered to form an audio signal and the processing unit 220 is arranged to add the selected audio effects to this audio signal.
[0098] The processing unit 220 is further arranged to transmit the processed audio signal to each loudspeaker 112 for the purpose of broadcasting a sound signal in the environment 100.
[0099] By "to each loudspeaker 112", it is meant here that the output of the processing unit 220 is in fact transmitted to an electronic circuit capable of shaping the processed audio signal to allow each loudspeaker 112 to produce the desired sound signal.
[0100] It should be noted that the core of the invention relates to the selection of suitable audio effects by identifying an audio ambiance. Consequently, such an electronic circuit is not detailed here and is not shown in the drawings. It may nevertheless be noted that, typically, the processed audio signal output from the processing unit 220 is a digital audio signal and that the electronic circuit includes at least one digital-to-analog converter (DAC) arranged to convert the digital audio signal into an analog audio signal and an amplifier arranged to amplify the analog audio signal and transmit it to each loudspeaker 112. The shaping of a digital audio signal—whether or not it is augmented with audio effects—is part of the general knowledge of a person skilled in the art.
[0101] With further reference to [Fig. 1], it appears that the audio processing system 114 is part of the sound distribution system 110, which is integrated into the car. However, it should be understood that some of the components of the audio processing system 114 shown in [Fig. 2] may be located outside the car.
[0102] A method for selecting audio effects implemented by device 210 will now be described with reference to [Fig.3].
[0103] A typical context for implementing such an audio effects selection method is as follows: a passenger in a vehicle, for example, a passenger in the car shown in [Fig. 1], uses a user terminal to connect to a digital distribution platform. Such a user terminal is, for example, a smartphone or a touchscreen integrated into the vehicle. The user terminal allows the passenger to connect to a digital distribution platform via a website, software, or an application. Once connected, the passenger consults the catalog offered by the digital distribution platform and selects the audio content of their choice, for example, music or a podcast. The audio file corresponding to the selected audio content is retrieved by the audio source 200, along with the associated metadata.
[0104] It should be noted that this is only one possible context and that, alternatively, the audio files and associated metadata may already be stored in memory 212 so that the audio effects selection process can be implemented locally, thus without the need to connect to the Internet and access the digital distribution platform. In such a case, the audio files and associated metadata may have been downloaded beforehand.
[0105] It should also be noted that the car in [Fig.1] is only one example of environment 100 and that the sound diffusion system 110 can refer to any audio system comprising at least one loudspeaker.
[0106] During an operation 300, memory 212 stores data sets each associated with a respective audio ambiance.
[0107] As detailed previously, each dataset includes a reference value vector, a weighting coefficient vector, and audio effects.
[0108] It should be noted that operation 300, which allows memory 212 to be configured with the data necessary for selecting suitable audio effects for an audio file, does not have to be implemented every time such audio effects are to be selected.
[0109] Memory 212 can be updated regularly, for example to refine the reference value vector and the weighting coefficient vector of each of the audio ambiences. Such refinement can be achieved by enriching memory 212 with the value vectors associated with newly labeled audio files, that is to say, audio files corresponding respectively to audio content for which an appropriate audio ambiance has been chosen.
[0110] During an operation 310, metadata associated with an audio file is received by the device 210, and more precisely by the computer 214. As explained previously, the relevant information contained in this metadata is in fact the vector of values f, each value corresponding to an audio component.
[0111] In the example of [Fig.2], the vector of values f is supplied to the device 210 by the audio source 200.
[0112] As specified above, the value vector ' is associated with an audio file corresponding to audio content that a passenger in the vehicle wishes to listen to and that he has selected, for example by means of a Human-Machine Interface (HMI) of the user terminal used for this purpose.
[0113] During an operation 320, the calculator 214 identifies the audio ambiance corresponding to the value vector ', and therefore to the audio file to which the value vector is associated.
[0114] To do this, the calculator 214 calculates a respective distance between the audio file to which the vector of values f is associated and each audio ambience. The calculator 214 calculates a distance for each audio ambience.
[0115] In the example developed above, five possible audio environments are mentioned: dynamic, acoustic, cinema, surround sound, and natural. In such a case, five distances must therefore be calculated.
[0116] The distance between the audio file and an audio environment is calculated based on several differences. A difference corresponds to the difference between, on the one hand, a value associated with the audio file—that is, a value in the value vector associated with the audio file—and therefore corresponding to an audio component, and, on the other hand, the reference value associated with the audio environment—that is, a reference value in the reference value vector of the dataset associated with the audio environment—and corresponding to the same audio component. Each difference thus corresponds to an audio component. Furthermore, each difference is weighted by the weighting coefficient associated with the audio environment—that is, a weighting coefficient in the weighting coefficient vector of the dataset associated with the audio environment—and corresponding to the same audio component.
[0117] For example, the distance between the audio file and an audio ambience is calculated as follows by calculator 214:
[0118] where: O i is an index corresponding to an audio ambiance, O 1 is the vector of values associated with the audio file, OJ is the number of audio components, O wiJ is the reference value corresponding to the j-th audio component of the audio ambience i, O Pjj is the weighting coefficient corresponding to the j-th audio component of the audio ambience!, and O tj is the value corresponding to the j-th audio component of the value vector
[0119] Finally, during an operation 330, the computer 214 selects, in memory 212, the audio effects from the dataset associated with the audio ambiance for which the distance is the smallest.
[0120] Consequently, the audio effects set selected by the calculator 214 to process the audio signal produced from the audio data of the audio file is the audio effects set AEk corresponding to the audio ambiance k;
[0121] k = argmin,e[y]1d(i,t)
[0122] It can be noted that, on Spotify (registered trademark), the metadata associated with an audio file corresponding to a podcast includes information about this. It is possible to define a podcast audio mood and store in memory 212 a dataset associated with the podcast audio mood that includes suitable audio effects. If information is detected within the metadata of an audio file indicating that the audio file in question corresponds to a podcast, the calculator 214 can directly select the audio effects from the dataset associated with the podcast audio mood. Consequently, it is not necessary for the calculator 214 to perform the calculations mentioned above relating to the value vector. It is understood that, unlike other audio moods, the podcast audio mood may not include a reference value vector or a weighting coefficient vector.Operation 320 described above can then include a preliminary sub-operation of searching, in the metadata associated with an audio file, for information indicating that the audio file in question corresponds to a podcast in order to, if necessary, directly select the audio effects of the podcast audio ambiance.
[0123] An audio processing method implemented by the audio processing system 114 will now be described with reference to [Fig.4].
[0124] Such an audio processing method typically falls within the same context as the audio effects selection method described above.
[0125] During an operation 400, the audio source 200 provides, on the one hand, audio data from an audio file to the processing unit 220 and, on the other hand, the vector of values 1 associated with the audio file to the device 210.
[0126] The reception of the value vector 1 by the device 210 corresponds to operation 310 of the audio effects selection process described above. Operation 400 This corresponds, in addition to the provision of audio data - or, equivalently, of the audio signal formed by the audio data - to the processing unit 220, at least to operations 310, 320 and 330 of the audio effect selection process of [Fig.3].
[0127] At the end of operation 400, the computer 214 transmits the selected audio effects to the processing unit 220, which then has the audio data and the selected audio effects.
[0128] During operation 410, the processing unit 220 applies the selected audio effects to the audio signal. The output of the processing unit 220 is therefore a processed audio signal.
[0129] The processed audio signal is then transmitted to each loudspeaker 112 of the sound diffusion system 110. As mentioned above, it should be understood here that the processed audio signal is received by an electronic circuit capable of shaping the processed audio signal to allow each loudspeaker 112 to produce the desired sound signal and to diffuse it within the environment 100 which, in the example of [Fig.1], is a car.
Claims
Demands
1. Audio effect selection device (210) for processing an audio signal comprising: - a memory (212) arranged to store datasets each associated with a respective audio ambience, each dataset comprising a vector of reference values each corresponding to a respective audio component, a vector of weighting coefficients each corresponding to a respective audio component, and audio effects, and - a calculator (214) arranged to receive a vector of values associated with an audio file, each value corresponding to a respective audio component, for calculating a respective distance between said audio file and each audio ambience, which distance is calculated based on respective differences between a value associated with said audio file and the reference value associated with the considered audio ambience corresponding to the same audio component,each difference being weighted by the weighting coefficient associated with said audio environment in question corresponding to said same audio component, and to select, as audio effects to be applied to an audio signal taken from said audio file, the audio effects from the dataset associated with the audio environment for which the distance is the smallest.
2. Device (210) according to claim 1, wherein the calculator (214) is arranged to calculate the distance between the audio file and an audio ambience as follows: where: O i is an index corresponding to an audio ambience, O 1 is the vector of values associated with the audio file, OJ is the number of audio components, O wiJ is the reference value corresponding to the j-th audio component of the audio ambience i, O is the weighting coefficient corresponding to the j-th audio component of the audio ambience i, and Oh is the value corresponding to the j-th audio component of the vector of values l.
3. Device (210) according to claim 1 or 2, wherein the sum of The weighting coefficient of an audio environment is equal to 1.
4. Device (210) according to any one of the preceding claims, wherein the reference value vector and the weighting coefficient vector of a given audio ambiance are constructed from a plurality of value vectors associated respectively with audio files tagged with said given audio ambiance, each reference value corresponding to an audio component depending on an average of the values associated with said audio files, each corresponding to said audio component, each weighting coefficient corresponding to an audio component depending on the inverse of the average deviation of the values associated with said audio files, each corresponding to said audio component.
5. Device (210) according to any one of claims 1 to 3, wherein the reference value vector and the weighting coefficient vector of a given audio ambiance are constructed by supervised learning from training data comprising a plurality of value vectors associated respectively with audio files tagged with said given audio ambiance.
6. Device (210) according to any one of the preceding claims, wherein the audio components are selected from: danceability, energy, loudness, speech, acoustics, instrumentality, livability, valence, tempo and duration.
7. A method for selecting audio effects for processing an audio signal implemented by computer means and comprising the following operations: - storing (300) datasets each associated with a respective audio ambience, each dataset comprising a vector of reference values each corresponding to a respective audio component, a vector of weighting coefficients each corresponding to a respective audio component, and audio effects, - receiving (310) a vector of values associated with an audio file, each value corresponding to a respective audio component, - calculating (320) a respective distance between said audio file and each audio ambience, which distance is calculated as a function of respective differences between a value associated with said audio file and the reference value associated with the considered audio ambience corresponding to the same audio component,each difference being weighted by the weighting coefficient associated with said audio environment in question corresponding to said same audio component, and, - select (330), as audio effects to apply to an audio signal taken from said audio file, the audio effects from the dataset associated with the audio ambience for which the distance is the smallest.
8. Audio processing system (114) comprising: - a device (210) according to any one of claims 1 to 6 arranged to select audio effects to be applied to an audio signal taken from an audio file, and - a processing unit (220) arranged to apply said selected audio effects to said audio signal.
9. Audio processing method implemented by computer means and comprising the following operations: - selecting (400) audio effects to be applied to an audio signal taken from an audio file by implementing the method according to claim 7, - applying (410) said selected audio effects to said audio signal.
10. Computer program comprising instructions the execution of which, by at least one processor, results in the implementation of the method according to claim 7 and / or 9.