Device and method for selecting audio effects for processing an audio signal
The device and method automate the selection of audio effects by calculating distances between audio file components and ambiance characteristics, addressing the challenge of diverse audio content and providing a personalized listening experience.
Patent Information
- Application Number
- FR2024000174
- Authority / Receiving Office
- FR · FR
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-01-09
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2044-01-09
AI Technical Summary
The selection of suitable audio effects is difficult due to the diversity of audio content, making it unlikely for users to perform audio processing on a case-by-case basis, especially when using simple smartphones to access digital distribution platforms.
A device and method for selecting audio effects that utilize a memory to store data sets associated with audio ambiances, each containing reference values, weighting coefficients, and audio effects, and a calculator to determine the smallest distance between an audio file's values and ambiances, thereby selecting appropriate effects based on these calculations.
Automates the selection of audio effects, ensuring a personalized listening experience without requiring users to manually choose effects, by identifying the most suitable audio ambiance and applying corresponding effects to the audio signal.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
Title of the invention: Device and method for selecting audio effects for processing an audio signal
[0001] The field of the invention relates to audio processing, and more precisely to the addition of audio effects to an audio signal.
[0002] Audio effects refer to various techniques used in audio processing that aim to alter or improve the characteristics of an audio signal. Such techniques are usually carried out with electronic devices - themselves called "audio effects" or more simply "effects" - which allow the audio signal to be manipulated, and more precisely to modify its dynamics, temporality or frequency response. Audio effects find application in various contexts, including video games, film post-production and, of course, music.
[0003] Dynamic effects include, among others, the de-esser, which limits certain sibilant frequencies in a recorded voice, the compressor, which reduces dynamic differences by attenuating the gain when the sound level exceeds a certain threshold, or the Y expander, which increases dynamic differences by attenuating the gain when the sound level is below a certain threshold.
[0004] Temporal effects include, among others, delay, which reproduces a sound signal with a delay so as to simulate an echo, or reverberation (sometimes abbreviated as "reverb"), which creates a persistence of sound similar to that of a concert hall and thus gives an impression of depth.
[0005] Equalization (commonly referred to as "EQ") consists of modifying the frequency response of an audio signal, i.e., attenuating or amplifying different frequency bands. Equalization allows you to act on the harmonic content of an audio signal, i.e., all the frequency peaks around the fundamental frequency of a sound, to remove unwanted high frequencies or accentuate the bass, for example. It is possible, thanks to equalization, to restore a tonal balance by preventing one frequency component from dominating the others.
[0006] Audio effects make it possible to enrich the listening experience of music or audio program on demand (“podcast” in English). Such audio processing is all the more interesting today since the recent development of digital distribution platforms, such as Spotify (registered trademark), Deezer (registered trademark) or Napster (registered trademark), now makes it easy to access audio files for downloading or streaming.
[0007] However, the selection of suitable audio effects is made difficult by the diversity of audio content and it is quite unlikely that the audio processing will be carried out on a case-by-case basis by the user, particularly when the user uses a simple smartphone to access a digital distribution platform.
[0008] There is therefore a need to automate audio processing without requiring users to do more than just choose the audio content they wish to listen to.
[0009] The present invention improves the situation.
[0010] In this respect, the invention relates to a device for selecting audio effects for processing an audio signal comprising: - a memory arranged to store data sets each associated with a respective audio ambiance, each data set comprising a vector of reference values each corresponding to a respective audio component, a vector of weighting coefficients each corresponding to a respective audio component, and audio effects, and - a calculator arranged to receive a vector of values associated with an audio file, each value corresponding to a respective audio component, to calculate a respective distance between the audio file and each audio ambiance, which distance is calculated as a function of respective differences between a value associated with the audio file and the reference value associated with the audio ambiance considered corresponding to the same audio component, each difference being weighted by the weighting coefficient associated with the audio ambiance considered corresponding to the same audio component, and to select, as audio effects to be applied to an audio signal taken from the audio file, the audio effects of the data set associated with the audio ambiance for which the distance is the smallest.
[0011] In one or more embodiments, the calculator is arranged to calculate the distance between the audio file and an audio ambience as follows: 1) -
[0012] where: O i is an index corresponding to an audio ambiance, O 7 is the vector of values associated with the audio file, OJ is the number of audio components, O wij is the reference value corresponding to the j-th audio component of the audio environment i, O is the weighting coefficient corresponding to the j-th audio component of the audio environment i, and O 0 is the value corresponding to the j-th audio component of the value vector T.
[0013] In one or more embodiments, the sum of the weighting coefficients of an audio ambiance is equal to 1.
[0014] In one or more embodiments, the vector of reference values and the vector of weighting coefficients of a given audio ambiance are constructed from a plurality of vectors of values respectively associated with audio files labeled with the given audio ambiance, each reference value corresponding to an audio component depending on an average of the values associated with the audio files each corresponding to the audio component, each weighting coefficient corresponding to an audio component depending on the inverse of the average deviation of the values associated with the audio files each corresponding to the audio component.
[0015] Alternatively, the vector of reference values and the vector of weighting coefficients of a given audio ambiance are constructed by supervised learning from training data comprising a plurality of vectors of values respectively associated with audio files labeled with the given audio ambiance.
[0016] In one or more embodiments, the audio components are selected from: danceability, energy, loudness, speech, acoustics, instrumentality, livability, valence, tempo and duration.
[0017] The invention also relates to a method for selecting audio effects for processing an audio signal implemented by computer means and comprising the following operations: - storing data sets each associated with a respective audio ambiance, each data set comprising a vector of reference values each corresponding to a respective audio component, a vector of weighting coefficients each corresponding to a respective audio component, and audio effects, - receive a vector of values associated with an audio file, each value corresponding to a respective audio component, - calculate a respective distance between the audio file and each audio ambiance, which distance is calculated based on respective differences between a value associated with the audio file and the reference value associated with the audio ambiance considered corresponding to the same audio component, each difference being weighted by the weighting coefficient associated with the audio ambiance considered corresponding to the same audio component, and - select, as audio effects to be applied to an audio signal taken from the audio file, the audio effects from the dataset associated with the audio ambiance for which the distance is the smallest.
[0018] The invention also relates to an audio processing system comprising: - a device as described previously arranged to select audio effects to be applied to an audio signal taken from an audio file, and - a processing unit arranged to apply the selected audio effects to the audio signal.
[0019] The invention further relates to an audio processing method implemented by computer means and comprising the following operations: - selecting audio effects to be applied to an audio signal taken from an audio file by implementing the audio effects selection method described above, - apply the selected audio effects to the audio signal.
[0020] Finally, the invention relates to a computer program comprising instructions whose execution, by at least one processor, results in the implementation of the audio effects selection method and / or audio processing method described previously.
[0021] Other characteristics, details and advantages will appear on reading the detailed description below, and on analyzing the attached drawings in which:
[0022] [Fig.l] illustrates a sound diffusion system;
[0023] [Fig.2] illustrates an audio processing system according to the invention;
[0024] [Fig.3] illustrates a method of selecting audio effects according to the invention; and
[0025] [Fig.4] illustrates an audio processing method according to the invention.
[0026] [Fig.l] illustrates an environment 100 within which a sound diffusion system 110 is installed.
[0027] The sound broadcasting system 110 is arranged to broadcast audio content, for example music or a podcast, to one or more individuals present in the environment 100.
[0028] In the example of [Fig.l], the environment 100 is a car and the sound diffusion system 110 makes it possible to diffuse audio content to the passengers present in the passenger compartment.
[0029] The environment 100 may be a vehicle other than a car. It should also be noted that the environment 100 is not necessarily a vehicle and may, for example, be a home audio system, while the sound system 110 may include an audio headset. Generally, the sound system 110 may refer to any type of audio system comprising at least one loudspeaker, and the environment 100 may refer to any environment in which such an audio system may be installed.
[0030] The sound diffusion system 110 comprises at least one loudspeaker 112 and an audio processing system 114.
[0031] The loudspeaker 112 is arranged to produce an audio signal from an electrical signal. More particularly here, the loudspeaker 112 is arranged to convert an audio signal processed by the audio processing system 114 into an audio signal.
[0032] In the example of [Fig.l], the sound diffusion system 110 comprises four speakers 112. However, the sound diffusion system 110 may also comprise only a single speaker 112.
[0033] The audio processing system 114 is arranged to perform audio processing, i.e. to add audio effects to an audio signal. Such audio processing makes it possible to modify the dynamics, temporality and / or frequency response of the audio signal. To do this, the audio processing system 114 can use various audio effects such as dynamic effects (de-esser, compressor, expander, etc.), temporal effects (delay, reverb, etc.) and / or filters (equalization, etc.).
[0034] Furthermore, the audio processing system 114 is arranged to transmit the processed audio signal to each loudspeaker 112 for the purpose of broadcasting a sound signal in the environment 100.
[0035] To do this, the audio processing system 114 can communicate with each speaker 112 using wired technology. Alternatively, the audio processing system 114 can communicate with each speaker 112 using short-range wireless communication technology. Among the wireless technologies, near-field communication (better known by the English acronym NFC for “near-field communication”), ZigBee (registered trademark), Bluetooth (registered trademark) or Wi-Fi can for example be envisaged.
[0036] [Fig.2] schematically illustrates the audio processing system 114.
[0037] The audio processing system 114 comprises an audio source 200, an audio effects selection device 210 and a processing unit 220.
[0038] The audio source 200 is arranged to store audio files. Each audio file contains the audio data necessary for producing an audio signal. Furthermore, each audio file is accompanied by metadata. The set formed by the audio data and the metadata can be referred to as an audio track.
[0039] Typically, the audio source 200 has the necessary means to connect to a digital distribution platform and access the catalog offered by the latter to retrieve audio files and the associated metadata. For example, the audio source 200 is capable of connecting to a digital distribution platform via a wide area network (WAN), for example the Internet. The communication between the audio source 200 and the digital distribution platform can be part of a client-server or peer-to-peer architecture (P2P).
[0040] In particular, the audio source 200 can connect to the digital distribution platform by means of a player created using a software development kit (also known by the English acronym SDK for “software development kit”), which can take the form of an application programming interface (also known by the English acronym API for “application programming interface” made available by the digital distribution platform.
[0041] The audio source 200 is configured to retrieve the audio files and the associated metadata upon request from a user, who can consult the catalog of the digital distribution platform via a website, software or an application on a user terminal such as a smartphone.
[0042] The digital distribution platform is for example Spotify (registered trademark), Amazon Music (registered trademark), Apple Music (registered trademark), Tidal (registered trademark), Napster (registered trademark), YouTube Music or Deezer (registered trademark).
[0043] The audio source 200 is arranged to communicate with the device 210 and with the processing unit 220. More particularly, the audio source 200 is arranged to transmit the audio data of an audio file to the processing unit 220 and to transmit the metadata associated with the audio file to the device 210.
[0044] In the context of the invention, the metadata associated with an audio file comprises at least one vector of values. Each value corresponds to a respective audio component. In a paradigm in which the set of audio contents is assimilated to a vector space, the family of audio components can be seen as a basis of such a vector space.
[0045] For example, Spotify (registered trademark) provides a value for each of the following audio features: danceability, energy, loudness, speechiness, acoustics, instrumentality, liveness, valence, tempo, and duration.
[0046] In the remainder of the description, the vector of values is noted ' and the value corresponding to the j-th audio component is noted 0- Finally, it is considered that a number J of audio components are used. Consequently:
[0047] ...
[0048] The device 210 is arranged to select audio effects from the metadata associated with an audio file, and more precisely from the vector of values
[0049] Such audio effects are intended to be added to the audio signal produced from the audio data of the audio file with which the metadata comprising the vector of values f received by the device 210 are associated.
[0050] As illustrated in [Fig.2], the device 210 comprises a memory 212 and a cal- culator 214.
[0051] The memory 212 is arranged to store data sets each associated with a respective audio ambiance.
[0052] In the context of the invention, an audio ambiance is characterized by a vector of reference values, each reference value corresponding to a respective audio component. According to the paradigm described above, the vector of reference values of an audio ambiance corresponds to the representation of this audio ambiance in the vector space of the audio contents.
[0053] The Applicant has defined the following audio ambiances for the purposes of configuring the device 210: dynamic, acoustic, cinema, surround and natural. Of course, other audio ambiances can be defined.
[0054] To determine the vector of reference values for each audio ambiance, it is possible to use the metadata associated with audio files available on one or more digital distribution platforms.
[0055] We consider a given audio ambiance, for example the cinema audio ambiance. The Applicant has identified, in the Spotify (registered trademark) catalog, the following four pieces of music corresponding to this audio ambiance: H. Zimmer, JN Howard. (2012). A Dark Knight (hereinafter ^); H. Jackman. (2011). First Class (hereinafter O); L Williams. (1977). Star Wars (Main Title) (hereinafter / 3); and D. Arnold. (2006). The Name's Bond... James Bond (hereinafter £4).
[0056] The same audio components are used to determine the vector of reference values for each of the audio ambiances. In the case described here, the values recovered for each of the music are those corresponding to the following audio components: danceability, energy, speech, acoustics, instrumentality, livability and valence.
[0057] We extract the following value vectors:
[0058] 11 = [ 0.277; 0.182; 0.034; 0.482; 0.782; 0.125; 0.0357]
[0059] t2 = [0.534; 0.487; 0.0302; 0.74; 0.885; 0.109, 0.237]
[0060] = [0.245; 0.321; 0.0379; 0.865; 0.855; 0.156; 0.142]
[0061] = [ 0.264; 0.57; 0.0338; 0.21; 0.706; 0.0783; 0.105]
[0062] It should be noted that Spotify (registered trademark) provides normalized values, which explains why all the values of the vectors t2, and ^4 are between 0 and 1.
[0063] This information is gathered in the table below: Music A Dark Knight First Class Star Wars (Main Title) The Name 's Bond... James Bond Audio Health Compo Danceability 0.277 0.534 0.245 0.264 Energy 0.182 0.487 0.321 0.57 Speech 0.034 0.0302 0.0379 0.0338 Acoustics 0.482 0.74 0.865 0.21 Instrumentality 0.782 0.885 0.855 0.706 Liveability 0.125 0.109 0.156 0.0783 Valence 0.0357 0.237 0.142 0.105
[0064] The vector of reference values of the cinema audio ambiance can be constructed by calculating an average of the vectors 1], h and ^4. More particularly, the reference value corresponding to an audio component is calculated as a function of the average of the values of the vectors 1[, U, ^3 and / 4 corresponding to this audio component.
[0065] We obtain the following vector of reference values for the cinema audio ambiance (rounding to the nearest 105):
[0066] tref = [0.33000; 0.39000; 0.03398; 0.57425; 0.80700,0.11708; 0.12993]
[0067] In the example given here, the average used is the arithmetic average. However, it is possible to use a geometric average or a harmonic average.
[0068] Such a vector of reference values can be refined by using more music labeled with the cinema audio ambiance.
[0069] By proceeding in this way for the other audio ambiances cited above, the Applicant was thus able to construct the following table in which each column contains the reference values of the corresponding audio ambiance: Audio Ambiances Dynamic Acoustics Natural Surround Cinema Audio components Danceability 0.68480 0.45725 0.33000 0.77400 0.63500 Energy 0.75960 0.11743 0.39000 0.59950 0.81420 Speech 0.07426 0.05710 0.03398 0.19680 0.12084 Acoustics 0.03357 0.90750 0.57425 0.21620 0.08786 Instrumentality 0.10801 0.67450 0.80700 0.00156 0.03888 Liveability 0.19814 0.11280 0.11708 0.91750 0.16792 Valence 0.54080 0.22935 0.12993 0.67400 0.62540
[0070] It is possible to proceed otherwise than by calculating the average. For example, the vector of reference values of a given audio ambiance can be constructed by supervised learning from training data comprising a plurality of vectors of values associated with audio files labeled with this given audio ambiance.
[0071] We can thus define a matrix W = { where / is the number of atmospheres audio and wi.j is the reference value corresponding to the j-th audio component of the audio ambiance i. Each row of the matrix W thus corresponds to the vector of reference values of an audio ambiance.
[0072] In addition to a vector of reference values, each data set associated with an audio ambiance also includes a vector of weighting coefficients, each weighting coefficient corresponding to a respective audio component.
[0073] The weighting coefficients make it possible to prioritize the respective contributions of the audio components for a given audio ambiance.
[0074] Again, to determine the vector of weighting coefficients for each audio ambiance, it is possible to use the metadata associated with audio files available on one or more digital distribution platforms.
[0075] We consider again the cinema audio ambiance and the four musics mentioned above for which the respective value vectors 1], h and ^4 were extracted.
[0076] The vector of weighting coefficients of the cinema audio ambiance can be constructed by calculating this time the inverse of the average deviation (“mean absolute difference” in English) of the vectors t], O, ^3 and ^4. More particularly, the weighting coefficient corresponding to an audio component is calculated as a function of the inverse of the average deviation of the values of the vectors h and £4 corresponding to this audio component.
[0077] We obtain the following vector of weighting coefficients for the cinema audio ambiance (rounding to the nearest 105):
[0078]
[0079]
[0080]
[0081]
[0082]
[0083]
[0084]
[0085] = [9.80392:7.22022; 506.32911; 4.38116:15.87302; 42.68943; 16.78556] Again, such a vector of weighting coefficients can be refined by using more music labeled with the cinema audio ambiance. By proceeding in this way for the other audio ambiances cited above, the Applicant was thus able to construct the following table in which each column contains the weighting coefficients of the corresponding audio ambiance: Audio Ambiances Dynamic Acoustics Natural Surround Cinema Audio Components Danceability 8.51499 13.74570 9.80392 9.80392 8.03859 Energy 15.94388 17.21911 7.22022 9.04977 12.3517 8 Speech 58.16659 30.34901 506.3291 1 6.65779 10.7610 2 Acoustics 29.95734 20.61856 4.38116 6.63130 8.11330 Instrumentality 5.92436 2.96516 15.87302 643.05373 18.6405 5 Viability 13.60766 54.94505 42.68943 2000.00000 9.79777 Valence 6.13648 12.87830 16.78556 20.00000 4.60236 It is possible to proceed in another way than by calculating the inverse of the average deviation. For example, the vector of weighting coefficients of a given audio ambiance can be constructed by supervised learning from training data comprising a plurality of value vectors associated with audio files labeled with this given audio ambiance. We can thus define a matrix , where 7 is the number of atmospheres audio and Pj j is the weighting coefficient corresponding to the j-th audio component of the audio ambiance i. Each row of the matrix B thus corresponds to the vector of weighting coefficients of an audio ambiance. Once calculated, the weighting coefficients of an audio ambience can be normalized as follows: where: . is the normalized weighting coefficient corresponding to the j-th audio component of the audio ambience i.
[0086] It is also possible, for a given audio ambiance, not to take into account one or more audio components by assigning a zero value to the respective weighting coefficient of this or these audio components.
[0087] Finally, in addition to a vector of reference values and a vector of weighting coefficients, each data set associated with an audio ambiance also includes audio effects.
[0088] In other words, for a number / of audio ambiances, a number / of audio effect sets AE^, ■ ■ ■, AE, are stored in the memory 212.
[0089] As explained previously, such audio effects may be dynamic effects, temporal effects and / or filters. The audio effects AE^, ■ ■ ■, AEl may be stored in the memory 212 in the form of data or coefficients suitable for conversion by the processing unit 220 into a specific audio processing. Alternatively, the audio effects AE^, ■ ■ ■, AEj may be stored in the memory 212 in the form of an identifier from which the processing unit 220 can find the corresponding audio effects in a specific storage medium.
[0090] The memory 212 may designate any data storage medium arranged to receive and store digital data, for example a hard disk, a solid-state drive (SSD) or more generally any computer hardware allowing the storage of data on flash memory. The memory 212 may also be a random access memory or a magneto-optical disk. A combination of several storage media may also be envisaged. In the latter case, each data set associated with an audio ambiance may be divided into several sub-data sets distributed between different storage media.
[0091] Typically, the vector of reference values and the vector of weighting coefficients of a data set are stored in a first storage medium while the audio effects of this data set are stored in a second storage medium. The memory 212 then designates the combination of the first and second storage media.
[0092] Furthermore, the memory 212 can also be arranged to store instructions whose execution, by the computer 214, results in the operation of the device 210.
[0093] The computer 214 is arranged to receive a vector of values 1 and to identify, using the data sets stored in the memory 212 and each associated with a respective audio ambiance, the audio ambiance corresponding to the audio signal. The identification of the appropriate audio ambiance allows the computer 214 to select the appropriate audio effects, i.e. the audio effects of the data set associated with the identified audio ambiance.
[0094] The operation of the calculator 214 will be described in more detail below, with reference to [Fig.3].
[0095] The calculator 214 can be produced in any known manner, for example in the form of a microprocessor, a programmable logic circuit (better known by the acronym PLD for “Programmable Logical Device”) or a dedicated chip of the FPGA type (acronym for “Field Programmable Gate Array”) or SoC (acronym for “System on Chip”), a grid of computing resources, a microcontroller or any other specific form having the computing power necessary for the selection of audio effects. One or more of these elements can also be produced in the form of specialized electronic circuits of the ASIC type (acronym for “Application-Specific Integrated Circuit”). A combination of processors and electronic circuits can also be envisaged.
[0096] Finally, the processing unit 220 is arranged to receive, on the one hand, the audio data of an audio file and, on the other hand, the audio effects selected by the device 210 from the vector of values { associated with the audio file, and to produce an audio signal presenting the selected audio effects.
[0097] Equivalently, it can be considered that the audio data of an audio file received by the processing unit 220 form an audio signal and that the processing unit 220 is arranged to add the selected audio effects to this audio signal.
[0098] The processing unit 220 is further arranged to transmit the processed audio signal to each loudspeaker 112 for the purpose of broadcasting a sound signal in the environment 100.
[0099] By "to each loudspeaker 112" is meant here that the output of the processing unit 220 is in fact transmitted to an electronic circuit capable of shaping the processed audio signal to allow each loudspeaker 112 to produce the desired sound signal.
[0100] It should be noted that the core of the invention relates to the selection of suitable audio effects by identifying an audio ambiance. Therefore, such an electronic circuit is not detailed here and is not shown in the drawings. It may nevertheless be noted that, typically, the processed audio signal at the output of the processing unit 220 is a digital audio signal and that the electronic circuit comprises at least one digital-to-analog converter (DAC) arranged to convert the digital audio signal into an analog audio signal and an amplifier arranged to amplify the analog audio signal and transmit it to each loudspeaker 112. The shaping of a digital audio signal - whether it is augmented with audio effects or not - is part of the general knowledge of a person skilled in the art.
[0101] Referring again to [Fig.l], it appears that the audio processing system 114 is part of the sound diffusion system 110, which is integrated into the car. However, it should be understood that some of the entities of the audio processing system 114 shown in [Fig.2] may be remote from the car.
[0102] A method of selecting audio effects implemented by the device 210 will now be described with reference to [Fig.3].
[0103] A context for implementing such an audio effects selection method is typically the following: a passenger of a vehicle, for example a passenger of the car shown in [Fig.l], uses a user terminal to connect to a digital distribution platform. Such a user terminal is for example a smartphone or a touch screen integrated into the vehicle. The user terminal allows the passenger to connect to a digital distribution platform via a website, software or an application. Once connected, the passenger consults the catalog offered by the digital distribution platform and selects the audio content of his choice, for example music or a podcast. The audio file corresponding to the selected audio content is retrieved by the audio source 200 as well as the associated metadata.
[0104] It should be noted that this is only one possible context and that, alternatively, the audio files and associated metadata may already be stored in the memory 212 so that the audio effects selection method can be implemented locally, thus without the need to connect to the Internet and access the digital distribution platform. In such a case, the audio files and associated metadata may have been downloaded beforehand.
[0105] It should further be noted that the car of [Fig.l] is only one example of environment 100 and that the sound diffusion system 110 can designate any audio system comprising at least one loudspeaker.
[0106] During an operation 300, the memory 212 stores data sets each associated with a respective audio ambiance.
[0107] As detailed previously, each dataset includes a vector of reference values, a vector of weighting coefficients and audio effects.
[0108] It should be noted that operation 300, which allows memory 212 to be configured with the data necessary for selecting suitable audio effects for an audio file, does not have to be implemented each time such audio effects are to be selected.
[0109] The memory 212 may be updated regularly, for example to refine the vector of reference values and the vector of weighting coefficients of each of the audio ambiances. Such refinement may be achieved by enriching the memory 212 with the vectors of values associated with newly labeled audio files, i.e. audio files corresponding to audio content for which an appropriate audio ambiance has been chosen.
[0110] During an operation 310, metadata associated with an audio file is received by the device 210, and more precisely by the calculator 214. As explained previously, the relevant information contained in this metadata is in fact the vector of values f, each value corresponding to an audio component.
[0111] In the example of [Fig.2], the vector of values f is provided to the device 210 by the audio source 200.
[0112] As specified above, the vector of values ' is associated with an audio file corresponding to audio content that a passenger of the vehicle wishes to listen to and that he has selected, for example by means of a Human-Machine Interface (HMI) of the user terminal used for this purpose.
[0113] During an operation 320, the calculator 214 identifies the audio ambiance corresponding to the vector of values ', and therefore to the audio file with which the vector of values is associated.
[0114] To do this, the calculator 214 calculates a respective distance between the audio file with which the vector of values f is associated and each audio ambiance. The calculator 214 calculates a distance for each audio ambiance.
[0115] In the example developed above, five possible audio ambiances are mentioned: dynamic, acoustic, cinema, ambiophonic and natural. In such a case, five distances must therefore be calculated.
[0116] The distance between the audio file and an audio ambiance is calculated based on several differences. A difference corresponds to the difference between, on the one hand, a value associated with the audio file - that is to say a value of the vector of values ' associated with the audio file - and therefore corresponding to an audio component and, on the other hand, the reference value associated with the audio ambiance - that is to say a reference value of the vector of reference values of the data set associated with the audio ambiance - and corresponding to the same audio component. Each difference therefore corresponds to an audio component. Furthermore, each difference is weighted by the weighting coefficient associated with the audio ambiance - that is to say a weighting coefficient of the vector of weighting coefficients of the data set associated with the audio ambiance - and corresponding to the same audio component.
[0117] For example, the distance between the audio file and an audio ambiance is calculated as follows by the calculator 214:
[0118] where: O i is an index corresponding to an audio ambiance, O 1 is the vector of values associated with the audio file, OJ is the number of audio components, O wiJ is the reference value corresponding to the j-th audio component of the audio environment i, O Pjj is the weighting coefficient corresponding to the j-th audio component of the audio ambiance !, and O tj is the value corresponding to the j-th audio component of the value vector
[0119] Finally, during an operation 330, the computer 214 selects, in the memory 212, the audio effects of the data set associated with the audio ambiance for which the distance is the smallest.
[0120] Therefore, the set of audio effects selected by the computer 214 to process the audio signal produced from the audio data of the audio file is the set of audio effects AEk corresponding to the audio ambiance k;
[0121] k = argmin,e[y]1d(i,t)
[0122] It may be noted that, on Spotify (registered trademark), the metadata associated with an audio file corresponding to a podcast includes information on this subject. It is possible to define a podcast audio ambiance and to store in the memory 212 a data set associated with the podcast audio ambiance which includes suitable audio effects. In the event of detection, within the metadata of an audio file, of information according to which the audio file in question corresponds to a podcast, the calculator 214 can directly select the audio effects from the data set associated with the podcast audio ambiance. It is therefore not necessary for the calculator 214 to carry out the calculations mentioned above and relating to the vector of values. It is understood that, unlike other audio ambiances, the audio podcast ambiance may not include a vector of reference values or a vector of weighting coefficients.The operation 320 described above can then include a prior sub-operation of searching, in the metadata associated with an audio file, for information according to which the audio file in question corresponds to a podcast in order, where appropriate, to directly select the audio effects of the podcast audio ambiance.
[0123] An audio processing method implemented by the audio processing system 114 will now be described with reference to [Fig.4].
[0124] Such an audio processing method typically falls within the same context as the audio effects selection method described above.
[0125] During an operation 400, the audio source 200 provides, on the one hand, audio data from an audio file to the processing unit 220 and, on the other hand, the vector of values 1 associated with the audio file to the device 210.
[0126] The reception of the vector of values 1 by the device 210 corresponds to the operation 310 of the audio effects selection method described previously. The operation 400 thus corresponds, in addition to the provision of the audio data - or, in an equivalent manner, of the audio signal formed by the audio data - to the processing unit 220, at least to the operations 310, 320 and 330 of the audio effects selection method of [Fig.3].
[0127] At the end of operation 400, the calculator 214 transmits the selected audio effects to the processing unit 220, which then has the audio data and the selected audio effects.
[0128] During an operation 410, the processing unit 220 applies the selected audio effects to the audio signal. The output of the processing unit 220 is therefore a processed audio signal.
[0129] The processed audio signal is then transmitted to each loudspeaker 112 of the sound diffusion system 110. As specified above, it must be understood here that the processed audio signal is received by an electronic circuit capable of shaping the processed audio signal to allow each loudspeaker 112 to produce the desired sound signal and to diffuse it within the environment 100 which, in the example of [Fig.l], is a car.
Claims
Claims
1. Device (210) for selecting audio effects for processing an audio signal comprising: - a memory (212) arranged to store data sets each associated with a respective audio ambiance, each data set comprising a vector of reference values each corresponding to a respective audio component, a vector of weighting coefficients each corresponding to a respective audio component, and audio effects, and - a calculator (214) arranged to receive a vector of values associated with an audio file, each value corresponding to a respective audio component, to calculate a respective distance between said audio file and each audio ambiance, which distance is calculated as a function of respective differences between a value associated with said audio file and the reference value associated with the audio ambiance considered corresponding to the same audio component,each difference being weighted by the weighting coefficient associated with said audio ambiance considered corresponding to said same audio component, and to select, as audio effects to be applied to an audio signal taken from said audio file, the audio effects of the data set associated with the audio ambiance for which the distance is the smallest.,
2. Device (210) according to claim 1, wherein the calculator (214) is arranged to calculate the distance between the audio file and an audio ambiance as follows: where: O i is an index corresponding to an audio ambiance, O 1 is the vector of values associated with the audio file, OJ is the number of audio components, O wiJ is the reference value corresponding to the j-th audio component of the audio ambiance i, O is the weighting coefficient corresponding to the j-th audio component of the audio ambiance i, and O h is the value corresponding to the j-th audio component of the vector of values l.
3. Device (210) according to claim 1 or 2, wherein the sum of the weighting coefficients of an audio ambiance is equal to 1.
4. Device (210) according to one of the preceding claims, in which the vector of reference values and the vector of weighting coefficients of a given audio ambiance are constructed from a plurality of vectors of values respectively associated with audio files labeled with said given audio ambiance, each reference value corresponding to an audio component depending on an average of the values associated with said audio files each corresponding to said audio component, each weighting coefficient corresponding to an audio component depending on the inverse of the average deviation of the values associated with said audio files each corresponding to said audio component.
5. Device (210) according to one of claims 1 to 3, in which the vector of reference values and the vector of weighting coefficients of a given audio ambiance are constructed by supervised learning from training data comprising a plurality of vectors of values associated respectively with audio files labeled with said given audio ambiance.
6. Device (210) according to one of the preceding claims, in which the audio components are chosen from: danceability, energy, loudness, speech, acoustics, instrumentality, livability, valence, tempo and duration.
7. A method for selecting audio effects for processing an audio signal implemented by computer means and comprising the following operations: - storing (300) data sets each associated with a respective audio ambiance, each data set comprising a vector of reference values each corresponding to a respective audio component, a vector of weighting coefficients each corresponding to a respective audio component, and audio effects, - receiving (310) a vector of values associated with an audio file, each value corresponding to a respective audio component, - calculating (320) a respective distance between said audio file and each audio ambiance, which distance is calculated as a function of respective differences between a value associated with said audio file and the reference value associated with the audio ambiance considered corresponding to the same audio component,each difference being weighted by the weighting coefficient associated with said audio ambiance considered corresponding to said same audio component, and, - selecting (330), as audio effects to be applied to an audio signal taken from said audio file, the audio effects from the data set associated with the audio ambiance for which the distance is the smallest.
8. Audio processing system (114) comprising: - a device (210) according to one of claims 1 to 6 arranged to select audio effects to be applied to an audio signal taken from an audio file, and - a processing unit (220) arranged to apply said selected audio effects to said audio signal.
9. Audio processing method implemented by computer means and comprising the following operations: - selecting (400) audio effects to be applied to an audio signal taken from an audio file by implementing the method according to claim 7, - applying (410) said selected audio effects to said audio signal.
10. Computer program comprising instructions whose execution, by at least one processor, results in the implementation of the method according to claim 7 and / or 9.
Citation Information
Patent Citations
Audio similarity detection method and device, medium and computing equipment
CN114464214A
Sound effect determination method and device, equipment and storage medium
CN117093741A