Device and method for selecting audio effects for processing an audio signal
The device and method automate the selection of audio effects by calculating distances between audio file metadata and ambiance data, ensuring appropriate processing is applied, thus enhancing the listening experience without user intervention.
Patent Information
- Application Number
- PCT/FR2025/050021
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-09
- Filing Date
- 2025-01-08
- Publication Date
- 2025-07-17
AI Technical Summary
The selection of suitable audio effects for audio processing is difficult due to the diversity of audio content, making it unlikely that users will perform audio processing on a case-by-case basis, especially when using simple devices like smartphones to access digital distribution platforms.
A device and method for selecting audio effects that utilize a memory to store data sets associated with audio ambiances, each containing reference values, weighting coefficients, and audio effects, and a calculator to determine the smallest distance between an audio file's values and ambiance data, automatically selecting appropriate effects based on metadata from digital platforms.
Automates the selection of audio effects, enhancing the listening experience by applying the most suitable audio processing without requiring users to manually choose effects for each audio content.
Smart Images

Figure FR2025050021_17072025_PF_FP_ABST
Abstract
Description
[0001] Description
[0002] Title: Device and method for selecting audio effects for processing an audio signal
[0003] The field of the invention relates to audio processing, and more precisely to the addition of audio effects to an audio signal.
[0004] Audio effects refer to various techniques used in audio processing that aim to alter or enhance the characteristics of an audio signal. Such techniques are usually achieved with electronic devices—themselves called "audio effects" or simply "effects"—that allow the audio signal to be manipulated, and more specifically, to change its dynamics, temporality, or frequency response. Audio effects find application in a variety of contexts, including video games, film postproduction, and, of course, music.
[0005] Dynamic effects include, among others, the de-esser, which limits certain sibilant frequencies in a recorded voice, the compressor, which reduces dynamic differences by attenuating the gain when the sound level exceeds a certain threshold, or the expander, which increases dynamic differences by attenuating the gain when the sound level is below a certain threshold.
[0006] Temporal effects include, among others, delay, which reproduces a sound signal with a delay in order to simulate an echo, or reverberation (sometimes abbreviated as "reverb"), which creates a persistence of sound similar to that of a concert hall and thus gives an impression of depth.
[0007] Equalization (commonly referred to as "EQ") involves modifying the frequency response of an audio signal, i.e., attenuating or amplifying different frequency bands. Equalization allows you to act on the harmonic content of an audio signal, i.e., all the frequency peaks around the fundamental frequency of a sound, to remove unwanted high frequencies or accentuate the bass, for example. Equalization can restore tonal balance by preventing one frequency component from dominating the others.
[0008] Audio effects can enrich the listening experience of music or podcast lovers. Such audio processing is all the more interesting today given the recent development of digital distribution platforms, such as Spotify (registered trademark), Deezer (registered trademark) or Napster (registered trademark), which now make it easy to access audio files for downloading or streaming.
[0009] However, the selection of suitable audio effects is made difficult by the diversity of audio content and it is quite unlikely that audio processing will be carried out on a case-by-case basis by the user, particularly when the user uses a simple smartphone to access a digital distribution platform.
[0010] There is therefore a need to automate audio processing without requiring users to do more than just choose the audio content they want to listen to.
[0011] The present invention improves the situation.
[0012] In this respect, the invention relates to a device for selecting audio effects for processing an audio signal comprising:
[0013] - a memory arranged to store data sets each associated with a respective audio ambiance, each data set comprising a vector of reference values each corresponding to a respective audio component, a vector of weighting coefficients each corresponding to a respective audio component, and audio effects, and
[0014] - a calculator arranged to receive a vector of values associated with an audio file, each value corresponding to a respective audio component, to calculate a respective distance between the audio file and each audio ambiance, which distance is calculated as a function of respective differences between a value associated with the audio file and the reference value associated with the audio ambiance considered corresponding to the same audio component, each difference being weighted by the weighting coefficient associated with the audio ambiance considered corresponding to the same audio component, and to select, as audio effects to be applied to an audio signal taken from the audio file, the audio effects of the data set associated with the audio ambiance for which the distance is the smallest.
[0015] In one or more embodiments, the calculator is arranged to calculate the distance between the audio file and an audio ambience as follows: d(i, t) = A where: oi is an index corresponding to an audio ambiance, ot is the vector of values associated with the audio file, oj is the number of audio components, o wt j is the reference value corresponding to the j-th audio component of the audio ambiance i, o Pi j is the weighting coefficient corresponding to the j-th audio component of the audio ambiance i, and o tj is the value corresponding to the j-th audio component of the vector of values t.
[0016] In one or more embodiments, the sum of the weighting coefficients of an audio ambience is equal to 1.
[0017] In one or more embodiments, the vector of reference values and the vector of weighting coefficients of a given audio ambiance are constructed from a plurality of vectors of values respectively associated with audio files labeled with the given audio ambiance, each reference value corresponding to an audio component depending on an average of the values associated with the audio files each corresponding to the audio component, each weighting coefficient corresponding to an audio component depending on the inverse of the average deviation of the values associated with the audio files each corresponding to the audio component.Alternatively, the vector of reference values and the vector of weighting coefficients of a given audio ambiance are constructed by supervised learning from training data comprising a plurality of vectors of values respectively associated with audio files labeled with the given audio ambiance.
[0018] In one or more embodiments, the audio components are selected from: danceability, energy, loudness, speech, acoustics, instrumentality, livability, valence, tempo, and duration.
[0019] The invention also relates to a method for selecting audio effects for processing an audio signal implemented by computer means and comprising the following operations:
[0020] - storing data sets each associated with a respective audio ambiance, each data set comprising a vector of reference values each corresponding to a respective audio component, a vector of weighting coefficients each corresponding to a respective audio component, and audio effects,
[0021] - receive a vector of values associated with an audio file, each value corresponding to a respective audio component,
[0022] - calculate a respective distance between the audio file and each audio ambiance, which distance is calculated based on respective differences between a value associated with the audio file and the reference value associated with the audio ambiance considered corresponding to the same audio component, each difference being weighted by the weighting coefficient associated with the audio ambiance considered corresponding to the same audio component, and
[0023] - select, as audio effects to be applied to an audio signal taken from the audio file, the audio effects from the dataset associated with the audio ambiance for which the distance is the smallest.
[0024] The invention also relates to an audio processing system comprising:
[0025] - a device as previously described arranged to select audio effects to be applied to an audio signal taken from an audio file, and - a processing unit arranged to apply the selected audio effects to the audio signal.
[0026] The invention further relates to an audio processing method implemented by computer means and comprising the following operations:
[0027] - selecting audio effects to be applied to an audio signal taken from an audio file by implementing the audio effects selection method described above,
[0028] - apply the selected audio effects to the audio signal.
[0029] Finally, the invention relates to a computer program comprising instructions whose execution, by at least one processor, results in the implementation of the audio effects selection method and / or audio processing method described previously.
[0030] Other characteristics, details and advantages will appear on reading the detailed description below, and on analyzing the attached drawings in which:
[0031] [Fig. 1] illustrates a sound diffusion system;
[0032] [Fig. 2] illustrates an audio processing system according to the invention;
[0033] [Fig. 3] illustrates a method of selecting audio effects according to the invention; and
[0034] [Fig. 4] illustrates an audio processing method according to the invention.
[0035] [Fig. 1] illustrates an environment 100 within which a sound diffusion system 110 is installed.
[0036] The sound broadcasting system 110 is arranged to broadcast audio content, for example music or a podcast, to one or more individuals present in the environment 100.
[0037] In the example of [Fig. 1], the environment 100 is a car and the sound diffusion system 110 makes it possible to diffuse audio content to the passengers present in the passenger compartment.
[0038] The environment 100 may be a vehicle other than a car. It should also be noted that the environment 100 is not necessarily a vehicle and may, for example, be a home audio system, while the sound system 110 may include a headset. Generally, the sound system 110 may refer to any type of audio system comprising at least one loudspeaker, and the environment 100 may refer to any environment in which such an audio system may be installed.
[0039] The sound diffusion system 110 comprises at least one loudspeaker 112 and an audio processing system 114.
[0040] The loudspeaker 112 is arranged to produce an audio signal from an electrical signal. More particularly here, the loudspeaker 112 is arranged to convert an audio signal processed by the audio processing system 114 into an audio signal.
[0041] In the example of [Fig. 1], the sound diffusion system 110 comprises four loudspeakers 112. However, the sound diffusion system 110 may also comprise only a single loudspeaker 112.
[0042] The audio processing system 114 is arranged to perform audio processing, i.e. to add audio effects to an audio signal. Such audio processing makes it possible to modify the dynamics, temporality and / or frequency response of the audio signal. To do this, the audio processing system 114 can use various audio effects such as dynamic effects (de-esser, compressor, expander, etc.), temporal effects (delay, reverb, etc.) and / or filters (equalization, etc.).
[0043] Furthermore, the audio processing system 114 is arranged to transmit the processed audio signal to each loudspeaker 112 for the purpose of broadcasting a sound signal in the environment 100.
[0044] To do this, the audio processing system 114 can communicate with each speaker 112 using wired technology. Alternatively, the audio processing system 114 can communicate with each speaker 112 using short-range wireless communication technology. Among the wireless technologies, near-field communication (NFC), ZigBee (registered trademark), Bluetooth (registered trademark) or Wi-Fi can for example be considered. [Fig. 2] schematically illustrates the audio processing system 114.
[0045] The audio processing system 114 comprises an audio source 200, an audio effects selection device 210 and a processing unit 220.
[0046] The audio source 200 is arranged to store audio files. Each audio file contains the audio data necessary to produce an audio signal. Furthermore, each audio file is accompanied by metadata. The set formed by the audio data and the metadata can be referred to as an audio track.
[0047] Typically, the audio source 200 has the necessary means to connect to a digital distribution platform and access the catalog offered by the latter to retrieve audio files and the associated metadata. For example, the audio source 200 is capable of connecting to a digital distribution platform via a wide area network (WAN), for example the Internet. The communication between the audio source 200 and the digital distribution platform may be part of a client-server or peer-to-peer architecture (P2P).
[0048] In particular, the audio source 200 can connect to the digital distribution platform by means of a player created using a software development kit (also known by the English acronym SDK for “software development kit”), which can take the form of an application programming interface (also known by the English acronym API for “application programming interface”), made available by the digital distribution platform.
[0049] The audio source 200 is configured to retrieve the audio files and associated metadata upon request from a user, who can consult the catalog of the digital distribution platform via a website, software or an application on a user terminal such as a smartphone. The digital distribution platform is for example Spotify (registered trademark), Amazon Music (registered trademark), Apple Music (registered trademark), Tidal (registered trademark), Napster (registered trademark), YouTube Music or Deezer (registered trademark).
[0050] The audio source 200 is arranged to communicate with the device 210 and with the processing unit 220. More particularly, the audio source 200 is arranged to transmit the audio data of an audio file to the processing unit 220 and to transmit the metadata associated with the audio file to the device 210.
[0051] In the context of the invention, the metadata associated with an audio file comprises at least one vector of values. Each value corresponds to a respective audio component. In a paradigm in which the set of audio contents is assimilated to a vector space, the family of audio components can be seen as a basis of such a vector space.
[0052] For example, Spotify (registered trademark) provides a value for each of the following audio features: danceability, energy, loudness, speechiness, acousticness, instrumentality, liveness, valence, tempo, and duration.
[0053] In the following description, the vector of values is denoted t and the value corresponding to the j-th audio component is denoted t7. Finally, we consider that a number J of audio components are used. Therefore:
[0054] The device 210 is arranged to select audio effects from the metadata associated with an audio file, and more precisely from the vector of values t. Such audio effects are intended to be added to the audio signal produced from the audio data of the audio file with which the metadata comprising the vector of values t received by the device 210 are associated.
[0055] As illustrated in [Fig. 2], the device 210 comprises a memory 212 and a computer 214.
[0056] The memory 212 is arranged to store data sets each associated with a respective audio ambiance.
[0057] In the context of the invention, an audio ambiance is characterized by a vector of reference values, each reference value corresponding to a respective audio component. According to the paradigm described above, the vector of reference values of an audio ambiance corresponds to the representation of this audio ambiance in the vector space of the audio contents.
[0058] The Applicant has defined the following audio ambiances for the purpose of configuring the device 210: dynamic, acoustic, cinema, surround and natural. Of course, other audio ambiances can be defined.
[0059] To determine the reference value vector for each audio ambiance, it is possible to use the metadata associated with audio files available on one or more digital distribution platforms.
[0060] We consider a given audio ambiance, for example the cinema audio ambiance. The Applicant has identified, in the Spotify (registered trademark) catalog, the following four pieces of music corresponding to this audio ambiance: H. Zimmer, JN Howard. (2012). A Dark Knight (hereinafter t); H. Jackman. (2011). First Class (hereinafter t2) ! J- Williams. (1977). Star Wars (Main Title) (hereinafter t3); and D. Arnold. (2006). The Name's Bond... James Bond (hereinafter t4).
[0061] The same audio components are used to determine the reference value vector for each of the audio ambiances. In the case described here, the values retrieved for each of the music are those corresponding to the following audio components: danceability, energy, speech, acoustics, instrumentality, livability and valence.
[0062] We extract the following value vectors: ti = [0.277; 0.182; 0.034; 0.482; 0.782; 0.125; 0.0357] t2= [0.534; 0.487; 0.0302; 0.74; 0.885; 0.109; 0.237] t3= [0.245; 0.321; 0.0379; 0.865; 0.855; 0.156; 0.142] t4= [0.264; 0.57; 0.0338; 0.21; 0.706; 0.0783; 0.105]
[0063] It should be noted that Spotify (registered trademark) provides normalized values, which explains why all vector values t2, t3 and t4 are between 0 and 1.
[0064] This information is gathered in the table below:
[0065] The reference value vector of the cinema audio ambiance can be constructed by calculating an average of the vectors t , t2, t3 and t4. More particularly, the reference value corresponding to an audio component is calculated based on the average of the values of the vectors t , t2, t3 and t4 corresponding to this audio component.
[0066] We obtain the following vector of reference values for the cinema audio ambiance (rounded to 10' 5 near): t re / = [0.33000; 0.39000; 0.03398; 0.57425; 0.80700; 0.11708; 0.12993]
[0067] In the example given here, the average used is the arithmetic mean. However, it is possible to use a geometric mean or a harmonic mean. Such a vector of reference values can be refined by using more music labeled with the cinema audio ambiance.
[0068] By proceeding in this way for the other audio ambiances cited above, the Applicant was thus able to construct the following table in which each column contains the reference values of the corresponding audio ambiance:
[0069] It is possible to proceed in an alternative way to calculating the average. For example, the reference value vector of a given audio ambiance can be constructed by supervised learning from training data comprising a plurality of value vectors associated with audio files labeled with this given audio ambiance.
[0070] We can thus define a matrix W = is the number of ambiances i<;<; audio and w £j - is the reference value corresponding to the j-th audio component of the audio ambiance i. Each row of the matrix W thus corresponds to the vector of reference values of an audio ambiance.
[0071] In addition to a vector of reference values, each dataset associated with an audio ambiance also includes a vector of weighting coefficients, each weighting coefficient corresponding to a respective audio component.
[0072] The weighting coefficients make it possible to prioritize the respective contributions of the audio components for a given audio ambiance.
[0073] Again, to determine the vector of weighting coefficients for each audio ambiance, it is possible to use the metadata associated with audio files available on one or more digital distribution platforms. We again consider the cinema audio ambiance and the four musics mentioned above for which the respective value vectors t2, t3 and t4 were extracted.
[0074] The vector of weighting coefficients of the cinema audio ambiance can be constructed by calculating this time the inverse of the average deviation ("mean absolute difference" in English) of the vectors t , t2, t3 and t4. More specifically, the weighting coefficient corresponding to an audio component is calculated according to the inverse of the average deviation of the values of the vectors t , t2, t3 and t4 corresponding to this audio component.
[0075] We obtain the following vector of weighting coefficients for the cinema audio ambiance (rounding to 10' 5 close) :
[0076] P = [9.80392; 7.22022; 506.32911; 4.38116; 15.87302; 42.68943; 16.78556]
[0077] Again, such a vector of weighting coefficients can be refined by using more music labeled with the cinema audio ambiance.
[0078] By proceeding in this way for the other audio ambiances cited above, the Applicant was thus able to construct the following table in which each column contains the weighting coefficients of the corresponding audio ambiance:
[0079] It is possible to proceed in a different way than by calculating the inverse of the average deviation. For example, the vector of weighting coefficients for a given audio ambiance can be constructed by supervised learning from training data comprising a plurality of value vectors associated with audio files labeled with this given audio ambiance. We can thus define a matrix B where I is the number of audio moods i<;<; and / ? i ; - is the weighting coefficient corresponding to the j-th audio component of the audio ambiance i. Each row of the matrix B thus corresponds to the vector of weighting coefficients of an audio ambiance.
[0080] Once calculated, the weighting coefficients of an audio ambience can be normalized as follows: where: pj is the normalized weighting coefficient corresponding to the j-th audio component of the audio environment i.
[0081] It is also possible, for a given audio ambiance, to not take into account one or more audio components by assigning a zero value to the respective weighting coefficient of this or these audio components.
[0082] Finally, in addition to a vector of reference values and a vector of weighting coefficients, each dataset associated with an audio ambiance also includes audio effects.
[0083] In other words, for a number I of audio ambiances, a number I of audio effect sets AE , ■■■ . AEj are stored in the memory 212.
[0084] As explained previously, such audio effects may be dynamic effects, temporal effects and / or filters. The audio effects AE , ■■■ . AEj may be stored in the memory 212 in the form of data or coefficients suitable for conversion by the processing unit 220 into a specific audio processing. Alternatively, the audio effects AE , ■■■ .AEj may be stored in the memory 212 in the form of an identifier from which the processing unit 220 can find the corresponding audio effects in a specific storage medium.
[0085] The memory 212 may designate any data storage medium designed to receive and store digital data, for example a hard disk, a solid-state drive (SSD) or more generally any computer hardware allowing the storage of data on flash memory. The memory 212 may also be a random access memory or a magneto-optical disk. A combination of several storage media may also be envisaged. In the latter case, each data set associated with an audio ambiance may be divided into several sub-data sets distributed between different storage media.
[0086] Typically, the vector of reference values and the vector of weighting coefficients of a data set are stored in a first storage medium while the audio effects of this data set are stored in a second storage medium. The memory 212 then designates the combination of the first and second storage media.
[0087] Furthermore, the memory 212 can also be arranged to store instructions whose execution, by the computer 214, results in the operation of the device 210.
[0088] The computer 214 is arranged to receive a vector of values t and to identify, using the data sets stored in the memory 212 and each associated with a respective audio ambiance, the audio ambiance corresponding to the audio signal. The identification of the appropriate audio ambiance allows the computer 214 to select the appropriate audio effects, i.e. the audio effects from the data set associated with the identified audio ambiance.
[0089] The operation of the calculator 214 will be described in more detail below, with reference to [Fig. 3].
[0090] The calculator 214 can be implemented in any known manner, for example in the form of a microprocessor, a programmable logic device (PLD) or a dedicated chip of the FPGA (Field Programmable Gate Array) or SoC (System on Chip) type, a grid of computing resources, a microcontroller or any other specific form having the computing power necessary for the selection of audio effects. One or more of these elements can also be implemented in the form of specialized electronic circuits of the ASIC (Application-Specific Integrated Circuit) type. A combination of processors and electronic circuits can also be envisaged.
[0091] Finally, the processing unit 220 is arranged to receive, on the one hand, the audio data of an audio file and, on the other hand, the audio effects selected by the device 210 from the vector of values t associated with the audio file, and to produce an audio signal presenting the selected audio effects.
[0092] Equivalently, it can be considered that the audio data of an audio file received by the processing unit 220 forms an audio signal and that the processing unit 220 is arranged to add the selected audio effects to this audio signal.
[0093] The processing unit 220 is further arranged to transmit the processed audio signal to each loudspeaker 112 for the purpose of broadcasting a sound signal in the environment 100.
[0094] By "to each speaker 112" is meant here that the output of the processing unit 220 is in fact transmitted to an electronic circuit capable of shaping the processed audio signal to enable each speaker 112 to produce the desired sound signal.
[0095] It should be noted that the core of the invention relates to the selection of suitable audio effects by identifying an audio ambiance. Therefore, such an electronic circuit is not detailed here and is not shown in the drawings. It may nevertheless be noted that, typically, the processed audio signal at the output of the processing unit 220 is a digital audio signal and that the electronic circuit comprises at least one digital-to-analog converter (DAC) arranged to convert the digital audio signal into an analog audio signal and an amplifier arranged to amplify the analog audio signal and transmit it to each loudspeaker 112. The shaping of a digital audio signal - whether it is augmented with audio effects or not - is part of the general knowledge of those skilled in the art.
[0096] Referring again to [Fig. 1], it appears that the audio processing system 114 is part of the sound system 110, which is integrated into the car. However, it should be understood that some of the entities of the audio processing system 114 shown in [Fig. 2] may be remote from the car.
[0097] A method of selecting audio effects implemented by the device 210 will now be described with reference to [Fig. 3].
[0098] A context for implementing such an audio effects selection method is typically as follows: a passenger in a vehicle, for example a passenger in the car shown in [Fig. 1], uses a user terminal to connect to a digital distribution platform. Such a user terminal is for example a smartphone or a touchscreen integrated into the vehicle. The user terminal allows the passenger to connect to a digital distribution platform via a website, software or an application. Once connected, the passenger consults the catalog offered by the digital distribution platform and selects the audio content of his choice, for example music or a podcast. The audio file corresponding to the selected audio content is retrieved by the audio source 200 as well as the associated metadata.
[0099] It should be noted that this is only one possible context and that, alternatively, the audio files and associated metadata may already be stored in the memory 212 so that the audio effects selection method can be implemented locally, thus without the need to connect to the Internet and access the digital distribution platform. In such a case, the audio files and associated metadata may have been downloaded beforehand.
[0100] It should further be noted that the car of [Fig. 1] is only one example of environment 100 and that the sound diffusion system 110 can designate any audio system comprising at least one loudspeaker.
[0101] During an operation 300, the memory 212 stores data sets each associated with a respective audio ambiance.
[0102] As detailed previously, each data set includes a vector of reference values, a vector of weighting coefficients and audio effects. It should be noted that operation 300, which configures memory 212 with the data necessary for selecting suitable audio effects for an audio file, does not have to be implemented each time such audio effects are to be selected.
[0103] The memory 212 may be updated regularly, for example to refine the vector of reference values and the vector of weighting coefficients of each of the audio ambiances. Such refinement may be achieved by enriching the memory 212 with the vectors of values associated with newly labeled audio files, i.e. audio files corresponding respectively to audio contents for which an appropriate audio ambiance has been chosen.
[0104] During an operation 310, metadata associated with an audio file is received by the device 210, and more precisely by the calculator 214. As explained previously, the relevant information contained in this metadata is in fact the vector of values t, each value corresponding to an audio component.
[0105] In the example of [Fig. 2], the vector of values t is provided to the device 210 by the audio source 200.
[0106] As specified above, the vector of values t is associated with an audio file corresponding to audio content that a vehicle passenger wishes to listen to and that he has selected, for example by means of a Human-Machine Interface (HMI) of the user terminal used for this purpose.
[0107] During an operation 320, the calculator 214 identifies the audio ambiance corresponding to the vector of values t, and therefore to the audio file with which the vector of values t is associated.
[0108] To do this, the calculator 214 calculates a respective distance between the audio file to which the vector of values t is associated and each audio ambiance. The calculator 214 calculates a distance for each audio ambiance. In the example developed above, five possible audio ambiances are mentioned: dynamic, acoustic, cinema, surround sound and natural. In such a case, five distances must therefore be calculated.
[0109] The distance between the audio file and an audio ambiance is calculated based on several differences. A difference corresponds to the difference between, on the one hand, a value associated with the audio file - that is, a value of the vector of values t associated with the audio file - and therefore corresponding to an audio component and, on the other hand, the reference value associated with the audio ambiance - that is, a reference value of the vector of reference values of the dataset associated with the audio ambiance - and corresponding to the same audio component. Each difference therefore corresponds to an audio component. Furthermore, each difference is weighted by the weighting coefficient associated with the audio ambiance - that is, a weighting coefficient of the vector of weighting coefficients of the dataset associated with the audio ambiance - and corresponding to the same audio component.
[0110] For example, the distance between the audio file and an audio ambience is calculated as follows by the calculator 214: where: oi is an index corresponding to an audio ambiance, ot is the vector of values associated with the audio file, oj is the number of audio components, o wtj is the reference value corresponding to the j-th audio component of the audio ambiance i, o Pi j is the weighting coefficient corresponding to the j-th audio component of the audio ambiance i, and o tj is the value corresponding to the j-th audio component of the vector of values t. Finally, during an operation 330, the computer 214 selects, in the memory 212, the audio effects of the data set associated with the audio ambiance for which the distance is the smallest.
[0111] Therefore, the audio effects set selected by the computer 214 to process the audio signal produced from the audio data of the audio file is the audio effects set AE kcorresponding to the audio ambiance k: k = argmin ieM d(t, t)
[0112] It may be noted that, on Spotify (registered trademark), the metadata associated with an audio file corresponding to a podcast includes information on this subject. It is possible to define a podcast audio ambiance and to store in the memory 212 a data set associated with the podcast audio ambiance which includes suitable audio effects. In the event of detection, within the metadata of an audio file, of information according to which the audio file in question corresponds to a podcast, the calculator 214 can directly select the audio effects from the data set associated with the podcast audio ambiance. It is therefore not necessary for the calculator 214 to carry out the calculations mentioned above and relating to the vector of values t. It is understood that, unlike other audio ambiances, the podcast audio ambiance may not include a vector of reference values or a vector of weighting coefficients.The operation 320 described above can then include a prior sub-operation of searching, in the metadata associated with an audio file, for information according to which the audio file in question corresponds to a podcast in order, where appropriate, to directly select the audio effects of the podcast audio ambiance.
[0113] An audio processing method implemented by the audio processing system 114 will now be described with reference to [Fig. 4],
[0114] Such an audio processing method typically falls within the same context as the audio effects selection method described above. During an operation 400, the audio source 200 provides, on the one hand, audio data of an audio file to the processing unit 220 and, on the other hand, the vector of values t associated with the audio file to the device 210.
[0115] The reception of the vector of values t by the device 210 corresponds to the operation 310 of the audio effects selection method described previously. The operation 400 thus corresponds, in addition to the provision of the audio data - or, in an equivalent manner, of the audio signal formed by the audio data - to the processing unit 220, at least to the operations 310, 320 and 330 of the audio effects selection method of [Fig. 3].
[0116] At the end of operation 400, the computer 214 transmits the selected audio effects to the processing unit 220, which then has the audio data and the selected audio effects.
[0117] During an operation 410, the processing unit 220 applies the selected audio effects to the audio signal. The output of the processing unit 220 is therefore a processed audio signal.
[0118] The processed audio signal is then transmitted to each loudspeaker 112 of the sound diffusion system 110. As specified above, it must be understood here that the processed audio signal is received by an electronic circuit capable of shaping the processed audio signal to allow each loudspeaker 112 to produce the desired sound signal and to diffuse it within the environment 100 which, in the example of [Fig. 1], is a car.
Claims
Claims
1. Device (210) for selecting audio effects for processing an audio signal comprising: - a memory (212) arranged to store data sets each associated with a respective audio ambiance, each data set comprising a vector of reference values each corresponding to a respective audio component, a vector of weighting coefficients each corresponding to a respective audio component, and audio effects, and - a calculator (214) arranged to receive a vector of values associated with an audio file, each value corresponding to a respective audio component, to calculate a respective distance between said audio file and each audio ambiance, which distance is calculated as a function of respective differences between a value associated with said audio file and the reference value associated with the audio ambiance considered corresponding to the same audio component, each difference being weighted by the weighting coefficient associated with said audio ambiance considered corresponding to said same audio component, and to select, as audio effects to be applied to an audio signal taken from said audio file, the audio effects of the data set associated with the audio ambiance for which the distance is the smallest.
2. Device (210) according to claim 1, wherein the calculator (214) is arranged to calculate the distance between the audio file and an audio ambiance as follows: d(i, t) = where: oi is an index corresponding to an audio ambiance, ot is the vector of values associated with the audio file, oj is the number of audio components, o wtj is the reference value corresponding to the j-th audio component of the audio ambiance i, o Pi j is the weighting coefficient corresponding to the j-th component audio of the audio ambiance i, and where tj is the value corresponding to the j-th audio component of the value vector t.
3. Device (210) according to claim 1 or 2, wherein the sum of the weighting coefficients of an audio ambience is equal to 1.
4. Device (210) according to one of the preceding claims, in which the vector of reference values and the vector of weighting coefficients of a given audio ambiance are constructed from a plurality of vectors of values respectively associated with audio files labeled with said given audio ambiance, each reference value corresponding to an audio component depending on an average of the values associated with said audio files each corresponding to said audio component, each weighting coefficient corresponding to an audio component depending on the inverse of the average deviation of the values associated with said audio files each corresponding to said audio component.
5. Device (210) according to one of claims 1 to 3, in which the vector of reference values and the vector of weighting coefficients of a given audio ambiance are constructed by supervised learning from training data comprising a plurality of vectors of values associated respectively with audio files labeled with said given audio ambiance.
6. Device (210) according to one of the preceding claims, in which the audio components are chosen from: danceability, energy, loudness, speech, acoustics, instrumentality, livability, valence, tempo and duration.
7. A method of selecting audio effects for processing an audio signal implemented by computer means and comprising the following operations: - storing (300) data sets each associated with a respective audio ambiance, each data set comprising a vector of reference values each corresponding to a respective audio component, a vector of weighting coefficients each corresponding to a respective audio component, and audio effects, - receive (310) a vector of values associated with an audio file, each value corresponding to a respective audio component, - calculating (320) a respective distance between said audio file and each audio ambiance, which distance is calculated as a function of respective differences between a value associated with said audio file and the reference value associated with the audio ambiance considered corresponding to the same audio component, each difference being weighted by the weighting coefficient associated with said audio ambiance considered corresponding to said same audio component, and - selecting (330), as audio effects to be applied to an audio signal taken from said audio file, the audio effects from the data set associated with the audio ambiance for which the distance is the smallest.
8. An audio processing system (114) comprising: - a device (210) according to one of claims 1 to 6 arranged to select audio effects to be applied to an audio signal taken from an audio file, and - a processing unit (220) arranged to apply said selected audio effects to said audio signal.
9. A method of audio processing implemented by computer means and comprising the following operations: - selecting (400) audio effects to be applied to an audio signal taken from an audio file by implementing the method according to claim 7, - applying (410) said selected audio effects to said audio signal.
10. Computer program comprising instructions whose execution, by at least one processor, results in the implementation of the method according to claim 7 and / or 9.
Citation Information
Patent Citations
Audio similarity detection method and device, medium and computing equipment
CN114464214A
Sound effect determination method and device, equipment and storage medium
CN117093741A