Enclosure including an electronic wall detection device

The enclosure uses microphones and neural networks to detect and adapt to nearby walls, improving sound quality by minimizing reflections and enhancing user experience.

FR3153717B1Active Publication Date: 2026-04-10DEVIALET
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
FR · FR
Patent Type
Patents
Current Assignee / Owner
DEVIALET
Filing Date
2023-10-03
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

The acoustics of loudspeakers are degraded when positioned close to walls due to sound reflections, leading to a diminished user experience.

Method used

An enclosure equipped with microphones and an electronic wall detection device using neural networks to calculate spectrograms of audio signal arrival directions and energy levels to determine the presence and position of walls, allowing for adaptive sound diffusion modes.

Benefits of technology

Accurately detects the presence and position of walls, enhancing sound quality by adapting the loudspeaker's emission to minimize sound reflections.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000017_0000
    Figure 00000017_0000
  • Figure 00000018_0000
    Figure 00000018_0000
  • Figure 00000019_0000
    Figure 00000019_0000
Patent Text Reader

Abstract

Enclosure comprising an electronic wall detection device. The present invention relates to an enclosure (10) comprising: at least one loudspeaker (20) for emitting an audio signal, several microphones (25) for acquiring a received audio signal, whether a wall (15) is present in the enclosure's environment, the received audio signal including the audio signal reflected from the wall, and an electronic wall detection device (30) comprising a first computing module (35) for calculating a spectrogram of the arrival directions of the received audio signal. The detection device is characterized in that it further comprises: a second computing module (40) for calculating an energy level of the audio signal received from each microphone, and a determination module (45) for determining the presence or absence of the wall in the enclosure's environment using a neural network model. Figure for the abstract: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Title of the invention: Enclosure comprising an electronic wall detection device

[0001] The present invention relates to an enclosure comprising at least one loudspeaker suitable for emitting an audio signal, and several microphones each suitable for acquiring an audio signal received by the enclosure.

[0002] The present invention relates to the field of acoustic speakers, preferably portable acoustic speakers, also called nomadic acoustic speakers.

[0003] Such enclosures are generally suitable for being positioned in different environments with different structures.

[0004] However, the acoustics of the sounds emitted by a loudspeaker are strongly influenced by the environment, and in particular by the presence of obstacles such as walls. Indeed, when the loudspeaker is positioned close to a wall, for example less than 50 cm from it, the emitted sounds are reflected off the wall and propagate to a user with a delay compared to sounds following a direct path between the loudspeaker and the user.

[0005] This leads to a degraded experience for the user of the speaker.

[0006] There is therefore a need to determine the location of the walls in the environment of the acoustic enclosure in order to be able to adapt the emission of signals by the enclosure to limit the degradation of the user experience.

[0007] To this end, the invention relates to an enclosure further comprising at least one loudspeaker capable of emitting an audio signal, - several microphones, each capable of acquiring a received audio signal; if a wall is present in the enclosure's environment, the received audio signal will include the emitted audio signal that has been reflected off the wall, and - an electronic wall detection device connected to each microphone,

[0008] the electronic detection device comprising: • a first calculation module designed to calculate a spectrogram of the arrival directions of a component of the audio signal received by the speaker, based on the audio signal acquired by each of the microphones,

[0009] characterized in that the electronic detection device further comprises: • a second calculation module specifically designed to calculate the energy level of the received audio signal acquired from each microphone, and • a determination module capable of determining the presence or absence of a wall in the enclosure's environment by applying a neural network model with spectrogram of arrival directions and determined energy levels.

[0010] According to particular embodiments, the enclosure according to the invention comprises one or more of the following features, taken individually, or according to all technically possible combinations: - The neural network model includes a first convolutional neural network designed to receive the spectrogram of arrival directions, a second neural network designed to receive energy levels, a concatenation block designed to concatenate data from the first and second neural networks to form concatenated data, and a third neural network designed to process the concatenated data to determine the presence or absence of the wall, - the first neural network comprises a circular sliding window and at least one layer of neurons, the sliding window being suitable for application to the arrival direction spectrogram before the layer of neurons, - The neural network model is suitable for determining, if the wall is present in the environment of the enclosure, a position of the wall relative to the enclosure from among a plurality of predefined positions, based on the spectrogram of arrival directions and calculated energy levels. - at least one loudspeaker is designed to emit the audio signal according to a speaker's diffusion mode,

[0011] the electronic detection device further comprising an adaptation module capable of reconfiguring the diffusion mode of the enclosure according to the determination of the presence, or absence, of the wall, - The arrival direction spectrogram is a matrix comprising a plurality of values, each value corresponding to the power of the audio signal received by the speaker according to a predefined angular interval relative to a predefined reference frame centered on a center of the speaker, and according to a predefined frequency interval, - the emitted audio signal is part of an audio stream originating from a broadcast instruction from a user; and - the enclosure further includes its own accelerometer to determine if the enclosure is stationary and to issue a calculation instruction when the enclosure is stationary, the first and second calculation modules being specific to calculating the spectrogram of the directions of arrival and the energy levels following the reception of the calculation instruction from the accelerometer.

[0012] The invention also relates to a method for detecting a wall in an enclosed environment comprising at least one loudspeaker, several microphones and an electronic wall detection device connected to each microphone,

[0013] the process comprising the following steps: - emission of an audio signal from at least one loudspeaker, - acquisition, by each microphone, of a received audio signal, if a wall is present in the enclosure environment, the received audio signal including the audio signal that was reflected off the wall, - Calculation of a spectrogram of the arrival directions of a component of the audio signal received by the speaker, based on the audio signal acquired by each microphone.

[0014] characterized in that the process further comprises the following steps: - calculation of the energy level of the received audio signal acquired from each microphone, and - determination of the presence or absence of the wall in the environment of the enclosure by applying a neural network model to the spectrogram of the directions of arrival and the determined energy levels.

[0015] The invention also relates to a computer program product comprising software instructions which, when executed by a computer, implement a detection method as described above.

[0016] The invention will become clearer upon reading the following description, given solely by way of non-limiting example, and made with reference to the drawings in which: - [Fig.1] [Fig.1] is a schematic representation of an enclosure according to the invention; - [Fig. 2] [Fig. 2] is an example of a directional spectrogram arrival specific to be determined by an electronic detection device included within the enclosure of the [Fig.l]; - [Fig.3] [Fig.3] is an example of energy levels specific to being determined by the electronic detection device included in the enclosure of the [Fig.1]; - [Fig.4] [Fig.4] is a schematic representation of a model with neural network(s) implemented by the electronic detection device included in [Fig. 1]; and - [Fig. 5] [Fig. 5] is a flowchart of a detection process implemented work by the enclosure of the [Fig.l].

[0017] Figure 1 shows an enclosure 10 in an environment. The environment includes a wall 15 in the vicinity of the enclosure 10. "In the vicinity" means a distance less than or equal to 1 m and preferably less than or equal to 50 cm.

[0018] The enclosure 10 comprises at least one loudspeaker 20 suitable for emitting audio signals. Preferably, the enclosure 10 comprises several loudspeakers 20, preferably distributed around the periphery of the enclosure 10 to emit audio signals in several directions.

[0019] In the example of [Fig.1], the enclosure 20 comprises two loudspeakers 20.

[0020] The enclosure 10 further comprises a plurality of microphones 25, each adapted to acquire the audio signals present in the environment of the enclosure 10. In the example of [Fig. 1], the enclosure 10 comprises four microphones distributed over its outer surface. The microphones 25 are preferably angularly and regularly spaced from one another. Thus, in the example where the enclosure 10 comprises four microphones 25, the microphones 25 are angularly distributed with an angular displacement of 90° so as to acquire audio signals from directions as far apart from each other as possible.

[0021] The speaker 10 further comprises a control device 27, an accelerometer 28, and an electronic detection device 30. The control device 27 is connected to the speakers 20 and optionally to an audio content source (not shown), such as a digital player, an optical disc player, a vinyl record player, or a smartphone. The control device 27 is adapted to receive audio content from the source and to send an excitation command to the speakers 20 to play the received audio content. The excitation command then causes the speakers 20 to emit an audio signal, referred to as the emitted audio signal. Preferably, the control device 27 is adapted to send the excitation command according to a first playback mode of the speaker 10 selected from several predefined playback modes of the speaker 10.For example, the chosen diffusion mode is a neutral mode in which the 20 speakers are controlled synchronously.

[0022] It is then understood that the emitted audio signal is part of an audio stream originating from a broadcast instruction from a user. In other words, the emitted audio signal is, for example, any type of music or voice, and is not limited to a predefined test signal.

[0023] If the wall 15 is present in the environment of the enclosure 10, the emitted audio signal is at least partially reflected off the wall 15.

[0024] Each microphone 25 is then suitable for acquiring an audio signal received by the speaker 10 following the emission of the audio signal emitted by the loudspeakers 20.

[0025] The received audio signal includes a direct component which is the audio signal emitted along the direct path between the loudspeakers 20 and the microphones 25.

[0026] If the wall 15 is present in the environment of the enclosure 10, as in the example of [Fig.1], the received audio signal further includes a reflected component which is the audio signal emitted and reflected on the wall 15.

[0027] Optionally, the received audio signal also includes noise. In other words, the received audio signal includes: - the direct component and the noise, if no wall 15 is present in the environment of enclosure 10; or - the direct component, the reflected component, and the noise, if wall 15 is present in the environment of enclosure 10.

[0028] The accelerometer 28 is connected to the detection device 30. The accelerometer 28 is suitable for measuring an acceleration of the enclosure 10 and for sending a calculation instruction to the detection device 30 when the enclosure 10 is newly stationary.

[0029] The electronic detection device 30 includes a first module 35 for calculating a spectrogram of the directions of arrival, a second module 40 for calculating energy levels, a module 45 for determining the presence or absence of the wall 15, and optionally a module 50 for adapting a diffusion mode.

[0030] In the example of [Fig.1], the electronic detection device 30 is a computer formed for example of a memory 55 and a processor 60 associated with the memory 55.

[0031] In the example of [Fig. 1], the first calculation module 35, the second calculation module 40, and the determination module 45, as well as the optional adaptation module 50, are each implemented as a software program, or a software component, executable by the processor. The memory 55 of the electronic detection device 25 is then capable of storing a first calculation program, a second calculation program, and a determination program, as well as, optionally, an adaptation program. The processor is then capable of executing each of the following programs: the first calculation program, the second calculation program, and the detection program, as well as, optionally, the adaptation program.

[0032] In an alternative not shown, the first calculation module 35, the second calculation module 40 and the determination module 45, as well as the optional adaptation module 50, are each implemented as a programmable logic component, such as an FPGA (Field Programmable Gate Array), or as an integrated circuit, such as an ASIC (Application-Specific Integrated Circuit).

[0033] When the electronic detection device 30 is implemented in the form of one or more software programs, i.e., in the form of a computer program, also called a computer program product, it is further capable of being stored on a computer-readable medium (not shown). A computer-readable medium is, for example, a medium capable of storing electronic instructions and being connected to a bus of a computer system. By way of example, a readable medium is an optical disc, a magneto-optical disc, ROM, RAM, any type of non-volatile memory (e.g., FLASH or NVRAM), or a magnetic card. A computer program comprising software instructions is then stored on the readable medium.

[0034] The control device 27 is, for example, also a computer comprising a non-represented memory and a non-represented processor.

[0035] Alternatively, the detection device 30 and the control device 27 form a single computer comprising a single processor 60 and a single memory 55. The memory 55 then also stores control software.

[0036] The first calculation module 35 is suitable for acquiring the audio signal received from each microphone 25, for example following the receipt of the calculation instruction from the accelerometer 28. The first calculation module 35 is suitable for calculating a spectrogram of the directions of arrival on the speaker 10 of the received signal, from the audio signal received by the speaker 10 and acquired by each of the microphones 25. The spectrogram of the directions of arrival is calculated from all the components of the received audio signal.

[0037] To this end, the first calculation module 35 is, for example, suitable for applying the MUSIC (Multiple Signal Classification) algorithm, which is known from the prior art and described in particular on the following web page: https: / / fr.wikipedia.org / wiki / MUSIC_(algorithme). The MUSIC algorithm is configured, as is known, to detect a sound signal coming from a search direction. When it detects a sound signal coming from a particular direction, it is necessarily the reflected component because, by construction, the direct component arrives identically from all directions.

[0038] The use of the MUSIC algorithm then makes it possible to identify a preferred arrival direction for the received audio signal.

[0039] If the wall 15 is present in the environment of the enclosure 10, the identified component corresponds substantially to the reflected component of the received audio signal.

[0040] The MUSIC algorithm provides a matrix.

[0041] The columns of this matrix represent the arrival directions around the enclosure 10 according to a first sampling of the arrival directions. The columns therefore correspond to an angular interval of identical width from one column to another, in a predefined coordinate system centered on a center of the enclosure 10.

[0042] The rows of this matrix correspond to frequency bins. For example, the matrix comprises 257 rows.

[0043] Each value of the matrix is ​​associated with a column and a row, and corresponds to the power of the component identified by the MUSIC algorithm, in the angular interval of the associated column, and for the frequency interval of the associated row.

[0044] The first calculation module 35 is also optionally capable of performing frequency subsampling and smoothing in sixths of an octave on the frequency axis of the matrix to reduce the number of rows, for example to 25 rows. Alternatively, the smoothing is performed, for example, in twelfths of an octave, in thirds of an octave, or by octave.

[0045] Furthermore, the first calculation module 35 is optionally suitable for reducing the number of columns by averaging the values ​​over several angular intervals, for example so that the matrix comprises only 8 columns. In other words, after the averaging is performed, each column corresponds to an angular interval of 45°.

[0046] In addition, the first calculation module 35 is preferably suitable for normalizing the values ​​of the matrix obtained so that the highest value corresponds to the value 1 and the other values ​​correspond to percentages of the highest value.

[0047] Thus, the direction(s) of arrival for which the values ​​are highest are the most probable directions of arrival on the enclosure 10 of the reflected component. Therefore, this or these direction(s) of arrival are the most probable directions in which the wall 15 lies relative to the enclosure 10.

[0048] The matrix obtained by the first calculation module 35 is then the spectrogram of the arrival directions of the sound sources around the enclosure 10, including in particular the reflected component identified when present. An example of such a spectrogram is illustrated in grayscale in [Fig. 2]. In [Fig. 2], each cell corresponds to a value of the matrix obtained.

[0049] In [Fig. 2], the lighter a square is, the higher the power of the identified reflected component on enclosure 10 in the corresponding frequency and angular interval. Conversely, the darker a square is, the lower the power of the identified reflected component in the corresponding frequency and angular interval.

[0050] In [Fig.2], two zones 65 are shown. The zones 65 correspond to sets of frequency intervals for which the power of the identified reflected component is highest.

[0051] In the example of [Fig. 2], these zones correspond to the angular intervals between 90° and 135° on the one hand, and between 270° and 315° on the other. There is therefore an ambiguity of 180° concerning the direction of wall 15 relative to enclosure 10. This ambiguity arises from the fact that the MUSIC algorithm is designed to calculate a spectrogram of the arrival directions in the context of a point source and not of a reflection of an audio signal on a surface formed by wall 15.

[0052] Furthermore, the material of the wall 15 influences the frequency reflection of the emitted audio signal. Indeed, depending on the material of the wall 15, certain frequencies are predominantly absorbed while others are predominantly reflected.

[0053] The second calculation module 40 is also suitable for receiving, from each microphone 25, the audio signal received, for example following the reception of the calculation instruction from the accelerometer 28.

[0054] The second calculation module 40 is designed to calculate an energy level of the received audio signal, acquired from each microphone 25.

[0055] To this end, the second calculation module 40 is adapted to calculate the energy level of the received audio signal, acquired by each microphone 25 for a predefined duration, for example 5 seconds. Preferably, the calculation module 40 is adapted to calculate the energy levels only over a predefined frequency band, for example between 20 Hz and 20,000 Hz.

[0056] For example, the energy level is calculated using the following formula:

[0057] _ X^con^X^

[0058] where: Ej is the energy level received by the microphone \

[0059] T is the predefined duration,

[0060] f min and fmax are respectively the minimum and maximum frequencies of the band frequency,

[0061] X^t, f) is the audio signal received by microphone 1 at time t and at frequency f, in the frequency domain, and

[0062] conjQ is the complex conjugate function.

[0063] The second calculation module 40 is designed to normalize the energy levels so that the highest energy level is equal to 1 and the other energy levels are equal to percentages of the highest energy level.

[0064] Figure 3 shows an example of calculated energy levels. It is visible in Figure 3 that the first microphone 25 acquires the received audio signal with the highest energy level, i.e., a normalized level equal to 1. It is also visible that the other microphones 25 acquire the received audio signal with a lower energy level, i.e., a normalized level between 0 and 1.

[0065] Again with reference to [Fig.1], the determination module 45 is connected to the first 35 and second 40 calculation modules.

[0066] The determination module 45 is suitable for determining the presence of the wall 15 or the absence of the wall 15 in the environment of the enclosure 10 by applying a neural network model 70 to the spectrogram of the arrival directions and the determined energy levels.

[0067] Fig. 4 illustrates an architecture of the neural network model 70 implemented by the determination module 45.

[0068] The neural network model 70 comprises a first convolutional neural network 75 suitable for receiving the spectrogram of the arrival directions, a second neural network 80 suitable for receiving the energy levels, a concatenation block 83 for concatenating data from the first 75 and second 80 neural networks to form concatenated data, and a third neural network 85 suitable for processing the concatenated data to determine the presence or absence of the wall 15.

[0069] Each neural network comprises an ordered succession of layers of neurons, each of which takes its inputs from the outputs of the previous layer.

[0070] More precisely, each layer includes neurons taking their inputs from the outputs of the neurons of the previous layer, or from the input variables for the first layer.

[0071] Each neuron is also associated with an operation, that is to say a type of processing, to be carried out by said neuron within the corresponding processing layer.

[0072] Each layer is connected to the other layers by a plurality of synapses. A synaptic weight is associated with each synapse, and each synapse forms a link between two neurons. Each synaptic weight is preferably a real number, which takes both positive and negative values. In some cases, each synaptic weight is a complex number.

[0073] Each neuron is capable of performing a weighted sum of the value(s) received from the neurons of the preceding layer, each value being multiplied by the respective synaptic weight of each synapse, or connection, between said neuron and the neurons of the preceding layer, then applying an activation function, typically a non-linear function, to said weighted sum, and delivering at the output of said neuron, in particular to the neurons of the next layer connected to it, the value resulting from the application of the activation function. The activation function allows for the introduction of non-linearity in the processing performed by each neuron. The sigmoid function, the hyperbolic tangent function, the Heaviside function, the The Rectified Linear Unit function, also called the ReLU function (from the English, Rectified Linear Unit), or the softmax function, are examples of activation functions.

[0074] As an optional complement, each neuron is also capable of applying, in addition, a multiplicative factor, also called bias, to the output of the activation function, and the value delivered at the output of said neuron is then the product of the bias value and the value from the activation function.

[0075] A convolutional neural network is also sometimes called a convolutional neural network or by the acronym CNN (from the English, Convolutional Neural Networks).

[0076] In a convolutional neural network, each neuron in the same layer exhibits exactly the same connection pattern as its neighboring neurons, but at different input positions. The connection pattern is called the convolutional kernel or, more commonly, the "kernel" in reference to the corresponding English term.

[0077] A fully connected layer of neurons is a layer in which the neurons of said layer are each connected to all the neurons of the preceding layer.

[0078] Such a type of layer is more often referred to by the English term "fully connected", and sometimes designated by the name "dense layer".

[0079] The first neural network 75 comprises a circular padding sliding window and at least one layer of neurons. Preferably, the first neural network 75 comprises three layers of neurons.

[0080] The sliding window comprises a kernel of width M and height N, where M is less than the number of columns in the matrix forming the spectrogram, and N is less than the number of rows in said matrix. The window is designed to move within the spectrogram to form submatrices of dimensions M by N. The submatrices are provided to at least one layer of neurons.

[0081] The circular type of the sliding window is such that the window is suitable for forming sub-matrices comprising values ​​from the first and last columns. In other words, the sliding window moves across the spectrogram as if the first column of the corresponding matrix were contiguous with the last column, i.e., as if the spectrogram formed a cylinder.

[0082] The sliding window is suitable for being applied to the arrival direction spectrogram, calculated by the first computing module 35, before the or each layer of neurons.

[0083] Thus, the spectrogram of the arrival directions is treated as an image by the first neural network 75.

[0084] The first neural network 75 provides, at its output, initial data.

[0085] The second neural network 80 is a fully connected neural network. The second neural network 80 is designed to take as input a vector comprising the normalized energy level of each microphone 25 calculated by the second computing module 40.

[0086] The second neural network 80 provides second data in its output.

[0087] The concatenation block 83 is connected to the output of the first 75 and second 80 neural networks and is designed to concatenate the first and second data into a single vector to form concatenated data.

[0088] The third neural network 85 is connected to the concatenation block 83. The third neural network 85 is, for example, of the fully connected type.

[0089] The third neural network 85 is preferentially suited to classifying concatenated data into a plurality of classes, for example nine classes.

[0090] For example, each class except one corresponds to a predefined position of the wall 15 relative to the enclosure 10, the last class corresponding to the absence of wall 15. The position of the wall 15 is defined here as the direction in which the wall 15 lies relative to the enclosure 10 in a predefined coordinate system centered on the enclosure 10. Thus, each class corresponds to a predefined position of the wall 15, defined by an angular interval in which the wall 15 lies relative to the enclosure 10.

[0091] If the plurality of classes comprises nine classes, each of the first eight classes corresponds to an angular interval of 45°, the ninth class corresponding to the absence of wall 15.

[0092] To this end, the third neural network 85 is designed to provide an output vector comprising nine components. The activation function of the neurons in the last layer of the third neural network 85 is preferably the softmax function. Thus, the value of each component of the output vector is between 0 and 1, and the sum of the values ​​of all the components of said vector is equal to 1. The value of each component is then a probability of belonging to the corresponding class.

[0093] The determination module 45 is for example suitable for determining the absence of wall 15 if the component of the output vector having the highest value is the component associated with the class corresponding to the absence of wall 15. The determination module 45 is then suitable for determining the presence of wall 15 otherwise.

[0094] As an optional complement, the determination module 45 is further able to determine the position of the wall 15 as being the position defined by the class whose output component has the highest value in the output vector.

[0095] Optionally, the determination module 45 is configured to compare said highest value to a predefined threshold. If the highest value is lower than the predefined threshold, then the determination module 45 is capable of determining that the wall detection 15 is not conclusive and to order the reiteration of calculations by the first 35 and second 40 calculation modules from second audio signals acquired at a later time.

[0096] The adaptation module 50 is suitable for receiving the position of the wall 15, or information on the absence of wall 15, from the determination module 45.

[0097] The adaptation module 50 is suitable for reconfiguring the diffusion mode of the enclosure 10 according to the determination of the presence or absence of the wall 15, and more preferably according to the position of said wall 15 where applicable.

[0098] In the example of [Fig. 1], a diffusion mode is associated with each of the following positions of the wall 15: opposite one of the loudspeakers 20, opposite the other loudspeaker 20, at 90° to a line joining the two loudspeakers 20 in one direction, and at 90° to the line joining the two loudspeakers 20 in the opposite direction. Thus, four diffusion modes are defined by these positions.

[0099] For each of these four diffusion modes, a delay is for example applied by the control device 27 to one of the loudspeakers 20 to compensate for the presence of the wall 15.

[0100] A fifth diffusion mode is defined as a neutral diffusion mode in which the loudspeakers 20 are controlled synchronously, i.e., without delay. The adaptation module 50 is, for example, suitable for associating the fifth diffusion mode with all other positions and for determining the absence of a wall 15.

[0101] In the fifth diffusion mode, an equalization of the different channels, in phase and in amplitude, is preferably carried out.

[0102] In the first, second, third and fourth diffusion modes, other processing is preferentially also carried out to further improve the sound rendering for a user located on the opposite side of the wall from the speaker 10. For example, this processing includes phase and amplitude equalization, and decomposition of the signal into components to be diffused in front of and behind the speaker 10.

[0103] The adaptation module 50 is adapted to send the reconfigured diffusion mode to the control device 27 so that it controls the loudspeakers 20 in accordance with this reconfigured diffusion mode.

[0104] The operation of enclosure 10 will now be described with reference to [Fig.5] illustrating a flowchart of a process for determining the presence or absence of wall 15.

[0105] During an emission step 110, the control device 27 receives audio content from the source, and sends an excitation command to the speakers 20 to emit the audio content.

[0106] At a certain instant, the accelerometer 28 detects that the speaker has become stationary following movement. The accelerometer then sends an instruction to the detection device 30 indicating that the speaker 10 has become stationary. The audio content emitted by the loudspeakers 20 from this instant onward is the emitted audio signal.

[0107] During an acquisition step 120, each microphone 25 acquires the received audio signal.

[0108] Then, in a first calculation step 130, the first calculation module 35 calculates the spectrogram of the arrival directions of the component of the audio signal received by the speaker 10, from the audio signal received, acquired by each of the microphones 25, for example as described previously.

[0109] During a second calculation step 140, the second calculation module 40 calculates the energy level of the received audio signal, acquired from each microphone 25, for example as explained previously.

[0110] Preferably, the first 130 and second 140 calculation steps are implemented simultaneously.

[0111] Then, during a determination step 150, the determination module 45 determines the presence of the wall 15 or the absence of the wall 15 in the environment of the enclosure 10, by applying the neural network model 70 to the spectrogram of the directions of arrival and the calculated energy levels.

[0112] For example, the determination module 45 applies the first neural network 75 to the spectrogram to obtain the first data and the second neural network 80 to the energy levels to obtain the second data. The determination module 45 then applies the concatenation block 83 to the first and second data to form the concatenated data. The determination module 45 then provides the concatenated data to the third neural network 85 to determine the output vector.

[0113] According to this optional complement, the determination module 45 determines the position of the wall 15 relative to the enclosure 10 or the absence of wall 15, as previously indicated, from the output vector.

[0114] Optionally, during an adaptation step 160, the adaptation module 50 reconfigures the diffusion mode of the enclosure 10 according to the position of the wall 15 relative to the enclosure 10, or the absence of a wall 15, for example by sending the reconfigured diffusion mode to the control device 27.

[0115] With the enclosure 10 according to the invention, the presence of the wall 15, or its absence, is detected accurately.

Claims

Demands

1. Enclosure (10) comprising: - at least one loudspeaker (20) capable of emitting an audio signal, - several microphones (25), each capable of acquiring a received audio signal, if a wall (15) is present in an environment of the enclosure (10), the received audio signal includes the emitted audio signal that has been reflected off the wall (15), and - an electronic wall detection device (30) connected to each microphone, the electronic detection device comprising: • a first calculation module (35) designed to calculate a spectrogram of the arrival directions of the audio signal received by the speaker (10), from the audio signal received by each of the microphones (25), characterized in that the electronic detection device further comprises: • a second calculation module (40) designed to calculate an energy level of the received audio signal acquired from each microphone (25), and • a determination module (45) suitable for determining the presence of the wall (15) or the absence of the wall (15) in the environment of the enclosure (10) by applying a neural network model (70) to the spectrogram of the arrival directions and the determined energy levels.

2. Enclosure (10) according to the preceding claim, wherein the neural network model (70) comprises a first convolutional neural network (75) adapted to receive the spectrogram of the arrival directions, a second neural network (80) adapted to receive the energy levels, a concatenation block (83) adapted to concatenate data from the first (75) and second (80) neural networks to form concatenated data, and a third neural network (85) and suitable for processing concatenated data to determine the presence or absence of the wall (15).

3. Enclosure (10) according to claim 2, wherein the first neural network (75) comprises a circular sliding window and at least one layer of neurons, the sliding window being suitable for being applied to the arrival direction spectrogram before the layer of neurons.

4. Enclosure (10) according to any one of the preceding claims, wherein the neural network model (70) is suitable for, if the wall (15) is present in the environment of the enclosure (10), determining a position of the wall (15) relative to the enclosure (10) among a plurality of predefined positions, from the spectrogram of the arrival directions and calculated energy levels.

5. Enclosure (10) according to any one of the preceding claims, wherein at least one loudspeaker (20) is suitable for emitting the audio signal emitted according to a diffusion mode of the enclosure (10), the electronic detection device (30) further comprising an adaptation module (50) suitable for reconfiguring the diffusion mode of the enclosure (10) according to the determination of the presence, or absence, of the wall (15).

6. Enclosure (10) according to any one of the preceding claims, wherein the arrival direction spectrogram is a matrix comprising a plurality of values, each value corresponding to the power of the audio signal received by the enclosure (10) over a predefined angular interval with respect to a predefined reference frame centered on a center of the enclosure (10), and over a predefined frequency interval.

7. Speaker (10) according to any one of the preceding claims, wherein the emitted audio signal is included in an audio stream from a broadcast instruction from a user.

8. Enclosure (10) according to any one of the preceding claims, further comprising an accelerometer (28) suitable for determining whether the enclosure (10) is stationary and for issuing a calculation instruction when the enclosure (10) is stationary, the first (35) and second (40) calculation modules being suitable for calculating the arrival direction spectrogram and energy levels following receipt of the calculation instruction from the accelerometer (28).

9.

10. Method for detecting a wall (15) in an enclosure environment (10) comprising at least one loudspeaker (20), several microphones (25) and an electronic wall detection device (30) connected to each microphone (25), the method comprising the following steps: - emission (110) of an audio signal emitted by at least one loudspeaker (20), - acquisition (120), by each microphone (25), of a received audio signal, if the wall (15) is present in an environment of the enclosure (10), the received audio signal including the audio signal which has been reflected off the wall (15), - calculation (130) of a spectrogram of the arrival directions of the audio signal received by the speaker (10), from the received audio signal acquired by each microphone (25), characterized in that the process further comprises the following steps: - calculation (140) of an energy level of the received audio signal acquired from each microphone (25), and - determination of the presence of the wall (15) or the absence of the wall (15) in the environment of the enclosure (10) by application of a neural network model (70) to the spectrogram of the directions of arrival and the determined energy levels. Product computer program comprising software instructions which, when executed by a computer, implement a detection method according to the preceding claim.