Method and device for configuring an audio system
The method and device utilize a camera to overlay graphic representations of listening zones onto a video feed, addressing the complexity of existing audio system configuration methods by providing an intuitive and equipment-free solution for configuring and switching between listening zones.
Patent Information
- Application Number
- EP2024213498
- Authority / Receiving Office
- EP · EP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-04
- Filing Date
- 2024-11-18
- Publication Date
- 2025-06-11
- Estimated Expiration
- 2044-11-18
AI Technical Summary
Existing audio system configuration methods require additional equipment such as microphones and complex software processing, or in-depth knowledge of acoustics, making it difficult and cumbersome to determine and switch between multiple listening zones.
A method and device using a camera to visually select and configure a listening area by overlaying graphic representations of potential listening zones onto a video feed, allowing users to intuitively choose and adjust the geometric parameters of the listening area in real-time.
Enables simple and intuitive configuration of listening zones without the need for additional equipment or specialized knowledge, allowing users to easily switch between different listening configurations by providing real-time visual feedback.
Smart Images

Figure IMGAF001_ABST
Abstract
Description
Domaine technique
[0001] A method and device for configuring an audio system are described, using a camera to select a listening area. The method and device can be used in particular for configuring a home audio system. Arrière-plan technique
[0002] An audio system, for example a home audio system, can be associated with a listening zone, which represents a sound diffusion location in which a listener benefits from sound that meets desired criteria. Determining such a zone may require the implementation of various additional equipment. For example, it may be necessary to implement a microphone to make sound recordings in the room in which the audio system is placed, or to use a mobile phone. The recorded sound must then be processed by dedicated software to identify the listening zone. Other solutions may require in-depth knowledge of acoustics and sound reinforcement. The problem increases if several listening zones are envisaged, these listening zones corresponding for example to different configurations of the sound processing by the audio system.Switching from one zone to another requires performing the operations mentioned above again.
[0003] An intuitive audio system configuration solution is offered, allowing simple determination of the listening zone(s). Résumé
[0004] 1. One or more embodiments relate to a method for configuring an audio system implemented by a device comprising a processor, the method comprising configuring an audio processing processor for generating one or more audio signals feeding one or more respective audio sources as a function of one or more geometric parameters defining a listening area chosen by a user, the method further comprising: obtaining, from a camera, a video of a location where sound is broadcast by the one or more audio sources; generating a signal representative of a video to be displayed on a display device, the video to be displayed comprising all or part of the video obtained by the camera;the overprinting, in the video to be displayed, of at least one graphic representation of a listening area among several listening areas, the video with overprint displayed on the screen constituting real-time visual feedback adapted to help the user configure the location where the listening area is located; obtaining data representative of a choice, by the user, of a displayed listening area, for the configuration of the audio processing processor.;
[0005] The visual feedback obtained in real time using the camera helps the user with the configuration of the premises and / or the choice of a listening area and its geometric parameters. The video to be displayed may include only a part of the video obtained by the camera, in the sense that only a part of the camera's field of vision is reproduced in the video to be displayed.
[0006] According to one or more exemplary embodiments, the geometric parameter(s) comprising at least one of: a distance between the listening area and the audio source(s); a width of the listening area.
[0007] According to one or more exemplary embodiments, a listening zone is an area in which the sound produced by the audio source(s) meets one or more quality criteria.
[0008] According to one or more exemplary embodiments, the method comprises superimposing, in the video to be displayed, a menu comprising several choices of value of a parameter, highlighting a current choice of parameter value, as well as the graphical representation of a listening zone corresponding only to a current choice of parameter value.
[0009] According to one or more exemplary embodiments, the method comprises displaying an invitation to a user to move a piece of furniture in the broadcast location based on the graphical representation.
[0010] According to one or more exemplary embodiments, the audio source(s) are orientable, the orientation of the audio source(s) comprising one of: the display of a message inviting a user to orient one or more audio sources according to a chosen listening area; the generation of a control signal for a motor for orienting one or more audio sources according to a chosen listening area.
[0011] According to one or more exemplary embodiments, a geometric parameter is defined by a value, or a range of values, or an open range.
[0012] One or more exemplary embodiments relate to a device comprising a camera, at least one audio source, a processor and a memory comprising software code, the processor, when executing the software code, causing the device to implement one of the methods described.
[0013] According to one or more exemplary embodiments, the at least one audio source can be oriented either manually or by means of an actuator. Brève description des figures
[0014] Other features and advantages will become apparent when reading the detailed description which follows, for the understanding of which reference will be made to the attached drawings, among which: there figure 1 is a schematic diagram of an audio system according to one or more embodiments; the figure 2 is a schematic diagram of a top view of a room in which the system of the figure 1 ; there figure 3 is a flowchart of a configuration method according to one or more embodiments; the figure 4 is a schematic representation of a video image that can be displayed with, superimposed, a graphical representation of a listening area; the figure 5 is a first schematic representation of a video image illustrating the screen display at various stages; the figure 6 is a second schematic representation of a video image illustrating the screen display at various stages; figure 7 is a third schematic representation of a video image illustrating the screen display at various stages; figure 8 is a fourth schematic representation of a video image illustrating the screen display at various stages; figure 9 is a fifth schematic representation of a video image illustrating the screen display at various stages; figure 10 is a sixth schematic representation of a video image illustrating the screen display at various stages; figure 11 is a schematic diagram of a portion of the device according to one or more embodiments illustrating examples of possible orientations of an audio source; the figure 12 is a schematic representation of a video image illustrating the display on the screen of menu items having the orientations illustrated by the figure 11 , as well as the selection of a first orientation; the figure 13 is a schematic representation of a video image illustrating the display on the screen of menu items having the orientations illustrated by the figure 11 , as well as the selection of a second orientation; figure 14 is a schematic diagram illustrating a method of determining the position of a point in the image as a function of a distance; figure 15 is a schematic diagram that shows the positioning of certain elements of the figure 14 in the displayed video image. Description détaillée
[0015] In the following description, identical, similar or analogous elements will be designated by the same reference numerals. The block diagrams, flowcharts and message sequence diagrams in the figures illustrate the architecture, functionalities and operation of systems, devices, methods and computer program products according to one or more exemplary embodiments. Each block of a block diagram or each phase of a flowchart may represent a module or a portion of software code comprising instructions for implementing one or more functions. According to certain implementations, the order of the blocks or phases may be changed, or the corresponding functions may be implemented in parallel.The process blocks or phases may be implemented using circuitry, software, or a combination of circuitry and software, in a centralized manner or in a distributed manner for all or some of the blocks or phases. The systems, devices, processes, and methods described may be modified, added to, and / or deleted within the scope of this disclosure. For example, the components of a device or system may be integrated or separated. Also, the described functions may be implemented using more or fewer components or phases, or with other components or through other phases. Any suitable data processing system may be used for the implementation. For example, a suitable data processing system or device includes a combination of software code and circuitry, such as a processor, controller, or other circuitry suitable for executing the software code.When the software code is executed, the processor or controller causes the system or device to implement all or part of the functionalities of the blocks and / or phases of the processes or methods according to the exemplary embodiments. The software code may be stored in non-volatile memory or a readable non-volatile storage medium (USB key, memory card or other medium) directly or through an interface adapted by the processor or controller.
[0016] There figure 1 is a schematic diagram of an audio system illustrating one or more embodiments in a non-limiting manner. The system of the figure 1 comprises a configuration device 100 and a screen 101. The device 100 can be controlled by a user 102 using a remote control 103. The device 100 comprises a camera 104, a processor 105, a non-volatile memory 106 comprising software code, and an audio processing processor 107. The various components of the device 100 are connected by an internal bus 110. The audio processing processor 107 receives audio data as input and generates audio signals intended to be reproduced by one or more audio sources, represented here by two loudspeakers 108 and 109. The device 100 further comprises an interface (not illustrated) by which it is connected to the screen 101. This interface is for example an HDMI interface. The device 100 is adapted to generate a video signal for display on the screen 101. The generation of the video signal is for example carried out by the processor 104. The system of the figure 1 is given for illustrative purposes for the clarity of presentation of the exemplary embodiments and a current implementation may obviously differ. Furthermore, the device 100 may comprise a single speaker or more than two speakers. Depending on the implementations, the speakers 108 and 109 are fixed or orientable. In the case of orientable speakers, the orientation may be adjusted manually or via actuators such as electric motors. According to certain implementations, position sensors are associated with the manually orientable speakers to allow the device 100 to determine the orientation.
[0017] The device 100 integrates for example a video receiver and decoder functionality in addition to the audio and camera functionalities (product designated under the name 'video and sound box' or 'video sound box' in English). It should also be noted that both the functionality of the processor 105 and that of the processor 107 can be implemented with more than one component, or even be jointly implemented by one or more components. For example, a single processor can be used to implement both functionalities. In the example of the figure 1 , the device 100 comprises two audio sources and the camera is placed on an axis of symmetry 111 of these audio sources. This is however not obligatory.
[0018] There figure 2 a schematic diagram illustrating an example of system placement figure 1 in a room 200. The camera 104 of the device 100 has a horizontal viewing angle AdV_H. According to the example of the figure 2 , the room also includes furniture such as a sofa 202. A listening zone 201 is schematically represented by an ellipse, but its actual shape depends on the implementation, the number of loudspeakers, etc. A listening zone is a sound diffusion zone in which a listener benefits from sound meeting desired criteria, for example quality criteria. Such a criterion is, for example, that the phase shift between the sounds produced by the different loudspeakers is below a certain threshold, so that the stereophonic or spatial perception is of a desired level of quality. The geometric parameters defining a listening zone include, for example, the distance of a listening zone from the audio sources, or the width of the zone. The audio processing receives at least these geometric parameters as input and generates corresponding audio signals.
[0019] According to one or more exemplary embodiments, a graphical representation of a listening area is displayed on the screen, superimposed with a video image obtained by the camera 104 of the room in which the device 100 is located. This visual feedback helps the user in real time with the configuration of the premises and / or the choice of a listening area. This choice may for example include the selection of a listening area from a list, the selection of one or more parameters defining a listening area or any other action or series of actions resulting in a determination of a listening area by the user. The configuration of the premises includes for example the movement of furniture so that a listener can easily position themselves in the listening area. In the specific context of the example of the figure 2 , this can be achieved by placing for example the sofa 201 in the listening area. The choice of a listening area makes it possible to obtain a new listening area, for example more or less wide or more or less distant from the device 100. This choice of a new listening area is followed by a configuration of the audio processing carried out by the device 100 so that the characteristics of the sound produced correspond to the new listening area. The display is updated on the fly according to the user's choice.
[0020] There figure 3 is a flowchart of a configuration method 300 according to one or more exemplary embodiments. In a first step (301), a signal representative of the video of the room in which the audio system is broadcasting is obtained from the camera. In a second step (302), a graphical representation of the currently configured listening area is added as an overlay on this video. A corresponding video signal is generated for display on the screen 101.
[0021] The user 102 may then decide to change the listening area. Various ways of making this choice may then be envisaged, some of which are illustrated in the figures 5 à 10 described further.
[0022] The audio processing is adapted to the new listening area and the new listening area is displayed as an overlay at 302, allowing the user to both appreciate the new sound processing and visualize the new listening area.
[0023] According to an alternative embodiment, the user is offered the option of selecting a listening area for overlay display, but only actually configuring the audio processing following additional validation of a selected listening area. This alternative allows the user to view a listening area (or several listening areas consecutively) without modifying the audio processing configuration parameters.
[0024] If the user decides not to - or no longer to - change the current listening zone in 303, he is asked to exit the configuration mode in 305. If this is the case, the configuration mode is stopped in 306. If not, the process loops back to 303.
[0025] There figure 4 is a schematic representation of a video image that may be displayed with, superimposed thereon, a graphical representation of a listening area 401.
[0026] THE figures 5 à 10 are schematic representations of video images illustrating the display on the screen at various stages of the process. Some of these figures illustrate in particular particular implementations for the configuration or selection of a listening zone. Some implementations use menus, it being noted that these implementations are given only as examples and that other ways of configuring a listening zone can be implemented, in particular by explicitly entering parameter values for a desired listening zone.
[0027] There figure 5 illustrates an initial positioning of the listening area visually delineated by the graphical representation 401, with a sofa 202 placed to the left of the room, outside the current listening area. The user can then either choose another listening area and / or move their sofa. The overlay video displayed on the screen provides immediate, real-time visual feedback helping the user to correctly move their furniture to a location in the room within the current listening area.
[0028] According to one or more embodiments, the user can choose a listening area by configuring a distance between the device 100 and the listening area. figure 5 illustrates a possible example of choices offered at 501, with a limited number of predefined choices presented on the screen as individually selectable menu items. In the figure 5 , three choices are displayed, namely a distance of less than three meters, a distance between 3 and 5 meters and a distance of more than 5 meters. Note that the distance is defined here by a range, but it is quite possible to define a distance by a single value. The distance range indicates for example the distance frame for which the audio processing input parameters will not be modified as long as the user's distance remains within the frame. In the example of the figure 5 , the intermediate distance is selected, the corresponding element is visually highlighted, and the corresponding graphical representation of the listening area is displayed superimposed on the video image. The user can either change the location of his furniture or vary the listening area. The user can, for example, align his sofa with the displayed listening area and / or indicate at what distance from the device he wishes to place the listening area. In this context, the figure 6 illustrates the alignment. Then the user chooses the distance corresponding to the location (the furthest distance in the example). This last case is illustrated by the figure 7 .
[0029] According to one or more embodiments, the user can also configure the width of the listening area. figure 8 illustrates an example of width choices, with a limited number of predefined choices presented on screen as individually selectable menu items. In the figure 8 , three choices are displayed, namely reduced width, standard width and extended width. In the example of the figure 8 , the standard width has been selected and the corresponding element is visually highlighted. figures 9 And 10 respectively represent the video image with the graphical representations of reduced width and extended width areas.
[0030] According to one or more embodiments, the orientation of the sound source(s) may be modified. A listening area is associated with each orientation of the sound source(s), as well as a corresponding graphical representation.
[0031] According to a particular embodiment, the device 100 then allows the user to preview the graphic representation of a listening zone for a particular orientation.
[0032] According to another particular embodiment, a user first modifies the orientation of the sound sources, and the device 100 then generates a display, a graphical representation of the listening area corresponding to this orientation. The device determines the orientation of the sources either automatically using suitable sensors or based on information provided by the user. The use of sensors also makes it possible to verify that the orientations of several audio sources correspond to an authorized configuration and thus to warn the user accordingly. For example, the user may have oriented two symmetrical sources with non-symmetrical orientations.
[0033] According to another particular embodiment that can be combined with one of the two embodiments above, the orientation of the audio source(s) is adjustable by one or more actuators. An actuator can be controlled by the processor 105 according to a command entered by the user. This command is for example a choice of a particular orientation. The adjustment of the orientation can be carried out as appropriate once the user is satisfied with a previewed listening area, or immediately according to a viewed listening area.
[0034] According to one or more embodiments, the orientation of the audio sources is a configuration parameter for processing the audio data by the audio processor 107.
[0035] There figure 11 is a schematic diagram of a portion of the device 100 in the case where this device comprises two audio sources which, combined, generate stereo sound. The figure 11 illustrates the left part of the device, the audio source can take three predefined orientations. The default orientation is for example an angle α of 45° relative to an axis parallel to the axis of symmetry of the device, with the other two positions varying for example by ± β° relative to the default position. Of course other positions can be chosen.
[0036] In the example of the figure 11 , the three possible orientations can be arbitrarily labeled, for example from 1 to 3, with orientation 2 being the default orientation. This numbering can optionally be reported on the housing of the device 100 for easy identification of a current orientation of the audio source(s).
[0037] According to other embodiments, the orientation can take more or less distinct values. According to still other embodiments, the orientation of an audio source is continuously adjustable.
[0038] There figure 12 is a schematic representation of a video image illustrating the display on the screen of three menu items having the orientations illustrated by the figure 11 A graphical representation of the listening area corresponding to the current orientation is displayed as an overlay. The corresponding element is visually highlighted.
[0039] Similarly, the figure 13 illustrates the case where the first orientation is selected, giving a wider listening area.
[0040] There figure 14 is a schematic diagram illustrating a method for determining the position of a point in the image as a function of a distance. This method can be applied to place a graphical representation at a given distance from the device 100.
[0041] The diagram of the figure 14 represents a side view of the device 100 placed on a piece of furniture 1401 just in front of the screen 101. The camera of the device 100 has a vertical viewing angle AdV_V, in degrees. It is placed at a height h1 from the ground. The video return plane 1402 of the camera has a height h in number of pixels. The ground represented by the bottom L1 of this plane is located at a distance d1 from the camera. The camera can be inclined at an angle 'inc' relative to the horizontal.
[0042] We also consider resV, the vertical resolution of the camera sensor in number of pixels, and ρ, the angle corresponding to a pixel, in degrees, and nbpx, the number of pixels between the bottom of the camera's video return plane, i.e. line L1, and a line L2 of this plane, line L2 corresponding in plane 1402 to a line A in the room and for which we wish to determine the distance d2 from the camera. nbpx therefore represents a number of pixels on a vertical line in the image, between the two lines L1 and L2. α and ß represent respectively the angle between the horizontal and the straight line passing through the camera and line L2, and the angle between the horizontal and the straight line passing through the camera and L1. Line L2 is the projection of line A onto plane 1402 along the straight line passing through the camera and line A.
[0043] There figure 15 is a schematic diagram showing the placement of lines L1 and L2 in the displayed image. Lines L2 and A overlap in this image.
[0044] We can ask: d 1 = h 1 tan β et d 2 = h 1 tan α β = AdV _ V 2 − inc ρ = AdV _ V resV β − α = nbpx × ρ α = AdV _ V 2 − inc − nbpx × AdV _ V resV d 1 = h 1 tan AdV _ V 2 − inc d 2 = h 1 tan AdV _ V 2 − inc − nbpx × AdV _ V resV
[0045] We thus obtain d2 for a given value of nbpx, and conversely, we can determine nbpx for a given value of d2. Thus, it is possible to place a graphical representation of a listening zone in the image according to the desired distance, and in particular between two values of the distance d2. For example, if we consider AdV_V=46°, inc=6°, resV=1080px, h1=0.8m and nbpx chosen at 270 pixels, then d1=2.6m and d2=8.3m.
[0046] Regarding the different widths of listening areas, we can simply define the width of a listening area as a fraction of the camera's viewing angle. Returning to the figure 2for example, an extended width ZE3 corresponds to ¾ of the camera's viewing angle AdV_H, a standard width ZE2 corresponds to 1 / 2 of this viewing angle, while a reduced width ZE1 corresponds to ¼ of this viewing angle. If the horizontal resolution of the camera sensor is resH, then it is sufficient to apply a coefficient of ¾, ½, or ¼ to obtain the width in pixels of the listening area.
Claims
1. A method of configuring (300) an audio system implemented by a device (100) comprising a processor (105), the method comprising - configuring (304) an audio processing processor (107) for generating one or more audio signals feeding one or more respective audio sources (108, 109) as a function of one or more geometric parameters defining a listening area chosen by a user, the method further comprising: - obtaining (301), from a camera, a video of a location where sound is broadcast by the audio source(s) (108, 109); - generating a signal representative of a video to be displayed on a display device (101), the video to be displayed comprising all or part of the video obtained by the camera;- the overprinting (302), in the video to be displayed, of at least one graphic representation (401) of a listening zone among several listening zones, the video with overprinting displayed on the screen constituting a visual feedback in real time adapted to help the user in the configuration of the place where the listening zone is located; - the obtaining (303) of data representative of a choice, by the user, of a displayed listening zone, for the configuration (304) of the audio processing processor (107).; 2. Method according to claim 1, the geometric parameter(s) comprising at least one of: - a distance between the listening area and the audio source(s); - a width of the listening area.
3. Method according to one of claims 1 or 2, in which a listening zone is an area in which the sound produced by the audio source(s) meets one or more quality criteria.
4. Method according to one of claims 1 to 3, comprising the overprinting, in the video to be displayed, of a menu (501) comprising several choices of value of a parameter, the highlighting of a current choice of parameter value, as well as the graphic representation (401) of a listening zone corresponding only to a current choice of parameter value.
5. Method according to one of claims 1 to 4, comprising displaying an invitation to a user to move a piece of furniture in the broadcast location based on the graphic representation.
6. Method according to one of claims 1 to 5, in which the audio source(s) are orientable, the orientation of the audio source(s) comprising one of: - the display of a message inviting a user to orient one or more audio sources according to a chosen listening area; - the generation of a control signal for a motor for orienting one or more audio sources according to a chosen listening area.
7. Method according to one of claims 1 to 6, a geometric parameter being defined by a value, or a range of values, or an open range.
8. Device (100) comprising a camera (104), at least one audio source (108, 109), a processor (105) and a memory (106) comprising software code, the processor, when it executes the software code, causing the device to implement a method according to one of claims 1 to 7.
9. Device according to claim 8, the at least one audio source being adjustable either manually or by means of an actuator.
Citation Information
Patent Citations
Loudspeaker System
US20150373452A1
Vehicle Audio System Interface
US20140096003A1
Audio system with configurable zones
US20170374465A1
System and method for differentially locating and modifying audio sources
US20180088900A1
Information processing device, information processing method, and information processing system
US20210266692A1