Method and device for configuring an audio system

The method and device utilize a camera to overlay graphic representations of listening areas onto a live video feed, addressing the complexity of existing audio system configuration methods by enabling intuitive and equipment-free zone selection and adjustment.

FR3156221A1Pending Publication Date: 2025-06-06SAGEMCOM BROADBAND SAS
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
FR2023013491
Authority / Receiving Office
FR · FR
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-12-04
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

Existing audio system configuration methods require additional equipment such as microphones and complex software processing, and are cumbersome when switching between multiple listening areas.

Method used

A method and device using a camera to visually select and configure a listening area by overlaying graphic representations of potential zones onto a live video feed, allowing users to choose and adjust listening areas intuitively.

Benefits of technology

Enables simple and intuitive configuration of listening areas without the need for additional equipment, providing real-time visual feedback and efficient switching between multiple zones.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

A method (300) for configuring an audio system is described, comprising: obtaining (301), from a camera, an image of a place where sound is broadcast by one or more audio sources, each audio source being configured to be supplied by a respective audio signal; generating a signal representative of an image to be displayed on a display device; superimposing (302), in the image to be displayed, at least one graphical representation (401) of a listening area from among several listening areas, each listening area being a function of one or more geometric parameters; obtaining (303) data representative of a choice of a listening area displayed by a user; configuring (304) an audio processing processor for generating the respective audio signal(s) as a function of the parameter(s) of the chosen listening area. A corresponding device is also described. Figure for the abstract: 3
Need to check novelty before this filing date? Find Prior Art

Description

Title of the invention: Method and device for configuring an audio system Technical field

[0001] A method and device for configuring an audio system are described, using a camera for selecting a listening area. The method and device can be used in particular for configuring a home audio system. Technical background

[0002] An audio system, for example a home audio system, may be associated with a listening area, which represents a sound diffusion location in which a listener benefits from sound meeting desired criteria. Determining such an area may require the implementation of various additional equipment. For example, it may be necessary to implement a microphone to make sound recordings in the room in which the audio system is placed, or to use a mobile phone. The recorded sound must then be processed by dedicated software to identify the listening area. Other solutions may require in-depth knowledge of acoustics and sound systems. The problem increases if several listening areas are envisaged, these listening areas corresponding for example to different configurations of the sound processing by the audio system.Switching from one zone to another requires performing the operations mentioned above again.

[0003] An intuitive audio system configuration solution is proposed, allowing simple determination of the listening zone(s). Summary

[0004] One or more embodiments relate to a method of configuring an audio system implemented by a device comprising a processor, the method comprising: - obtaining, from a camera, an image of a place where sound is broadcast by one or more audio sources, each audio source being configured to be supplied by a respective audio signal; - the generation of a signal representative of an image to be displayed on a display device, the image to be displayed comprising all or part of the image obtained by the camera; - the overprinting, in the image to be displayed, of at least one graphic representation of a listening zone among several listening zones, each listening zone being a function of one or more geometric parameters; - obtaining data representative of a choice of a listening area displayed by a user; - the configuration of an audio processing processor for the generation of the respective audio signal(s) according to the parameter(s) of the chosen listening area.

[0005] The visual feedback obtained in real time using the camera helps the user with the configuration of the premises and / or the choice of a listening area and its geometric parameters.

[0006] According to one or more exemplary embodiments, the geometric parameter(s) comprising at least one of: - a distance between the listening area and the device; - a width of the listening area.

[0007] According to one or more exemplary embodiments, a listening zone is an area in which the sound produced by the audio source(s) meets one or more quality criteria.

[0008] According to one or more exemplary embodiments, the method comprises the overprinting, in the image to be displayed, of a menu comprising several choices of value of a parameter, the highlighting of a current choice of parameter value, as well as the graphical representation of a listening zone corresponding only to a current choice of parameter value.

[0009] According to one or more exemplary embodiments, the method comprises displaying an invitation to a user to move a piece of furniture in the broadcast location based on the graphical representation.

[0010] According to one or more exemplary embodiments, the audio source(s) are orientable, the orientation of the audio source(s) comprising one of: - the display of a message inviting a user to orient one or more audio sources according to a chosen listening area; - the generation of a controller signal from a motor for orienting one or more audio sources according to a chosen listening area.

[0011] According to one or more exemplary embodiments, a geometric parameter is defined by a value, or a range of values, or an open range.

[0012] One or more exemplary embodiments relate to a device comprising a camera, at least one audio source, a processor and a memory comprising software code, the processor, when it executes the software code, causing the device to implement one of the methods described.

[0013] According to one or more exemplary embodiments, the at least one audio source can be oriented either manually or by means of an actuator.

[0014] Other characteristics and advantages will become apparent when reading the detailed description which follows, for the understanding of which reference will be made to the attached drawings among which:

[0015] [Fig-1] - [Fig.l] is a schematic diagram of an audio system according to one or several embodiments;

[0016] [Fig.2] - [Fig.2] is a schematic diagram of a top view of a part in which the system of [Fig.l] is placed;

[0017] [Fig.3] - [Fig.3] is a flowchart of a configuration method according to one or several embodiments;

[0018] [Fig.4] - [Fig.4] is a schematic representation of a video image that can be displayed with, superimposed, a graphic representation of a listening area;

[0019] [Fig.5] - [Fig.5] is a first schematic representation of a video image illustrating the screen display at various stages;

[0020] [Fig.6] - [Fig.6] is a second schematic representation of a video image illustrating the screen display at various stages;

[0021] [Fig.7] - [Fig.7] is a third schematic representation of a video image illustrating the screen display at various stages;

[0022] [Fig.8] - [Fig.8] is a fourth schematic representation of a video image illustrating the screen display at various stages;

[0023] [Fig.9] - [Fig.9] is a fifth schematic representation of a video image illustrating the screen display at various stages;

[0024] [Fig. 10] - [Fig. 10] is a sixth schematic representation of a video image illustrating the screen display at various stages;

[0025] [Fig. 11] - [Fig.l 1] is a schematic diagram of a part of the device according to one or more embodiments illustrating examples of possible orientations of an audio source;

[0026] [Fig. 12] - [Fig. 12] is a schematic representation of a video image illustrating the display on the screen of the menu items having the orientations illustrated by [Fig.l 1], as well as the selection of a first orientation;

[0027] [Fig. 13] - [Fig. 13] is a schematic representation of a video image illustrating the display on the screen of menu items having the orientations illustrated by [Fig.l 1], as well as the selection of a second orientation;

[0028] [Fig. 14] - [Fig. 14] is a schematic diagram illustrating a method of determining the position of a point in the image as a function of a distance;

[0029] [Fig. 15] - [Fig. 15] is a schematic diagram showing the positioning of certain elements of [Fig. 14] in the displayed video image.

[0030] In the following description, identical, similar or analogous elements will be designated by the same reference numerals. The block diagrams, algorithms and message sequence diagrams in the figures illustrate the architecture, functionalities and operation of systems, devices, methods and computer program products according to one or more exemplary embodiments. Each block of a block diagram or each phase of an algorithm may represent a module or a portion of software code comprising instructions for implementing one or more functions. According to certain implementations, the order of the blocks or phases may be changed, or the corresponding functions may be implemented in parallel.The process blocks or phases may be implemented using circuitry, software, or a combination of circuitry and software, in a centralized manner or in a distributed manner for all or some of the blocks or phases. The systems, devices, processes, and methods described may be modified, added to, and / or deleted within the scope of this disclosure. For example, the components of a device or system may be integrated or separated. Also, the described functions may be implemented using more or fewer components or phases, or with other components or through other phases. Any suitable data processing system may be used for the implementation. For example, a suitable data processing system or device includes a combination of software code and circuitry, such as a processor, controller, or other circuitry suitable for executing the software code.When the software code is executed, the processor or controller causes the system or device to implement all or part of the functionalities of the blocks and / or phases of the processes or methods according to the exemplary embodiments. The software code may be stored in non-volatile memory or a readable non-volatile storage medium (USB key, memory card or other medium) directly or through an interface adapted by the processor or controller.

[0031] [Fig.l] is a schematic diagram of an audio system illustrating one or more embodiments in a non-limiting manner. The system of [Fig.l] comprises a configuration device 100 and a screen 101. The device 100 can be controlled by a user 102 using a remote control 103. The device 100 comprises a camera 104, a processor 105, a non-volatile memory 106 comprising software code, and an audio processing processor 107. The various components of the device 100 are connected by an internal bus 110. The audio processing processor 107 receives audio data as input and generates audio signals intended to be reproduced by one or more audio sources, represented here by two loudspeakers 108 and 109. The device 100 further comprises an interface (not illustrated) by which it is connected to the screen 101. This interface is for example a HD MI interface. The device 100 is adapted to generate a video signal for display on the screen 101. The generation of the video signal is for example carried out by the processor 104. The system of [Fig.l] is given for illustrative purposes for the clarity of presentation of the exemplary embodiments and a current implementation may obviously differ. Furthermore, the device 100 may comprise a single speaker or more than two speakers. Depending on the implementations, the speakers 108 and 109 are fixed or orientable. In the case of orientable speakers, the orientation may be carried out manually or by means of actuators such as electric motors. According to certain implementations, position sensors are associated with the manually orientable speakers to allow the device 100 to determine the orientation.

[0032] The device 100 integrates for example a video receiver and decoder functionality in addition to the audio and camera functionalities (product designated under the name 'video and sound box' or 'video sound box' in English). It should also be noted that both the functionality of the processor 105 and that of the processor 107 can be implemented with more than one component, or even be jointly implemented by one or more components. For example, a single processor can be used to implement both functionalities. In the example of [Fig.l], the device 100 comprises two audio sources and the camera is placed on an axis of symmetry 111 of these audio sources. This is not, however, obligatory.

[0033] [Fig. 2] a schematic diagram illustrating an example of placement of the system of [Fig. 1] in a room 200. The camera 104 of the device 100 has a horizontal viewing angle AdV_H. According to the example of [Fig. 2], the room also includes furniture such as a sofa 202. A listening zone 201 is schematically represented by an ellipse, but its actual shape depends on the implementation, the number of loudspeakers, etc. A listening zone is a sound diffusion zone in which a listener benefits from sound meeting desired criteria, for example quality criteria. Such a criterion is, for example, that the phase shift between the sounds produced by the different loudspeakers is below a certain threshold, so that the stereophonic or spatial perception is of a desired level of quality.Geometric parameters defining a listening area include, for example, the distance of a listening area from audio sources, or the width of the area. Audio processing receives at least these geometric parameters as input and generates corresponding audio signals.

[0034] According to one or more exemplary embodiments, a graphic representation of a listening area is displayed on the screen, superimposed on a video image obtained by the camera 104 of the room in which the device 100 is located. This visual feedback helps the user in real time with the configuration of the premises and / or the choice of an area. listening area. This choice may, for example, include selecting a listening area from a list, selecting one or more parameters defining a listening area, or any other action or series of actions resulting in the user determining a listening area. The configuration of the premises includes, for example, moving furniture so that a listener can easily position themselves in the listening area. In the specific context of the example of [Fig.2], this can be achieved by placing, for example, the sofa 201 in the listening area. The choice of a listening area makes it possible to obtain a new listening area, for example more or less wide or more or less distant from the device 100. This choice of a new listening area is followed by a configuration of the audio processing carried out by the device 100 so that the characteristics of the sound produced correspond to the new listening area. The display is updated on the fly according to the user's choice.

[0035] [Fig.3] is a flowchart of a method 300 of configuration according to one or several exemplary embodiments. In a first step (301), a signal representative of the video of the room in which the audio system is broadcasting is obtained from the camera. In a second step (302), a graphic representation of the currently configured listening area is added as an overlay on this video. A corresponding video signal is generated for display on the screen 101.

[0036] The user 102 can then decide to change listening zone. Various ways of making this choice can then be envisaged, some of which are illustrated in figures 5 to 10 described later.

[0037] The audio processing is adapted to the new listening area and the new listening area is displayed as an overlay at 302, which allows the user both to appreciate the new sound processing and to visualize the new listening area.

[0038] According to an alternative embodiment, the user is offered the option of selecting a listening area for the overlay display, but of actually configuring the audio processing only following additional validation of a selected listening area. This alternative allows the user to view a listening area (or several listening areas consecutively) without modifying the audio processing configuration parameters.

[0039] In the case where the user decides not to - or no longer to - change the current listening zone in 303, he is offered to leave the configuration mode in 305. If this is the case, the configuration mode is stopped in 306. If not, the process loops back to 303.

[0040] [Fig.4] is a schematic representation of a video image that can be displayed with, superimposed, a graphic representation of a listening area 401.

[0041] Figures 5 to 10 are schematic representations of video images illustrating the display on the screen at various stages of the process. Some of these figures illustrate including specific implementations for configuring or choosing a listening area. Some implementations use menus, it being noted that these implementations are given only as examples and that other ways of configuring a listening area can be implemented, including explicitly entering parameter values ​​for a desired listening area.

[0042] [Fig.5] illustrates an initial positioning of the listening area visually delimited by the graphical representation 401, with a sofa 202 placed to the left of the room, outside the current listening area. The user can then either choose another listening area and / or move his sofa. The overlay video displayed on the screen constitutes immediate and real-time visual feedback helping the user to correctly move his furniture to a location in the room located within the current listening area.

[0043] According to one or more embodiments, the user can choose a listening area by configuring a distance between the device 100 and the listening area. [Fig. 5] illustrates a possible example of choices proposed at 501, with a limited number of predefined choices presented on the screen in the form of individually selectable menu items. In [Fig. 5], three choices are displayed, namely a distance of less than three meters, a distance between 3 and 5 meters and a distance of more than 5 meters. Note that the distance is here defined by a range, but it is entirely possible to define a distance by a single value. The distance range indicates for example the distance frame for which the input parameters of the audio processing will not be modified as long as the user's distance remains within the frame. In the example of [Fig.5], the intermediate distance is selected, the corresponding element is visually highlighted, and the corresponding graphic representation of the listening area is displayed superimposed on the video image.

[0044] The user can either change the location of his furniture or vary the listening area.

[0045] The user can, for example, align his sofa with the displayed listening area and / or indicate at what distance from the device he wishes to place the listening area. In this context, [Fig.6] illustrates the alignment. Then the user chooses the distance corresponding to the location (the furthest distance in the example). This latter case is illustrated by [Fig.7].

[0046] According to one or more embodiments, the user can also configure the width of the listening area. [Fig. 8] illustrates an example of width choices, with a limited number of predefined choices presented on the screen as individually selectable menu items. In [Fig. 8], three choices are displayed, namely a reduced width, a standard width and an extended width. In the example In [Fig.8], the standard width has been selected and the corresponding element is visually highlighted. Figures 9 and 10 show the video image with the graphical representations of reduced width and extended width areas respectively.

[0047] According to one or more embodiments, the orientation of the sound source(s) may be modified. A listening area is associated with each orientation of the sound source(s), as well as a corresponding graphical representation.

[0048] According to a particular embodiment, the device 100 then allows the user to preview the graphic representation of a listening zone for a particular orientation.

[0049] According to another particular embodiment, a user first modifies the orientation of the sound sources, and the device 100 then generates a display, a graphical representation of the listening area corresponding to this orientation. The device determines the orientation of the sources either automatically using suitable sensors or based on information provided by the user. The use of sensors also makes it possible to verify that the orientations of several audio sources correspond to an authorized configuration and thus to warn the user accordingly. For example, the user may have oriented two symmetrical sources with non-symmetrical orientations.

[0050] According to another particular embodiment that can be combined with one of the two embodiments above, the orientation of the audio source(s) is adjustable by one or more actuators. An actuator can be controlled by the processor 105 according to a command entered by the user. This command is for example a choice of a particular orientation. The adjustment of the orientation can be carried out as appropriate once the user is satisfied with a previewed listening area, or immediately according to a viewed listening area.

[0051] According to one or more embodiments, the orientation of the audio sources is a configuration parameter for processing the audio data by the audio processor 107.

[0052] [Fig. 11] is a schematic diagram of a part of the device 100 in the case where this device comprises two audio sources which, combined, generate stereo sound. [Fig.l 1] illustrates the left part of the device, the audio source being able to take three predefined orientations. The default orientation is for example an angle α of 45° relative to an axis parallel to the axis of symmetry of the device, with the other two positions varying for example by ± B° relative to the default position. Of course other positions can be chosen.

[0053] In the example of [Fig. 11], the three possible orientations can be arbitrarily labeled, for example from 1 to 3, with orientation 2 being the default orientation. This numbering may optionally be reported on the housing of the device 100 for easy identification of a current orientation of the audio source(s).

[0054] According to other embodiments, the orientation can take more or less distinct values. According to still other embodiments, the orientation of an audio source is continuously adjustable.

[0055] [Fig. 12] is a schematic representation of a video image illustrating the display on the screen of three menu items having the orientations illustrated by [Fig. 11]. A graphical representation of the listening area corresponding to the current orientation is displayed as an overlay. The corresponding item is visually highlighted.

[0056] Similarly, [Fig. 13] illustrates the case where the first orientation is selected, giving a wider listening area.

[0057] [Fig. 14] is a schematic diagram illustrating a method for determining the position of a point in the image as a function of a distance. This method can be applied to place a graphical representation at a given distance from the device 100.

[0058] The diagram in [Fig. 14] represents a side view of the device 100 placed on a piece of furniture 1401 just in front of the screen 101. The camera of the device 100 has a vertical viewing angle AdV_V, in degrees. It is placed at a height hl from the ground. The video return plane 1402 of the camera has a height h in number of pixels. The ground represented by the bottom L1 of this plane is located at a distance dl from the camera. The camera can be inclined at an angle 'inc' relative to the horizontal.

[0059] We also consider resV, the vertical resolution of the camera sensor in number of pixels, and p, the angle corresponding to a pixel, in degrees, and nbpx, the number of pixels between the bottom of the video return plane of the camera, i.e. line L1, and a line L2 of this plane, line L2 corresponding in plane 1402 to a line A in the room and for which we wish to determine the distance d2 relative to the camera, nbpx therefore represents a number of pixels on a vertical line in the image, between the two lines L1 and L2. a and B respectively represent the angle between the horizontal and the straight line passing through the camera and line L2, and the angle between the horizontal and the straight line passing through the camera and LL. Line L2 is the projection of line A onto plane 1402 along the straight line passing through the camera and line A.

[0060] [Fig. 15] is a schematic diagram showing the placement of lines L1 and L2 in the displayed image. Lines L2 and A are superimposed in this image.

[0061] We can ask:

[0062] [Equation 1]

[0064] [Equation 2]

[0065]

[0066] [Equation 3]

[0068] [Equation 4]

[0069] fi-a- nbpx xp

[0070] [Equation 5]

[0071] a - _ jnc _ nbpx x

[0072] [Equation 6]

[0073] d 1 =----------- tan ( ———inc J

[0074] [Equation 7]

[0075] d2 =.............----nrr- tan('1 2~; -inc-nbpxxA'reiÿ ]

[0076] We thus obtain d2 for a determined value of nbpx, and conversely, we can determine nbpx for a determined value of d2. Thus, it is possible to place a graphical representation of a listening zone in the image as a function of the desired distance, and in particular between two values ​​of the distance d2. For example, if we consider AdV_V=46°, inc=6°, resV=1080px, hl=0.8m and nbpx chosen at 270 pixels, then dl=2.6m and d2=8.3m.

[0077] Regarding the different widths of the listening areas, we can simply define the width of a listening area as a fraction of the camera's viewing angle. Returning to [Fig.2] for example, an extended width ZE3 corresponds to ¾ of the camera's viewing angle AdV_H, a standard width ZE2 corresponds to 1 / 2 of this viewing angle, while a reduced width ZE1 corresponds to ¾ of this viewing angle. If the horizontal resolution of the camera sensor is resH, then it is sufficient to apply a coefficient of ¾, / 2, or to obtain the width in pixels of the listening area.

Claims

Claims

1. Method for configuring (300) an audio system implemented by a device (100) comprising a processor (105), the method comprising: - obtaining (301), from a camera, an image of a place where sound is broadcast by one or more audio sources (108, 109), each audio source being configured to be supplied by a respective audio signal; - generating a signal representative of an image to be displayed on a display device (101), the image to be displayed comprising all or part of the image obtained by the camera; - superimposing (302), in the image to be displayed, at least one graphical representation (401) of a listening area among several listening areas, each listening area being a function of one or more geometric parameters; - obtaining (303) data representative of a choice of a listening area displayed by a user;- configuring (304) an audio processing processor (107) for generating the respective audio signal(s) as a function of the parameter(s) of the selected listening area;

2. Method according to claim 1, the geometric parameter(s) comprising at least one of: - a distance between the listening area and the device; - a width of the listening area.

3. A method according to either of claims 1 or 2, wherein a listening area is an area in which the sound produced by the audio source(s) meets one or more quality criteria.

4. Method according to one of claims 1 to 3, comprising the overprinting, in the image to be displayed, of a menu (501) comprising several choices of value of a parameter, the highlighting of a current choice of parameter value, as well as the graphical representation (401) of a listening zone corresponding only to a current choice of parameter value.

5. A method according to one of claims 1 to 4, comprising displaying of an invitation to a user to move a piece of furniture in the broadcast location based on the graphic representation.

6. Method according to one of claims 1 to 5, in which the audio source(s) are orientable, the orientation of the audio source(s) comprising one of: - the display of a message inviting a user to orient one or more audio sources according to a chosen listening area; - the generation of a controller signal of a motor for orienting one or more audio sources according to a chosen listening area.

7. Method according to one of claims 1 to 6, a geometric parameter being defined by a value, or a range of values, or an open range.

8. Device (100) comprising a camera (104), at least one audio source (108, 109), a processor (105) and a memory (106) comprising software code, the processor, when executing the software code, causing the device to implement a method according to one of claims 1 to 7.

9. Device according to claim 8, the at least one audio source being orientable either manually or by means of an actuator.

Citation Information

Patent Citations

  • Vehicle Audio System Interface

    US20140096003A1

  • Loudspeaker System

    US20150373452A1

  • Information processing device, information processing method, and information processing system

    US20210266692A1

  • Angular sensing for optimizing speaker listening experience

    US20210385604A1

  • Loudspeaker system and control

    US20230239646A1