FUSION OF MULTIMODAL MULTI-USER INTERACTIONS

DE602023013947T2Active Publication Date: 2026-03-25ORANGE SA
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
DE · DE
Patent Type
Patents
Current Assignee / Owner
Filing Date
2023-05-17
Publication Date
2026-03-25

AI Technical Summary

Technical Problem

Current digital services struggle to manage multimodal interactions from multiple users simultaneously, particularly when users are distributed across different sites, leading to complex and inefficient command processing.

Method used

A method and system for generating an enriched request based on the interpretation of multimodal human-machine interactions from multiple users, utilizing user identification, contextual information, and temporal coordination to synchronize and merge signals from different users, enabling the formulation of explicit and inclusive commands.

Benefits of technology

Enhances usability by effectively managing multimodal interactions from multiple users, allowing natural and spontaneous behavior, and facilitating coordinated actions across distributed environments.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader
Need to check novelty before this filing date? Find Prior Art

Description

Domaine technique

[0001] This disclosure falls within the domain of human-computer interaction.

[0002] More specifically, this disclosure relates to a process for handling a request from a digital service, a computer program including instructions for executing such a process, and a terminal for implementing such a process. Technique antérieure

[0003] Multimodality in current digital services allows users to interact with a service using one or more modalities from among those supported by the service. For example, for incoming interactions, the user can choose touch or voice to submit their requests.

[0004] In more advanced digital services, processes for merging interaction modalities allow several interaction modalities to cooperate at the same time for the same and unique transfer of information, i.e. a user command, and this by taking into account different types of cooperation of modalities, such as complementarity, redundancy, or competition.

[0005] These interaction methods can be a gesture, pressing keys, voice, gaze, ... It is known, in current digital services, to interpret certain methods, and to decide whether to merge them or not in order to deduce a single whole command, or several separate commands, to be executed.

[0006] For example, as illustrated on the figure 1 A user can give a voice command 100 to turn on a lamp, and simultaneously give a gesture command 110 by extending their arm towards the living room lamp. An incoming interaction interpretation module 1, within the digital service, merges these two interactions, received from two input interfaces 10 and 11, to determine a single action: to turn on the living room lamp alone. This action is expressed as a command 900 issued through an output interface 90.

[0007] Thus, as illustrated on the figure 2 It is known to provide an incoming interaction interpretation module 1 which allows the collection of different modalities 100, 110, 120, 130 from a single user and received for example by a plurality of input interfaces 10, 11, 12, 13. The incoming interaction interpretation module 1 is capable of combining certain incoming interactions and deciding on distinct actions to be implemented by issuing independent commands 900, 910 through output interfaces 90, 91.

[0008] As illustrated on the figure 3 The analysis of direct incoming interactions can, in certain digital services, be enriched with contextual information from previous direct or indirect interactions. An indirect interaction is defined as the detection by a digital service of information related to the user without the user being aware of this interaction: for example, detection of presence in a room, detection of a leg in a cast, noise level, brightness level, etc.

[0009] The processing of cooperation of incoming modalities from the same user with the fusion process is a rather complex operation because each individual is unique and therefore interacts differently, depending on the context, even when using the same interaction modalities for the same command.

[0010] Therefore, in current digital services, each direct incoming interaction, also called an intentional incoming interaction—that is, one emitted with the intention of triggering a command—is associated, in a simple case, with a single command. Even in more complex cases, where a request is formulated through a cooperation of incoming modalities, this cooperation is individual and local.

[0011] That is to say, as illustrated on the figure 4 , cooperation is associated with a single user, who is the author of all the modalities considered, and with a single place, which corresponds to the location of this user.

[0012] In other words, when initial incoming modalities are identified as originating from a first user and subsequent incoming modalities are identified as originating from a second user distinct from the first user, notably through an identification module (not shown) receiving the first and second incoming modalities, it is known to provide a first interpretation module 1 to interpret the incoming modalities originating from the first user and a second interpretation module 2 to interpret the incoming modalities originating from the second user, the two interpretation modules operating separately and independently. Each interpretation module is dedicated to interpreting queries expressed by a specific user and uses a given set of input and output interfaces for this purpose.

[0013] Thus, according to the figure 4 The first interpretation module 1 receives as input, from input interfaces 10, 11, 12, 13, different first incoming interactions 100, 110, 120, 130 with different first modalities, interprets them and outputs commands 900, 910 through output interfaces 90, 91. Similarly, the second interpretation module 2 receives as input, from input interfaces 20, 21, 22, 23, different second incoming interactions 200, 210, 220, 230 with different second modalities, interprets them and outputs commands 800, 810 through output interfaces 80, 81.

[0014] Prior art documents US 2018 / 046851 A1 and WO 2012 / 066557 A1 describe multi-modal, multi-user interactions.

[0015] However, there is a need for an interactive environment capable of interacting simply, naturally, and ergonomically, not with a single user at a time, but with multiple users collectively. It would also be desirable for such an interactive environment to be compatible with users being distributed across different sites. Résumé

[0016] This disclosure improves the situation.

[0017] A method is proposed for processing a request from a digital service in an interactive environment, the method comprising: generating an enriched request based on an interpretation of signals received from a multimodal set of human-machine interactions resulting from the coordinated actions of a plurality of distinct users in the interactive environment, each human-machine interaction, considered individually, being indicative only of a simple request, the enriched request triggering the emission of at least one command.

[0018] The proposed process improves the usability of the digital service because it effectively manages multimodal interactions from different users. This helps to encourage natural and spontaneous behavior among users of the digital service.

[0019] The expression "interactive environment" refers to a set of human-machine interfaces enabling interaction between the digital service and its users.

[0020] These human-machine interfaces include input and output interfaces. Input interfaces encompass, for example, a keyboard, a pointing device, and / or a touchscreen. Output interfaces can include, for example, a display or an audio output coupled with a speech synthesis module. A head-mounted display (HUD) is an example of a terminal comprising a set of motion sensors as an input interface and a display device as an output interface.

[0021] A camera coupled with an automatic gesture recognition module, or a voice-activated interface coupled with an automatic speech recognition module, are other examples of input interfaces. The interactive environment can thus be integrated, for example, into a video conferencing application using the audio and video input interfaces of meeting participants. In this context, the digital service can leverage the interactive environment to automatically interpret different types of signals corresponding to instructions jointly provided by the participants. These signals might, for example, indicate a gesture from the first participant, a spoken instruction from the second participant, a third participant's interaction with a button, and so on.

[0022] A so-called "simple" request is individual in that it can be determined from a single human-computer interaction originating from a single user. Conversely, a so-called "enriched" request is joint in that it can only be determined from a plurality of human-computer interactions originating collectively from distinct users. The proposed method includes detection to ensure that distinct users are the source of the received signals.

[0023] Considering, for example, a rich query designed to send an email to a group of users, an email subject line can be automatically suggested based on the analysis of a direct interaction, such as text or voice, from a first user, while the email content can be determined based on the analysis of another interaction from a second user. Thus, a rich query is constructed from two simple queries originating from two distinct users.

[0024] The process may further include an interpretation of the received signals, determining the enriched query based on at least one of the following: simple queries extracted from the received signals; the users who originated each received signal; or the result of detecting that distinct users originated the received signals. The interpretation can be implemented by a logical entity, referred to here as the analyzer or signal interpretation module. The analyzer may also include several elementary modules, each exclusively interpreting signals that have been previously associated with a user identified as the source of said signals by the aforementioned identification module. The analyzer allows a generator to generate one or more enriched queries. These enriched queries automatically trigger commands corresponding to them.

[0025] In one example, the process involves identifying, for each received signal, a user who originated the received signal.

[0026] Identification can be part of interpreting a received signal; for example, voice signals can be analyzed to find a person's voiceprint. Identification can also be based on a user identifier associated with the origin of a received signal, such as an IP address identifying a user individually connected to the digital service. Identification allows a given signal to be interpreted differently depending on the user identified as the source of that signal. For example, an enriched request can be formulated based on both an initial signal indicating a proposal to send a file to interested users and a series of second signals from distinct users, each expressing interest in the proposal. A series of "send file to X" commands can thus be issued, where X designates each identified user associated with at least one second signal.

[0027] Furthermore, when multiple received signals are each associated with an identified user, the interpretation of a combination of these received signals can be based on the identification of the originating users. Thus, the merging of incoming modalities can be performed differently depending on whether these incoming modalities originate from a single identified user or from several distinct identified users. In particular, consider an example where several related modalities from the same identified user, along with one or more additional modalities from one or more distinct identified users, contribute to the generation of an enriched query by a generator.To determine, in this example, that the modalities emanating from the same identified user are linked, it is possible, for example, to base oneself on a history of modalities emanating from this identified user and to extract from it a habit of temporal coordination constituting a characteristic specific to this identified user.

[0028] In one example, at least one of said signals includes a user identifier and generating the rich query includes at least one of the following steps: - generate the rich query based on the interpretation of the received signals and the associated user identifiers, - generate the rich query by adapting, based on the user identifiers, the rich query obtained by the interpretation of the received signals.

[0029] In one example, at least one of said signals contains contextual information, and the generation of the enriched query adapts the enriched query based on an interpretation of this contextual information.

[0030] Considering, for example, an enriched request to turn on a lamp from a set of lamps, the precise identification of the lamp to be turned on can be performed based on a low level of brightness detected in a given area. The low level of brightness is here an example of contextual information that implicitly complements an enriched request formulated by merging incoming modalities, and thus contributes to the interpretation of the received signals and the adaptation of the enriched request by the digital service.

[0031] For example, considering a user in a noisy environment, it might be possible to isolate the user's voice from the ambient noise so that only the user's own words are considered as potential voice commands, and not those of another person nearby who is not also a user of the digital service. The ambient noise level is an example of contextual information associated with an input interface receiving human-computer interaction. This element is interpreted, in this specific case, to differentiate the signals received by this input interface and thus, more generally, to adapt the enriched request(s) resulting from the coordination of incoming modalities.

[0032] Isolating sounds from each user can also allow, for example, finding multiple sounds from multiple users using the same sound sensor connected to the same input interface. In other words, it is possible to isolate several sounds within an incoming audio stream, each of which could contain a simple query, while simultaneously identifying the source of each isolated sound, in order to generate the enriched query.

[0033] Contextual information can be associated with a user and can refer, for example, to a preference or characteristic specific to that user, or even to an indirect interaction between the user and the digital service. Contextual information associated with an input interface can refer, for example, to a characteristic specific to a terminal and independent of the user; that is, common to the signals received by the input interface from that terminal, regardless of whether these signals originate from a single user or from several distinct users. Generally speaking, contextual information is digital data that can be used to aid in the interpretation of a direct interaction between the user and the digital service.

[0034] The digital data mentioned in the examples above, i.e. the user identifier and contextual information, are not necessarily included directly in one or more received signals but can more generally be obtained in any technically feasible way, whether by separate transmission or by determination, for example by interpretation of received signals, such determination being able to be implemented locally, or centrally, or even with the help of cloud computing means.

[0035] In one example, the proposed process further includes timestamping at least the start times of reception of said signals, interpreting time differences between the respective start times of reception of said signals, and adapting the enriched query according to the interpretation of these time differences.

[0036] The temporal coordination of linked modalities, such as a gesture instruction coupled with a voice instruction, varies from one user to another. These differences can relate both to the duration of each modality, depending on the user implementing it, and to the synchronization of the modalities. Despite these differences, it is possible to define a threshold below which two modalities, particularly two different modalities, at least in that they originate from distinct users, are considered linked, that is, relating to the generation of the same enriched request. As an alternative, it is possible to associate a time gap between the start of incoming interaction modalities with a probability level that these modalities are linked.Thus, the same final command can be obtained from a combination of an indicative signal of a gestural instruction with an indicative signal of a voice instruction, regardless of the user who issued each instruction.

[0037] For example, at least some of the received signals may originate from a plurality of input interfaces distributed across different sites.

[0038] Thus, the proposed process allows users distributed across different sites, each with at least one input interface, to act in a coordinated manner to formulate their intentions. The signals received by the input interfaces translate these intentions, which are associated with distinct users. These combined signals are used by the digital service to formulate an enriched request corresponding to the users' intentions and to issue the command corresponding to the formulated enriched request.

[0039] For example, the proposed process may also include: timestamp the timestamps of the start of reception of said signals, correct the respective start times of reception of said signals according to a latency between said sites, interpret corrected time differences between corrected time-stamped times, and adapt the enriched query according to the interpretation of the corrected time differences.

[0040] Thus, the proposed method takes into account latency in communication between different sites. By synchronizing the time reference systems of each user, the proposed method allows for better identification of coordinated modalities as being linked.

[0041] According to another aspect, a computer program is proposed containing instructions for implementing the proposed process, when said instructions are executed by a processor.

[0042] According to another aspect, a device for processing a request from a digital service in an interactive environment is proposed, the processing device comprising at least one generator of an enriched request based on an interpretation of signals received from a multimodal set of human-machine interactions resulting from the coordinated actions of a plurality of distinct users in the interactive environment, each human-machine interaction, considered individually, being indicative only of a simple request, the enriched request triggering the emission of at least one command.

[0043] The processing device may also include an analyzer capable of interpreting received signals to determine the enriched query based on at least one of the following elements: simple queries extracted from received signals; for each of the received signals, the users who originated the received signals.

[0044] The processing device may also include a user identifier capable of identifying, for each received signal, a user who was the originator of the received signal.

[0045] The processing device can be implemented, for example, in at least one of the following devices: a terminal, a connected object manager capable of controlling at least one object connected to a communication network by means of at least one command corresponding to the enriched request destined for at least one connected object.

[0046] For example, the processing device may include a processing circuit comprising a storage memory storing the proposed computer program and a processor connected to the storage memory and to one or more terminal interfaces.

[0047] Alternatively, the processing device may lack such a processing circuit, which is then remote and configured to exchange with the processing device via a communication channel. Brève description des dessins

[0048] Other features, details, and advantages will become apparent upon reading the detailed description below and analyzing the attached drawings, on which: Fig. 1 [ Fig. 1 ] illustrates a known processing of a request arising from actions emanating from a single user of a digital service, where two modes of interaction are merged to construct a command. Fig. 2 [ Fig. 2 ] illustrates a known variant of the treatment of the figure 1 , where several interaction modalities are handled to construct several commands. Fig. 3 [ Fig. 3 ] illustrates a known variant of the treatment of the figure 2 where contextual data is taken into account to interpret one or more incoming interactions. Fig. 4 [ Fig. 4 ] illustrates side by side two identical treatments known according to the figure 1 . The processes each concern a unique user and are carried out in a completely isolated manner from each other, each process relating to a distinct user arising solely from interactions which originate from that user and constructing a distinct command. Fig. 5 [ Fig. 5 ] illustrates a processing of requests resulting from coordinated actions of several users of a digital service, according to an example implementation. Fig. 6 [ Fig. 6 ] illustrates a variant of the treatment of the figure 5 where contextual data is taken into account to interpret one or more incoming interactions. Fig. 7 [ Fig. 7 ] illustrates the differences in coordination of related modalities for a given individual query depending on their author. Fig. 8 [ Fig. 8 ] illustrates an example of coordinating related modalities for a given collective request. Fig. 9 [ Fig. 9 ] illustrates a particular example of the processing of specific interaction modalities, according to the general principle of the figure 5 , to build a specific order. Fig. 10 [ Fig. 10 ] illustrates another specific example of the treatment of particular interaction modalities, still according to the general principle of the figure 5 , to build a specific order. Description des modes de réalisation

[0049] The principle of the present invention is to allow several users to interact together, each according to the modality or modalities of their choice, with a digital service.

[0050] To achieve this, a digital service manages one or more input interfaces capable of collecting signals relating to a group of users, specifically those who are subscribers to the digital service. These users may all be in the same location or, conversely, located far apart. The composition of the user group is managed by the digital service itself according to its own rules and contexts. In one example, it could be the residents of a house using a home automation service. In another example, it could be a group of people participating in a meeting using a video conferencing service.

[0051] Each input interface is configured to collect one or more signals relating to one or more users, so that the set of signals collected forms a multimodal set of interactions between distinct users and the digital service.

[0052] The digital service manager includes a detector designed to verify that the users originating the received signals are distinct. The detector may include a user identifier to identify the users originating the received signals. Such detection or identification can be performed either before or after the signals are received.

[0053] There are many possible ways to perform such detection or identification.

[0054] For example, it is possible to obtain data directly at the input interface that identifies the users responsible for different received signals. Such data is referred to in this document by the generic term "user identifier." This could include, for example, user account identifiers in cases where each user has their own account to access the digital service.

[0055] Alternatively, multiple received signals can be extracted, for example, from the same audio or video stream. Individual analysis of the extracted signals can be performed to identify the users who originated them. A comparative analysis of the extracted signals can also be performed using a specific algorithm to detect that the users originating the extracted signals are distinct, without necessarily identifying the users in question.

[0056] A physical location is another example of data that can be associated with a received signal and can be used, at least in some cases, for the same purpose of detecting that the users originating different signals are distinct, again without necessarily allowing a given user to be formally identified on its own.

[0057] Alternatively, it is possible to compare data identifying the users responsible for different received signals, and to obtain only the result of this comparison at an input interface of the digital service. This makes it possible to detect that different received signals originate not from a single user, but rather from distinct users, based solely on the comparison result, without requiring the digital service manager to identify the users.

[0058] The digital service is capable of interpreting the collected signals by detecting a cooperation of direct, successive, or simultaneous incoming interaction modalities from multiple users. In this way, the digital service can generate an enriched request from multiple users by merging interaction modalities, each containing a simple request that is a fragment of this enriched request. Once the collective request is generated, a corresponding complete command is issued by the digital service through an output interface.

[0059] In general, and according to the principle described, the present invention offers an enriched user experience through multi-user coordination capabilities, even remotely, mutual assistance, and potentially greater inclusivity for different user profiles. Thus, an enriched, or even explicit, incoming command, also called a complete or full incoming command, can be formulated using one or more interaction methods, expressed by one or more users. The incoming command is considered explicit if the merging of requests resolves one or more ambiguities related to one of the simple requests. For example, a voice command from a first user such as "turn on the light" might be ambiguous, but merging it with a request from a second user resulting from that user's gesture pointing at a lamp will produce a command to turn on the designated lamp, which is therefore explicit.However, if the lamp has a light dimmer, the command is not explicit since no simple request allows the expected light intensity level to be determined; it is therefore an enriched but not explicit command.

[0060] The applications of multimodal interaction between multiple users and a digital service are numerous, both in the public and professional spheres. For example, one possible application in the workplace is to trigger commands from several people in a meeting, whether in person or remotely. A possible application in the security sector is to enforce synchronized and multimodal interactions from multiple actors for a single command. A possible application in the disability sector is to assist multiple users in challenging situations or to enable them to cooperate in initiating a command together.

[0061] Reference is now being made to the figure 5 , which illustrates a logical architecture of a digital service manager according to an example of an embodiment of the invention.

[0062] The digital service manager is understood to designate at least one processing circuit comprising a processor, storage memory, at least one input interface, and at least one output interface. The digital service manager can be implemented, for example, at the terminal level and / or at the level of a connected object manager capable of controlling at least one object connected to a communication network by means of at least one command corresponding to the enriched request addressed to at least one connected object.

[0063] Computer program instructions are stored in storage memory. When read by the processor, these instructions cause the processor to implement a process for handling a request from a multimodal set of incoming human-machine interactions.

[0064] The digital service manager includes a generator (not shown) for an enriched request. This is a logic module capable of generating an enriched request based on an interpretation of signals received from a multimodal set of human-computer interactions resulting from the coordinated actions of multiple distinct users in the interactive environment. Each human-computer interaction, considered individually, is indicative only of a simple request. The enriched request triggers the transmission of one or more output commands through an output interface. Upstream of this generator, the digital service manager may also include an analyzer, or incoming interaction interpretation module. This is a logic module configured to process signals received by the input interface(s). These signals are indicative of the multimodal set of incoming human-computer interactions.

[0065] The input interface(s) contribute, from the point of view of a plurality of users, to forming an interactive environment.

[0066] This means that several users are each able to act on the operation of the digital service through at least one input interface.

[0067] For example, an interactive element might be displayed on a touchscreen, and it might be expected that detecting a user's touch on this interactive element would lead to the reception of a signal by an input interface. It might also be expected that a user will look in a specific direction to interact with the digital service in another way, and that detecting this gaze direction, by a suitable sensor, would lead to the reception of another signal by an input interface. Furthermore, it might be expected that a device worn by a user could be geolocated, and that a particular location of the device would trigger the reception of yet another signal by an input interface.

[0068] Generally, users may be equipped with terminals which can be, for example, a computer, a mobile phone, a watch, a television, etc. Optionally, a user may not have their own equipment, so that the same terminal can be used by several users, simultaneously or successively.

[0069] Each terminal can include one or more human-machine interfaces as a source of signals transmitted to an input interface associated with the analyzer 3. In this case, a terminal's human-machine interface is, from the user's perspective, considered a simple proxy for an input interface in the interactive environment. Depending on its human-machine interfaces and its use, a single terminal may be limited to acquiring only one type of signal indicating the actions of a single user, or it may, conversely, acquire several signals or types of signals indicating the actions of one or more users. When several users are the source of different received signals, these signals are interpreted to detect that the originating users are distinct. Such interpretation is particularly appropriate when these signals are received by the same human-machine interface.

[0070] At least one terminal can also be configured to receive and execute a command from an output interface associated with the generator.

[0071] The invention is not limited in the number or nature of the interfaces considered. It is only necessary that at least two signals be received, the first signal indicating a human-machine interaction with a first user according to a first modality, and the second signal indicating a human-machine interaction with a second user according to a second modality. Thus, using, for example, a single camera simultaneously recording two users, it is possible to receive an audio track containing a spoken instruction from the first user and a video track containing a movement from the second user. The audio and video tracks thus form a set of received signals that is both multimodal and related to several users.In this example, an analysis of the video track and / or the audio track can be implemented in order to identify the two users, particularly in the case where they do not interact with a connected terminal identifying them, or in order to confirm an identification of the two users in the opposite case.

[0072] For example, the figure 5 Several input interfaces 30, 31, 32, 33 each receive an indicative signal of an incoming human-machine interaction 300, 310, 320, 330 from the same first user. Other input interfaces 34, 35, 36, 37 each receive a signal of an incoming human-machine interaction 340, 350, 360, 370 from the same second user. For the sake of simplicity, it is assumed that the modalities of the different signals received are all different. Of course, the figure 5 is only illustrative, and in general several signals from the same or several users can be received on the same interfaces.

[0073] Analyzer 3 identifies, among the received signals, human-machine interactions resulting from coordinated actions of a plurality of users.

[0074] There are different ways to perform such an identification.

[0075] The identification of coordinated actions can be based exclusively on useful data contained in the received signals, that is, without taking into account any metadata.

[0076] In this scenario, analyzer 3 is unaware that the human-computer interactions originate from multiple users. Processing the received signals therefore involves merging different interaction modalities, regardless of the user initiating each, to formulate a request and issue a command. In other words, in this case, a command resulting from a merged modality will always be determined identically by analyzer 3, regardless of whether the modalities in question originate from a single user or from multiple users.

[0077] In one example, the digital service manager includes a user identifier (not shown) capable of identifying, for each received signal, the user who originated that signal. This identification can be performed by analyzing each received signal, for example, through speech recognition. A received signal can also be identified by retrieving a user identifier. Specifically, received signals may contain a user identifier. This can be explicitly stated in the content itself (text, audio, image format, etc.), or via the transport channel itself, which is associated with a user identifier (IP address, for example). The user identifier has the additional effect of allowing the digital service manager to determine not only whether a set of received signals relates to the same single user or, conversely, to distinct users.

[0078] Analyzer 3 can therefore be configured to formulate a query in a differentiated way depending on whether the query is simple or enriched.

[0079] As indicated on the figure 5 Several incoming interactions 300, 310, 330, 340 within a multimodal set of human-computer interactions can be identified by the analyzer 3 as resulting from coordinated actions of several users of the digital service. The analyzer merges these interactions so that the generator produces a collective request from these users, or an enriched request. The enriched request thus generated then triggers the preparation of a command 700 corresponding to this collective request, and this command is issued through an output interface 70.

[0080] Similarly, analyzer 3 can estimate that, still within the same multimodal set of human-computer interactions, several other incoming interactions 320, 350 also result from coordinated actions and correspond to another collective request from several users. In this case, a command 710 is issued separately, either through the same output interface as before, or through a separate output interface 71.

[0081] The analyzer 3 can also determine that, within the same multimodal set of human-computer interactions, other incoming interactions 360, 370 result from coordinated actions by a single user and correspond to an individual request from that user. In this case, the analyzer 3 prepares a command 720 corresponding to this individual request and issues this command through any appropriate output interface 72.

[0082] As already mentioned, the formulation of collective or individual requests, in general, does not necessarily require taking into account any metadata.

[0083] Nevertheless, the timing of interactions, their context of reception and the identification of their authors are all examples of data that can be included in the received signals and whose automatic processing can make it possible to refine the formulation of queries and to formulate possible additional queries.

[0084] To present these different data simply and how their processing can influence the formulation of queries, we now consider a particular example of modalities and the resulting enriched query.

[0085] In this particular example, a first user verbally requests that the heating be turned on in a room, while a second user gestures towards a radiator. A microphone records an audio stream containing the first user's request, and a camera records a video stream containing the second user's request. These two streams are received by an input interface and processed by analyzer 3.

[0086] Without additional data, Analyzer 3 can only formulate a rough request to activate a heating system in a room, without knowing precisely which room. Analyzer 3 can optionally overcome this limitation by prompting one or more users to refine their request, for example, by naming the room or selecting it from a list of suggestions.

[0087] Contextual information is a special type of additional data whose automatic processing can influence query formulation. Thus, at least one of the received signals may contain contextual information 60 which is interpreted by the analyzer 3 to refine the formulation of a query.

[0088] Continuing with this specific example, it is possible to assume that the location of the camera filming the second user is known beforehand. This location could, for example, take the form of a room identifier, which is previously associated, in a database of received signal sources, with a camera identifier.

[0089] Thus, the interpretation, by analyzer 3, of a received signal containing the video stream and the camera identifier, may include an interaction with the database of received signal sources based on the camera identifier contained in the received signal.

[0090] Such an interaction allows analyzer 3 to obtain the ID of the room where users collectively wish to activate the heating. This room ID can then be used to refine the enriched request, specifically by correctly identifying one or more particular heating devices to activate. Additionally, the direction of a user's outstretched arm, as analyzed in the incoming video stream, constitutes a simple request for identifying a specific heating device to activate, thus contributing to the generation of the enriched request as a fusion of two simple requests: one in the form of a voice signal, the other in the form of a gesture signal.

[0091] In general, contextual information 60 as represented in the figure 6 may be related to one or more aspects including sound level, brightness, temperature or any other parameter relating to any characteristic of a user, their environment, or a data source used by the user and used to acquire one or more signals received by analyzer 3.

[0092] Thus defined, contextual information 60, 61 obtained by analyzer 3 translates a set of indirect interactions 600, 610 with the digital service, as opposed to direct interactions which are intentionally manifested by users.

[0093] Since users are distinct and their locations at the time of interactions are also potentially different, different signals received can naturally contain different contextual information.

[0094] Unlike known devices and processes implementing fusions of modalities from a single user, it is useful here that the analyzer 3 can integrate contextual information 60, 61 differentiated by user, in order to be able to correctly interpret the contexts of obtaining different modalities from distinct users.

[0095] User profiles are another special type of additional data whose automatic processing can influence the formulation of queries.

[0096] It may also be envisaged to use a history of requests, in particular provided to analyzer 3, and / or generated by analyzer 3 itself, in particular by keeping in memory all or part of the signals and associated requests.

[0097] A query history might include, for example: on the one hand, sets of received signals, each containing at least one signal corresponding to a user action and having made it possible to formulate at least one request, and on the other hand, for each set of received signals, one or more commands issued downstream on the basis of the formulated request.

[0098] Analyzer 3 can be configured to search, in such a history, for signals similar to one of the commonly received signals, and to consult, for each of these signals present in the history, the commands actually issued downstream.

[0099] Such an approach, the practical implementation of which involves training an artificial intelligence according to known general principles, can make it possible to refine the formulation of a query, for example by choosing to reuse or adapt specific formulations of queries that have already been used in the past from similar received signals.

[0100] The process may also include obtaining an element of a profile associated with a user who generated one of the human-machine interactions and / or with at least one input interface receiving one of the human-machine interactions.

[0101] For example, the query history described above can be broken down by user in order to customize how previous queries are taken into account when formulating a current query.

[0102] Taking up the previous example of two users issuing the collective request to heat a room, analyzer 3 can also obtain or infer tastes or preferences of at least one of the users concerned in order to refine the formulation of the request by opting for example for a personalized temperature in the room.

[0103] In general, a user profile element refers to any form of past interaction, direct or indirect, between the user and the digital service, which can be interpreted by the analyzer 3 to refine the formulation of a query.

[0104] The user ID, already mentioned, is a specific example of a user profile element. A potential use of a user ID in this context is to refine the formulation of a vague query. For example, during a video conference, a query from the presenters, automatically identified by their user IDs, might be to mute the microphones of the other participants. The presenters' user IDs allow the parser to refine the query and deduce, through a process of elimination, the list of participants whose microphones should be muted.

[0105] The temporality of interactions, which can include, for example, the duration of each interaction or the time gap between the beginnings of two interactions, is an additional element likely to affect the detection of related interactions and the formulation of queries.

[0106] Taking this temporality into account is all the more complex as each user does not act with the same speed or the same precision.

[0107] Let's take the example of a command to turn on the living room lamp, detected by the cooperation of the following two interaction modalities: Voice command: "turn on the lamp" Gesture command: user's outstretched arm towards the living room lamp.

[0108] The merging of these two modes of interaction alone allows us to define the command to be executed.

[0109] However, one user may be faster than another, with a different coordination of related modalities, yet the command to be retrieved is indeed the same. This phenomenon is illustrated on the figure 7 , where the duration and synchronization of linked modalities 100, 110 from a first user are different from those 200, 210 from a second user. Modalities 100, 200 are vocal while modalities 200, 210 are gestural. They are represented on the figure 7 in the form of time arrows, a modality begins at a time corresponding to the origin of the arrow and ends at a later time corresponding to the tip of the arrow. Although the two modalities implemented by different users do not have the same duration or synchronization, the expected final command is identical in both cases.

[0110] Modalities from multiple users in parallel do not necessarily have to be taken into account by analyzer 3 in the same way as linked modalities from a single user.

[0111] Indeed, repeated interaction patterns by the same user are not equivalent to identical interaction patterns from multiple users. The former suggests redundancy or insistence, while the latter implies the same involvement or a shared request among several users.

[0112] For example, as shown on the figure 8 , when two users perform at the same time, or in close succession, the two modalities 500, 510 in cooperation: in voice “turn on the lamp” and with the gesture of the arm extended towards the lamp in the meeting room, this does not mean turning on this lamp twice, but turning it on only once, and as this is desired by two people, it can also signify an insistence or a real shared need.

[0113] For example, received signals containing voice information from one user and gesture information from another user may each include at least a start timestamp of interaction and optionally an end timestamp of interaction or equivalently a duration of the interaction.

[0114] In general, by interpreting time-stamped signals, in particular by interpreting time gaps between the beginning or end of one interaction and the beginning or end of another interaction, it is possible to determine, for example, whether these interactions are overlapping, close together or far apart in time, and on this basis to determine whether these interactions are linked or not.

[0115] Time is also a crucial element in determining the type of cooperation between several interaction modalities, particularly complementarity or redundancy. Thus, the processing by analyzer 3 of time-stamped moments relating to linked interactions offers the possibility of interpreting a given interaction modality in light of another interaction modality that precedes it.

[0116] By considering several related interaction modalities, the first interaction modality, in terms of timing, can, for example, define an initial, possibly incomplete, user request, while each subsequent request refines its meaning, transforming it into a collective request involving several distinct users. Thus, the consideration of the temporality of interactions, and in particular their order, by analyzer 3 helps to reconstruct the logical flow of users and thereby accurately transcribe certain collective requests.

[0117] When received signals originate from different sites because the users themselves are located at different sites, the latency between these sites can affect timing discrepancies. In this case, it may be necessary to synchronize the time references with respect to the same event (for example, following a triggered event such as a ping) in each time reference. This synchronization may require correcting timestamps performed separately at each site before interpreting the time-stamped signals.

[0118] Conversely, processing a fusion resulting from the interaction modalities of several users can pool, for multiple commands, one of the modalities from one user to be combined with each of the modalities from the other users. For example, in the same room, when one user says "turn on these lights," while a second user points to the ceiling light, and a third points to the table lamp, two commands must result from the fusion of the three modalities: a first command to turn on the ceiling light and a second command to turn on the table lamp.

[0119] Two specific examples of applications of the invention are now described.

[0120] In a first specific example illustrated on the figure 9 Pierre is in a meeting with his colleagues. Using his laser pointer, he points to one of the filenames displayed on the current slide and says "launch this file". The system combines the two interaction methods—"name written on the board and pointed to" and the spoken "launch this file"—and deduces the command "launch file xxx", xxx being read from the board using a camera that detects the targeted name.

[0121] The following week, Patrick, who had broken his arm, was at home to participate in the remote meeting. His colleague, present at the meeting, manipulated the laser pointer and directed it to the file to be shown at the same time as Pierre said "launch this file." This time, the service merged the two different interaction methods from the two meeting participants to deduce the entire command "launch file xxx."

[0122] At the end of the meeting, he adds verbally, "Send the file by email to those who want it," while his colleague points to the file name again. Some people raise their hands, others say "yes, me" or "yes, I'd like that." The system detects the raised hands of the different users using the video feeds from the various PC cameras connected remotely or present in the room. The digital service also detects, using the microphones of the connected computers and the microphone and camera in the meeting room, those who have communicated something equivalent to "yes," regardless of the method used.

[0123] The service thus merges the modalities of direct incoming interactions from the group of meeting participants and deduces the command to send the pointed-to file to the list of people who responded in one way or another in the affirmative.

[0124] In a second specific example illustrated on the figure 10 Jeanne and Alexis are using a digital tool to design their future home. Jeanne says "put this" and then clicks on an object on the screen, and Alexis follows by saying "here" and clicking on a location on the house plan. Tonight they are together in the living room, but tomorrow they can continue their digital design, he from his work, she from the house, while maintaining the same opportunities for cooperation and interaction.

Claims

1. Method for processing a request for a digital service in an interactive environment, the method comprising: generating an enriched request as a function of an interpretation of signals received from a multimodal set of human-machine interactions resulting from the human-machine interactions (300, 310, 320, 330, 340, 350, 360, 370) arising from coordinated actions of a plurality of distinct users in the interactive environment, each human-machine interaction, considered individually, being indicative only of a simple request, each simple request being a fragment of the enriched request, the enriched request triggering a transmission of at least one corresponding command (700, 710), characterized in that the method further comprises: detecting that distinct users are originating the received signals.

2. Method according to Claim 1, wherein the method comprises an interpretation of the received signals determining the enriched request as a function of at least one of the following elements: - simple requests extracted from the received signals; - for each of the received signals, users originating the received signals; - the result of a detection that distinct users are originating the received signals.

3. Method according to either of Claims 1 and 2, wherein the method comprises an identification, for each received signal, of a user originating the received signal.

4. Method according to any of the preceding claims, wherein at least one of said signals comprises a user identifier and generating the enriched request comprises at least one of the following steps: generating the enriched request as a function of the interpretation of the received signals and of the associated user identifiers, and generating the enriched request by adapting, as a function of the user identifiers, the enriched request obtained by the interpretation of the received signals.

5. Method according to any of the preceding claims, wherein, when at least one of said signals comprises contextual information (50), the generation of the enriched request adapts the enriched request according to an interpretation of this contextual information.

6. Method according to any of the preceding claims, further comprising: time stamping at least start times of receptions of said signals, interpreting time differences between the start times of respective receptions of said signals, and adapting the enriched request according to the interpretation of these time differences.

7. Method according to any of the preceding claims, wherein at least a part of said received signals comes from a plurality of input interfaces (30, 31, 32, 33, 34, 35, 36, 37) distributed over different sites.

8. Method according to Claim 7, further comprising: time stamping at least start times of receptions of said signals, correcting instants at which respective receptions of said signals start as a function of a latency between said sites, interpreting corrected time differences between corrected time-stamped instants, and adapting the enriched request according to the interpretation of the corrected time differences.

9. Computer program comprising instructions for the implementation of the processing method according to any of Claims 1 to 8, when said instructions are executed by a processor.

10. Device for processing a request for a digital service in an interactive environment, the processing device comprising at least: one generator of an enriched request as a function of an interpretation of signals received from a multimodal set of human-machine interactions resulting from the human-machine interactions (300, 310, 320, 330, 340, 350, 360, 370) arising from coordinated actions of a plurality of distinct users in the interactive environment, each human-machine interaction, considered individually, being indicative only of a simple request, each simple request being a fragment of the enriched request, the enriched request triggering a transmission of at least one corresponding command (700, 710), characterized in that the device further comprises a detector capable of detecting that distinct users are originating the received signals.

11. Processing device according to Claim 10, wherein the processing device comprises an analyser (3) capable of interpreting received signals determining the enriched request as a function of at least one of the following elements: - simple requests extracted from the received signals; - for each of the received signals, users originating the received signals.

12. Processing device according to Claim 11, wherein the processing device comprises a user identifier capable of identifying, for each received signal, a user originating the received signal.

13. Processing device according to any of Claims 10 to 12, wherein the processing device is implemented in at least one of the following devices: - a terminal, - a connected object manager capable of controlling at least one object connected to a communication network by means of at least one command corresponding to the enriched request intended for at least one connected object.