Method and apparatus for initiating an action - Patents.com

JP2024537866A5Pending Publication Date: 2025-09-12KONINKLIJKE PHILIPS NV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024520952
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2021-10-07
Filing Date
2022-09-08
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

Current human-machine interaction systems are inefficient and unreliable in dynamic environments with multiple people and devices, lacking flexibility and robustness, particularly in scenarios where multiple interactions occur simultaneously.

Method used

An apparatus and method that utilize multiple sensors and processors to detect real-world orientations and bidirectional information exchange links between entities, enabling the initiation of actions based on these links through audiovisual communication.

Benefits of technology

Enhances user experience and operational efficiency by adaptively managing interactions between multiple people and devices, mimicking human-like interaction in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

The apparatus includes a first sensor 101 and a second sensor 103 for determining a first and a second set of characteristics of an entity, which may be a device or a person, the sets being determined according to different sensor modalities. A first processor 105 determines a direction between the entities and a second processor 107 determines an orientation of the entities. A first detector 109 detects a bidirectional information exchange link between the first person and another entity from among a plurality of possible bidirectional information exchange links between the entities in response to the direction and at least one orientation. An initiator 111 initiates an action in response to the detection of the bidirectional information exchange link. The detection of the bidirectional information exchange link is in response to the first set of characteristics and the second set of characteristics.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to an apparatus and method for initiating actions based on detection of links between entities, and particularly, but not exclusively, to initiating actions based on man-machine interaction. [Background technology]

[0002] Man-machine interaction is becoming more and more prevalent, and many new applications based on or utilizing human-machine interaction are being developed. For human-machine interaction, voice control is becoming increasingly important and popular as it provides more efficient and user-friendly interaction in many practical situations. In critical environments such as hospital environments, it is often necessary to provide contactless operation and a simpler user interface (fewer physical buttons).

[0003] As an example, an increasing number of devices that interact with humans are becoming part of the home and professional environment. Indeed, homes and offices increasingly include numerous virtual or voice assistants with which users can interface by using voice commands and queries. Examples include home assistant devices such as Amazon's Alexa, Apple's Siri, Microsoft's Cortana, and Google's Assistant, which are prevalent in many homes and offices. Additionally, voice assistants or direct human interfaces may be implemented in home appliances such as televisions, radios, and other devices. Such devices are often operated and accessed by different people at different times and are often used in environments where multiple people are present at the same time, such as in a home or office space.

[0004] Another example is the medical industry, where there are often many devices in the same room. For example, in an operating room, many devices are present and used to monitor the health and biological status of patients and provide information to medical professionals, such as surgeons, specialists, nurses, etc. Furthermore, a relatively large number of people may be dynamically interacting with the devices, and the interactions between people and devices often change rapidly and significantly. Summary of the Invention [Problem to be solved by the invention]

[0005] It is therefore becoming important in many scenarios that human-machine interaction be efficient, robust and practical when used in dynamic environments where multiple people and devices may be present. However, most of the current systems tend to focus on a direct link between one person and one device, where the device itself directly interacts with one person to detect commands or requests. However, while such an approach may be efficient in many scenarios and applications, it has several drawbacks and may not be optimal in environments with multiple devices and people.

[0006] Improved approaches would be advantageous in many scenarios, particularly approaches that allow for improved operation, increased flexibility, reduced complexity, easier implementation, improved user experience, more reliable and robust interactions or operations, reduced computational burden, greater applicability, easier operation, and / or improved performance or operation.

[0007] SUMMARY OF THE DISCLOSURE Accordingly, the Invention seeks to preferably mitigate, alleviate or eliminate one or more of the above mentioned disadvantages singly or in any combination. [Means for solving the problem]

[0008] According to an aspect of the present invention, an apparatus is provided, the apparatus including: a first sensor that determines a first set of characteristics of a plurality of entities in a real-world environment, the first set of characteristics being determined according to a first sensor modality, each entity of the plurality of entities being a person or a device; a second sensor that determines a second set of characteristics of the plurality of entities, the second set of characteristics being determined according to a second sensor modality, the second sensor modality being different from the first sensor modality; a first processor that determines a real-world direction between entities of the plurality of entities, the real-world direction between the two entities being a direction from one of the two entities to another of the two entities in the real-world environment; and a first processor that determines a real-world direction between entities of the plurality of entities in response to the first set of characteristics. the first detector detects a real-world bidirectional information exchange link that exists between a first person of the plurality of entities and another entity of the plurality of entities from among a plurality of possible real-world bidirectional information exchange links between the entities of the plurality of entities in response to the real-world direction and the at least one real-world orientation, the real-world bidirectional information exchange link being a real-world audiovisual communication link enabling information exchange from the first person to the other entity and information exchange from the other entity to the first person; and an initiator that initiates an action in response to detection of the real-world bidirectional information exchange link, the first detector detecting the real-world bidirectional information exchange link in response to the first set of characteristics and the second set of characteristics.

[0009] The present invention provides improved user experience and / or enhanced functionality and / or performance in many applications and scenarios. For example, in many embodiments, an improved human-machine interface is provided. Typically, improved operation / performance / user experience is enabled, especially in environments where multiple people and / or devices interact with each other. This approach allows, for example, to initiate actions that reflect communication interactions between people and / or between people and devices.

[0010] In many embodiments, a more robust, reliable and flexible user interaction is achieved by improving user control, for example based on voice or gesture control. In many scenarios, the device acts as an "intermediary" or "middleman" that detects an interaction between two entities and adapts the behavior of the system / device, in fact one of the entities where a two-way information exchange link is detected, in response to the detection. Actions similar to those adopted unconsciously by humans in sometimes complex and non-uniform environments are obtained.

[0011] In some embodiments, the second processor determines an orientation of at least one of the entities of the plurality of entities in response to both the first set of characteristics and the second set of characteristics.

[0012] In some embodiments, the second processor determines an orientation of at least one of the entities in the plurality of entities in response to the first set of characteristics, and the first processor determines a direction between the entities in the plurality of entities in response to the second set of characteristics.

[0013] In some embodiments, the first processor determines the direction in response to the first set of characteristics. In some embodiments, the first processor determines the direction in response to the second set of characteristics. In some embodiments, the first processor determines the direction in response to both the first set of characteristics and the second set of characteristics.

[0014] A combined set of directions and at least one orientation is generated by the first processor and the second processor depending on both the first set of characteristics and the second set of characteristics. The combined set of directions and at least one orientation may be generated by the first processor and the second processor depending on both the first and second sensor modality data.

[0015] The first set of characteristics and / or the second set of characteristics may include at least one characteristic selected from the group of: a power state of the device, a position of the entity, an orientation of the entity, an orientation of the person's head, a pose of the person's eyes, a direction of the person's gaze, a gesture of the person, a direction of movement of the entity, a user action, a sound emitted by the entity, and an utterance from the person.

[0016] An initiator initiates an action by generating an action start command or message and forwarding / sending it to a processor.

[0017] A real-world audiovisual communication link is one that supports the communication of information / data using sound and / or light.

[0018] In an optional feature of the invention, the at least one real world orientation includes a real world orientation of the first person.

[0019] This provides improved performance in many scenarios, and in particular improved detection of two-way information exchange links in many scenarios, applications and embodiments, allowing for efficient operation and in many cases enhanced functionality.

[0020] In an optional feature of the invention, the first detector determines a real-world two-way information exchange link in response to a real-world orientation between the first person and another entity.

[0021] This provides improved performance in many scenarios, and in particular improved detection of two-way information exchange links in many scenarios, applications and embodiments, allowing for efficient operation and in many cases enhanced functionality.

[0022] In an optional feature of the invention, the first detector determines a real-world bidirectional information exchange link in response to detecting that a real-world orientation of the first person is aligned with a real-world direction between the first person and another entity.

[0023] This provides a particularly advantageous operation in many embodiments, allowing efficient detection of two-way information exchange links that are particularly well suited for adapting operations and initiating actions to enhance the experience of users, particularly those involving first persons, in many scenarios.

[0024] The first detector evaluates the alignment criteria by comparing a directional vector reflecting the orientation with a directional vector reflecting a direction between the first person and another entity, and alignment is deemed to exist if the directions are sufficiently parallel, e.g., have a minimum angle that does not exceed a given threshold, or if the normalized dot product of the vectors is not below a given threshold.

[0025] In an optional feature of the invention, the at least one real world orientation includes an orientation of another entity.

[0026] This provides a particularly advantageous operation in many embodiments, in particular allowing for improved detection of appropriate two-way information exchange links for initiating an action.

[0027] In an optional feature of the invention, the first detector determines a real-world bidirectional information exchange link in response to criteria including a requirement that a direction of information emission from another entity to the first person is aligned with a direction between the first person and the other entity.

[0028] This provides a particularly advantageous operation in many embodiments, in particular allowing for improved detection of suitable two-way information exchange links for initiating actions. The projection direction is in particular the main or central direction for projecting information from another entity. In particular it may be a direction perpendicular to the display plane or the central axis of a speaker or speaker arrangement.

[0029] In an optional feature of the invention, the first detector determines a real-world bidirectional information exchange link in response to criteria including a requirement that a view direction of the first person is aligned with a direction between the first person and another entity.

[0030] This provides a particularly advantageous operation in many embodiments, in particular allowing for improved detection of appropriate two-way information exchange links for initiating an action.

[0031] In an optional feature of the invention, the device further includes a second detector for detecting a triggering action by the first person, the initiator initiating an action in response to the triggering action.

[0032] This allows for improved usability, performance, and / or user experience in many embodiments.

[0033] In an optional feature of the invention, a second detector detects the triggering action as a communication by the first person over the real-world two-way information exchange link.

[0034] This allows for improved usability, performance, and / or user experience in many embodiments.

[0035] In an optional feature of the invention, the first sensor modality is a visual modality and the second sensor modality is an auditory modality.

[0036] This allows for improved usability, performance, and / or user experience in many embodiments.

[0037] In an optional feature of the invention, the other entity is a person.

[0038] This allows for improved usability, performance, and / or user experience in many embodiments, and in particular allows for adaptation of the system, such as adapting the human-machine interface specifically based on detection of a two-way information exchange link between people.

[0039] In an optional feature of the invention, the first detector detects a real-world bidirectional information exchange link in response to detecting that a real-world pose of the first person and a real-world pose of the other entity satisfy a matching criterion and sounds from at least one of the first person and the other entity satisfy the criterion.

[0040] This provides improved performance in many scenarios, and in particular improved detection of two-way information exchange links in many scenarios, applications and embodiments, allowing for efficient operation and in many cases enhanced functionality.

[0041] In an optional feature of the invention, the action is an action of another entity.

[0042] In an optional feature of the invention, the first sensor modality and the second sensor modality are different modalities selected from the group of visual, auditory, tactile, ultrasonic, infrared, radar, and tag detection.

[0043] In some embodiments, the other entity is a device.

[0044] This allows for improved usability, performance, and / or user experience in many embodiments.

[0045] In some embodiments the two-way information exchange link includes an audiovisual communication link from the first person to another entity.

[0046] This allows for improved usability, performance, and / or user experience in many embodiments.

[0047] In some embodiments the two-way information exchange link includes an audiovisual communication link from another entity to the first person.

[0048] This allows for improved usability, performance, and / or user experience in many embodiments.

[0049] In some embodiments, the first sensor includes multiple sensor elements at different locations in the environment.

[0050] This allows for improved usability, performance, and / or user experience in many embodiments.

[0051] In some embodiments, the apparatus further includes a user output for generating user instructions in response to detecting the two-way information exchange link.

[0052] This allows for improved usability, performance, and / or user experience in many embodiments, allowing for improved adaptation to provide more feedback to the user and for combined user and device adaptation.

[0053] In some embodiments, the initiator determines an identity designation of the first person, and the initiation of the action depends on the identity designation.

[0054] This allows for improved usability, performance, and / or user experience in many embodiments, which allows for improved user adaptation and optimization for individual users.

[0055] According to an aspect of the invention, there is provided a method of initiating an action, the method including the steps of: determining a first set of characteristics of a plurality of entities in a real-world environment, the first set of characteristics being determined according to a first sensor modality, each entity of the plurality of entities being a person or a device; determining a second set of characteristics of the plurality of entities, the second set of characteristics being determined according to a second sensor modality, the second sensor modality being different from the first sensor modality; determining a real-world direction between entities of the plurality of entities, the real-world direction between the two entities being a direction from one of the two entities to another of the two entities in the real-world environment; and, in response to the first set of characteristics, determining a real-world direction between the entities of the plurality of entities. determining a real-world orientation of at least one of the entities of the plurality of entities; detecting a real-world bidirectional information exchange link between a first person of the plurality of entities and another entity of the plurality of entities from among a plurality of possible real-world bidirectional information exchange links between the entities of the plurality of entities in response to the real-world orientation and the at least one real-world orientation, the real-world bidirectional information exchange link being a real-world audiovisual communication link enabling information exchange from the first person to the other entity and information exchange from the other entity to the first person; and initiating an action in response to detecting the real-world bidirectional information exchange link, the detection of the real-world bidirectional information exchange link being responsive to the first set of characteristics and the second set of characteristics.

[0056] These and other aspects, features and advantages of the invention will be apparent from and elucidated with reference to the embodiments described hereinafter. [Brief description of the drawings]

[0057] Embodiments of the invention will now be described, by way of example only, with reference to the drawings in which:

[0058] [Figure 1] FIG. 1 illustrates example components of an apparatus according to some embodiments of the present invention. [Diagram 2] Figure 2 shows an example of people interacting in an environment. [Diagram 3] FIG. 3 illustrates an example of the use of beamforming audio sensors to detect audio sources in an environment. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0059] The presence of devices interacting with humans is becoming increasingly common in everyday life, and the amount of human-machine interaction is growing rapidly and becoming ubiquitous.

[0060] As an example, a house or a room in a house, or an office, may contain a relatively large number of devices that can be controlled by a human, such as by audio actions (e.g., verbal commands), visual actions (e.g., gestures), etc. Multiple people may be in the room and may wish to interact with various devices, such as voice assistants, displays, audio playback systems, etc.

[0061] As another example, in a medical environment, such as an ambulance, a hospital ward, or an operating room, multiple medical professionals may wish to interact with various devices, such as via voice or gesture commands.

[0062] However, man-machine communication in an environment with many entities (each entity being a person or a device) tends to become very difficult to manage, and ensuring that the correct (device) action is initiated based on appropriate characteristics of a user action tends to become very difficult in such an environment.

[0063] In the following, an approach is described that supports the initiation of an action in a real-world environment (in the following, the term "environment" is used for brevity to refer to the real-world environment) in which there are multiple entities, specifically multiple people and devices in the same environment, such as the same room, in many embodiments and scenarios. The approach is based on an apparatus that detects a bidirectional information exchange link in the real world between at least a first person in the real-world environment and another entity (specifically, another person and / or another device in the environment). A bidirectional information exchange link in the real world is a link in which information is exchanged from the first person to the other entity and information is exchanged from the entity to the first person. A bidirectional information exchange link in the real world allows the first person to receive information from the first entity and provide information to another entity. This link is a real-world communication link that allows data and information to be exchanged bidirectionally between the first person and another device / person. For example, a bidirectional information exchange link exists between a person and a display, where the person controls the display by voice commands and the display provides an image that the user can see. Based on the detection of the link between the first person and the other entity, an action is initiated.

[0064] The link may be a visual link, for example a two-way information exchange link exists between the person and the display, allowing the person to control the part of the image that is magnified by the focus of their eyes.

[0065] The link may be a gesture link, for example a two-way information exchange link exists between a person and a robotic arm device, allowing the person to control the robotic arm with hand gestures.

[0066] 1 illustrates an example of a device that can initiate actions based on properties detected in an environment. The environment is a real-world environment, and properties of real-world entities and real-world interactions are determined. Similarly, directions and orientations of people and entities are determined in the real world, and are real-world directions and / or real-world orientations, such as directions and orientations in the real world. Thus, references below to such properties may be considered to refer to real-world properties (as opposed to virtual properties).

[0067] The device includes a first sensor 101 and a second sensor 103 that sense the real-world environment. These two sensors 101, 103 use different modalities. For example, the first sensor 101 is a camera that captures visual information of the environment, and the second sensor 103 is a microphone array that captures audio information of the environment. The sensors can determine characteristics of the environment based on the capture of the environment. For example, the first sensor 101 detects features in the captured image that correspond to possible entities in the environment, such as image objects that correspond to faces, humans, displays, etc. For example, the second sensor determines various audio sources in the environment and also determines, for example, the direction of arrival, the volume, and the type of audio (speech, music, etc.) for the different audio sources.

[0068] The apparatus further includes a first processor 105 for determining real-world directions between entities among a plurality of entities in the real-world environment, each entity being a person or a device. A direction from a first entity to a second entity (and therefore, equivalently, a direction from a second entity to the first entity) is determined and expressed, for example, as a vector. In many embodiments, the first processor determines the positions of different entities, such as the positions of people and devices in the environment, and determines the direction from this position. The first processor 105 determines the positions and directions based on sensor inputs, or determines the direction and / or positions based on, for example, direct user input for one or more entities. For example, a user can input the positions of static devices, such as displays or medical monitoring equipment, and the first processor 105 can estimate the positions of people in the environment based on video and audio capture. As another example, the devices can communicate their positions via GPS or based on fixtures to a known grid, such as a mounting rack. The directions between all entities can then be determined from the determined positions.

[0069] The first sensor 101 and the second sensor 103 are further coupled to a second processor 107 that determines at least one orientation of the entity based on characteristics determined by at least one of the first sensor 101 and the second sensor 103.

[0070] The second processor 107 typically determines / estimates a real-world orientation of at least one person within the real-world environment. The orientation is determined to reflect the direction the person is facing, head orientation, gaze direction, eye tracking direction, etc. Thus, the orientation of a person indicates the direction of the person's focus.

[0071] In many embodiments, the second processor 107 determines the orientation of multiple entities in response to the sensor data. For example, the second processor 107 detects devices in the environment and determines the direction they are facing. For example, the direction a display is facing is estimated based on the video captured by the first sensor 101.

[0072] The first processor 105 and the second processor 107 dynamically update the direction and / or orientation to reflect changes to at least one of the entities. In particular, as the person moves or turns in the environment, the first processor 105 and the second processor 107 update the direction and / or orientation to the entity to reflect the person's movement.

[0073] The first processor 105 and the second processor 107 are coupled to a first detector 109. The first detector 109 evaluates the determined direction and orientation in order to detect a bidirectional information exchange link in the real-world environment between the (first) person and another entity (hereafter referred to as a target entity). The target entity may be another person or a device such as a display or a medical instrument or device (e.g., electrocardiogram machine, dialysis machine, etc.).

[0074] A real-world bidirectional information exchange link can be considered as a real-world link between two real-world entities that is ready for bidirectional information exchange. The actual information can be either unidirectional (e.g., only a command with no response from the receiving entity) or bidirectional (e.g., a command followed by a response (e.g., display of an image) from the receiving entity). It is also possible that no information exchange takes place after the bidirectional information exchange link is established, e.g., a doctor checking the current status of a display.

[0075] A real-world two-way information exchange link is an information exchange / transmission link formed in the real world and existing / through real-world space. It may be a communication link existing through the air of the environment. Specifically, a real-world two-way information exchange link is a link formed / supported by the transmission of light and / or sound between / to / from people and entities. A real-world two-way information exchange link may include one or more optical / light-based communication / information links and / or one or more audio / acoustic-based communication / information links.

[0076] The detector is coupled to an initiator 111 that initiates an action in response to detecting a bidirectional information exchange link. Specifically, the initiator 111 generates an action initiation command and sends it to a receiver, specifically another device or entity. The action initiation command includes a command to initiate an action. For example, initiating an action is initiating a process, initiating a software program / routine, etc. In many embodiments, the initiator 111 can initiate an action by sending an action initiation command to a remote device. The action is specifically an action of a (device) entity of the multiple entities (for which a direction and / or orientation has been determined). The initiator 111 can initiate an action of a device of the multiple entities by sending an action initiation command / message to that device.

[0077] In some embodiments, the initiator 111 initiates an action when it is detected that a two-way information exchange link exists. Detection of the two-way information exchange link triggers the initiation of the action. For example, when it is detected that a two-way information exchange link exists between a given first person and one of the devices, the initiator 111 can generate and send an action initiation command to the device with which the two-way information exchange link is formed (e.g., triggered by eye contact with the device).

[0078] In other embodiments, the action is initiated when a trigger is detected. In such embodiments, the device includes a second detector 113 that detects a trigger action by the first person. The trigger action is typically determined based on data / characteristics generated by a sensor. Specifically, the trigger action is detected by detecting a specific sound or movement of the first person using output from an audio or video sensor.

[0079] The second detector 113 is coupled to the initiator 111 to which an indication of the detected trigger action is provided. In response to detecting the trigger action, the initiator 111 proceeds to initiate the action, such as by sending an action trigger command to an appropriate external device.

[0080] In some embodiments, an example of a trigger action is an audible cue, such as a "wake word" trigger. For example, if a person is looking at a given medical device in an operating room, it is detected that there is a bidirectional information exchange link with that medical device that provides the person (e.g., doctor) with information about a health condition, etc. However, even if a bidirectional information exchange link is detected, the device may not take any action. If at some point the person / doctor issues an appropriate audible command (e.g., "start the examination"), the second detector 113 detects this cue and the initiator 111 sends an action start command to the medical device. This causes the medical device to perform a diagnostic examination procedure. Depending on the detected bidirectional information exchange link, the device determines which device to send the action start command to. For example, when the expression "start the examination" is detected, the device determines which of multiple possible devices should start the examination based on the detected bidirectional information exchange link. That is, it determines which other device the person skilled in the art wants to instruct.

[0081] In some embodiments, the destination of the action initiation command varies depending on the two-way information exchange link, for example, specifically, the destination is a device that forms the two-way information exchange link.

[0082] In some embodiments, the action initiated and / or its characteristics vary depending on the bidirectional information exchange link. For example, the command sent varies depending on the device with which the bidirectional information exchange link is formed, so that the command is adapted to the destination device. For example, upon detecting an audible cue with the expression "start inspection" being issued by the first person, the device evaluates the device with which the bidirectional information exchange link is formed and generates an action initiation command that is specifically suited to this device. Thus, different commands can be sent to different devices without the user having to worry about explicitly defining the intended recipient of the command.

[0083] In some embodiments, the action and / or its characteristics that are initiated vary depending on the triggering action. For example, the second detector 113 can detect various trigger expressions, and depending on the trigger detected, different action initiation commands are generated and transmitted. For example, depending on whether the user utters expressions such as "start test 1," "start test 2," "start test 3," etc., the initiator 111 generates different action initiation commands that instruct the medical device to perform different diagnostic tests.

[0084] It will be appreciated that different trigger actions are used in different embodiments and that these may be detected based on different modalities (particularly based on data from the first sensor 101, the second sensor 103, or both). For example, in some embodiments, gestures or movements are detected as trigger actions based on visual capture of the environment.

[0085] Many different algorithms and approaches for detecting user input actions and commands will be known to those skilled in the art, and these will not be described in detail herein for the sake of brevity and clarity, although it will be appreciated that any suitable approach for detecting triggering actions by a user / person may be used without compromising the principles and approaches described.

[0086] In many embodiments, the second detector 113 detects the triggering action as a communication by the first person over a two-way information exchange link, particularly an action that is part of an interaction from the first person to a device or person with which the first person has been detected to have formed a two-way information exchange link.

[0087] For example, when detecting the existence of a two-way information exchange link between a first person and another person by the people having a conversation, the second detector 113 monitors speech from the first person captured by the microphone array of the second sensor, and if, for example, a predefined phrase is detected in the speech, a trigger action is deemed to be detected.

[0088] As another example, if a two-way information exchange link is detected between a first person and a display that displays information that the first person can read and detects, for example, gestures from the first person, the second detector 113 monitors the first person to detect such gestures. Upon detecting a gesture, the device initiates an action. For example, the device can generate an action initiation command that is sent to another device to cause an action to be taken / performed on this device.

[0089] In some embodiments, the device detects a two-way information exchange link between a person and another entity, and initiates an action in another device in response to information from the first person to the other entity. For example, a person may give a gesture or audio command to a display, causing the different device to generate an audio indication. In such a scenario, the device can effectively "listen in" on the information exchange between a person and the device or other person, and take action if a particular triggering action is detected.

[0090] The apparatus helps to provide efficient man-machine interface and interaction in many scenarios and environments where multiple entities coexist. For example, there are environments where multiple people are present and attempt to interact with multiple devices using man-machine interfaces based on audio or visual commands and instructions, etc. For example, in an operating room or intensive care unit, there is a relatively large number of different equipments used to monitor and evaluate the health status of a patient (or multiple patients). Furthermore, there may be a large number of devices to provide different information to medical personnel, for example several displays providing different information targeted to different specialists. Furthermore, there may be multiple medical personnel, such as surgeons, consultants / specialists, nurses, etc., interacting with these devices.

[0091] In such an environment, it is difficult to ensure reliable, efficient and user-friendly device control based on interactions between people and devices or between different people. Indeed, traditional approaches such as each individual device detecting speech and / or gestures are insufficient and unreliable, and are usually prone to uncertainty as to which action or device is intended to be activated.

[0092] The described approach facilitates and / or enables improved operation, for example in scenarios with many devices and people. The approach effectively acts as a "middleman" that can detect connections / links between entities and initiate actions based on the specific links that are formed. The apparatus can monitor the environment, particularly continuously, and adapt the operation of one or more of the devices (or another device) depending on the specific links that are established.

[0093] As a specific example, the device is implemented in an operating room, for example with a camera capturing the entire room and one or more microphone arrays installed at various positions in the room. For example, the device detects multiple faces in the room and determines the corresponding head orientations. In this way, it is detected that the medical personnel is facing a voice-controlled medical device having a display facing the medical personnel. For example, this may be based on manually input information regarding the position or orientation of the display, or based on automatic detection of the display based on camera sensor data, etc. Based on the audio sensor, it is further detected that the medical personnel is speaking while facing the display, and based on the camera sensor, for example, it is detected that the medical device is turned on and displays information. Thus, it is detected by the device that a two-way information exchange link is formed between the medical personnel and the medical device, and an action is initiated accordingly on the device itself or on another device, usually specifically the medical device. For example, the device sends a command to the medical device indicating that a two-way information exchange link has been detected and that the medical device should respond to the voice command. As another example, another action may be initiated, such as turning on a light that highlights an area associated with the medical device (e.g., a complex user interface).

[0094] As another example, a depth sensor and skeletal detection software (e.g., Microsoft's Skeletal Tracking™, known in the art) may determine the head-gaze vector of each person detected in the environment. For example, such a vector is pre-computed as a function of the 3D ear, nose, and eye positions using a variation of a human head graphic model. Given the detected skeletal data, the device calculates this head-gaze vector. All people's head positions and this vector are the output of the head pose estimation stage. Multi-microphone signal analysis provides sound data as a rough function of orientation relative to the capture device. In some embodiments, the depth camera and microphones are located in a single device, and the device associates sound snippets with head poses based on their orientation relative to the capture device. After such a correspondence estimation step, the device detects a two-way information exchange link by analyzing the relative positions and directions of the pose vectors. A simple criterion can be used to establish a two-way information exchange link based on both sound and image data.

[0095] As an example, the following approach can be used to detect a two-way information exchange link between two people: 1. Input: 2 head pose vectors: 2. Input: Sound from each head 3. Both heads emit sounds over a given past time interval (e.g., 5 seconds). 4. When the two head pose vectors are sufficiently parallel in opposite directions (e.g., the angle difference is less than 30 degrees) 5. There is communication between the two people

[0096] As another example, an active two-way information exchange link between a person and a display can be detected by following the following approach: 1. Input: Multiple head pose vectors 2. Input: Sound from each head 3. Input: Display position and orientation 4. Find the head pose vector that is most parallel to the display normal vector and in the opposite direction. 5. When the angle difference is less than a threshold (e.g., 10 degrees) 6. There is communication between the person and the display

[0097] In response to detecting this two-way information exchange link, the display may, for example, show a small picture or name of the person with whom communication is ongoing, rotate the display, or the 3D view may be optimized for the person's orientation.

[0098] Another application scenario is a hybrid communication scenario where multiple people are physically in the same room and available to multiple other people via audio / visual links. The device detects a two-way information exchange link corresponding to someone in the same space talking to someone else in the space. The device then initiates an action to send information of this direct person-to-person two-way information exchange link to people who are not in the same space but are in contact using communication via audio / video links. In this way, people who are not in the same location can be informed of direct person-to-person communication (as well as communication links). A graphical representation is presented that indicates who is talking to whom in other locations, for example (e.g., by changing the pose of a rendered avatar). This significantly increases the sense of immersion for people who are not in the same location.

[0099] This approach in particular improves performance and operation in many embodiments. The device is believed to provide similar effects and operations in some scenarios as humans achieve in an environment where a group of humans interact with each other. In some embodiments, the device acts as a "middleman" between people / devices, providing additional information regarding the interaction / communication / information exchange between various entities in the environment. This allows for an acknowledged adaptation to the current scenario, for example, allowing for improved man-machine interface in complex scenarios.

[0100] For example, the main advantages of using voice control of equipment are contactless operation and a simpler user interface (fewer physical buttons), especially in a hospital environment. However, while important steps have been taken to improve the performance of voice control engines, human-machine interaction is far from the immersive experience of human-human interaction and tends not to offer the same ease, adaptability, and reliability of interaction, especially in complex scenarios. In particular, use cases where multiple people interact with one or more devices (simultaneously) are challenging. The described apparatus helps in such instances by monitoring or surveying the environment in a manner similar to the way a human present in the environment uses multiple pieces of information to evaluate the scenario.

[0101] In some embodiments, this approach can reduce the difference between machine-to-human interaction in order to increase the immersiveness of human-to-machine interaction. That is, specifically, it allows a user to interact with a device in many scenarios in a way that is more natural and similar to the way a user interacts with another human. For example, in many practical environments, especially hospital environments, a person in the middle needs to interpret and execute certain commands issued by a doctor or the like. Also, a headset or other on-body device may be required for reliable interaction with the device.

[0102] In real-life "human-to-human" communication, various visual and audible cues are used unconsciously. Figure 2 shows an exemplary scenario of human interaction. In such a scenario, an important visual cue is "eye contact". When people B and C look at person A, they indicate that they are ready to receive information from person A. When person A also looks at person B, a two-way communication channel is usually established and person A interacts directly with person B. After a two-way communication channel is established between two people A and B, for example through eye contact, another person C can still receive information from person A (and person B), but person C implicitly knows that he should not interfere while there is still communication between people A and B.

[0103] Devices sense their environment using multiple different modalities to observe and detect two-way information exchange links between people, or between people and devices. This can provide operations similar to how a human "intermediary" uses their senses (eyes and ears) to determine who is talking to whom. Devices can detect such two-way information exchange links in their environment and initiate actions such as informing another device of the presence of the link and its characteristics (e.g., which entities are involved). This may be used by the device to adapt its operation, for example by providing information to a particular device that it is part of a detected two-way information exchange link and therefore should respond to voice commands, etc.

[0104] This approach is further enhanced by the device detecting a trigger, which may be an audible or visual cue, particularly from the first person. The trigger is taken into account when generating the action, e.g., the action is not performed until a triggering action is detected, or the characteristics of the action vary depending on the triggering action. The triggering action may be an audible cue, such as a "wake word" trigger.

[0105] As an example, in a human scenario, person C gets person A's attention by calling person A's name. In a typical person-to-person communication, whether and when person C emits an audible cue depends on the state of the visual cues, specifically whether person A is still in eye contact with person B, or whether person A or person B are still actively communicating. The described apparatus allows for accomplishing similar operations, where various modalities are evaluated to detect, for example, whether a two-way information exchange link exists and whether a trigger action has been taken to perform an action.

[0106] The particular algorithms and criteria for detecting a two-way information exchange link will vary from embodiment to embodiment, depending on the particular operation and performance desired.

[0107] In general, detection of a bidirectional information exchange link is based on the direction between the entities and the orientation of at least one (and often multiple or all) of the entities.

[0108] In many embodiments, the first processor first determines the current location of all devices and / or people considered. This is based on knowledge of sensor locations, such as camera and microphone locations, microphone arrays, and object detection that provides directions from the sensors to various items detected. For example, in the case of a camera, image object detection can be performed to detect objects corresponding to the person and associated devices (e.g., displays or other devices identified by optical properties, including the case of attaching optically detectable stickers to devices to aid in detection and identification). Similarly, in the case of a microphone array, audio sources are detected. Directions to these can be detected from beamforming weights. Based on the directions estimated from known sensor locations, an estimated location of the person and / or device can be estimated.

[0109] In some embodiments, some or all of the location is determined in response to user input, such as by the user explicitly inputting location data. In some embodiments, a dedicated location determination process is performed. For example, a display may show a particular image that facilitates detection and identification, or a particular sound may be output that is easily detectable.

[0110] Based on the positions, a direction between the devices is determined. These are represented, for example, by vectors. A (real-world) direction between two entities is a direction (in the real-world environment) from one of the two entities to another of the two entities. In some embodiments, the positions and / or distances between the entities are also taken into account, represented, for example, by the start point and length of a vector, respectively.

[0111] Locations may be static or dynamically updated. Indeed, in many embodiments, some entities (e.g., static devices) are represented by static locations, while other entities (e.g., people) are represented by locations that are constantly changing. Correspondingly, some directions (e.g., between static devices) are constant, while other directions between entities may dynamically change and be dynamically updated.

[0112] Additionally, the second processor 107 determines the orientation of one or more entities in the real-world environment. For example, all human orientations are determined based on face recognition and / or skeleton tracking, etc. In many embodiments, device orientations are taken into account. These may be estimated, for example, based on sensor inputs, or may be estimated automatically, for example, for some or all devices. For example, the size and shape of image objects of the display in an image captured by the camera of the first sensor 101 may be evaluated to determine the orientation of the display, and therefore in which direction the display is projecting an image.

[0113] (Real world) orientation and direction are determined with reference to a coordinate system that applies to the real world environment, so that the coordinates / orientations of the coordinate system correspond directly to the coordinates / orientations in the real world environment.

[0114] In some embodiments, the first detector 109 detects the bidirectional information exchange link using criteria that considers, for example, the direction between the first person and the devices not involved in the bidirectional information exchange link, for example, if the first person is looking at a particular marker (placeholder) in the room, its position does not correspond to the position of the actual device with which the bidirectional information exchange link is established (the device may not be visible to the first person).

[0115] In many embodiments, the first detector 109 determines a bidirectional information exchange link in response to an orientation between the first person and a device or person. In particular, the first detector 109 detects whether an orientation between the first person and another device matches an orientation of the first person.

[0116] For example, the criterion determines the direction from the first person to all devices in the room. Then the face orientation / direction and / or eye gaze orientation / direction are determined. Then the directions to the various devices are evaluated to see if any of the directions align with the orientation of the first person. If the alignment requirement is met, in particular if the angular difference between the direction to one of the devices and the orientation direction of the first person is below a given threshold, a bidirectional information exchange link between the first person and the device can be considered to be detected. Such a requirement basically corresponds to the detection that the first person is facing the device, and in particular corresponds to the detection that the first person is focusing on the device, e.g. looking at or talking to the device. Thus, a bidirectional information exchange link is considered to be detected between the first person and the device in the face / eye gaze direction. The orientation of the first person is usually the view direction. The view direction is an indication / estimation of the direction in which the person is looking.

[0117] Specifically, if the orientation of the first person is consistent with the orientation between the first person and the first device / person (e.g., the angle between them is less than a threshold (5°, 10°, 15°, etc.)), then a two-way information exchange link is considered to be detected between the first person and the first device / person. Further requirements may be included, such as detecting audio or gestures from the first person (and / or from the second person) and / or detecting audio emanating from or displaying images from the second device.

[0118] In some embodiments, it may be detected that the other entity is oriented towards the first person, i.e. that the device or person associated with the first person is aligned with the orientation between the first person and the other device or person. The orientation of the second person is in particular the facing orientation, i.e. the head facing orientation and / or the gaze direction. The orientation of the second device is in particular the (information) projection or emission orientation. For example, in case of a display, the orientation is the direction perpendicular to the display plane, in case of a speaker, the orientation is the main sound emission direction.

[0119] Specifically, if the orientation of the second device or person is consistent with the orientation between the second device / person and the first person (e.g., the angle between them is less than a threshold (5°, 10°, 15°, etc.)), then it is considered that a two-way information exchange link has been detected between the first person and the second device / person. Further requirements may be included, such as detecting audio or gestures from the first person (and / or from the second person) and / or detecting audio emanating from or displaying images from the second device.

[0120] In many embodiments, both the orientation of the first person and the orientation of the second person / device must be aligned with the direction between the first person and the second person / device.

[0121] In some embodiments, the first detector 109 detects the bidirectional information exchange link in response to detecting that the posture of the first person and the posture of the second person or device meet a matching criterion, and optionally that sounds from the first person and at least one of the other entities in the bidirectional information exchange link meet the criterion (and often both).

[0122] In some embodiments, the matching criteria is such that the angle between the vectors representing the orientations of the two entities is less than a given amount (e.g., 5°, 10°, 15°) but in opposite directions. For example, the matching criteria corresponds to the requirement that the normalized dot product between the vectors representing the orientations is negative and has a magnitude not below a given threshold (0.8, 0.9, 0.95, etc.).

[0123] The sound criterion is that the first person emits sound at a volume not below a given threshold (e.g., average over a given time) and / or the second person / device emits sound at a volume not below a given threshold (e.g., average over a given time).

[0124] In many scenarios, such approaches and detection provide a reliable indication that two people, or indeed one person and one device, are actively engaged with one another, that is, that the two entities are facing each other, engaged in a focused interaction, and exchanging information.

[0125] When considering alignment between two directions or between a direction and an orientation, in some embodiments the device only considers how parallel the directions / orientations are, while in other embodiments it also considers whether the assessed directions / orientations are pointing in the same direction. In some embodiments the alignment is determined based on the smallest angle formed between the directions / orientations, while in other embodiments the alignment also considers whether they are pointing in the same direction. In some embodiments a parameter that depends monotonically on the magnitude of the dot product between the two directions or between a direction and an orientation is considered, while in other embodiments a parameter that depends monotonically on the (signed) dot product between the two directions or between a direction and an orientation is considered alternatively or additionally.

[0126] As mentioned above, the device uses different modalities to determine various characteristics to detect a bidirectional information exchange link. As mentioned above, the (at least) two sensor modalities can be a visual modality and an audio modality. In particular, one or more cameras are used to detect the position and orientation of people and / or devices in the environment.

[0127] For example, human pose estimation can be used to detect a person's focus direction from a 2D image. For example, the 3D position of the head can be determined from the known size of a typical human head. Part of the head's orientation with respect to the optical axis of the camera can be determined by the relative image positions of facial features and their relationship to a typical 3D head model. Finally, a "virtual head" can be placed at an approximately correct location and orientation in 3D space using the known (e.g., pre-calibrated) pose of the capture camera with respect to the world coordinate system. Alternatively, the same can be achieved using a depth sensor such as that found in Azure Kinect. For example, the Azure Kinect Body Tracking SDK outputs the positions of human skeletal joints in 3D space. Knowing the pose of the Azure Kinect's depth sensor with respect to world space, a point in world space can be directly represented. A human's gaze vector can be constructed orthogonal to the positions of the two ears, pointing outwards from the middle of the line connecting the two ears in the direction of the nose. Similarly, gestures made by a person can be recognized based on camera detections.

[0128] Microphones (especially microphone arrays and directional microphones) are used to monitor sounds in the environment and to specifically identify who is speaking (e.g., by sound patterns that distinguish different voices, or by detecting the direction of audio reception in conjunction with vision-based location detection).

[0129] For example, in the case of an audio / microphone capture modality, beamforming can be used to isolate the person or people speaking and optionally determine the 2D or 3D direction from the microphone and therefore the location. This is used to detect the person speaking, which can then be annotated (with name, title, etc.). Such an approach is based, for example, on an initial registration procedure. Audio modalities and sensors can be used to track sounds in a particular area, which allows, for example, the use of exclusion and priority zones.

[0130] For example, by strategically placing a collection of microphone sensors, the direction of a coherent source can be determined using time / phase differences. Using adaptive delay-and-sum beamforming (DSB) or filter-and-sum beamforming (FSB), a sensitive beam can be directed towards a speaker that may be moving. Figure 3 shows an example of a multi-beam FSB solution that can simultaneously direct focused tracking beams towards multiple sources and also simultaneously scan the area for new sources by using a fast tracking beam. Based on the filter coefficients of the FSB filters, the angle of each beam relative to the microphone shape can be determined.

[0131] Another example modality is depth from a given sensor position. For example, depth sensors can be used to detect 3D skeletal points of multiple people in real time. Using one or more such systems, the head pose of a person can be inferred in real time. Such detection and estimation can then be used to infer the communicative intent of different people and detect bidirectional information exchange links. For example, the estimated head pose is derived from the skeletal data and evaluated with respect to the known / estimated positions of various devices (e.g., displays).

[0132] Other examples of modalities include ultrasound, infrared, radar, and tag detection.

[0133] The exact approaches used by, and the characteristics determined by, different sensors using different modalities will vary according to the preferences and requirements of a particular embodiment, as will the approaches for detecting bidirectional information exchange links, e.g., for determining direction and position, according to the preferences and requirements of a particular embodiment.

[0134] For example, the characteristics from the sensors may include one or more of a device power state, an entity's position, an entity's orientation, a person's head orientation, a person's eye pose, a person's gaze direction, a person's gestures, an entity's direction of movement, a user action, a sound emitted by the entity, and a speech from the person, and / or an approach for detecting a bidirectional information exchange link may take one or more of these into account.

[0135] In many embodiments, the two-way information exchange link includes an audiovisual communication link from a first person to a device / other person forming the two-way information exchange link. An audiovisual link is a link in which audio and / or visual information is exchanged in two directions. An audiovisual link is formed in a real-world environment. An audiovisual link is formed directly by sound and / or light propagating in the real-world environment.

[0136] A typical example is for a person to speak commands into a display, which responds by adapting the displayed information seen by the first person. Another example is for a first person to use gesture input to control, for example, a medical device, and the control device responds with a sound or speech that the first person can hear. In many embodiments, the bidirectional information exchange link is a communications link that transmits both audio and video in at least one direction. For example, the display can emit sound, and the medical device can detect both gestures and voice commands by the first person.

[0137] Thus, in many embodiments, the two-way information exchange link includes an audiovisual communication link from the first person to the device / other person and / or from the device / other person to the first person. For example, when a two-way information exchange link reflects two people speaking directly to each other, typically both audio and visual information is exchanged in both directions.

[0138] In some embodiments, at least one of the sensors includes multiple sensor elements at different locations in the environment. The sensor elements may be spaced apart from one another, for example positioned at a minimum distance of at least 1 meter, 2 meters, 5 meters, or in some applications 10 meters.

[0139] For example, microphones (or microphone arrays) can be placed in different locations in a room and audio signals captured from the different microphones can be used to detect a two-way information exchange link. For example, the location of a user can be determined based on the audio captured at different locations, e.g., by simply detecting which microphone detects the loudest audio, or by triangulating the beams of an adaptive beamforming microphone, etc.

[0140] As another example, cameras may be placed in various positions around a room to increase visibility of all areas and to reduce the risk of a person or device blocking another person or device, for example, a person may be tracked by multiple cameras and the orientation of the person's head may be determined based on the camera in which the head is most clearly visible.

[0141] In some embodiments, the device includes a user output for generating a user indication in response to detecting a bidirectional information exchange link. For example, when a new bidirectional information exchange link is detected, the device generates an alert. The alert may simply indicate that a bidirectional information exchange link has been detected, or may indicate some characteristics of the bidirectional information exchange link, such as, among other things, which other people or entities are part of forming the bidirectional information exchange link. The alert may, for example, audio the names and types of devices detected in the operating room, and the device may indicate that a particular medical device has sent a command that receives an audio input. For example, the device may issue the statement, "ECG monitor, command ready." In other examples, the device may simply generate an audio or light alert.

[0142] The device therefore includes a feedback mechanism to provide information to the user in the room. For example, it may provide an audible acknowledgment or display a visual indication on a screen using a logo or LEDs (red, orange, green, etc.). In Healthcare (HC) environments, which tend to be noisy environments, visual indications are often preferred.

[0143] After the feedback acknowledgement, in many embodiments, the person need not maintain focus on that particular device, but rather, the device can be considered enabled to receive voice input after a two-way information exchange link has been detected, with this status being maintained, for example, until a new two-way information exchange link is detected for the same person.

[0144] In some embodiments, the initiator 111 determines an identity indication of the first person included in the two-way information exchange link. For example, based on a camera, face detection is applied to identify the person with whom the two-way information exchange link is detected. Alternatively or additionally, the initiator 111 may detect the identity of the speaker from the audio captured by one or more microphones based on the captured audio. For example, the initiator 111 compares a signature determined from the video or audio with a stored signature linked to a particular identity. If the match between the signatures is sufficiently accurate, the first person is deemed to be identified as the person linked to the stored signature.

[0145] The initiator 111 further initiates an action depending on the determined identity indication. In some embodiments, the action is initiated only if the identity indication indicates a person previously selected as a person who may perform the action. Thus, in some embodiments, the action is initiated only if the identified user is a user who is qualified to perform the action. For example, a particular medical device can only be operated by a particular medical professional / consultant. In this case, when a first person is detected forming a bidirectional information exchange link with the medical device, an action is sent to the medical device, but only if the first person is that particular medical professional / consultant (previously approved).

[0146] In some embodiments, the action to be initiated is adapted in response to the identity indication. As a specific example, the action is person-modified, e.g., the volume of the display is adapted to the preferences of the particular person identified. As another example, the action-initiating command is sent to another entity, particularly a device with which a two-way information exchange link has been formed. For example, the action-initiating command includes data indicative of the detected identity, which allows the device to adapt operations to the particular user.

[0147] All issues and terms, including terms such as direction, orientation, entity, environment, and two-way information exchange links, may be replaced with real-world terms, including real-world direction, real-world orientation, real-world entity, real-world environment, and real-world two-way information exchange links.

[0148] It will be appreciated that for clarity, the above description describes embodiments of the invention with reference to various functional circuits, units, and processors. However, it will be apparent that functionality may be distributed between various functional circuits, units, or processors as appropriate without detracting from the invention. For example, functionality described as being performed by separate processors or controllers may be performed by the same processor or controller. Thus, references to specific functional units or circuits are to be regarded merely as references to suitable means for providing the described functionality, rather than indicative of a strict logical or physical structure or organization.

[0149] The invention can be implemented in any suitable form including hardware, software, firmware or any combination of these. The invention may optionally be implemented at least partly as computer software running on one or more data processors and / or digital signal processors. The elements and components of the embodiments of the invention may be physically, functionally and logically implemented in any suitable way. Indeed functionality may be implemented in a single unit, in several units or as part of other functional units. Thus, the invention may be implemented in a single unit or may be physically and functionally distributed between different units, circuits and processors.

[0150] Although the present invention has been described in connection with some embodiments, it is not intended to be limited to the specific form set forth herein. Rather, the scope of the present invention is limited only by the appended claims. Moreover, although certain features may appear to be described in connection with a particular embodiment, those skilled in the art will recognize that various features of the described embodiments may be combined in accordance with the present invention. In the claims, the term "comprising" does not exclude the presence of other elements or steps.

[0151] Moreover, although individually listed, a plurality of means, elements, circuits or method steps may be implemented by, for example, one circuit, unit or processor. Moreover, although individual features may be included in different claims, these features may also be advantageously combined, and the inclusion in various claims does not imply that the combination of features is not feasible and / or advantageous. Moreover, the inclusion of a feature in one category of claims does not imply a limitation of this category, but rather indicates that the feature may be applied to other claim categories as well, where appropriate. Moreover, the order of features in the claims does not imply a particular order in which the features must function, and in particular the order of individual steps in method claims does not imply that the steps must be performed in this order. Rather, the steps may be performed in any suitable order. Moreover, a reference in the singular does not exclude a reference in the plural. Thus, a singular element does not exclude a plural. Reference signs in the claims are provided only as examples for clarity, and these examples should not be construed as limiting the scope of the claims in any way.

[0152] Generally, example devices and methods are illustrated in the following embodiments.

[0153] Embodiments: 1. A first sensor (101) for determining a first set of characteristics of a plurality of entities in an environment, the first set of characteristics being determined according to a first sensor modality, each entity of the plurality of entities being a person or a device; a second sensor (103) for determining a second set of characteristics of the plurality of entities, the second set of characteristics being determined according to a second sensor modality, the second sensor modality being different from the first sensor modality; a first processor (105) for determining a direction between entities of a plurality of entities; a second processor (107) responsive to the first set of characteristics to determine an orientation of at least one of the entities of the plurality of entities; a first detector (109) responsive to the direction and at least one orientation to detect a bidirectional information exchange link between a first person of the plurality of entities and another entity of the plurality of entities from among a plurality of possible bidirectional information exchange links between the entities of the plurality of entities; an initiator (111) for initiating an action in response to detecting a bidirectional information exchange link; a first detector (109) for detecting a bidirectional information exchange link in response to the first set of characteristics and the second set of characteristics. 2. The apparatus of claim 1, wherein the at least one orientation includes an orientation of a first person. 3. The apparatus of claim 1 or 2, wherein the first detector (109) determines a two-way information exchange link in response to a direction between the first person and another entity. 4. The apparatus of any one of claims 1 to 3, wherein the first detector (109) determines a bidirectional information exchange link in response to detecting that an orientation of the first person is aligned with a direction between the first person and another entity. 5. The apparatus of claim 1 or 2, wherein at least one orientation includes an orientation of another entity. 6. An apparatus as claimed in any one of claims 1 to 5, wherein the first detector (109) determines the bidirectional information exchange link in response to criteria including a requirement that a projection orientation from the other entity to the first person is aligned with a direction between the first person and the other entity. 7. An apparatus as described in any one of claims 1 to 6, wherein the first detector (109) determines a bidirectional information exchange link in response to criteria including a requirement that a viewing direction of the first person is aligned with a direction between the first person and another entity. 8. The apparatus of any one of claims 1 to 7, further comprising a second detector (113) for detecting a triggering action by a first person, the initiator (111) initiating an action in response to the triggering action. 9. The apparatus of claim 8, wherein the second detector (113) detects the trigger action as a communication by the first person over the two-way information exchange link. 10. The apparatus of any one of claims 1 to 9, wherein the first sensor modality is a visual modality and the second sensor modality is an auditory modality. 11. An apparatus according to any one of claims 1 to 10, wherein the other entity is a person. 12. The apparatus of any one of claims 1 to 11, wherein the first detector (109) detects a bidirectional information exchange link in response to detecting that a posture of the first person and a posture of the other entity satisfy a matching criterion and that sounds from at least one of the first person and the other entity satisfy the criterion. 13. An apparatus according to any one of claims 1 to 12, wherein the action is an action of another entity. 14. The first sensor modality and the second sensor modality include Vision, Hearing, Touch, Ultrasound, Infrared, radar, The device of claim 1, wherein the different modalities are selected from the group of tag detection. 15. A method of initiating an action, comprising: determining a first set of characteristics of a plurality of entities in the environment, the first set of characteristics being determined according to a first sensor modality, each entity of the plurality of entities being a person or a device; determining a second set of characteristics of the plurality of entities, the second set of characteristics being determined according to a second sensor modality, the second sensor modality being different from the first sensor modality; determining a direction between entities of the plurality of entities; determining an orientation of at least one of the plurality of entities in response to the first set of characteristics; detecting a bidirectional information exchange link between a first person of the plurality of entities and another entity of the plurality of entities from among a plurality of possible bidirectional information exchange links between entities of the plurality of entities in response to the direction and at least one orientation; initiating an action in response to detecting the two-way information exchange link; Including, A method, wherein the detection of the bidirectional information exchange link is responsive to a first set of characteristics and a second set of characteristics.

Claims

1. a first sensor that determines a first set of characteristics of a plurality of entities in a real-world environment, the first set of characteristics being determined according to a first sensor modality, and each entity of the plurality of entities is a person or a device; a second sensor that determines a second set of characteristics of the plurality of entities, the second set of characteristics being determined according to a second sensor modality, the second sensor modality being different from the first sensor modality; and a first processor that determines a real-world direction between entities of the plurality of entities, the real-world direction between two entities being a direction from one of the two entities to another of the two entities in the real-world environment; a second processor responsive to the first set of characteristics to determine a real-world orientation of at least one of the entities; a first detector responsive to the real-world direction and the at least one real-world orientation to detect a real-world bidirectional information exchange link that exists between a first person of the plurality of entities and another entity of the plurality of entities from among a plurality of possible real-world bidirectional information exchange links between entities of the plurality of entities, the real-world bidirectional information exchange link being a real-world audiovisual communication link that enables information exchange from the first person to the other entity and from the other entity to the first person; an initiator that initiates an action in response to said detection of said real-world bidirectional information exchange link; Including, The apparatus, wherein the first detector detects the real-world bidirectional information exchange link in response to the first set of characteristics and the second set of characteristics.

2. The device of claim 1 , wherein the at least one real-world orientation comprises a real-world orientation of the first person.

3. The apparatus of claim 1 or 2, wherein the first detector determines the real-world two-way information exchange link in response to a real-world orientation between the first person and the other entity.

4. 4. The apparatus of claim 1, wherein the first detector determines the real-world bidirectional information exchange link in response to detecting that a real-world orientation of the first person is aligned with a real-world direction between the first person and the other entity.

5. The apparatus of claim 1 or 2, wherein the at least one real-world direction comprises an orientation of the other entity.

6. 6. The apparatus of claim 1, wherein the first detector determines the real-world two-way information exchange link in response to criteria including a requirement that a view direction of the first person is aligned with a direction between the first person and the other entity.

7. The device of claim 1 , further comprising a second detector that detects a triggering action by the first person, the initiator initiating the action in response to the triggering action.

8. The apparatus of claim 7 , wherein the second detector detects the triggering action as a communication by the first person over the real-world two-way information exchange link.

9. The apparatus of claim 1 , wherein the first sensor modality is a visual modality and the second sensor modality is an auditory modality.

10. The apparatus of claim 1 , wherein the other entity is a person.

11. 11. The apparatus of claim 1, wherein the first detector detects the real-world bidirectional information exchange link in response to detecting that a real-world posture of the first person and a real-world posture of the other entity meet a match criterion and that sounds from at least one of the first person and the other entity meet the criterion.

12. The apparatus of claim 1 , wherein the action is an action of the other entity.

13. The first sensor modality and the second sensor modality are Vision, hearing, tactile, ultrasound, Infrared, radar, The device of claim 1 , wherein the different modalities are selected from the group of tag detection.

14. 1. A method for initiating an action, the method comprising: determining a first set of characteristics of a plurality of entities in a real-world environment, the first set of characteristics being determined according to a first sensor modality, and each entity of the plurality of entities being a person or a device; determining a second set of characteristics of the plurality of entities, the second set of characteristics being determined according to a second sensor modality, the second sensor modality being different from the first sensor modality; determining a real-world direction between entities of the plurality of entities, wherein a real-world direction between two entities is a direction from one of the two entities to another of the two entities in the real-world environment; determining a real-world orientation of at least one of the plurality of entities in response to the first set of characteristics; detecting, in response to the real-world direction and the at least one real-world orientation, from among a plurality of possible real-world bidirectional information exchange links between entities of the plurality of entities, a real-world bidirectional information exchange link between a first person of the plurality of entities and another entity of the plurality of entities, the real-world bidirectional information exchange link being a real-world audiovisual communication link enabling information exchange from the first person to the other entity and from the other entity to the first person; initiating an action in response to said detecting said real-world bidirectional information exchange link; Including, The method, wherein the detection of the real-world bidirectional information exchange link is responsive to the first set of characteristics and the second set of characteristics.