Mixed reality video conferencing across multiple locations
The system addresses unnatural avatar positioning in mixed reality videoconferencing by mapping attendee locations and translating interactions, ensuring avatars face each other naturally, enhancing the mixed reality experience.
Patent Information
- Application Number
- JP2023503109
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-07-23
- Filing Date
- 2021-07-19
- Publication Date
- 2026-01-07
- Estimated Expiration
- 2041-07-19
AI Technical Summary
Conventional mixed reality videoconferencing systems fail to consider the positioning of remote attendees, leading to unnatural interactions as avatars do not face each other correctly, despite the actual physical layout of the locations.
A computer-based videoconferencing system that maps the locations of remote attendees, translating interactions to mimic natural behavior by orienting avatars based on the relative positions of participants, regardless of their actual locations, using computer vision and machine learning techniques to superimpose avatars onto real-world environments.
Enables natural interactions among attendees by correctly positioning avatars, efficiently utilizing space, and maintaining realistic interactions across multiple locations.
Smart Images

Figure 0007795266000001 
Figure 0007795266000002 
Figure 0007795266000003
Abstract
Description
[Technical Field]
[0001] The present invention relates generally to computer processing systems, and more particularly to computer processing systems that enable mixed-reality teleconferencing across multiple locations. [Background technology]
[0002] A computer system can generate a mixed-reality user experience by overlaying digitally created objects onto a real-world environment on a user's computer display. A mixed-reality experience allows the user to interact with both the digitally created and real-world objects. Mixed reality differs from virtual reality and augmented reality, which are primarily based on the level of interaction available between the user and the displayed environment. Virtual reality is a display technology in which a computing system creates a fully simulated environment. The user can interact with objects in the simulated environment but cannot interact with real-world objects. Augmented reality is a display technology in which a computer system creates an augmented environment by overlaying computer-generated perceptual information on top of the real-world environment. The user can interact with real-world objects but cannot interact with the computer-generated perceptual information. Summary of the Invention
[0003]
[0003] Embodiments of the present invention are directed to mixed reality videoconferencing across multiple locations. A non-limiting exemplary computer-implemented method includes overlaying a graphical representation of a first user over a real-world image of a first location, the first user being physically located at a second location different from the first location; detecting an interaction between the first user and a second user, the second user being physically located at the first location; determining a current location of the second user within the first location; and orienting the graphical representation of the first user toward the current location of the second user.
[0004] Other embodiments of the present invention implement features of the aforementioned methods in computer systems and computer program products.
[0005] Other technical features and advantages are realized by the techniques of the present invention. Embodiments and aspects of the present invention are described in detail herein and are considered a part of the claimed subject matter. For a better understanding, refer to the detailed description and drawings.
[0006] The particulars of the proprietary rights set forth herein are particularly pointed out and distinctly claimed in the claims at the conclusion of the specification. The foregoing and other features and advantages of embodiments of the present invention will become apparent from the following detailed description taken in conjunction with the accompanying drawings. [Brief explanation of the drawings]
[0007] [Figure 1] FIG. 1 illustrates a block diagram of components of a system for generating mixed reality video conferences in accordance with one or more embodiments of the present invention. [Figure 2] FIG. 1 illustrates a block diagram of components of a system for generating mixed reality video conferences in accordance with one or more embodiments of the present invention. [Figure 3] FIG. 1 illustrates a flow diagram of a process for initiating a mixed reality video conference in accordance with one or more embodiments of the present invention. [Figure 4]FIG. 1 illustrates a flow diagram of a process for initiating a mixed reality video conference in accordance with one or more embodiments of the present invention. [Figure 5] FIG. 1 illustrates a flow diagram of a process for generating a mixed reality video conference in accordance with one or more embodiments of the present invention. [Figure 6] FIG. 1 illustrates a flow diagram of a process for mixed reality videoconferencing in accordance with one or more embodiments of the present invention. [Figure 7] FIG. 1 illustrates a cloud computing environment in accordance with one or more embodiments of the present invention. [Figure 8] FIG. 1 illustrates an abstract model layer in accordance with one or more embodiments of the present invention. [Figure 9] FIG. 1 illustrates a block diagram of a computer system for use in implementing one or more embodiments of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0008] The diagrams shown herein are exemplary. There may be many variations of the diagrams or operations described herein without departing from the scope of the invention. For example, operations may be performed in a different order, or operations may be added, deleted, or modified. Also, the term "coupled" and variations thereof indicate that a communication path exists between two elements, and do not imply a direct connection between elements without an intervening element / connection between them. All of these variations are considered part of this specification.
[0009] One or more embodiments of the present invention provide a computing system that generates a mixed reality video conference for attendees at multiple locations. Attendees can use their computing devices to view three-dimensional graphical representations (e.g., avatars) of other attendees who are not physically present. The system takes into account the positioning of each attendee at each location, so that when an attendee at one location interacts with an attendee at another location, the computer system controls the attendees' avatars to face each other, regardless of the attendees' actual physical positioning.
[0010] Conventional computer-based videoconferencing systems can generate a videoconferencing space in which three-dimensional representations (e.g., avatars) of remote attendees who are not physically present are superimposed on an image of the real-world videoconferencing location. This mixed reality image can be viewed by attendees via a mixed reality display on a mobile computing device (e.g., a head-mounted display (HMD), a tablet display, or a smartphone display). However, conventional mixed reality systems simply have the avatars directly mimic the attendees' actual movements without considering the positioning of other attendees. Conventional mixed reality systems do not consider whether the avatar movements between two videoconferencing attendees follow natural interactions.
[0011] For example, a first attendee at a first location may interact with a second attendee at a second location. Each attendee is represented by an avatar at a location other than where the attendee is physically present. In this situation, the first attendee may turn to the right and rotate their device to view the second attendee's avatar. In a conventional system, if the first attendee turns to the right and views the second attendee's avatar display, the second attendee using a mobile computing device will see the first attendee's avatar facing right. However, if the second attendee is actually to the left of the first attendee's avatar at the second location, the first attendee's avatar will be facing the wrong direction, and the interaction will not appear natural.
[0012] One or more embodiments of the present invention address one or more of the aforementioned shortcomings by providing a computer-based videoconferencing system that maps the locations of remote attendees without limitations on the physical layout of each location's mixed reality space. The computer-based system translates interactions between virtual and real attendees, mimicking the interactions as if each attendee were present at each location, regardless of the attendees' relative locations. The computer-based videoconferencing system enables efficient use of space at each location while maintaining natural interactions.
[0013] Referring now to FIG. 1, a system 100 for generating a mixed reality video conference in accordance with one or more embodiments of the present invention is generally illustrated. Generally, the system 100 operates to create a mixed reality video conference for attendees at remote locations. The system 100 includes a first location spatial unit 102 and a second location spatial unit 104. Each of the first and second location spatial units 102, 104 is operable to receive topological, geometric, or geographic features of each location and identify dimensions and spatial relationships between objects at each location. Each of the first and second location spatial units 102, 104 is further operable to superimpose computer-generated graphical representations (avatars) of attendees onto a visual display of each real-world location. Real-world locations include spaces where solids, liquids, and gases exist. An example is a conference room in an office building where attendees are physically present. The system 100 also includes an interaction converter 106 for receiving and analyzing data to determine which attendees are interacting with each other. System 100 also includes an avatar model updater 108 for receiving data from interaction converter 106 and generating input data that causes the avatar's behavior to conform to natural behavior during interaction. For the following discussion, FIG. 1 illustrates system 100 in operative communication with first and second attendee devices 110, 112 at a first location and third and fourth attendee devices 114, 116 at a second location. However, system 100 can accommodate as many remote locations as are participating in the videoconference. System 100 and first attendee device 110 are described in further detail below with reference to FIG. 2. The description of first attendee device 110 is applicable to any of second, third, and fourth attendee devices 112, 114, 116.
[0014] Referring to FIG. 2 , the system 100 includes a first-location spatial unit 102 operable to receive and aggregate data from multiple attendee devices 110, 112. The first-location spatial unit 102 is further operable to employ computer vision techniques on the aggregated data for object detection. Object detection includes both image classification and object localization. Image classification involves predicting the class of one or more objects in an image. To perform image classification, the first-location spatial unit 102 receives an image as input and outputs a class label in the form of one or more integer values mapped to a class value. Object localization involves identifying the location of one or more identified objects in the image. To perform object localization, the first-location spatial unit 102 can process the received image and output one or more bounding boxes that define the spatial relationships of objects in the image. Through object detection, the first-location spatial unit 102 builds a three-dimensional spatial model for identifying objects in the videoconferencing location. For example, the spatial unit 102 at the first location can use the model to identify attendees, furniture, and visual materials.
[0015] In some embodiments of the present invention, the first location spatial unit 102 can apply machine learning techniques to perform object detection. In an exemplary embodiment, the first location spatial unit 102 uses a trained artificial neural network (e.g., a region-based convolutional neural network (R-CNN)). R-CNN typically operates in three stages. First, R-CNN analyzes an image to extract independent regions within the image and draws the regions as candidate bounding boxes. Second, R-CNN extracts features from each region, for example, using a deep convolutional neural network. Third, a classifier (e.g., a support vector machine (SVM)) is used to analyze the features within the regions and predict one or more object classes. In other embodiments, the first location spatial unit 102 is a different form of neural network than R-CNN.
[0016] As used herein, "machine learning" broadly refers to the ability of an electronic system to learn from data. A machine learning system, engine, or module may include machine learning algorithms that can be trained, such as in an external cloud environment (e.g., cloud computing environment 50), to learn functional relationships between currently unknown inputs and outputs. In one or more embodiments, machine learning capabilities may be implemented using artificial neural networks (ANNs) that have the ability to be trained to perform currently unknown functions. In machine learning and cognitive science, ANNs are a family of statistical learning models inspired by biological neural networks in animals and, in particular, the brain. ANNs can be used to estimate or approximate systems and functions that depend on a large number of inputs.
[0017] An ANN can be embodied as a so-called "neuromorphic" system of interconnected processor elements that function as simulated "neurons" and exchange "messages" with each other in the form of electronic signals. Similar to the so-called "plasticity" of synaptic neurotransmitter connections that transmit messages between biological neurons, the connections within an ANN that transmit electronic messages between simulated neurons are provided with numerical weights that correspond to the strength or weakness of a particular connection. These weights can be adjusted based on experience, allowing the ANN to adapt to inputs and learn. For example, an ANN for handwritten character recognition is defined by a set of input neurons that can be activated by pixels in an input image. The activation of these input neurons is weighted and transformed by a function determined by the network designer, and then passed on to other downstream neurons, often called "hidden" neurons. This process is repeated until an output neuron is activated, which determines which character has been read.
[0018] The spatial unit 102 at the first location is further operable to receive an avatar model for each attendee and merge this model with a local spatial model. The spatial unit 102 at the first location inputs location features into the avatar model. These features are extracted from images from each attendee's device and the position and observation angle of each attendee's device. The spatial unit 102 at the first location inputs this data to combine the avatar with objects detected in the real-world image and overlay the avatar on the real-world image. The appearance of each avatar is based at least in part on the position and observation angle of each attendee's device. For example, a top view, a bottom view, or a side view of the avatar may be displayed based on the observation angle of the attendee's device.
[0019] The spatial unit 102 at the first location further determines the position of each avatar based at least in part on the physical layout of the real-world location where the avatars are displayed. For example, at the first location, a first and second attendee are seated at a first square table facing each other. At the second location, a third and fourth attendee are also seated at a second square table, but on adjacent sides of the table. At the first location, the first location unit superimposes the avatars of the third and fourth attendees facing each other on the empty sides of the square table. At the second location, the spatial unit 104 at the second location superimposes the avatars of the first and second attendees on the empty adjacent sides of the table. Thus, the positioning of the avatars is converted to natural positioning during the videoconference, regardless of the attendees' actual positioning.
[0020] The interaction converter 106 is operable to receive interaction data from the first attendee's device 110. The interaction data is used to determine which attendees are interacting with which attendees. The interaction data may include data representing an avatar of the first attendee that may be displayed on a display of the second attendee's device. The interaction data may further include data representing an avatar of the second attendee that may be displayed on a display of the first attendee's device. Based on a determination that the time interval during which each avatar remains on the other attendee's display exceeds a threshold time, the interaction converter 106 may conclude that the two attendees are interacting. This determination may also be based on the time interval during which a single attendee's avatar is displayed on another attendee's device exceeding a threshold time.
[0021] The interaction data can include the position and orientation of the first attendee device. The position and orientation data can include data representing the x-, y-, and z-directions of the first attendee device. The interaction converter 106 analyzes the position and orientation data and data from the other attendee devices to determine which attendee the user of the first attendee device 110 is interacting with and the location of this user. For example, the first attendee device 110 can include a magnetic sensor and a gravity field sensor that can calculate the direction of magnetic north and the downward direction at the location of the first attendee device. The interaction converter 106 can use global positioning system (GPS) coordinates of the location of the first attendee device 110 and a world map of the angle between true north and magnetic north, and the interaction converter 106 calculates the correction angle needed for true north. The interaction converter 106 uses this calculation to determine the location of the first attendee device 110 and, as a result, the direction the first attendee device 110 is facing. Based on the direction the first attendee's device is facing, the interaction converter 106 applies a three-dimensional spatial model to determine which attendee's avatar should be positioned toward the first attendee's device 110 .
[0022] The avatar model updater 108 is operable to translate the decisions of the interacting attendees and generate inputs for the avatar models of the interacting attendees. The avatar model updater 108 sends the inputs to the spatial units 102, 104 at the first and second locations, which update the avatar models of the interacting attendees, respectively. In response, the avatar models provide outputs for translating the inputs into movements for each avatar.
[0023] The first attendee device 110 may be a mobile computing device and may be physically guided by the first attendee. The first attendee device 110 may include a head-mounted display (HMD). An HMD is a display device that can be attached to an attendee's head. The first attendee device 110 may also include a smartphone, tablet, or other computing device. The first attendee device 110 includes a local spatial unit 200. The local spatial unit 200 is operable to receive data captured from multiple sensors 206 and use computer vision techniques for object detection. The local spatial unit 200 is similar to the spatial unit 102 at the first location. However, the local spatial unit 200 receives data from multiple sensors 206 rather than from multiple attendee devices 110, 112. Furthermore, the local spatial unit 200 may aggregate data from the sensors 206 and generate a three-dimensional spatial model of the videoconference location. Meanwhile, the spatial unit 102 at the first location can aggregate data from multiple attendee devices 110, 112 to generate a shared three-dimensional spatial model of the videoconference location.
[0024] The first attendee device 110 includes a display 202. The display 202 is a small display device operable to display computer-generated images combined with real-world images (e.g., the videoconference location). The image displayed on the display 202 is based on the field of view of the image capture sensor 206 of the first attendee device 110. The field of view is the extent of the real world observable through the display of the first attendee's 110 device. In operation, the display 202 displays the physical location, including all attendees physically present at the location, as well as avatars of attendees at other locations. As the user manipulates the position of the first attendee device 110, the displayed image changes. The videoconference attendees view the mixed reality videoconference via the display 202.
[0025] The first attendee's device 110 includes a detection unit 204. The detection unit 204 is operable to receive location and motion data and determine whether the location and motion data suggest that the attendee is interacting with another attendee. The detection unit 204 is further operable to determine that the motion does not indicate an interaction with another attendee. For example, a sneeze or stretch by an attendee is a transient motion that generates motion data but does not suggest an interaction with another attendee. The detection unit 204 may be trained to recognize motions associated with these transient motions. The detection unit 204 is further operable to receive audio data to determine whether the attendee is interacting with another attendee. The detection unit 204 may analyze audio using natural language processing techniques to determine whether the attendee is conversing with another attendee. For example, the detection unit 204 may compare a name spoken by an attendee with a name provided when the attendee joined the videoconference. By comparing the names, the detection unit 204 may determine which attendee is speaking with which attendee. The audio data may be combined with the motion and position data to determine whether an attendee is interacting with another attendee. If the detection unit 204 determines that an interaction between attendees is occurring, the detection unit 204 transmits the data to the interaction converter 106.
[0026] The sensors 206 include image capture sensors, motion capture sensors, depth capture sensors, inertial capture sensors, magnetic capture sensors, gravity field capture sensors, position capture sensors, and audio capture sensors. The sensors 206 include, but are not limited to, cameras, gyroscopes, accelerometers, and position-based circuitry for interfacing with a global positioning system. The sensors 206 are operable to collect data from the environment for three-dimensional analysis, including determining the location and dimensions of objects. To enable three-dimensional analysis, the sensors are operable to capture shadows, lighting, and reflectivity of the surface of objects. The sensors 206 continuously collect data during the videoconference to determine the attendee's field of view and their positioning via the display 202.
[0027] With reference to Figures 3-6, a method is described for various stages of creating a mixed reality videoconference in accordance with one or more embodiments of the present invention. For illustrative purposes, the method is described with reference to attendees A and B at a first location and attendees C and D at a second, different location. With reference to Figure 3, method 300 initiates a mixed reality videoconference session from the perspective of the attendees at the same location and from the perspective of a server. At block 302, attendee A connects to a server using a device. For example, attendee A can use the device to access a videoconference website and connect to the server. At block 304, attendee A requests a mixed reality videoconference S. After attendee A accesses the website, attendee A can enter identification information and a passcode and request that the server initiate the mixed reality videoconference.
[0028] At block 306, attendee A uploads or causes an avatar to be uploaded to the server. This avatar is a three-dimensional graphical representation of attendee A. This avatar can be a three-dimensional human-like representation and can resemble attendee A. The avatar can be pre-built or selected from a set of available avatar models.
[0029] At block 308, attendee A's device registers a first location where attendee A is located. The registration includes generating a three-dimensional spatial model of the first location using a local computing device. The server collects sensor-based data to map the structure of the videoconference location (e.g., room dimensions), as well as the furniture, visual materials, and all attendees. The sensor-based data may be collected via a remote device, such as an HMD worn by the attendee or sensors placed around the videoconference room. At block 310, attendee A's device uploads the three-dimensional spatial model of the first location to the server.
[0030] At block 312, Attendees B connects to the server using a computing device. For example, Attendees B may use the device to access the same videoconferencing website as Attendees A and connect to the server. Attendees B may enter identifying information and request access to the server. The server may authenticate Attendees B's credentials and grant or deny access to the server.
[0031] At block 314, attendee B joins videoconference S, and at block 316, attendee B uploads an avatar model to the server. At this point, the system can compare the avatar models uploaded by attendee A and attendee B to determine if they are the same avatar model. In this example, if the two avatar models are the same, the server can issue an alert to attendees A and B with suggested changes, such as a color change, a virtual name tag, etc. Attendees A and B can change their avatars so that attendee A is not confused with attendee B.
[0032] Attendees B join the videoconference S in block 318, and at block 320, attendee B downloads the three-dimensional spatial model of the first location to their computing device. Attendees B's computing device can now send sensor-based data to the server. For example, a chair is moved outside the image captured by attendee A's device, but attendee B's device captures the chair being moved. This data can be sent to the server to update the model. From the server's perspective, at block 322, in response to a request from attendee A, the server initiates the mixed reality videoconference.
[0033] At block 324, the server receives the three-dimensional spatial model from attendee A and initializes a shared three-dimensional spatial model of the first location. In some embodiments of the present invention, an attendee's device may generate the spatial model. For example, an HMD may collect data using an image sensor and a depth sensor and use this data to generate a local spatial model of the videoconference location where attendee A is located.
[0034] At block 326, the server introduces an avatar model of attendee A into the shared three-dimensional spatial model of the first location. This avatar model is a mathematical description of a three-dimensional illustration of attendee A. The avatar generated by the model can include a cartoon character, a realistic drawing, an animal, or other drawing. The avatar's movement is based on inputting sensor data representing attendee A's movement into the model and outputting the direction of the avatar's movement. For example, the HMD may use head tracking sensors to detect the user's position and orientation, hand tracking sensors to detect hands, eye tracking sensors to detect eye movements, face tracking sensors to detect facial expressions, and other sensors to detect legs and body. These detected movements may be mapped and input into the avatar model. The avatar model may then output corresponding movements with respect to the avatar's position. At block 328, in response to attendee B joining the videoconference, the server introduces an avatar model of attendee B into the shared three-dimensional spatial model of the first location.
[0035] 4, a method 400 for initiating a mixed reality videoconferencing session from the perspective of a co-located attendee and a server is shown. At block 402, attendee C connects to the server using a computing device, and at block 404, attendee C joins the mixed reality videoconference. For example, attendee C can use the device to access the same videoconferencing website as attendees A and B and connect to the server. Attendees C enter identifying credentials via the videoconferencing website, and the server can grant or deny permission to join the mixed reality videoconference.
[0036] At block 406, attendee C uploads an avatar model to the server. The avatar generated by this model can be a three-dimensional representation of attendee C. The avatar can be pre-built or selected from a set of available avatar models. The avatar can be stored on attendee C's computing device, a remote device, or a server.
[0037] At block 408, attendee C registers a second location, which includes generating a three-dimensional spatial model of the second location using a local computing device. When attendee C joins the videoconference, an avatar for attendee C is displayed on attendee A's and attendee B's devices. The avatar is positionally overlaid on the same real-world location on both attendee A's and attendee B's displays. For example, attendee A and attendee B can both see an avatar sitting in the same chair displayed on each device. However, this avatar appears different to each attendee based on the angle of their device and their distance from the chair.
[0038] At block 410, the three-dimensional spatial model of the second location is uploaded to the server. Attendees D connect to the server using a computing device at block 412. For example, Attendees D can use the device to access the same videoconferencing website as Attendees A, B, and C and connect to the server. Attendees D can similarly verify their identity via the videoconferencing website and join the mixed reality videoconference.
[0039] Attendees D join videoconference S in block 414, and at block 416, attendee D uploads an avatar model to a server. The avatar generated by attendee D's model may be pre-built or selected from a set of available avatar models. The avatar may be stored on attendee D's computing device, a remote device, or a server.
[0040] At block 418, attendee D joins videoconference S, and at block 420, attendee D downloads a three-dimensional spatial model of the first location. An avatar of attendee D is displayed to attendees A and B via their respective devices. When attendees A and B view the first location in the real world, attendee D's avatar appears embedded in the first location.
[0041] From the server's perspective, in block 422, in response to the request from attendee A, the server creates a shared location model of the second location. In some embodiments of the present invention, an attendee's device may generate the spatial model. For example, an HMD may generate the local spatial model by using image and depth sensors and using the data to generate a local spatial model of the videoconference location where the attendee is located.
[0042] At block 424, the server configures attendee A's avatar model in the three-dimensional spatial model of the second location. Attendees A's avatar model can now receive input related to the physical layout of the second location. In this sense, attendee A's avatar responds to the physical layout of the second location. For example, attendee A's avatar appears differently through attendees C and D's devices than it does to the physical objects at the second location.
[0043] At block 426, the server identifies attendee C's avatar as the target of the interaction. By setting attendee C as the target of the interaction, the avatar model is no longer updated based on interactions between the attendee and inanimate objects. Rather, each avatar model is updated based on interactions with other attendees.
[0044] At block 428, the server configures attendee B's avatar model in the shared three-dimensional spatial model of the second location. Attendees B's avatar model can now receive input related to the physical layout of the second location. Attendees B's avatar responds to the physical layout of the second location. For example, through attendees C and D's devices, attendee B's avatar appears different from the physical objects at the second location.
[0045] At block 430, the server identifies attendee D's avatar as the target of the interaction. This identification may be performed by encoding the avatar output from attendee D's avatar model to reflect that this avatar is the target of the interaction. Setting attendee D as the target of the interaction provides a focal point toward which another attendee's avatar can turn.
[0046] Referring to Figure 5, a method 500 for generating a shared mixed reality videoconferencing session is shown. At block 502, a server adds avatar models for attendees C and D to a shared first location model. The avatar models for attendees C and D are now able to receive input related to the physical layout of the first location. In this sense, attendee A's avatar is responsive to the physical layout of the second location.
[0047] At block 504, the server adjusts the placement of the avatars to conform to the placement of the first location. For example, in one case, if attendees C and D are sitting side by side at the second location, but only opposite seats are available at the first location, the server determines that the avatars of attendees C and D should be displayed facing each other at the first location.
[0048] At block 506, the server sets avatars for attendees C and D and fixes the positions of avatars C and D within the three-dimensional spatial model of the first location. Using the example above, the server sets avatars for attendees C and D to appear facing each other at the first location.
[0049] At block 508, the server adds avatar models for attendee A and attendee B to the three-dimensional spatial model of the second location. The avatar models for attendees A and B are now able to receive input related to the physical layout of the second location. At block 510, the server adjusts the placement of the avatars for attendees A and B to conform to the placement of the second location. As above, the server determines the positions of the avatars for attendees A and B to be displayed at the second location. At block 512, the server fixes the positions of the avatars for attendees A and B within the three-dimensional spatial model of the second location.
[0050] 6, a method 600 for generating interactions during a mixed reality video conference is shown. At block 602, a motion detector detects a movement of attendee A. For example, attendee A may turn their body 30 degrees to the left, and the system may detect this movement. At block 604, a local spatial model stored on attendee A's computing device is updated to reflect this movement.
[0051] In block 606, the server updates the shared location model to reflect the movements of attendee A. For example, a sensor located at attendee A, a sensor located at attendee B, or a sensor located at the video conference location may detect that attendee A has turned his head 15 degrees toward attendee C. The server inputs this data into attendee A's avatar model and rotates the avatar the appropriate number of degrees to face attendee C.
[0052] At block 608, attendee A's computing device determines, based on attendee A's movements, that attendee A is looking at attendee C, the subject of the interaction. Continuing the example above, as attendee A turns, attendee C's avatar becomes visible on attendee A's device, and then attendee A's avatar, turning toward attendee C, appears on the displays of attendees C and D's devices.
[0053] At block 610, attendee A's computing device sends data to the server that attendee A is looking at attendee C. Attendees C have been identified as the target of the interaction. Thus, the server causes attendee A's avatar to appear to be looking at attendee C, even if attendee C is placed near an inanimate object (e.g., a cabinet).
[0054] At block 614, the server updates the shared three-dimensional space model to reflect that attendee A is viewing attendee C. Thus, when attendee C views attendee A through attendee C's device, an avatar of attendee A looking at attendee C is displayed. At block 616, the location model stored on attendee B's computing device is synchronized with the updated three-dimensional space model stored on the server. This synchronization populates the location model to reflect that attendee A is viewing attendee C.
[0055] At blocks 618 and 620, the local spatial models stored on attendees C and D's devices, respectively, are synchronized to display attendee A's avatar facing toward attendee C. In other words, attendee A's avatar appears to face attendee C on each of attendee C's and D's devices.
[0056] In some embodiments, the avatar may further appear to respond to environmental stimuli at the videoconference location. Each attendee's device may be equipped with audio sensors, temperature sensors, and other suitable sensors. The sensors may be operable to sense environmental stimuli at the videoconference location, and the avatar displayed at the location may respond to the stimuli. For example, attendee A may be at a first location, and attendee A's avatar may be displayed on attendee B's device at a second location. Attendees B's device may be configured to sense environmental stimuli (e.g., a sudden drop in temperature, a loud noise, a sudden increase in sunlight). Environmental stimulus data from one location may be transmitted to a server, and the server may update the avatar model if the attendee is at another location and unaware of the environmental stimulus. For example, an attendee's avatar at a first location may squint in response to increased sunlight at a second location. The avatar may turn toward a loud sound, even if the actual attendee cannot hear the sound.
[0057] Although this disclosure includes detailed descriptions of cloud computing, it should be understood that implementation of the teachings presented herein is not limited to cloud computing environments. Embodiments of the present invention may be implemented in conjunction with any other type of computing environment now known or later developed.
[0058] Cloud computing is a service delivery model for enabling convenient, on-demand network access to a shared pool of configurable computational resources (e.g., networks, network bandwidth, servers, processing, memory, storage, applications, virtual machines, and services) and for rapidly provisioning and releasing these resources with minimal administrative effort or interaction with a service provider. This cloud model may include at least five characteristics, at least three service models, and at least four deployment models.
[0059] The features are as follows:
[0060] On-demand self-service: Cloud customers can unilaterally and automatically provision computing power, such as server time and network storage, as needed, without the need for human interaction with the service provider.
[0061] Broad network access: Capabilities are available over the network and can be accessed using standard mechanisms, facilitating use by heterogeneous thin-client or thick-client platforms (e.g., mobile phones, laptops, and PDAs).
[0062] Resource Pool: The provider's computing resources are pooled and offered to multiple consumers using a multi-tenant model, with various physical and virtual resources dynamically allocated and reallocated according to demand. There is a sense of location independence, where consumers typically have no control or knowledge regarding the exact location of the resources offered, although at a higher level of abstraction they may be able to specify a location (e.g., country, state, or data center).
[0063] Rapid Elasticity: Capacity is quickly and elastically provisioned, sometimes automatically, and can be quickly scaled out and quickly released to quickly scale in. The capacity available for provisioning often appears unlimited to the consumer, and any amount can be purchased at any time.
[0064] Metered Services: Cloud systems leverage metering capabilities to automatically control and optimize resource usage at a level of abstraction appropriate to the type of service (e.g., storage, processing, bandwidth, and active user accounts). Resource usage can be monitored, controlled, and reported, providing transparency to both providers and consumers of the services used.
[0065] The service model is as follows:
[0066] SaaS (Software as a Service): Consumers are offered the ability to use a provider's applications running on a cloud infrastructure. Those applications are accessible from a variety of client devices through thin-client interfaces such as web browsers (e.g., web-based email). The consumer does not manage or control the underlying cloud infrastructure, including the network, servers, operating systems, storage, or individual application functionality, with the possible exception of limited user-specific application configuration settings.
[0067] PaaS (Platform as a Service): The ability offered to a consumer is to deploy applications they create or acquire, written using programming languages and tools supported by the provider, onto a cloud infrastructure. The consumer does not manage or control the underlying cloud infrastructure, including the network, servers, operating systems, or storage, but does have control over the deployed applications and, in some cases, the configuration of the application hosting environment.
[0068] Infrastructure as a Service (IaaS): The ability provided to a customer is to provision processing, storage, network, and other basic computing resources upon which the customer can deploy and run any software, which may include operating systems and applications. The customer does not manage or control the underlying cloud infrastructure, but has control over the operating system, storage, deployed applications, and in some cases, limited control over selected network components (e.g., host firewalls).
[0069] The deployment model is as follows:
[0070] Private Cloud: This cloud infrastructure is operated solely for the organization, can be managed by the organization or a third party, and can reside on-premise or off-premise.
[0071] Community Cloud: This cloud infrastructure is shared by multiple organizations to support a specific community with shared interests (e.g., mission, security requirements, policy, and compliance considerations). It can be managed by these organizations or a third party and can reside on-premises or off-premises.
[0072] Public cloud: This cloud infrastructure is available for use by the general public or large industry organizations and is owned by an organization that sells cloud services.
[0073] Hybrid cloud: This cloud infrastructure is a composition of two or more clouds (private, community, or public) that remain distinct and are joined together by standardized or proprietary technologies that allow for the portability of data and applications (e.g., cloud bursting to balance load between clouds).
[0074] A cloud computing environment is a service-oriented environment that emphasizes statelessness, loose coupling, modularity, and semantic interoperability. At the heart of cloud computing is an infrastructure that contains a network of interconnected nodes.
[0075] Referring now to FIG. 7 , an exemplary cloud computing environment 50 is shown. As shown, the cloud computing environment 50 includes one or more cloud computing nodes 10 with which local computing devices used by cloud consumers (e.g., a personal digital assistant (PDA) or mobile phone 54A, a desktop computer 54B, a laptop computer 54C, and / or an automobile computer system 54N) can communicate. The nodes 10 may communicate with each other. The nodes 10 may be physically or virtually grouped in one or more networks (not shown), such as a private cloud, community cloud, public cloud, or hybrid cloud, or combinations thereof, as previously described herein. This enables the cloud computing environment 50 to provide an infrastructure, platform, and / or SaaS that does not require cloud consumers to maintain resources on their local computing devices. The types of computing devices 54A-N shown in FIG. 7 are intended to be illustrative only, and it is understood that computing node 10 and cloud computing environment 50 can communicate with any type of computer-controlled device via any type of network and / or network-addressable connection (e.g., a connection using a web browser).
[0076] Referring now to Figure 8, a set of functional abstraction layers provided by cloud computing environment 50 (Figure 7) is shown. It should be understood in advance that the components, layers, and functions shown in Figure 7 are intended to be illustrative only, and that embodiments of the present invention are not limited thereto. As shown, the following layers and corresponding functions are provided:
[0077] Hardware and software layer 60 includes hardware and software components. Examples of hardware components include mainframe 61, RISC (Reduced Instruction Set Computer) architecture-based server 62, server 63, blade server 64, storage device 65, and network and network components 66. In some embodiments, software components include network application server software 67 and database software 68.
[0078] The virtualization layer 70 comprises an abstraction layer capable of providing virtual entities such as virtual servers 71, virtual storage 72, virtual networks including virtual private networks 73, virtual applications and operating systems 74, and virtual clients 75.
[0079] By way of example, the management layer 80 may provide the following functions: Resource provisioning 81 dynamically procures computing and other resources used to execute tasks within the cloud computing environment; Metering and pricing 82 tracks costs as resources are utilized within the cloud computing environment and sends bills or invoices for the utilization of those resources; by way of example, those resources may include application software licenses; Security verifies the identity of cloud users and tasks and protects data and other resources; User portal 83 provides users and system administrators with access to the cloud computing environment; Service level management 84 allocates and manages cloud computing resources to meet required service levels; and Service Level Agreement (SLA) planning and execution 85 proactively prepares and procures cloud computing resources in accordance with SLAs in anticipation of future demand.
[0080] The Workload Layer 90 illustrates examples of functionality available in a cloud computing environment. Examples of workloads and functionality that may be provided from this layer include mapping and navigation 91, software development and lifecycle management 92, virtual classroom instruction delivery 93, data analytics processing 94, transaction processing 95, and creation of mixed reality video conferencing environments 96.
[0081] It will be appreciated that the present disclosure may be implemented in conjunction with any other type of computing environment now known or later developed. For example, FIG. 9 illustrates a block diagram of a processing system 900 for implementing the techniques described herein. In the example, the processing system 900 includes one or more central processing units (processors) 921a, 921b, 921c, etc. (collectively or generally referred to as processors 921 and / or processing devices). In embodiments of the present disclosure, each processor 921 may include a reduced instruction set computer (RISC) microprocessor. The processors 921 are coupled to system memory (e.g., random access memory (RAM) 924) and various other components via a system bus 933. A read-only memory (ROM) 922 is coupled to the system bus 933 and may include a basic input / output system (BIOS) that controls certain basic functions of the processing system 900.
[0082] Also shown are an input / output (I / O) adapter 927 and a network adapter 926 coupled to the system bus 933. The I / O adapter 927 may be a small computer system interface (SCSI) adapter that communicates with a hard disk 923 and / or storage device 925, or any other similar components. The I / O adapter 927, hard disk 923, and storage device 925 are collectively referred to herein as mass storage 934. An operating system 940 for execution on the processing system 900 may be stored on the mass storage 934. The network adapter 926 interconnects the system bus 933 with an external network 936, enabling the processing system 900 to communicate with other such systems.
[0083] A display (e.g., a display monitor) 935 is connected to the system bus 933 by a display adapter 932, which may include a graphics adapter to improve performance for graphics-intensive applications and a video controller. In one aspect of the present disclosure, adapters 926, 927, and / or 932 may be connected to one or more I / O buses, which are connected to the system bus 933 through an intermediate bus bridge (not shown). I / O buses suitable for connecting peripherals such as hard disk controllers, network adapters, and graphics adapters typically include common protocols such as Peripheral Component Interconnect (PCI). Other input / output devices are shown connected to the system bus 933 via a user interface adapter 928 and a display adapter 932. Input devices 929 (e.g., keyboard, microphone, touchscreen, etc.), input pointer 930 (e.g., mouse, trackpad, touchscreen, etc.), and / or speakers 931 may be interconnected to the system bus 933 via a user interface adapter 928, which may, for example, include a super I / O chip that integrates multiple device adapters into a single integrated circuit.
[0084] In some embodiments of the present disclosure, the processing system 900 includes a graphics processing unit 937. The graphics processing unit 937 is a specialized electronic circuit designed to manipulate and modify memory to speed up the creation of images in a frame buffer for output to a display. In general, the graphics processing unit 937 is very efficient at operating in computer graphics and image processing, and has a highly parallel structure that makes it more effective than a general-purpose CPU for algorithms where the processing of large blocks of data is performed in parallel.
[0085] Thus, as configured herein, processing system 900 includes processing capability in the form of processor 921, storage capability including system memory (e.g., RAM 924) and mass storage 934, input means such as keyboard 929 and mouse 930, and output capability including speaker 931 and display 935. In some aspects of the present disclosure, a portion of the system memory (e.g., RAM 924) and mass storage 934 collectively store an operating system 940 for coordinating the functioning of the various components illustrated in processing system 900.
[0086] Various embodiments of the present invention are described herein with reference to the associated drawings. Alternate embodiments of the present invention may be devised without departing from the scope of the present invention. In the following description and in the drawings, various connections and relationships (e.g., above, below, adjacent, etc.) between elements are shown. These connections and / or relationships may be direct or indirect unless otherwise specified, and the present invention is not intended to be limited in this respect. Thus, coupling of entities may refer to direct or indirect coupling, and relationships between entities may be direct or indirect. Furthermore, various operations and process steps described herein may be combined into a more comprehensive procedure or process that includes additional steps or functions not specifically described herein.
[0087] One or more of the methods described herein may be implemented using any one or combination of techniques, such as discrete logic circuits including logic gates for implementing logic functions on data signals, application specific integrated circuits (ASICs) including appropriate combinatorial logic gates, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), and the like, each of which is well known in the art.
[0088] For purposes of brevity, prior art related to making and using aspects of the present invention may or may not be described in detail herein. In particular, various aspects of computing systems and particular computer programs for implementing various technical features described herein are well known. Thus, for purposes of brevity, many conventional implementation details are only briefly described or omitted entirely herein, without providing details of known systems and / or processes.
[0089] In some embodiments, various functions or operations may be performed at a particular location, or in conjunction with the operation of one or more devices or systems, or both. In some embodiments, a portion of a particular function or operation may be performed at a first device or location, and the remainder of the function or operation may be performed at one or more additional devices or locations.
[0090] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting. As used herein, the singular forms "a," "an," and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms "comprises" and / or "comprising," when used herein, specify the presence of stated features, integers, steps, operations, elements, or components, or combinations thereof, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, or groups thereof, or combinations thereof.
[0091] Corresponding structures, materials, acts, and equivalents of all means or steps and functional elements in the claims below are intended to include any structure, material, or act for performing the function in combination with other claimed elements as specifically claimed. This disclosure has been presented for purposes of illustration and description, but is not intended to be exhaustive or limited to the disclosed form. It will be apparent to those skilled in the art that many modifications and variations are possible without departing from the scope of the disclosure. The embodiments have been chosen and described to best explain the principles and practical application of the disclosure and to enable others skilled in the art to understand the disclosure in terms of various embodiments with various modifications as may be suited to the particular uses contemplated.
[0092] The diagrams shown herein are exemplary. There can be many variations of the diagrams or steps (or operations) described herein without departing from the scope of this disclosure. For example, operations can be performed in a different order, or operations can be added, deleted, or modified. Also, the term "coupled" indicates that a signal path exists between two elements, and does not imply a direct connection between elements with no intervening element / connection between them. All these variations are considered to be part of this disclosure.
[0093] The following definitions and abbreviations are used in interpreting the claims and this specification. As used herein, the terms "comprises," "comprising," "includes," "including," "has," "having," "containing," "containing," or any other variation thereof, are intended to cover a non-exclusive inclusion. For example, a composition, mixture, process, method, article, or apparatus containing a list of elements is not necessarily limited to only those elements and can include other elements not expressly stated or inherent in such composition, mixture, process, method, article, or apparatus.
[0094] Additionally, the term "exemplary" is used herein to mean "serving as an example, instance, or illustration." Any embodiment or design described herein as "exemplary" is not necessarily to be construed as preferred or advantageous over other embodiments or designs. The terms "at least one" and "one or more" are understood to include any integer number greater than or equal to one (i.e., 1, 2, 3, 4, etc.). The term "plurality" is understood to include any integer number greater than or equal to two (i.e., 2, 3, 4, 5, etc.). The term "connected" can include both an indirect "connected" and a direct "connected."
[0095] The terms "about," "substantially," "approximately," and variations thereof are intended to include the degree of error associated with measurement of the particular quantity based on the equipment available at the time of filing this application. For example, "about" can include a range of ±8%, or 5%, or 2% of the particular value.
[0096] The present invention may be a system, method, or computer program product, or any combination thereof, at any possible level of technical detail of integration. The computer program product may include a computer-readable storage medium containing computer-readable program instructions for causing a processor to perform aspects of the present invention.
[0097] A computer-readable storage medium can be a tangible device that can hold and store instructions for use by an instruction execution device. A computer-readable storage medium can be, for example, but is not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination thereof. A non-exhaustive list of more specific examples of computer-readable storage media includes portable floppy disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disc read-only memory (CD-ROM), digital versatile disk (DVD), memory stick, floppy disk, mechanically encoded devices such as punch cards or ridge structures in grooves on which instructions are recorded, and any suitable combination thereof. As used herein, a computer-readable storage medium should not itself be construed as a transitory signal such as an electric wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., a light pulse passing through a fiber optic cable), or an electrical signal transmitted over a wire.
[0098] The computer-readable program instructions described herein may be downloaded from a computer-readable storage medium to each computing / processing device or to an external computer or storage device over a network (e.g., the Internet, a local area network, a wide area network, and / or a wireless network). This network may include copper transmission cables, optical fiber transmissions, wireless transmissions, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface within each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage on a computer-readable storage medium within each computing / processing device.
[0099] Computer-readable program instructions for carrying out the operations of the present invention may be either assembler instructions, instruction-set-architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, configuration data for integrated circuits, or source or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk®, C++, and procedural programming languages such as the “C” programming language or similar programming languages. The computer-readable program instructions may execute entirely on the user's computer, partially on the user's computer as a standalone software package, partially on the user's computer and on a remote computer, or entirely on a remote computer or server. In the latter scenario, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection may be to an external computer (e.g., via the Internet using an Internet Service Provider). In some embodiments, to carry out aspects of the present invention, electronic circuitry including, for example, programmable logic circuits, field programmable gate arrays (FPGAs), or programmable logic arrays (PLAs), may execute computer readable program instructions to customize the electronic circuitry by utilizing state information of the computer readable program instructions.
[0100] Aspects of the present invention are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0101] These computer-readable program instructions may be provided to a processor of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus to create a machine, where the instructions, executed by the processor of the computer or other programmable data processing apparatus, create means for performing the functions / acts specified in the blocks of the flowcharts and / or block diagrams. These computer-readable program instructions may be stored on a computer-readable storage medium and capable of directing a computer, programmable data processing apparatus, or other device, or combination thereof, to function in a particular manner, such that the computer-readable storage medium on which the instructions are stored comprises an article of manufacture containing instructions for performing aspects of the functions / acts specified in the blocks of the flowcharts and / or block diagrams.
[0102] Computer-readable program instructions may be loaded into a computer, other programmable data processing apparatus, or other device such that the instructions, which execute on the computer, other programmable apparatus, or other device, perform the functions / acts specified in the flowchart and / or block diagram blocks, thereby causing a series of operable steps to be performed on the computer, other programmable apparatus, or other device to produce a computer-implemented process.
[0103] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of instructions, comprising one or more executable instructions for implementing the specified logical function(s). In some alternative implementations, the functions shown in the blocks may occur out of the order shown in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently or in the reverse order, depending on the functionality involved. It should also be noted that each block in the block diagrams and / or flowchart diagrams, and combinations of blocks included in the block diagrams and / or flowchart diagrams, may be implemented by a special-purpose hardware-based system that performs the specified function(s) or operation(s), or executes a combination of special-purpose hardware and computer instructions.
[0104] The description of various embodiments of the present invention is presented for illustrative purposes, but is not intended to be exhaustive or limited to the disclosed embodiments. Many changes and modifications will be apparent to those skilled in the art without departing from the scope of the described embodiments. The terms used herein are selected to best explain the principles, practical applications, or technical improvements of the embodiments beyond those found in the market, or to enable others skilled in the art to understand the embodiments described herein.
Claims
1. superimposing, by a processor, a graphical representation of a first user onto a real-world image of a first location, the first user being physically located at a second location different from the first location; detecting, by the processor, an interaction between the first user and a second user, the second user being physically located at the first location; determining, by the processor, a current location of the second user at the first location; causing, by the processor, the graphical representation of the first user to face in the direction of the current location of the second user; detecting, by the processor, an environmental stimulus to light or temperature at the first location; causing, by the processor, the graphical representation of the first user to respond to the detected environmental stimulus at the first location, wherein causing the graphical representation of the first user to respond to the detected environmental stimulus at the first location occurs even if the first user is unaware of the environmental stimulus.
11. A computer-implemented method comprising:
2. detecting, by the processor, the first user located at a physical table (hereinafter referred to as a second table) at a second location; detecting, by the processor, the second user located at a physical table (hereinafter referred to as a first table) at a first location different from the second location; further comprising the superimposing superimposes an avatar of the first user onto an empty position on the first table; The computer-implemented method of claim 1 .
3. The computer-implemented method of claim 2 , wherein the superimposing positions the avatar of the first user at the empty location regardless of the first user's actual positioning at the empty location.
4. overlaying a graphical representation of the second user onto a real-world image of the second location; determining a current location of the first user within the second location; causing the graphical representation of the second user to face in the direction of the current location of the first user; and The computer-implemented method of any one of claims 1 to 3, further comprising:
5. 5. The computer-implemented method of claim 1, further comprising: placing the graphical representation of the first user over the real-world image of the first location based at least in part on a physical layout of the first location.
6. 6. The computer-implemented method of claim 5, further comprising placing the graphical representation of the second user over the real-world image of the second location based at least in part on a physical layout of the second location, wherein the physical layout of the first location is different from the physical layout of the second location.
7. Detecting an interaction between the first user and the second user includes: calculating a position and field of view of the first user's computing device; Detecting a location of the second user's computing device; comparing the position and field of view of the first user's computing device with the position of the second user's computing device; The computer-implemented method of any one of claims 1 to 6, comprising:
8. Detecting an interaction between the first user and the second user includes: detecting the graphical representation of the first user on a display of a computing device of the second user; determining whether the time interval during which the graphical representation of the first user was displayed exceeds a threshold time; The computer-implemented method of any one of claims 1 to 7, comprising:
9. 9. The computer-implemented method of claim 1, further comprising sharing the model of the second location with a computing device of a third user, the third user being physically located at the second location.
10. a memory containing computer readable instructions; one or more processors for executing the computer-readable instructions; wherein the computer readable instructions cause the one or more processors to: A system for executing the steps of the computer-implemented method according to any one of claims 1 to 9.
11. A computer program product for causing a processor to execute the steps of the computer-implemented method according to any one of claims 1 to 9.
12. A storage medium storing the computer program according to claim 11.
Citation Information
Patent Citations
Semiconductor storage device
JP1989010346A
Animated electronic conference room and video conference system and method
JP1995255044A
Avatar communication system
JP2001034787A
Conversation system, conversation method, and computer program using virtual space
JP2011039860A
Conference system, conference method and program
JP2014225801A