Creating mixed reality conference room

By generating and aligning the feature point cloud of physical objects, the alignment problem of remote physical locations is solved, and an immersive and interactive mixed reality conference room is realized, combining the advantages of VR and AR.

CN120345005APending Publication Date: 2025-07-18INTER IKEA SYST
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380087957.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-12-22
Filing Date
2023-12-20
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The prior art is unable to achieve precise alignment and placement of two remote physical locations, resulting in insufficient immersion and interactivity in VR and AR conference room experiences.

Method used

Feature point clouds are generated by scanning physical objects in different physical environments, feature point detection algorithms are used to determine candidate anchor parts, and compared with reference feature point clouds, common anchor parts are derived, and feature point clouds are aligned to render the visual representation of a mixed reality conference room.

Benefits of technology

Achieving precise alignment and placement of two remote physical locations and users, delivering an immersive and interactive mixed reality meeting room experience combining the benefits of VR and AR.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120345005A_ABST
    Figure CN120345005A_ABST
Patent Text Reader

Abstract

A computer-implemented method (100) for creating a mixed reality (MR) conference room (70) is provided. Two MR devices (10, 30) are configured to scan respective physical objects (22, 42) located in respective physical environments (40) in order to generate feature point clouds (24, 44) thereof. A feature point detection algorithm is applied to determine candidate anchoring portions (55). The candidate anchor portion is compared to a reference feature point cloud (64) and at least one common anchor portion (50) is derived. The feature point clouds (24, 44) are aligned relative to the common anchor portion (50), and the MR device (10, 30) renders a visual representation of the MR meeting room (70) based on the aligned feature point clouds (24, 44).
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure generally relates to mixed reality. More specifically, the present disclosure relates to a computer-implemented method for creating a mixed reality conference room. The present disclosure also relates to related computerized systems, non-transitory computer-readable storage media, and computing devices. Background Art

[0002] Mixed reality (MR) is the reality where the physical world meets the digital world. The MR spectrum covers human-environment-computer interaction that extends from virtual reality (VR) to physical reality, including the fields of augmented reality (AR) and augmented virtual (AV) technology. Although the term "mixed reality" was coined as early as the 1990s, advances in hardware and software technology in recent years have driven groundbreaking innovations in the MR field. For example, in the digital world known as the "metaverse", people can meet and interact in the virtual world even if they are physically far apart. With the help of the human sensory system and perceptual rules, MR provides users with an immersive environment in which the real-time coordination between physical positions and gestures is virtually reflected.

[0003] Despite some commonalities, as mentioned above, different areas of MR require very different technical solutions to function properly. This is particularly evident when determining how objects in the environment should be viewed and placed (e.g., where they should be located and at what angle / direction) and how the avatar (i.e., the user's virtual representation in the virtual world) should be placed and oriented relative to said objects. Existing solutions for creating MR meeting rooms currently have many shortcomings in both VR and AR.

[0004] In VR, users cannot see the physical world. Therefore, in a VR meeting room, users can be placed anywhere to participate in the experience. Users are usually presented as virtual avatars and participate in the virtual world together. For example, in a reference virtual world, two users can be placed in a virtual reference position through their virtual avatars simultaneously or successively. The virtual reference position can be any suitable surface, such as a floor surface, so that the virtual avatar can be placed standing (i.e., roughly perpendicular to the floor surface).

[0005] The disadvantage of VR conference rooms is that participating users will directly feel that the environment is artificial, which will have a negative impact on the immersion of the experience. Therefore, it is desirable to provide a virtual environment that is as close as possible to the physical environment of the room in which the user is actually located, both in terms of spatial environment (e.g., objects in the room) and spatial dimensions (e.g., physical boundaries defined by the room).

[0006] With AR technology, it is easier to provide an environment that closely approximates the physical environment of the user's room. This is because AR is built around the real world, which is composed of computer-generated images presented realistically. To enable AR functionality, AR devices and their software platforms need to identify certain objects in the physical world and build the AR experience based on this. Specifically, in an AR experience, the placement location of virtual content must be determined. For example, marker-based AR involves providing multiple different types of markers that can be used to place content, such as planar markers, image markers, face markers, or 3D object markers. These markers ensure that the content fixed to the marker maintains its original position and orientation in the virtual space, thus helping to maintain the illusion of virtual objects being placed in the physical world. Markers are typically used in combination with an inward-facing visual inertial SLAM tracking technology (viSLAM), which can achieve stable tracking performance even in frames where markers cannot be reliably detected. However, in marker-based AR, AR content can only be directly associated and displayed with fiducial markers, which have a pattern that uniquely identifies each marker. This allows for simple tracking procedures, but only if the marker is within the camera's field of view.

[0007] Similar to VR meeting rooms, AR meeting rooms also have some drawbacks. In particular, with currently known technologies, it is technically impossible to automatically achieve the precise alignment and placement of two remote physical locations, all the associated physical objects, and virtual avatars.

[0008] The present inventors have recognized the above-mentioned deficiencies in the prior art and hereby propose an improvement to provide an immersive and interactive MR meeting room experience. Summary of the Invention

[0009] Accordingly, the present inventors have proposed a method for creating a fully immersive MR conference room in which two participants from two physically different environments can meet. According to a first aspect of the present disclosure, there is provided a computer-implemented method for creating a mixed reality (MR) conference room. The method includes: scanning a first physical object located in a first physical environment by a first MR device; scanning a second physical object located in a second physical environment by a second MR device, wherein the first physical environment and the second physical environment are spatially different; during the scanning, generating first and second feature point clouds of the first and second physical objects respectively, each of the first and second feature point clouds including a plurality of feature points, each feature point having a unique spatial coordinate with respect to the associated physical environment; applying a feature point detection algorithm to determine candidate anchor portions / anchor regions of each of the first and second feature point clouds; comparing the candidate anchor portions with a reference feature point cloud of a virtual representation of a reference physical object; in response to a case where a matching condition of the comparison is satisfied, deriving at least one common anchor portion among the first, second, and reference feature point clouds from the candidate anchor portions; aligning the first and second feature point clouds such that they coincide with each other with respect to the at least one common anchor portion; and rendering a visual representation of the MR conference room by each of the first and second MR devices based on the aligned feature point clouds, such that a virtual representation of a user of the first MR device is visualized in the MR conference room with respect to the second physical object or its virtual representation, and such that a virtual representation of a user of the second MR device is visualized in the MR conference room with respect to the first physical object or its virtual representation.

[0010] For the sake of enhancing the coherence of the following disclosure, the term "MR enabled object" is introduced. An MR enabled object should be understood as a physical object around which the MR conference room is fully constructed in terms of layout and orientation, physical and virtual objects to be included in the MR conference room, and associated virtual avatars. Thus, an MR enabled object can be used as an object anchor that is spatially unrestricted and from different physical locations. For this purpose, by simply physically placing the respective MR enabled objects in their respective physical environments and storing the corresponding reference physical objects digitally, an immersive and interactive MR conference room experience can be enjoyed, which benefits from the concept of the first aspect of the present invention.

[0011] According to a first aspect, the present invention effectively utilizes the general concept that the feature point cloud representations of similar MR-enabled objects will be similar to each other after processing their respective feature points. By comparing a candidate anchoring portion with the feature point cloud of the virtual representation of a reference MR-enabled object, subsequent alignment procedures can be performed. The candidate anchoring portion of a physical object can be identified with high precision, aiming to provide an immersive and interactive MR conference room experience in which almost any number of users can participate simultaneously. Thus, while enjoying the advantages of a VR conference room experience in terms of precise object orientation and alignment, the realistic effects of an AR conference room experience can also be enjoyed.

[0012] The use cases for such applications are numerous. Advantageously, any type of object that can be found in a home or office environment can represent an MR-enabled object as long as the corresponding MR-enabled object is provided as a reference object. For example, people who want to sit together on a couch to watch TV, play video games, or just have a normal face-to-face conversation can use the couch as an MR-enabled object. In another example, people who want to participate in a group training session can use a carpet or a yoga mat as an MR-enabled object. As another example, people may want to play a board game, such as chess or a similar game, together on a table. In this case, either the table or the board game layout itself can be used as an MR-enabled object. Those skilled in the art will appreciate that, by virtue of the concepts according to the first aspect, a variety of different MR conference experiences can be additionally or alternatively achieved. The embodiments / examples / claims described in the present disclosure can achieve further technical advantages.

[0013] In one or more embodiments, the reference physical object shares at least one physical property with the first and second physical objects.

[0014] In one or more embodiments, the at least one physical property is texture, size, pattern, furniture type, or another property that enables the feature point cloud representation thereof to be distinguishable from the feature point cloud representations of other physical properties of the reference physical object.

[0015] In one or more embodiments, matching conditions are determined based on corresponding feature descriptor data of at least one physical property.

[0016] In one or more embodiments, the feature point detection algorithm is a salient feature point detection algorithm.

[0017] In one or more embodiments, users of the first and second MR devices participate in the MR conference room simultaneously.

[0018] In one or more embodiments, the method further includes rendering a visualization of virtual representations of one or more additional physical objects such that the one or more additional physical objects are visualized relative to the first physical object or the second physical object or their virtual representations in the MR meeting room.

[0019] In one or more embodiments, the MR meeting room is a virtual reality (VR) meeting room, where users of the first and second MR devices are virtually rendered as avatars and virtually participate in the VR meeting room through their avatars, and the avatars are visualized relative to the virtual representations of the first physical object or the second physical object in the VR meeting room.

[0020] In one or more embodiments, the MR meeting room is an augmented reality (AR) meeting room, where the user of the first MR device is virtually rendered as an avatar and virtually participates in the AR meeting room through the avatar, and where the user of the second MR device is physically participating / actually participating in the AR meeting room.

[0021] In one or more embodiments, the rendering of the virtual representation of the user is performed based on the spatial relationship between the user and at least one common anchoring portion, so as to visualize the virtual representation of the user at a spatial position relative to the common anchor point in the MR meeting room.

[0022] In one or more embodiments, at least one common anchoring portion defines the spatial position and / or orientation of the MR meeting room.

[0023] In one or more embodiments, the first and second physical objects are selected from the group consisting of: household furniture, household appliances, household equipment, office furniture, office appliances, and office equipment.

[0024] In one or more embodiments, the scanning is performed by near real-time 3D scanning technology implemented by the MR device.

[0025] In one or more embodiments, the feature points are voxels, and each voxel has its own unique 3D coordinates.

[0026] In one or more embodiments, the MR meeting room is persistent.

[0027] According to a second aspect, there is provided a computerized system for creating a mixed reality (MR) meeting room. The system includes a first MR device, a second MR device, and a backend service, where: the first MR device is configured to scan a first physical object located in a first physical environment; the second MR device is configured to scan a second physical object located in a second physical environment, the first physical environment and the second physical environment being spatially different; wherein each MR device and / or the backend service is further configured to: during the scanning, generate first and second point clouds of features of the first and second physical objects respectively, each of the first and second point clouds of features containing a plurality of feature points, each feature point having a unique spatial coordinate relative to the relevant physical environment; apply a feature point detection algorithm to determine candidate anchoring portions of each of the first and second point clouds of features; compare the candidate anchoring portions with a reference point cloud of features of a virtual representation of a reference physical object; in response to a case where a matching condition of the comparison is satisfied, derive at least one common anchoring portion between the first, second, and reference point clouds of features from the candidate anchoring portions; and align the first and second point clouds of features such that they coincide with each other with respect to the at least one common anchoring portion; and wherein each of the first and second MR devices is further configured to render a visual representation of the MR meeting room based on the aligned point clouds of features, such that a virtual representation of a user of the first MR device is visualized in the MR meeting room relative to the second physical object or its virtual representation, and such that a virtual representation of a user of the second MR device is visualized in the MR meeting room relative to the first physical object or its virtual representation.

[0028] In one or more embodiments, the MR meeting room is persistent.

[0029] In one or more embodiments, the first and second MR devices are selected from the group consisting of: a head-mounted display (HMD), a cave automatic virtual environment (CAVE), or an augmented reality (AR) device.

[0030] In a third aspect, there is provided a non-transitory computer-readable storage medium. The non-transitory computer-readable storage medium contains instructions that, when executed by a processor device, cause the processor device to perform the function of one of the first or second MR devices in the method described in the first aspect.

[0031] In a fourth aspect, there is provided a computing device. The computing device includes a processor device configured to perform the function of one of the first or second MR devices in the method described in the first aspect.

[0032] It will be apparent to those of ordinary skill in the art that the above aspects, the appended claims, and / or the examples disclosed above and below can be appropriately combined with each other.

[0033] Additional features and advantages are disclosed in the following description, claims, and drawings, and some of the features and advantages will be apparent to those skilled in the art or will be recognized by practicing the disclosure described herein. Also disclosed herein are control units, computer-readable media, and computer program products associated with the above-described technical advantages. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] Aspects of the present disclosure, which are mentioned as examples, will now be described in more detail with reference to the drawings.

[0035] Figure 1 is an exemplary illustration showing the creation of an MR conference room.

[0036] Figure 2 is an exemplary schematic diagram illustrating data related to the creation of an MR conference room.

[0037] Figure 3A is an exemplary schematic diagram illustrating a computerized system that can be configured to create an MR conference room.

[0038] Figure 3B is for Figure 3A an exemplary computing environment of a computerized system.

[0039] Figure 4 is an exemplary flowchart method for creating an MR conference room.

[0040] Figure 5 is an exemplary computer-readable storage medium. DETAILED DESCRIPTION

[0041] The aspects set forth below represent the necessary information that enables those skilled in the art to implement the present disclosure.

[0042] Figure 1 is an exemplary illustration of creating an MR conference room 70. As can be seen in the figure, a first person is wearing a first MR device 10 in a first physical environment 20. A second person is also shown wearing a second MR device 30 in a second physical environment 40. The physical environments 20, 40 are spatially different, which means they are different environments that are physically far apart. For example, the first environment 20 can be the home / office of the first person, while the second environment 40 can be the home / office of the second person, and the two people live or work at different addresses, cities, or even countries.

[0043] MR devices 10, 30 scan respective physical objects 22, 42 located in respective physical environments 20, 40. As described in the Summary of the Invention section, physical objects 22, 42 are MR-enabled objects / objects with MR capabilities. Physical objects 22 and 42 are not limited to a specific type of object. In this example, physical objects 22, 42 are sofas. The Summary of the Invention section describes some other use cases, where physical objects 22, 42 are carpets, yoga mats, tables, or game boards. Any other object that is typically found in a suitable environment, such as a home, office, restaurant, hotel, school, coffee shop, park, or beach environment, etc., can alternatively be used as an MR-enabled object. In some examples, physical objects 22, 42 are selected from the group consisting of: household furniture, household appliances, household equipment, office furniture, office appliances, and office equipment. In this regard, furniture, appliances, and equipment should be understood in the broadest possible sense, i.e., bicycles, musical instruments, stuffed toys, curtains, fireplaces, lamps, etc., can all be considered candidates for physical objects 22, 42. Obviously, household and office appliances, furniture, and / or equipment may also be present in other places outside the home / office, such as restaurants, hotels, schools, coffee shops, parks, or beaches.

[0044] Physical objects 22, 42 may share at least one physical property. For example, the physical property may be texture, size, pattern, furniture type, or color, etc. The person operating MR devices 10, 30 can visually distinguish these physical properties.

[0045] The scanning can be based on any scanning technology known in the art, such as a near real-time 3D scanning tool. For example, such tools may include LiDAR (Light Detection and Ranging), stereo camera systems, depth camera systems, structured light projection systems, etc. Alternatively, the scanning can also be based on photogrammetry software.

[0046] During the said scanning, respective feature point clouds 24, 44 of the first physical object 22 and the second physical object 42 are generated. Thus, a feature point cloud is a known concept to those skilled in the art. Each feature point cloud 24, 44 includes a plurality of feature points 26, 46, and each feature point 26, 46 has unique spatial coordinates (e.g., Cartesian coordinates x, y, z) relative to the relevant physical environment 20, 40. Generating feature point clouds 24, 44 based on the scanning can be based on any feature point generation technology known in the art, such as the above-mentioned 3D scanning technology or photogrammetry technology. The scanning results are as Figure 1 shown, where each physical object 22, 42 is digitally represented as feature points 26, 46. Feature points 26, 46 can be represented as voxels, where each voxel contains respective unique 3D coordinates relative to the corresponding physical environment 20, 40.

[0047] Figure 1MR conference room 70 is further shown in. The visualization of MR conference room 70 is generated by each MR device 10, 30. Before rendering the visual representation of the above MR conference room 70, several operations need to be performed first, which will be described in more detail later in this disclosure in conjunction with Figure 2 For a more detailed description.

[0048] MR devices 10 and 30 will render different visual representations of MR conference room 70 according to different factors. In some examples, two users of MR devices 10 and 30 experience MR conference room 70 at the same time as if they are both being accessed by each other (i.e., both users are "interviewees"). Therefore, both users of MR devices 10 and 30 will see the other user appear in the form of an avatar, but located in the other user's own home. Thus, the two users can experience the virtual access of the other user at the same time. In some examples, one of the users is the interviewee and the other user is the visitor.

[0049] MR conference room 70 can be persistent. The persistence of MR conference room 70 refers to extending the existence time of digital content beyond the actual usage time of the system, so that the content has a permanent position in MR conference room 70. To this end, MR conference room 70 can contain multiple conference room sessions. This enables the users of MR conference room 70 to join or exit the virtual world without losing progress. For example, in a digital painting session using a real canvas as physical object 22 or 42, although the user may enter or exit a certain MR conference room session multiple times, the artist's current creation progress can be restored at any point in time. Alternatively, persistence can also be provided for any other object in the VR conference room (i.e., not necessarily an MR-enabled object).

[0050] In some examples, although Figure 1 not explicitly shown in, visual representations of virtual representations of one or more additional physical objects can be rendered. These virtual representations are visualized in MR conference room 70 with reference to the first physical object 22 or its virtual representation or the second physical object 42 or its virtual representation.

[0051] The MR meeting room 70 can be a VR meeting room. In an example where the MR meeting room 70 is a VR meeting room, both users are virtually rendered as virtual avatars and virtually participate in the VR meeting room through their respective virtual avatars. Thus, both users will see a completely virtual world that closely resembles the real physical environment where either user is located. Alternatively, the virtual world can include features common to the two corresponding physical worlds, such as the fusion of certain physical aspects. The visualization of the virtual avatar is based on the virtual representation of the first physical object 22 or the second physical object 42. Since the VR meeting room is completely virtual, the sofa is rendered as a "visually perfect" sofa (in a more general example, simply referred to as a "visually perfect object").

[0052] The MR meeting room 70 can be an AR meeting room. In an example where the MR meeting room 70 is an AR meeting room, depending on which of the MR devices 10, 30 is rendering the MR experience, the rendering of each user will be different. From the perspective of the MR device 10, the other user will be virtually rendered as a virtual avatar and will virtually participate in the AR meeting room through the virtual avatar; from the perspective of the MR device 30, it is the opposite. In other words, both users will experience the other user as a virtual avatar located in their "own" physical environment.

[0053] The MR meeting room 70 can be a meeting room that combines AR and VR. For this purpose, the user of the MR device 10 can experience a VR meeting room, while the user of the MR device 30 can experience an AR meeting room.

[0054] In view of the above, the MR meeting room 70 is "cross-rendered", that is, it is rendered by the corresponding MR devices 10, 30 for all participating users, but not necessarily in the same way.

[0055] Therefore, compared with the prior art, the disclosed MR meeting room 70 achieves the precise alignment and placement of two remote physical locations 20, 40 and physical objects 22, 42 and the users (or virtual avatars of the users) associated therewith.

[0056] Figure 2 A data schematic diagram related to the creation of the MR meeting room 70 is shown. Figure 1 At least conceptually explains how to scan the first physical object 22 and the second physical object 42 to create the corresponding feature point clouds 24, 44, each of which contains a set of feature points 26, 46, thereby creating the MR meeting room 70. More technical details and implications will now be further elaborated. As Figure 2As shown, by scanning another physical object 92 and generating corresponding subsequent data (including another feature point cloud 94, associated feature points 96, and another candidate anchoring part / anchoring region 55), any number of additional users can participate in the experience provided by the MR conference room 70. To this end, it can be understood that the interviewee in the MR conference room 70 can be accessed by any number of visitors.

[0057] The present disclosure determines the candidate anchoring part 55. This determination, in combination with the derivation of at least one common anchoring part 50, determines how to perform the anchoring process of the respective virtual representations of the physical objects 22, 42, 92 relative to each other. Thus, this enables the MR conference room 70 to be aligned and enables the participating users (i.e., the interviewee and any number of visitors) to enjoy the relevant MR experience.

[0058] The "parts" used with reference to the candidate anchoring part 55 and at least one common anchoring part 50 should be interpreted broadly. The anchoring parts 50, 55 can be the smallest parts distinguishable from other parts of the feature point clouds 24, 44, 94. The anchoring parts 50, 55 can also be larger parts of the feature point clouds 24, 44, 94, such as corresponding to the seat cushion, armrest, or neck pillow of a sofa, or even the entire sofa (if its significant information needs to be distinguished). The anchoring parts 50, 55 can also be of any size between the above-mentioned smallest and largest parts.

[0059] To determine the candidate anchoring part 55, the feature descriptor data of the feature point clouds 24, 44, 94 can be calculated. The feature descriptor data can include edges. The edges correspond to the recognizable boundaries between different image regions of the feature point clouds 24, 44, 94. The feature descriptor data can include corner points. The corner points correspond to the recognizable rapid changes in the direction of the image regions of the feature point clouds 24, 44, 94. The feature descriptor data can contain blobs. The blobs correspond to the local maxima of the image regions of the feature point clouds 24, 44, 94 or the center of gravity of the feature points therein. The feature descriptor data can include ridge lines. The ridge lines correspond to one-dimensional curves representing the symmetry axes within the feature point clouds 24, 44, 94. Other suitable feature descriptor data can alternatively be used to determine the candidate anchoring part 55.

[0060] Calculating the feature descriptor data can be accomplished by applying a feature point detection algorithm. The feature point detection algorithm is configured to calculate information about the feature point clouds 24, 44, 94 and determine whether specific feature points 26, 46, 96 correspond to an image feature of a given type. Such an image feature of a given type may correspond to significant information about the respective physical objects 22, 42, 92, i.e., information that is considered to be of interest. Thus, significant information can be obtained by applying a significant feature point detection algorithm. The significant information can indicate physical attributes such as texture, size, pattern, furniture type, or other attributes by which its feature point cloud representation can be distinguished from the feature point cloud representations of other physical attributes of the physical objects 22, 42, 92.

[0061] Although there is no widely recognized definition for what constitutes an object of interest for a feature point detection algorithm in the art, it is generally believed that some feature descriptor data is easier to detect than other data, hence the definitions related to ridges, corners, blobs, and edges, among others, as described above. Feature detection algorithms typically perform low-level image processing operations at the pixel level (or subsets of pixels). Thus, repeatability is preferably achieved at the pixel level, i.e., the ability to detect common features in different feature point clouds 24, 44, 94. Different algorithms known in the art handle this differently, and the present disclosure is not limited to a particular type of feature point detection algorithm. Some exemplary algorithms that can be applied can be one of SIFT (Scale-Invariant Feature Transform), edge detection, FAST (Features from Accelerated Segment Test), contour curvature, Canny edge detector, Sobel-Feldman operator, corner detection (e.g., Hessian intensity feature measurement, SUSAN, Harris corner detector, contour curvature, or Shi&Tomasi algorithm), blob detection (e.g., Laplacian of Gaussian, determinant of Hessian, or gray-level blob), difference of Gaussians, MSER (Maximally Stable Extremal Regions), ridge detection (e.g., principal curvature ridge), to name just a few examples.

[0062] As described herein, the physical objects 22, 42, 92 are typically objects commonly found in environments such as homes, offices, restaurants, hotels, schools, coffee shops, parks, or beaches, to name just a few exemplary environments. Thus, their feature point clouds 24, 44, 94 will vary depending on the object type and the corresponding detectability of the applied feature point detection algorithm. Thus, some algorithms may work better for certain types of objects than others, depending on a variety of different factors such as texture, pattern, furniture type, and / or other distinguishable physical attributes.

[0063] For example, solid wood textures are easier to distinguish compared to mirror textures because they contain more structural details determined by the wood itself. Therefore, compared to the feature descriptor data of the feature point cloud representation of a mirror, the corresponding feature descriptor data of the feature point cloud representation of a wooden table, for example, may contain more distinguishable ridge lines, corner points, patches, edges, and curves. Thus, when the physical objects 22, 42, 92 are wooden tables instead of mirrors, the determination accuracy of the candidate anchor part 55 is higher. Similar examples can be achieved with other physical properties. For example, compared to a sofa with a uniform color (where the feature descriptor data is more difficult to distinguish), a sofa with a floral pattern (where the feature descriptor data is easier to distinguish), or compared to a chair with smooth corners (where the feature descriptor data is more difficult to distinguish), a chair with a large number of sharp corners / edges (where the feature descriptor data is easier to distinguish). Therefore, it can be understood that for certain physical objects 22, 42, 92, the alignment of the MR conference room 70 will be more accurate than for other physical objects.

[0064] Once the candidate anchor part 55 is determined, one or more common points among all the candidate anchor parts 55 need to be determined as at least one common anchor part 50. This is achieved by comparing the candidate anchor part 55 with reference information. The reference information is the reference feature point cloud 64 of the virtual representation of the reference physical object 62. The virtual representation of the reference physical object 62 is a 3D model of a physical object similar to the physical objects 22, 42, 92, that is, a virtual representation of an MR-enabled object. The reference physical object 62 can share at least one physical property with the physical objects 22, 42, 92 and thus can contain corresponding distinguishable physical property feature descriptor data.

[0065] At least one common anchor part 50 is derived from the candidate anchor part 55 in response to the matching condition being satisfied. The matching condition can be based on the feature descriptor data of at least one physical property of the virtual representation of the reference physical object 62, that is, the significant features of the reference feature point cloud 64. The feature descriptor data of the reference feature point cloud 64 can be obtained from the model file of the reference physical object 62. The model file can be a CAD file (such as STEP, QIF, JT, or 3D PDF), a mesh representation file (such as OBJ, GLTF, etc.), a neural radiance field model, or other data-readable files (such as json, yaml, xml, etc.), just to name a few examples. The feature descriptor data of the reference feature point cloud 64 can also be retrieved as a previous scan sample of the physical object.

[0066] By using the comparison as described above to derive the candidate anchor part 55, a very high confidence can be provided in the matching. The matching conditions can be set accordingly. The matching conditions are configured to determine what conditions one or more candidate anchor parts 55 must meet to be regarded as common parts. The matching conditions can be set to any percentage value, such as 80%, 90%, 95%, etc. The setting of the matching conditions may vary depending on the physical properties being compared and the way they are compared. For example, a certain texture that produces a specific ridge line may require a 99% matching rate to establish a common part, while a physical object of a certain size that produces a specific corner / edge may only require a 60% matching rate to establish a common part. The matching conditions may depend on the number of candidate anchor parts 55 being compared. Thus, the larger the number of similar candidate anchor parts 55, the greater the likelihood of a match may be indicated, and thus the more confident the determination of the common anchor part 50 will be.

[0067] Once at least one common anchor part 50 has been derived, the feature point clouds 24, 44, 94 are aligned relative to the anchor part 50 so that they coincide with each other. The alignment can be performed by using a 3D rigid or affine geometric transformation algorithm to align the arrays of the feature point clouds 24, 44, 94 into an aligned feature point cloud. A box grid filter can be applied to the aligned feature point cloud of a 3D box with a specific size. The feature points 26, 46, 96 within the same box grid can be merged into one point. To achieve the alignment, a feature point cloud alignment algorithm can be applied. Some exemplary algorithms that can be applied can be one of ICP (Iterative Closest Point), CPD (Coherent Point Drift), NDT (Normal Distribution Transform), FCGF (Fully Convolutional Geometric Features), D3Feat (Joint Learning Algorithm for Dense Detection and Description of 3D Local Features), PREDATOR (Point Cloud Registration with Deep Attention to Overlapping Regions), PointNet-LK (PointNet Lucas & Kanade), PCRNet (Point Cloud Registration Network), DCP (Deep Closest Point), DGR (Deep Global Registration), PRNet (Partial Registration Network), RPMNet (Robust Point Matching Network), just to name a few.

[0068] The last step of the program is to render the MR conference room 70. The rendering of the MR conference room 70 can use any 3D graphics rendering software known in the art, but none of these software constitutes a limitation. Suitable rendering software includes Unity, ARKit, Unreal Engine, OctaneReader, 3DSMax, V-Ray, Corona Renderer, and MaxwellRay, just to name a few.

[0069] The rendering of the virtual representation of the user can be based on the spatial relationship between the user and at least one common anchor 50. Thus, the virtual representation of the user is spatially located within the MR conference room 70 relative to the common anchoring portion 50. The at least one common anchoring portion 50 can define the spatial position and / or orientation of the MR conference room 70. Thus, the MR world is constructed around the at least one common anchoring portion 50, and the virtual content of the world (e.g., including the virtual representations of the user and objects) is virtually placed in a spatial relationship with the common anchoring portion 50.

[0070] Figure 3A For an exemplary schematic diagram, a computerized system 200 configured to create an MR conference room is illustrated. The computerized system 200 includes a server-side platform 210 and two MR devices 10, 30.

[0071] In the example provided, the MR devices 10, 30 are head-mounted displays (HMDs). In other examples, the MR devices 10, 30 can be any type of computing device known in the art. The computing device can be an AR device, such as a smart tablet, smartphone, laptop, smartwatch, smart glasses, or a Neuralink chip. The computing device can be a CAVE virtual environment.

[0072] The server-side platform 210 can be deployed on a cloud-based server, which can be implemented using any known cloud computing platform technology, such as Amazon Web Services (AWS), Google Cloud Platform (GCP), Microsoft Azure, DigitalOcean, Oracle Cloud Infrastructure (OCI), IBM Bluemix, or Alibaba Cloud. The cloud-based server can be included in a widely open distributed cloud network or limited to internal enterprise use. Alternatively, in some embodiments, the cloud-based server can be locally managed as, for example, a centralized server unit. Other alternative server configurations can be implemented based on any type of client-server or peer-to-peer (P2P) architecture. Thus, the server configuration can involve any combination of, for example, a web server, database server, email server, web proxy server, DNS server, FTP server, file server, DHCP server, etc., to name just a few examples.

[0073] The server - side platform 210 includes computing resources 212 and storage resources 214 in an operable communication state. The computing resources 212 are configured to perform processing activities discussed in this disclosure, such as generating a feature point cloud, applying a feature point detection algorithm, comparing candidate anchor portions with a reference feature point cloud, deriving at least one common anchor portion, and aligning the feature point cloud. To this end, the computing resources 212 include one or more processor devices that are configured to process data related to creating an MR meeting room.

[0074] The storage resources 214 can be maintained by a cloud service and / or be configured as cloud - based services, either internal or external to the computing resources 110. A connection to the storage resources 214 can be established using DBaaS (Database as a Service). For example, the storage resources 214 can be deployed as an SQL data model, such as MySQL, PostgreSQL, or Oracle RDBMS. Alternatively, a deployment based on a NoSQL data model, such as MongoDB, Amazon DynamoDB, Hadoop, or Apache Cassandra, can also be used. DBaaS technology is typically included as a service in relevant cloud computing platforms.

[0075] The MR devices 10, 30 are configured to perform processing activities discussed in this disclosure, such as scanning a physical environment and rendering a visual representation of the MR meeting room. In some examples, the MR devices 10, 30 can also be configured to perform the functions described with reference to the computing resources 212. In yet other examples, the MR devices 10, 30 and the computing resources 212 can be configured to perform different functions as described herein.

[0076] Communication between the server - side platform 210 and the MR devices 10, 30 can be achieved by means of any short - range or long - range wireless communication standard known in the art. For example, the wireless communication can be implemented by technologies including but not limited to: IEEE 802.11, IEEE 802.15, ZigBee, WirelessHART, WiFi, Bluetooth®, BLE, RFID, WLAN, MQTT IoT, CoAP, DDS, NFC, AMQP, LoRaWAN, Z - Wave, Sigfox, Thread, EnOcean, Mesh communication, any form of short - range device - to - device wireless communication, LTE Direct, W - CDMA / HSPA, GSM, UTRAN, LTE, or Starlink.

[0077] As Figure 3BAs shown, the computerized system 200 may include multiple units known to those skilled in the art for implementing the functions described in the present disclosure. The computerized system 200 may include one or more computing units, which can include firmware, hardware, and / or execute software instructions to implement the functions described herein. The computerized system 200 may include one or more processors (which may also be referred to as control units) 230, one or more memories 235, and one or more buses 240. The processors 230 may be respectively included in the computing devices 222a-e and the computing resources 212. The computerized system 200 may include at least one computing device having a processor 230. The system bus 240 may provide an interface for system components including, but not limited to, the memory 235 and the processor 230. The processor 230 may include any number of hardware components for data or signal processing, or for executing computer code stored in the memory. For example, the processor 230 may include a general-purpose processor, a special-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a circuit including processing components, a group of distributed processing components, a group of distributed computers configured for processing, or other programmable logic devices, discrete gate or transistor logic, discrete hardware components, or any combination thereof, for performing the functions described herein. The processor 230 may also include computer-executable code for controlling the operation of programmable devices.

[0078] The system bus 240 can be any one of a variety of types of bus structures, which can also use any one of a variety of bus architectures to interconnect with a memory bus (with or without a memory controller), a peripheral bus, and / or a local bus. The memory 235 can be one or more devices for storing data and / or computer code for performing or facilitating the methods described herein. The memory 235 can include database components, object code components, script components, or other types of information structures for supporting the various activities described herein. Any distributed or local memory device can be used with the systems and methods of this specification. The memory 235 can be communicatively connected to the processor 230 (e.g., via circuitry or any other wired, wireless, or network connection), and can contain computer code for performing one or more of the processes described herein. The memory can include non-volatile memory (e.g., read-only memory (ROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), etc.) and volatile memory (e.g., random access memory (RAM)), or any other medium that can be used to carry or store the required program code (in the form of machine-executable instructions or data structures) and can be accessed by a computer or other machine equipped with a processor device. The basic input / output system (BIOS) can be stored in non-volatile memory and can contain basic routines that help transfer information between the various elements within a computer system.

[0079] The memory 245 can be operatively connected to the computerized system 200 via, for example, an I / O interface (e.g., a card, a device) 250 and an I / O port 255. The memory 245 can include, but is not limited to, devices such as disk drives, solid-state drives, optical disc drives, flash memory cards, USB drives, etc. The memory 245 can also include cloud-based servers implemented using any common cloud computing platform, as described above. The memory 245 or the memory 235 can store an operating system for controlling and allocating the resources of the computerized system 200.

[0080] The computerized system 200 can interact with the network device 260 via the I / O interface 250 or the I / O port 255. Through the network device 260, the computerized system 200 can interact with the network. Through the network, the computerized system 200 can be logically connected to a remote computer. As described above, the server-side platform 210 can communicate with the client-side platform 220 via the network. The networks with which the computerized system 200 can interact include, but are not limited to, local area networks (LANs), wide area networks (WANs), and other networks.

[0081] Figure 4An exemplary method 100 for creating an MR conference room 70 is shown. The method 100 includes: scanning a first physical object 22 located in a first physical environment 20 by a first MR device 10. The method 100 further includes: scanning a second physical object 42 located in a second physical environment 40 by a second MR device 30, wherein the first physical environment 20 and the second physical environment 40 are spatially different. The method 100 further includes: generating a first feature point cloud 24 and a second feature point cloud 44 of the first physical object 22 and the second physical object 42 respectively during the scans 110, 120. Each of the first feature point cloud 24 and the second feature point cloud 44 includes a plurality of feature points 26, 46, and each feature point has a unique spatial coordinate relative to the relevant physical environment 20, 40. The method 100 further includes applying a 140 feature point detection algorithm to determine a candidate anchoring portion 55 of each of the first feature point cloud 24 and the second feature point cloud 44. The method 100 further includes comparing 150 the candidate anchoring portion 55 with a reference feature point cloud 64 of a virtual representation of a reference physical object 62. The method 100 further includes deriving 160 at least one common anchoring portion 50 between the first, second, and reference feature point clouds 22, 42, 62 from the candidate anchoring portion 55 in response to the matching condition of the comparison 150 being satisfied. The method 100 further includes aligning 170 the first feature point cloud 24 and the second feature point cloud 44 such that they coincide with each other relative to at least one common anchoring portion 55. The method 100 further includes: each of the first MR device 10 and the second MR device 30 rendering 180 a visual representation of the MR conference room 70 based on the aligned feature point clouds 24, 44. The rendering 180 is completed such that a virtual representation of a user of the first MR device 10 is visualized in the MR conference room 70 relative to the second physical object 42 or its virtual representation, and such that a virtual representation of a user of the second MR device 30 is visualized in the MR conference room 70 relative to the first physical object 22 or its virtual representation.

[0082] Reference Figure 5 , according to an exemplary embodiment, a schematic diagram of a (non-transitory) computer-readable (storage) medium 300 is shown. The computer-readable medium 300 may be associated with or connected to a computerized system 200 as described herein and is capable of storing a computer program product 310. In the disclosed embodiment, the computer-readable medium 300 is a memory stick, such as a universal serial bus (USB) stick. The USB stick 300 includes a housing 330 having an interface (such as a connector 340) and a memory chip 320. In the disclosed embodiment, the memory chip 320 is a flash memory, i.e., an electrically erasable and reprogrammable non-volatile data memory. The memory chip 320 stores a computer program product 310, which is programmed with computer program code (instructions) that, when loaded into a processor device, will execute a method, such as referenceFigure 4 The method 100 described above. The USB storage stick 300 is arranged to be connected to and read by a reading device so as to load instructions into the processor device. It should be noted that the computer-readable medium can also be other media, such as optical discs, digital video discs, hard disk drives or other common storage technologies. The computer program code (instructions) can also be downloaded from the computer-readable medium via a wireless interface and loaded into the processing device.

[0083] The operation steps described in any exemplary aspect herein are intended to provide examples and discussions. These steps can be performed by hardware components, can be implemented in the form of machine-executable instructions to enable a processor to execute these steps, or can be performed by a combination of hardware and software. Although a specific order of method steps may have been shown or described, the order of the steps may be different. Additionally, two or more steps can be performed simultaneously or partially simultaneously.

[0084] The terms used herein are for the purpose of describing particular aspects only and are not intended to limit the present disclosure. As used herein, the singular forms "a", "an" and "the" are intended to include the plural forms as well, unless the context clearly dictates otherwise. As used herein, the term "and / or" includes any and all combinations of one or more of the associated listed items. It should be further understood that the terms "comprises", "comprising", "includes" and / or "having" as used herein specify the presence of the stated features, integers, steps, operations, elements and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or combinations thereof.

[0085] It should be understood that although terms such as "first", "second", etc. may be used herein to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish between the respective elements. For example, without departing from the scope of the present disclosure, a first element can be referred to as a second element, and similarly, a second element can also be referred to as a first element.

[0086] Relative terms such as "below", "above", "upper", "lower", "horizontal" or "vertical" may be used herein to describe the relationship between one element and another shown in the figures. It should be understood that these terms, as well as the terms discussed above, are intended to cover different orientations of the device other than the orientation shown in the figures. It should be understood that when an element is referred to as "connected" or "coupled" to another element, it can be directly connected or coupled to the other element, or there may be intervening elements. In contrast, when an element is referred to as "directly connected" or "directly coupled" to another element, there are no intervening elements.

[0087] Unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains. It should also be understood that the terms used herein should be interpreted as being consistent with their meaning in the context of this specification and the relevant art, and should not be interpreted in an idealized or overly formal sense unless expressly so defined herein.

[0088] It should be understood that the present disclosure is not limited to the aspects described above and shown in the drawings; rather, those skilled in the art should recognize that many changes and modifications can be made within the scope of the present disclosure and the appended claims. The aspects disclosed in the drawings and the specification are for illustrative purposes only and not for limiting purposes, and the scope of the inventive concept is set forth by the following claims.

Claims

1. A computer-implemented method (100) for creating a mixed reality (MR) conference room (70), comprising: scanning (110) a first physical object (22) located in a first physical environment (20) by a first MR device (10); scanning (120) a second physical object (42) located in a second physical environment (40) by a second MR device (30), the first and second physical environments (20, 40) being spatially different; during the scanning (110, 120), generating (130) first and second feature point clouds (24, 44) of the first and second physical objects (22, 42) respectively, each of the first and second feature point clouds (24, 44) comprising a plurality of feature points (26, 46), each feature point (26, 46) having a unique spatial coordinate relative to the associated physical environment (20, 40); applying (140) a feature point detection algorithm to determine candidate anchor portions (55) of each of the first and second feature point clouds (24, 44); comparing (150) the candidate anchor portions (55) with a reference feature point cloud (64) of a virtual representation of a reference physical object (62); in response to a matching condition of the comparison (150) being satisfied, deriving (160) at least one common anchor portion (50) between the first, second, and reference feature point clouds (22, 42, 62) from the candidate anchor portions (55); aligning (170) the first and second feature point clouds (24, 44) such that the first and second feature point clouds coincide with each other relative to the at least one common anchor portion (50); and rendering (180) a visual representation of the MR conference room (70) by each of the first and second MR devices (10; 30) based on the aligned feature point clouds (24, 44), such that a virtual representation of a user of the first MR device (10) is visualized in the MR conference room (70) relative to the second physical object (42) or its virtual representation, and such that a virtual representation of a user of the second MR device (30) is visualized in the MR conference room (70) relative to the first physical object (22) or its virtual representation.

2. The computer-implemented method (100) according to claim 1, wherein, The reference physical object (62) shares at least one physical property with the first and second physical objects (22, 42).

3. The computer-implemented method (100) according to claim 2, wherein, The at least one physical property is texture, size, pattern, furniture type, or another property such that the feature point cloud representation thereof can be distinguished from the feature point cloud representations of other physical properties of the reference physical object (62).

4. The computer-implemented method (100) according to claim 2 or 3, wherein, The matching condition is determined based on corresponding feature descriptor data of the at least one physical property.

5. The computer-implemented method (100) according to any one of the preceding claims, wherein, The feature point detection algorithm is a salient feature point detection algorithm.

6. The computer-implemented method (100) according to any one of the preceding claims, wherein, Users of the first and second MR devices (10, 30) participate in the MR conference room (70) simultaneously.

7. The computer-implemented method (100) according to any one of the preceding claims, further comprising rendering a visual representation of virtual representations of one or more additional physical objects such that the one or more additional physical objects are visualized in the MR meeting room (70) relative to the first physical object (22) or the second physical object (42) or their virtual representations.

8. A computer-implemented method (100) according to any one of the preceding claims, wherein, The MR meeting room (70) is a virtual reality (VR) meeting room, wherein users of the first and second MR devices (10, 30) are virtually rendered as avatars and virtually participate in the VR meeting room through their avatars, and the avatars are visualized in the VR meeting room relative to the virtual representations of the first physical object (22) or the second physical object (42).

9. The computer-implemented method (100) according to any one of claims 1 to 7, wherein, The MR meeting room (70) is an augmented reality (AR) meeting room, wherein the user of the first MR device (10) is virtually rendered as an avatar and virtually participates in the AR meeting room through the avatar, and wherein the user of the second MR device (30) is physically participating in the AR meeting room.

10. The computer-implemented method (100) according to any one of the preceding claims, wherein, Rendering (180) of the virtual representation of the user is performed based on the spatial relationship between the user and the at least one common anchoring portion (50), thereby visualizing the virtual representation of the user at a spatial position in the MR meeting room (70) relative to the common anchor point (50).

11. A computer-implemented method (100) according to any one of the preceding claims, wherein, The at least one common anchoring portion (50) defines the spatial position and / or orientation of the MR meeting room (70).

12. The computer-implemented method (100) according to any one of the preceding claims, wherein, The first and second physical objects (22, 42) are selected from the group consisting of household furniture, household appliances, household equipment, office furniture, office appliances, and office equipment.

13. The computer-implemented method (100) according to any one of the preceding claims, wherein, The scanning (110, 120) is performed by near real-time 3D scanning technology implemented by the MR devices (10, 30).

14. The computer-implemented method (100) according to any one of the preceding claims, wherein, The feature points (26, 46) are voxels, each voxel having its own unique 3D coordinates.

15. The computer-implemented method (100) according to any one of the preceding claims, wherein, The MR meeting room (70) is persistent.

16. A computerized system (200) for creating a mixed reality (MR) meeting room (70), the system (200) comprising a first MR device (10), a second MR device (30), and a backend service (210), wherein: The first MR device (10) is configured to scan a first physical object (22) located in a first physical environment (20); The second MR device (30) is configured to scan a second physical object (42) located in a second physical environment (40), the first physical environment and the second physical environment (20, 40) being spatially different; wherein each MR device (10, 30) and / or the backend service (210) is further configured to: - During the scanning, first and second feature point clouds (24, 44) of the first and second physical objects (22, 42) are respectively generated, and each of the first and second feature point clouds (24, 44) includes a plurality of feature points (26, 46), and each feature point (26, 46) has unique spatial coordinates relative to the associated physical environment (20, 40); - Apply a feature point detection algorithm to determine candidate anchor portions (55) of each of the first and second feature point clouds (24, 44); - Compare the candidate anchor portions (55) with a reference feature point cloud (64) of a virtual representation of a reference physical object (62); - In response to a case where a matching condition of the comparison is satisfied, derive at least one common anchor portion (50) between the first, second, and reference feature point clouds (22, 42, 62) from the candidate anchor portions (55); and - Align the first and second feature point clouds (24, 44) such that the first and second feature point clouds coincide with each other relative to the at least one common anchor portion (50); and wherein each of the first and second MR devices (10, 30) is further configured to render a visual representation of the MR conference room (70) based on the aligned feature point clouds (24, 44), such that a virtual representation of a user of the first MR device (10) is visualized in the MR conference room (70) relative to the second physical object (42) or its virtual representation, and such that a virtual representation of a user of the second MR device (30) is visualized in the MR conference room (70) relative to the first physical object (22) or its virtual representation.

17. The computerized system (200) according to claim 16, wherein, The MR conference room (70) is persistent.

18. The computerized system (200) according to claim 16 or 17, wherein, The first and second MR devices (10, 30) are selected from the group consisting of: a head-mounted display (HMD), a cave automatic virtual environment (CAVE), or an augmented reality (AR) device.

19. A non-transitory computer-readable storage medium, comprising instructions that, when executed by a processor device, cause the processor device to perform the functions of one of the first or second MR devices (10, 30) of the method (100) according to any one of claims 1 to 15.

20. A computing device, comprising a processor device configured to perform the functions of one of the first or second MR devices (10, 30) of the method (100) according to any one of claims 1 to 15.