Mixed reality processing system and mixed reality processing method
By using cameras, head-mounted displays, and trackers in a coordinated manner, and employing simultaneous localization and mapping (SLAM) technology to obtain and overlap the coordinates of the virtual and real worlds, the problem of inaccurate segmentation and coordinate overlap errors in mixed reality processing is solved, enabling the generation of high-precision mixed reality images without the need for a green screen.
Patent Information
- Application Number
- CN202211073892.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2022-08-05
- Filing Date
- 2022-09-02
- Publication Date
- 2026-08-25
- Estimated Expiration
- 2042-09-02
AI Technical Summary
In existing mixed reality processing systems, the segmentation of the user's two-dimensional image is not accurate enough, and the camera with built-in real-world coordinates needs to be aligned with the controller with built-in virtual-world coordinates, which makes operation inconvenient and may produce errors.
Using cameras, head-mounted displays, trackers, and processors, the system acquires the coordinates of physical objects in the virtual and real worlds through simultaneous localization and mapping (SLAM) technology. The processor then merges these coordinates to create mixed reality images by merging the virtual scene with the physical objects.
It can accurately acquire images of physical objects without the need for a green screen, avoiding coordinate overlap errors and improving the convenience and accuracy of mixed reality processing.
Smart Images

Figure CN116664801B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a processing system, and more particularly to a mixed reality processing system and a mixed reality processing method. Background Technology
[0002] Generally speaking, creating mixed reality images requires the application of green screen removal. Since green screen removal is an image technology that can completely separate the user from the green screen background, it allows users to experience virtual reality within the green screen area by completely separating the user from the green screen background.
[0003] However, segmenting the user's 2D image may not be precise enough; for example, it might capture a portion of the green screen or fail to capture the user completely. Furthermore, traditional methods require aligning a camera with built-in real-world coordinates and a controller with built-in virtual-world coordinates to match the real-world and virtual-world coordinates, thus replacing the scene outside the user with the correct virtual reality scene. This method is inconvenient, and the alignment of the two coordinates may introduce errors.
[0004] Therefore, how to make mixed reality processing systems more convenient for generating mixed reality images has become one of the problems to be solved in this field. Summary of the Invention
[0005] This invention provides a mixed reality processing system, including a camera, a head-mounted display (HMD), a tracker, and a processor. The camera captures a two-dimensional image containing a physical object. The HMD displays a virtual scene and obtains the virtual world coordinates of the physical object using a first simultaneous localization and mapping (SLAM) map. These virtual world coordinates are generated based on real-world coordinates. The tracker extracts the physical object from the two-dimensional image and obtains the real-world coordinates of the object using a second SLAM map. The processor combines the virtual world coordinates and the real-world coordinates, merging the virtual scene with the physical object to generate a mixed reality image.
[0006] In one embodiment, the physical object is a human body, and the tracker inputs the two-dimensional image into a segmentation model, which outputs a human body block, which is a part of the two-dimensional image.
[0007] In one embodiment, the tracker inputs the two-dimensional image into a skeleton model, which outputs multiple human skeleton points. The tracker generates a three-dimensional pose based on the multiple human skeleton points, and the three-dimensional pose is used to adjust the capture range of the human body block.
[0008] In one embodiment, the processor is located in the tracker or in an external computer. The processor is used to align the origins and axes of the virtual world coordinates and the real world coordinates to generate coincident coordinates. The processor then uses these coincident coordinates to overlay the human body region from the tracker onto the virtual scene from the head-mounted display device to generate the mixed reality image.
[0009] In one embodiment, the tracker and the camera are located outside the head-mounted display device, and outside-in tracking technology is used to track the position of the head-mounted display device.
[0010] This invention provides a mixed reality processing method, comprising: capturing a two-dimensional image containing a physical object using a camera; displaying a virtual scene using a head-mounted display (HMD) and obtaining virtual world coordinates of the physical object in the virtual world using a first simultaneous localization and mapping (SLAM) map; wherein the virtual world coordinates are generated based on real-world coordinates; extracting the physical object from the two-dimensional image using a tracker and obtaining the real-world coordinates of the physical object in the real world using a second SLAM map; and merging the virtual world coordinates and the real-world coordinates using a processor to merge the virtual scene and the physical object to generate a mixed reality image.
[0011] In one embodiment, where the physical object is a human body, the mixed reality processing method further includes: inputting the two-dimensional image into a segmentation model via the tracker, the segmentation model outputting a human body block, the human body block being a part of the two-dimensional image.
[0012] In one embodiment, the tracker inputs the two-dimensional image into a skeleton model, which outputs multiple human skeleton points. The tracker generates a three-dimensional pose based on the multiple human skeleton points, and the three-dimensional pose is used to adjust the capture range of the human body block.
[0013] In one embodiment, wherein the processor is located in the tracker or in an external computer, the mixed reality processing method further includes: aligning the origins and axes of the virtual world coordinates and the real world coordinates by the processor to generate a coincident coordinate; and superimposing the human body block from the tracker onto the virtual scene from the head-mounted display device based on the coincident coordinate by the processor to generate the mixed reality image.
[0014] In one embodiment, the tracker and the camera are located outside the head-mounted display device, and outside-in tracking technology is used to track the position of the head-mounted display device.
[0015] In summary, the embodiments of the present invention provide a mixed reality processing system and a mixed reality processing method. The system extracts images of physical objects from a two-dimensional image using a tracker, and uses a processor to overlap the simultaneous localization and mapping (SMR) data in the tracker and the SMR data in the head-mounted display device to achieve coordinate calibration. The system then merges the virtual scene with the images of physical objects to generate a mixed reality image.
[0016] Therefore, the mixed reality processing system and mixed reality processing method of the present invention can obtain images of physical objects without using a green screen, and there is no need to align the camera with the built-in real-world coordinates and the controller with the built-in virtual-world coordinates for coordinate calibration. Attached Figure Description
[0017] Figure 1 This is a schematic diagram of a mixed reality processing system according to an embodiment of the present invention.
[0018] Figure 2 This is a flowchart illustrating a mixed reality processing method according to an embodiment of the present invention.
[0019] Figure 3 This is a schematic diagram illustrating a mixed reality processing method according to an embodiment of the present invention.
[0020] Figure 4 This is a schematic diagram illustrating the application of a mixed reality processing method according to an embodiment of the present invention.
[0021] Symbol explanation:
[0022] 100: Mixed Reality Processing System
[0023] CAM: Camera
[0024] TR: Tracker
[0025] ɑ: included angle
[0026] HMD: Head-mounted display device
[0027] USR: User
[0028] CR: Controller
[0029] MT: Movement Track
[0030] 200: Mixed Reality Processing Methods
[0031] 210~240: Steps
[0032] IMG: Two-Dimensional Imagery
[0033] MP1: First Simultaneous Localization and Map Building
[0034] MP2: Simultaneous Localization and Map Building
[0035] SK: 3D pose
[0036] TR: Tracker
[0037] SGM: Segmentation Model
[0038] SKM: Skeleton Model
[0039] PR: Processor
[0040] MRI: Mixed Reality Imaging
[0041] DP: Display
[0042] EC: External Computer Detailed Implementation
[0043] The following description is a preferred embodiment of the invention and is intended to describe the basic spirit of the invention, but is not intended to limit the invention. The actual scope of the invention must be understood by referring to the claims that follow.
[0044] It must be understood that the terms "comprising" and "including" used in this specification are used to indicate the presence of specific technical features, values, method steps, work processes, elements and / or components, but do not preclude the addition of more technical features, values, method steps, work processes, elements, components, or any combination thereof.
[0045] The use of terms such as "first," "second," and "third" in the claims is to modify elements in the claims and is not intended to indicate a priority order, a prior relationship, or that one element precedes another, or the chronological order of the execution of method steps. It is only used to distinguish elements with the same name.
[0046] Please refer to Figures 1-2 , Figure 1 This is a schematic diagram of a mixed reality processing system 100 according to an embodiment of the present invention. Figure 2 This is a flowchart illustrating a mixed reality processing method 200 according to an embodiment of the present invention.
[0047] In one embodiment, the mixed reality processing system 100 can be applied to a virtual reality system and / or XR extended reality.
[0048] In one embodiment, the mixed reality processing system 100 includes a camera (CAM), a head-mounted display (HMD), a tracker (TR), and a processor (PR).
[0049] In one embodiment, the processor PR may be located in the tracker TR or in an external computer EC (e.g., Figure 4 As shown in the figure. In one embodiment, the processor PR is located in the tracker TR, and there is also a processor in the external computer EC. When the computational workload of the processor PR is too large, some data can be transferred to the processor of the external computer EC for processing.
[0050] In one embodiment, the processor PR located in the tracker TR is used to perform various operations and may be implemented by an integrated circuit such as a microcontroller, microprocessor, digital signal processor, application-specific integrated circuit (ASIC), or a logic circuit.
[0051] In one embodiment, the camera (CAM) comprises at least one charge-coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) sensor.
[0052] In one embodiment, the tracker TR and the camera CAM are located outside the head-mounted display device (HMD). The tracker TR and the camera CAM use outside-in tracking to track the position of the HMD. Outside-in tracking has high accuracy and, because it transmits less data, has low computational latency, which can reduce some of the errors caused by latency.
[0053] In one embodiment, the tracker TR is placed adjacent to the camera CAM, for example... Figure 1 As shown, the camera (CAM) is placed above the tracker (TR). An angle α exists between the CAM's shooting range and the TR's tracking range. Both the TR and CAM can independently adjust their operating range up, down, left, and right, ensuring that the user (USR) remains within the angle α even if they move. In one embodiment, the TR and CAM are integrated into one device, or the CAM is integrated into the TR.
[0054] In one embodiment, the tracker TR and the camera CAM can be placed along a movable trajectory MT to capture images of the user USR.
[0055] In one embodiment, however, this is only an example, and the placement of the tracker TR and the camera CAM is not limited to this, as long as both can capture or track the user USR, and there is an angle less than an angular threshold (e.g., angle α) at the intersection of the tracking range and the shooting range.
[0056] In one embodiment, Figure 1 The user (USR) holds a handheld controller (CR) to operate games or applications and interact with objects in the virtual reality or augmented reality world. This invention is not limited to using a controller (CR); any device capable of operating games or applications, or any method capable of controlling displayed indicator signals (e.g., using gestures or electronic gloves), can be applied.
[0057] In one embodiment, the mixed reality processing method 200 can be implemented using components of the mixed reality processing system 100. Please refer to Figures 1-3 together. Figure 3 This is a schematic diagram illustrating a mixed reality processing method according to an embodiment of the present invention.
[0058] In step 210, a camera (CAM) captures a two-dimensional image (IMG) containing a physical object.
[0059] In one embodiment, the tracker TR includes a storage device. In one embodiment, the storage device may be implemented as a read-only memory, flash memory, floppy disk, hard disk, optical disk, USB flash drive, magnetic tape, network-accessible database, or other storage media with similar functionality that are readily conceived by those skilled in the art.
[0060] In one embodiment, the storage device is used to store a segmentation model (SGM) and a skeleton model (SKM), and the processor (PR) of the tracker (TR) can access the segmentation model (SGM) and / or the skeleton model (SKM) for execution.
[0061] In one embodiment, the segmentation model SGM is a pre-trained model. In another embodiment, the segmentation model SGM can perform graph-based image segmentation using a convolutional neural network (CNN) model, a region-with-CNN (R-CNN) model, or other algorithms applicable to image segmentation. However, those skilled in the art will understand that the present invention is not limited to these models; any other neural network model capable of segmenting human body regions can also be applied.
[0062] In one embodiment, the physical object is a human body (e.g., a user, USR). The processor PR of the tracker TR inputs a 2D image into the segmentation model SGM, and the segmentation model SGM outputs a human body block, which is part of the 2D image IMG. More specifically, the segmentation model SGM can segment an image of the user, USR, from the 2D image IMG.
[0063] In one embodiment, the skeleton model SKM is a pre-trained model. The skeleton model SKM is used to annotate important key points of the human body (such as joints of the head, shoulders, elbows, wrists, waist, knees, ankles, etc.) to generate a skeleton, facilitating the analysis of human posture and movement. In one embodiment, if the application of skeleton points is further extended to continuous movements, it can be used for applications such as behavior analysis and action comparison. In one embodiment, the skeleton model SKM can be generated using a convolutional neural network model, a region-based convolutional neural network model, or other algorithms applicable to finding the human skeleton. However, those skilled in the art will understand that the present invention is not limited to these models; any other neural network model capable of outputting a human skeleton can also be applied.
[0064] In one embodiment, the tracker TR inputs a two-dimensional image IMG into a skeleton model SKM (or the tracker TR directly inputs the human body block output by the segmentation model SGM into the skeleton model SKM). The skeleton model SKM outputs multiple human body skeleton points. The processor PR of the tracker TR generates a three-dimensional pose SK based on these human body skeleton points. The three-dimensional pose SK is used to adjust the capture range of the human body block.
[0065] In one embodiment, the segmentation model SGM can segment the image of the user's USR from the 2D image IMG. The processor PR then inputs the image of the user's USR into the skeleton model SKM. The skeleton model SKM outputs multiple human skeleton points. The processor PR generates a 3D pose SK based on these human skeleton points and adjusts the capturing range of the user's USR image based on the 3D pose SK. This allows the processor PR to capture a more accurate image of the user's USR.
[0066] In step 220, a head-mounted display device (HMD) displays a virtual scene and obtains virtual world coordinates of the physical objects in the virtual world through a first simultaneous localization and mapping map (SLAM map) MP1; wherein the virtual world coordinates are generated based on a real world coordinate.
[0067] In one embodiment, the first simultaneous localization and mapping (MP1) acquires sensory information from the environment, incrementally constructs a map of the surrounding environment, and uses the map to achieve autonomous localization. In other words, this technology enables the head-mounted display (HMD) to determine its own location and generate an environmental map, thereby assessing the entire space. With the first MP1, the HMD knows its own spatial coordinates and then generates virtual world coordinates based on real-world coordinates. These virtual world coordinates allow for the creation of virtual scenes, such as game scenes. Therefore, the HMD can calculate the virtual world coordinates of physical objects (e.g., the user's USR) within this virtual world and transmit these coordinates to the tracker TR.
[0068] In step 230, a tracker TR extracts the physical object from the 2D image IMG and obtains the real-world coordinates of the physical object in the real world through a second simultaneous localization and mapping (SMR) map MP2.
[0069] In one embodiment, the second simultaneous localization and mapping (MP2) acquires sensory information from the environment, incrementally constructs a map of the surrounding environment, and uses the map to achieve autonomous localization. In other words, this technology enables the tracker TR to determine its own location and generate an environmental map, thereby assessing the entire space. With MP2, the tracker (TR) can know its own spatial coordinates and thus obtain its real-world coordinates.
[0070] Therefore, it can be seen that the head-mounted display device (HMD) can obtain virtual and real-world coordinates through Simultaneous Localization and Mapping (SLAM) technology and transmit these coordinates to the tracker (TR). Conversely, the tracker (TR) can independently obtain real-world coordinates using SLAM technology. Based on the concept of map sharing, the tracker (TR) aligns the origins and coordinate axes (such as the X, Y, and Z axes) of the first and second SLAM maps (MP1 and MP2) to complete coordinate calibration.
[0071] In one embodiment, the coordinate calibration calculation can be performed by the processor PR of the tracker TR or by the processor in an external computer EC.
[0072] In step 240, the processor PR coincides the virtual world coordinates with the real world coordinates and merges the virtual scene with the physical objects to generate a mixed reality image MRI.
[0073] Since virtual world coordinates are generated based on real world coordinates, each point in the virtual world coordinates can be mapped to a real world coordinate. The tracker TR only needs to align the origin and coordinate axes (such as the X-axis, Y-axis and Z-axis) of the virtual world coordinates and the real world coordinates to generate coincident coordinates.
[0074] In one embodiment, the tracker TR aligns the origins and coordinate axes (such as the X-axis, Y-axis, and Z-axis) of the virtual world coordinates with those of the real world coordinates to generate a coincident coordinate. The processor PR then overlays the human body region of the tracker TR onto the virtual scene from the head-mounted display device HMD based on the coincident coordinate to generate a mixed reality MRI image.
[0075] More specifically, the tracker TR can first store the calculated human body blocks (such as images of the user's USR) in a storage device. The virtual scene currently displayed by the head-mounted display (HMD) is also transmitted to the tracker TR, which stores this virtual scene in its storage device.
[0076] In one embodiment, the processor PR coincides the virtual world coordinates and the real world coordinates. After calculating the coincident coordinates, it can also calculate the position of the user's USR image at the coincident coordinates, and read the virtual scene from the storage device, or receive the virtual scene currently displayed from the head-mounted display device HMD in real time, and then paste the user's USR image onto the virtual scene to generate mixed reality MRI.
[0077] In one embodiment, in order for the human body block of the tracker TR to be superimposed on the virtual scene from the head-mounted display device HMD in the future, the optimal spatial correspondence between the virtual scene and the image of the user USR must be obtained through a spatial alignment step method (such as step 240), including size, rotation and displacement. This result can be used to calibrate the map data of the coincident coordinates, thereby calculating the coordinates of the image of the user USR placed in the virtual scene to achieve the purpose of accurate superposition.
[0078] Figure 4 This is a schematic diagram illustrating the application of a mixed reality processing method according to an embodiment of the present invention.
[0079] In one embodiment, the computation to generate mixed reality MRI images can be performed by the processor PR of the tracker TR or by the processor in an external computer EC.
[0080] In one embodiment, the external computer EC can be a laptop, server, or other electronic device with computing and storage functions.
[0081] In one embodiment, the tracker TR can transmit mixed reality MRI images to an external computer EC, which can then process subsequent applications, such as uploading the mixed reality MRI images to YouTube for live game streaming, or displaying the mixed reality MRI images on a display DP to showcase the latest games to viewers.
[0082] In summary, the embodiments of the present invention provide a mixed reality processing system and a mixed reality processing method. The system extracts images of physical objects from a two-dimensional image using a tracker, and uses a processor to overlap the simultaneous localization and mapping (SMR) data in the tracker and the SMR data in the head-mounted display device to achieve coordinate calibration. The system then merges the virtual scene with the images of physical objects to generate a mixed reality image.
[0083] Therefore, the mixed reality processing system and mixed reality processing method of the present invention can obtain images of physical objects without using a green screen, and there is no need to align the camera with the built-in real-world coordinates and the controller with the built-in virtual-world coordinates for coordinate calibration.
[0084] While the present invention has been disclosed above with reference to preferred embodiments, it is not intended to limit the scope of the invention. Any person skilled in the art may make some modifications without departing from the spirit and scope of the invention. Therefore, the scope of protection of the present invention shall be determined by the appended claims.
Claims
1. A mixed reality processing system, comprising: A camera used to capture a two-dimensional image containing a physical object; A head-mounted display device is used to display a virtual scene and obtain the virtual world coordinates of the entity object in the virtual world through a first simultaneous localization and mapping (SMR) map construction; wherein the virtual world coordinates are generated based on a real world coordinate. A tracker is used to extract the physical object from the two-dimensional image and obtain the real-world coordinates of the physical object by a second simultaneous localization and mapping (SMR) system. A processor is used to align the virtual world coordinates with the real world coordinates and merge the virtual scene with the physical object to generate a mixed reality image.
2. The mixed reality processing system as described in claim 1, wherein the physical object is a human body, the tracker inputs the two-dimensional image into a segmentation model, the segmentation model outputs a human body block, and the human body block is a part of the two-dimensional image.
3. The mixed reality processing system as described in claim 2, wherein the tracker inputs the two-dimensional image into a skeleton model, the skeleton model outputs multiple human skeleton points, the tracker generates a three-dimensional pose based on the multiple human skeleton points, and the three-dimensional pose is used to adjust the capture range of the human body block.
4. The mixed reality processing system of claim 2, wherein the processor is located in the tracker or in an external computer, and the processor is used to coincide the origin and coordinate axis of the virtual world coordinates and the real world coordinates to generate a coincident coordinate. in, The processor overlays the human body block from the tracker onto the virtual scene from the head-mounted display device based on the overlapping coordinates to generate the mixed reality image.
5. The mixed reality processing system of claim 1, wherein the tracker and the camera are located outside the head-mounted display device, and an outside-in tracking technique is used to track the position of the head-mounted display device.
6. A mixed reality processing method, comprising: Capture a two-dimensional image containing a physical object using a single camera; A virtual scene is displayed through a head-mounted display device, and a virtual world coordinate is obtained by a first simultaneous localization and mapping (SMR) system to determine the location of the physical object in the virtual world; wherein the virtual world coordinate is generated based on a real-world coordinate. The object is extracted from the 2D image using a tracker, and its real-world coordinates are obtained through a second simultaneous localization and mapping (SMR) system. A processor is used to align the virtual world coordinates with the real world coordinates and merge the virtual scene with the physical object to generate a mixed reality image.
7. The mixed reality processing method of claim 6, wherein the physical object is a human body, and the mixed reality processing method further comprises: The tracker inputs the two-dimensional image into a segmentation model, which outputs a human body block that is part of the two-dimensional image.
8. The mixed reality processing method as described in claim 7, further comprising: The tracker inputs the two-dimensional image into a skeleton model, which outputs multiple human skeleton points. The tracker then generates a three-dimensional pose based on these multiple human skeleton points, which is used to adjust the capture range of the human body region.
9. The mixed reality processing method of claim 7, wherein the processor is located in the tracker or in an external computer, and the mixed reality processing method further comprises: The processor aligns the origins and axes of the virtual world coordinates and the real world coordinates to generate a coincident coordinate system; and The processor overlays the human body block from the tracker onto the virtual scene from the head-mounted display device based on the overlapping coordinates to generate the mixed reality image.
10. The mixed reality processing method of claim 6, wherein the tracker and the camera are located outside the head-mounted display device, and an outside-in tracking technique is applied to track the position of the head-mounted display device.
Citation Information
Patent Citations
New pattern and method of virtual reality system based on mobile devices
US20170116788A1
Methods for simultaneous localization and mapping (SLAM) and related apparatus and systems
US20190178654A1
Blink-based calibration of an optical see-through head-mounted display
US20200363867A1