Methods and systems for tracking the position of moving objects and moving cameras
By combining environmental feature points and time constraints, the coordinates of the feature points are corrected, and the six degrees of freedom of the orientation of the moving object and the camera are calculated using a neural network. This solves the problem of synchronous tracking in mixed reality and realizes automatic control and privacy protection of the virtual screen.
Patent Information
- Application Number
- CN202110554564.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2021-04-21
- Filing Date
- 2021-05-20
- Publication Date
- 2025-10-31
- Estimated Expiration
- 2041-05-20
AI Technical Summary
Existing technologies cannot simultaneously track the six degrees of freedom orientation of a moving object and a moving camera, especially in mixed reality applications, where it is difficult to accurately determine the position and orientation of objects and cameras in dynamic environments.
By using a movable camera and a six-degree-of-freedom orientation calculation unit for objects, combined with environmental feature points and geometric and time constraints, the coordinates of feature points are corrected, the six-degree-of-freedom orientation of the movable object and the camera is calculated, and the feature points are inferred and corrected using a neural network to achieve synchronous tracking.
It achieves six-DOF orientation-synchronized tracking of moving objects and cameras in mixed reality, provides automatic control of the virtual screen and enhanced privacy protection, and improves the user experience.
Smart Images

Figure CN113920189B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a method and system for simultaneously tracking the six degrees of freedom orientation of a moving object and a moving camera. Background Technology
[0002] Existing tracking technologies, such as Simultaneous Localization and Mapping (SLAM), can track the six degrees of freedom orientation of a moving camera, but cannot simultaneously track moving objects. This is because a moving camera requires stable environmental feature points for localization, while the feature points of a moving object are unstable and are usually discarded, making them unusable for tracking.
[0003] On the other hand, technologies used to track moving objects ignore environmental features to avoid interference, so these technologies cannot track moving cameras.
[0004] Most neural networks learn features to distinguish object types, rather than to calculate the six degrees of freedom (DOF) orientation of objects. Some neural networks used for pose or gesture recognition can only output the 2D coordinates (x, y) of skeletal joints in the image plane. Even if the distance between the joints and the camera is estimated using depth sensing technology, it is not a true 3D coordinate in space, let alone the ability to calculate the six degrees of freedom orientation in space.
[0005] In motion capture systems, multiple fixed cameras are used to track joint positions. Markers are usually placed on the joints to reduce errors. There is no six-degree-of-freedom orientation tracking camera.
[0006] Therefore, with the technology currently known, there is no technology that can simultaneously track a moving object and a moving camera.
[0007] The rapid development of mixed reality (MR) has prompted researchers to develop technologies capable of simultaneously tracking the six degrees of freedom (DOF) of orientation of both moving cameras and moving objects. In MR applications, because the camera mounted on MR glasses moves with the user's head, knowing the camera's six DFO is crucial for determining the user's position and orientation. Objects interacting with the user also move, so knowing the object's six DFO is equally important for displaying virtual content in the appropriate location and orientation. Users wearing MR glasses can move freely indoors or outdoors, making it difficult to place markers in the environment. Furthermore, for a better user experience, no special markers are typically affixed to objects beyond their inherent characteristics.
[0008] While these conditions increase the difficulty of tracking six degrees of freedom orientation, we have developed a technology that can simultaneously track moving objects and moving cameras to solve these problems and meet more application needs. Summary of the Invention
[0009] The technology proposed in this invention can be applied, for example, to display one or more virtual screens next to the physical screen of a handheld device, such as a mobile phone, when the user is wearing MR glasses. The preset position, orientation, and size of the virtual screens are set according to the six degrees of freedom orientation of the camera on the mobile phone and MR glasses. Furthermore, through six degrees of freedom orientation tracking, the rotation and movement of the virtual screens can be automatically controlled to align with the viewing direction. This invention can provide the following benefits to the user: (1) expanding a small physical screen into a large virtual screen; (2) adding a single physical screen to multiple virtual screens to view more applications simultaneously; (3) ensuring that the content of the virtual screens is not spied on by others.
[0010] According to an embodiment of the present invention, a method for simultaneously tracking the six degrees of freedom (6 DoF poses) of a movable object and a movable camera is proposed, comprising the following steps: capturing a series of images with a movable camera, extracting several environmental feature points from these images, matching these environmental feature points to calculate several camera matrices of the movable camera, and then calculating the six degrees of freedom pose of the movable camera using these camera matrices; simultaneously, inferring several feature points of the movable object from these images captured by the movable camera, using the camera matrices corresponding to these images, as well as predefined geometric constraints and time constraints, correcting the coordinates of these feature points of the movable object, and then calculating the six degrees of freedom pose of the movable object using these corrected feature point coordinates and their corresponding camera matrices.
[0011] According to another embodiment of the present invention, a system for simultaneously tracking the six degrees of freedom (DOF) orientation of a movable object and a movable camera is proposed, comprising a movable camera, a movable camera six-DOF orientation calculation unit, and a movable object six-DOF orientation calculation unit. The movable camera is used to capture a series of images. The movable camera six-DOF orientation calculation unit is used to extract several environmental feature points from these images, match these environmental feature points to calculate several camera matrices of the movable camera, and then calculate the six-DOF orientation of the movable camera using these camera matrices. The movable object six-DOF orientation calculation unit is used to deduce several feature points of the movable object from the images captured by the movable camera, correct the coordinates of these feature points of the movable object using the camera matrices corresponding to each image, and predefined geometric and time constraints, and then calculate the six-DOF orientation of the movable object using these corrected feature point coordinates and their corresponding camera matrices.
[0012] To provide a better understanding of the above and other aspects of the present invention, specific embodiments are described below in conjunction with the accompanying drawings: Attached Figure Description
[0013] Figure 1A , 1B This illustration shows the application of the present invention's technique for simultaneously tracking moving objects and a moving camera compared to existing technologies;
[0014] Figure 2A A system and method for simultaneously tracking the six degrees of freedom orientation of a moving object and a moving camera, according to one embodiment, are illustrated.
[0015] Figure 2B The system and method for tracking the six degrees of freedom orientation of a moving object and a moving camera during the training phase are illustrated.
[0016] Figure 3A The diagram illustrates the correspondence between environmental feature points and movable object feature points in a series of images captured by a movable camera.
[0017] Figure 3B Examples illustrating the position and orientation of an object in space;
[0018] Figures 4A-4B Draw the feature points of the movable object to be repaired;
[0019] Figures 5A-5D The diagram illustrates the feature point definition and various training data using a mobile phone as an example.
[0020] Figure 6 Draw the structure of the neural network during the training phase;
[0021] Figure 7 The method for calculating the displacement of feature points between two adjacent images is illustrated.
[0022] Figure 8 The method for calculating and determining time limits is illustrated.
[0023] Figure 9 This illustrates situations where incorrect displacement occurs due to a lack of time constraints.
[0024] Figure 10 The system and method of tracking the six-DOF orientation of a moving object and a moving camera while incorporating incremental learning are illustrated.
[0025] Figure 11 A system and method for simultaneously tracking the orientation of a moving object and a moving camera, applied to MR glasses, are illustrated.
[0026] The meanings of the symbols in each of the attached figures are as follows:
[0027] 100, 200, 300: A six-DOF system for simultaneously tracking the orientation of a moving object and a moving camera.
[0028] 110: Portable camera
[0029] 120: Six-DOF orientation calculation unit for movable camera
[0030] 121: Environmental Feature Extraction Unit
[0031] 122: Camera Matrix Calculation Unit
[0032] 123: Camera orientation calculation unit
[0033] 130: Six-DOF orientation calculation unit for movable objects
[0034] 131: Object Feature Coordinate Calculation Unit
[0035] 132: Object Feature Coordinate Correction Unit
[0036] 133: Object Orientation Calculation Unit
[0037] 140: Training Data Generation Unit
[0038] 150: Neural Network Training Unit
[0039] 260: Automatic expansion unit
[0040] 270: Weight Adjustment Unit
[0041] 310: Azimuth Correction Unit
[0042] 311: Cross-alignment unit
[0043] 312: Correction Unit
[0044] 320: Azimuth Stabilizing Unit
[0045] 330: Linear Calculation Unit
[0046] 340: Screen Orientation Calculation Unit
[0047] 350: Stereoscopic Image Generation Unit
[0048] 351: Image generation unit
[0049] 352: Imaging Unit
[0050] 900: Movable object
[0051] CD: Six-DOF orientation for a movable camera
[0052] CM: Camera Matrix
[0053] d, d', d”: displacement
[0054] D1, D4: Physical screen
[0055] D2, D3: Virtual Screen
[0056] DD: Six degrees of freedom orientation of the virtual screen
[0057] EF: Environmental Feature Points
[0058] ET: Feature Extractor
[0059] FL: Feature Point Coordinate Prediction Layer
[0060] FV: Feature Vector
[0061] G1: MR Glasses
[0062] GC: Geometric Limitation
[0063] GCL: Geometric Confinement Layer
[0064] IM, IM': Images
[0065] LV: Loss Value
[0066] MD: Neural Network Inference Model
[0067] m: Average displacement
[0068] OD: Six degrees of freedom orientation of a movable object
[0069] OF, OF', OF*, OF**: feature points
[0070] OLV: Total Loss Value
[0071] P1: Mobile phone
[0072] s: standard deviation of displacement
[0073] ST1: Training Phase
[0074] ST2: Tracking Phase
[0075] ST3: Incremental Learning Phase
[0076] TC: Time Limit
[0077] TCL: Time Limitation Layer
[0078] , , , , , , , , , ,:coordinate
[0079] Displacement
[0080] Penalty value
[0081] PL: Optimal Plane
[0082] C: Center point
[0083] N: Normal vector Detailed Implementation
[0084] Please refer to Figure 1A , 1B The illustration depicts the application of the present invention's technique for simultaneously tracking moving objects and a moving camera, compared to existing technologies. The technology proposed in this invention can be applied, for example, to: Figure 1A As shown, when a user wears MR glasses G1 (which is equipped with a movable camera 110), one or more virtual screens can be displayed next to the real screen of a handheld device, such as a mobile phone P1 (i.e., a movable object 900). The preset positions, orientations, and sizes of the virtual screens D2 and D3 are set according to the six degrees of freedom orientation of the mobile phone P1 and the movable camera 110 on the MR glasses G1. "Movable" in "movable camera 110" refers to its position relative to a stationary object in three-dimensional space. Furthermore, through six-degree-of-freedom orientation tracking, the virtual screens D2 and D3 can be automatically controlled to rotate and move, aligning them with the viewing direction (e.g., ...). Figure 1B As shown), users can also adjust the position and angle of these virtual screens D2 and D3 according to their own preferences. The virtual screens displayed in the prior art will move with the MR glasses G1, but will not move with the six degrees of freedom orientation of the object. The technology of the present invention can provide users with the following benefits: (1) expanding the small physical screen D1 into a large virtual screen D2; (2) increasing a single physical screen D1 to multiple virtual screens D2 and D3 to view more applications at the same time; (3) the content of the virtual screens D2 and D3 will not be spied on by others. The above technology can also be applied to tablet computers or laptops, setting up virtual screens next to their physical screens. In addition to physical screens, the movable object 900 can also be other objects with defined characteristics, such as cars, bicycles, pedestrians, etc. The movable camera 110 is not limited to the camera on the MR glasses G1, but can also be a camera on an autonomous mobile robot and a vehicle.
[0085] Please refer to Figure 2A The illustration depicts simultaneous tracking of a movable object 900 (labeled as shown in an embodiment) according to one embodiment. Figure 1AA system 100 and method for six-DOF orientation of a movable camera 110. A movable object 900 is, for example, a... Figure 1A The P1 mobile phone; the 110 portable camera, for example, is Figure 1A The system 100, which simultaneously tracks the six degrees of freedom (DOF) orientation of a movable object 900 and a movable camera 110, includes a movable camera 110, a movable camera six-DOF orientation calculation unit 120, and a movable object six-DOF orientation calculation unit 130. The movable camera 110 is used to capture a series of images (IM). The movable camera 110 can be mounted on a head-mounted stereoscopic display, a mobile device, a computer, or a robot. The movable camera six-DOF orientation calculation unit 120 and / or the movable object six-DOF orientation calculation unit 130 are, for example, circuits, chips, circuit boards, program code, or storage devices storing program code.
[0086] The movable camera six-DOF orientation calculation unit 120 includes an environmental feature extraction unit 121, a camera matrix calculation unit 122, and a camera orientation calculation unit 123. Its implementation may be, for example, a circuit, a chip, a circuit board, program code, or a storage device storing program code. The environmental feature extraction unit 121 extracts several environmental feature points EF from these images IM. The camera matrix calculation unit 122 matches these environmental feature points EF to calculate several camera matrices CM of the movable camera 110. The camera orientation calculation unit 123 then calculates the six-DOF orientation CD of the movable camera 110 from the camera matrices CM.
[0087] The movable object six-DOF orientation calculation unit 130 includes an object feature coordinate estimation unit 131, an object feature coordinate correction unit 132, and an object orientation calculation unit 133. Its implementation can be, for example, a circuit, a chip, a circuit board, program code, or a storage device storing program code. The object feature coordinate estimation unit 131 is used to estimate several feature points OF of the movable object 900 from the images IM captured by the movable camera 110. These feature points OF are predefined and are compared with the images IM captured by the movable camera 110 to estimate the coordinates of these feature points OF. The movable object 900 is a rigid object.
[0088] Please refer to Figure 2BAnother embodiment illustrated, the method for simultaneously tracking the six-DOF orientation of a movable object 900 and a movable camera 110, includes a training stage ST1 and a tracking stage ST2. The object feature coordinate estimation unit 131 uses a neural network inference model MD to infer the coordinates of feature points OF of the movable object 900 from the images IM captured by the movable camera 110. The neural network inference model MD is pre-trained, and the training data is obtained through manual or automatic labeling. Geometric constraints GC and time constraints TC are incorporated during the training process.
[0089] The object feature coordinate correction unit 132 uses the camera matrix CM corresponding to each of the image IMs, as well as predefined geometric constraints GC and temporal constraints TC, to correct the coordinates of these feature points OF of the movable object 900. Specifically, the object feature coordinate correction unit 132 uses the camera matrix CM to project the two-dimensional coordinates of these feature points OF to their corresponding three-dimensional coordinates. Based on the geometric constraints GC, it deletes feature points OF with three-dimensional coordinate deviations greater than a predetermined value, or supplements the coordinates of undetected feature points OF using the coordinates of adjacent feature points OF based on the geometric constraints GC. Furthermore, the object feature coordinate correction unit 132 also compares the coordinate changes of these feature points OF in multiple consecutive image IMs based on the temporal constraints TC, and then uses the coordinates of these corresponding feature points OF in these consecutive image IMs to correct the coordinates of feature points OF' with coordinate changes greater than a predetermined value, thus obtaining the corrected coordinates of these feature points OF'.
[0090] Please refer to Figure 3A The example illustrates the correspondence between environmental feature points and movable object feature points in a series of images captured by a movable camera. For non-planar objects, orientation and position can be defined using the centroids of several selected feature points (OF). Please refer to [reference needed]. Figure 3B The example illustrates the position and orientation of an object in space. Feature points OF are fitted to find the optimal plane PL. The center point C of the optimal plane PL can represent the position (x, y, x) of the object in three-dimensional space, and the normal vector N of the optimal plane PL can represent the orientation of the object.
[0091] Geometric constraints (GC) are defined in three-dimensional space. For rigid objects, the distance between feature points (OF) should be fixed. After being projected onto the two-dimensional image plane by the camera matrix, the positions of all feature points (OF) must be constrained within a reasonable range.
[0092] Please refer to Figures 4A-4BThe example illustrates the correction of feature point OF coordinates. The camera matrix CM can not only be used to calculate the six-DOF orientation of the movable camera 110 and the movable object 900, but also to apply 3D geometric constraints GC to correct the coordinates of feature point OF* projected onto the 2D image plane (e.g., ...). Figure 4A (as shown) or add missing feature point OF** coordinates (e.g.) Figure 4B (As shown).
[0093] The object orientation calculation unit 133 then uses the corrected coordinates of these feature points OF' and their corresponding camera matrices CM to calculate the six-degree-of-freedom orientation OD of the movable object 900. For planar movable objects, the best-fit plane is calculated using these feature points OF'. The six-degree-of-freedom orientation OD of the movable object 900 is defined by the center point and normal vector of the plane. For non-planar movable objects, the six-degree-of-freedom orientation OD of the movable object 900 is defined by the centroid of the three-dimensional coordinates of these feature points OF'.
[0094] like Figure 2B As shown, the training stage ST1 of the system 100 that simultaneously tracks the six degrees of freedom orientation of the movable object 900 and the movable camera 110 includes a training data generation unit 140 and a neural network training unit 150, which can be implemented as a circuit, chip, circuit board, program code, or storage device for storing program code.
[0095] The neural network training unit 150 is used to train the neural network inference model MD. The neural network inference model MD is used to infer the positions and sequences of feature points OF of the movable object 900. In the training data generation unit 140, the training data can be images with manually labeled feature point positions and sequences, or automatically augmented labeled images. Please refer to [reference needed]. Figures 5A-5D The figures illustrate various training data, using a mobile phone as an example. In these figures, the feature points OF are defined by the four interior corners of the physical screen D4. When the physical screen D4 is positioned vertically, the four feature points OF are designated in a clockwise order from the top left corner to the bottom left corner. Figure 5A As shown, the four feature points OF have coordinates in sequence. ,coordinate ,coordinate ,coordinate Even when the physical D4 screen is rotated to landscape orientation, the order of the feature points OF remains unchanged (e.g., Figure 5B (As shown). In some cases, not all feature points (OF) can be captured. Therefore, the training data needs to include some similar... Figure 5C or Figure 5D This type of image lacks some feature points (OF). For example... Figure 5A and Figure 5DAs shown, the feature point labeling process can distinguish between the front (i.e., the screen) and back of the phone, but labeling is only performed on the front. To achieve higher accuracy, each image is magnified during feature point labeling (OF) until each pixel is clearly visible. Since manual labeling is very time-consuming, automatic augmentation is necessary to scale the training data to the millions of images. Methods for automatically augmenting manually labeled images include: scaling and rotating proportionally, mapping using perspective projection, converting to different colors, adjusting brightness and contrast, adding motion blur and noise, and adding other objects to obscure certain feature points (e.g., ...). Figure 5C and Figure 5D (As shown), change the content displayed on the screen, or replace the background, etc. Then, recalculate the positions of these manually marked feature points (OF) in the automatically expanded image according to the transformation relationship.
[0096] Please refer to Figure 6 The example illustrates that the main structure of a neural network during the training phase includes feature extraction and feature point coordinate prediction. The feature extractor (ET) can use a deep residual network like ResNet or other networks with similar functionality. The extracted feature vector (FV) is fed into the feature point coordinate prediction layer (FL) to calculate the coordinates of the feature points (OF) (e.g., the coordinates of the current image's feature points OF). The coordinates of the feature points (OF) of the previous image are represented by... (Represented). In addition to the feature point prediction layer, this embodiment also adds a geometrically constrained layer (GCL) and a temporally constrained layer (TCL) to reduce erroneous predictions. During the training phase, each layer calculates the loss value LV between the predicted value and the true value based on the loss function, and then accumulates these loss values and their respective weights to obtain the total loss value OLV.
[0097] Please refer to Figure 7 The example illustrates how the displacement of a feature point is calculated between two adjacent images. The coordinates of the feature point OF in the current image are... The coordinates of the same feature point OF in the previous image are: The displacement between them is defined as .
[0098] Unreasonable displacement is penalized. Impose restrictions. Penalty value. For example, it can be calculated according to the following formula (1).
[0099]
[0100] Where m is the average displacement calculated for each feature point OF across all training data, s is the standard deviation of displacement, and d is the displacement of the same feature point OF between the previous and current images. When d ≤ m, the displacement is within an acceptable range and there is no penalty (i.e., Please refer to... Figure 8 The examples illustrate the time limit (TC) and penalty value. The calculation and determination method. The center of the circle represents the coordinates of the feature point OF in the previous image. The area of the circle represents the acceptable displacement of the feature point OF in the current image. If the predicted coordinates of the feature point OF in the current image... If the displacement is within the circle (d'≤m), then the penalty value is... The value is zero. If the predicted coordinates of the feature point OF in the current image are zero... Outside the circle (displacement d > m), the penalty value is... for The greater the displacement exceeds the radius of the circle (i.e., m), the larger the penalty value will be during training. A larger loss value is used to limit the coordinates of feature points (OF) to a reasonable range.
[0101] Please refer to Figure 9 The example illustrates the situation where incorrect displacement occurs due to the lack of a time limit TC. Figure 9 The left image is the previous image, and the right image is the current image. In the previous image, elements with coordinates were identified. The feature points OF. However, in current imagery, identifying coordinates from reflected images is difficult. Feature point OF, coordinates With coordinates The displacement between them is greater than the range set by the time limit TC, therefore the coordinates can be determined. Incorrect.
[0102] like Figure 2B As shown, during the tracking phase ST2, the movable camera 110 captures a series of images IM. Several environmental feature points EF are extracted from these images and then used to calculate the corresponding camera matrix CM and six-degree-of-freedom orientation CD of the movable camera 110. Simultaneously, the coordinates of the feature points OF of the movable object 900 are also calculated by the neural network inference model MD and transformed and corrected by the camera matrix CM to obtain the six-degree-of-freedom orientation OD of the movable object 900.
[0103] Please refer to Figure 10 Its illustration shows the tracking of a moving object 900 (marked in ST3) while adding the incremental learning stage. Figure 1AThe system 200 and method for six-degree-of-freedom orientation of a movable camera 110 includes an automatic augmentation unit 260 and a weight adjustment unit 270, which are implemented, for example, as circuits, chips, circuit boards, program code, or storage devices storing program code.
[0104] exist Figure 10 In the embodiment, during the training phase, the training data of the neural network inference model MD consists of manual labeling and automatic augmentation; while during the incremental learning phase, the training data consists of automatic labeling and automatic augmentation.
[0105] While tracking the movable object 900, the neural network inference model MD performs incremental learning in the background. The training data for incremental learning includes: image IM captured by the movable camera 110 and image IM' automatically augmented by the automatic augmentation unit 260 based on image IM. The automatic augmentation unit 260 replaces the manual labels with the corrected feature point OF coordinates of the corresponding images IM and IM' as the ground truth of the feature point coordinates. The weight adjustment unit 270 adjusts the weights in the neural network inference model MD to update it to the neural network inference model MD', thereby adapting to the usage scenario to accurately track the six degrees of freedom orientation OD of the movable object 900.
[0106] In addition, please refer to Figure 11 The diagram illustrates a system 300 and method for simultaneously tracking the six degrees of freedom orientation of a movable object 900 and a movable camera 110 in MR glasses. The system includes an orientation correction unit 310, an orientation stabilization unit 320, a visual axis calculation unit 330, a screen orientation calculation unit 340, and a stereoscopic image generation unit 350. The implementation of this unit could be, for example, a circuit, a chip, a circuit board, program code, or a storage device for storing program code. The orientation correction unit 310 includes a cross-comparison unit 311 and a correction unit 312, and its implementation could be, for example, a circuit, a chip, a circuit board, program code, or a storage device for storing program code. The stereoscopic image generation unit 350 includes an image generation unit 351 and an imaging unit 352, and its implementation could be, for example, a circuit, a chip, a circuit board, program code, or a storage device for storing program code.
[0107] As the movable camera 110 and the movable object 900 move, their six-degree-of-freedom orientations (CD and OD) need to be cross-compared and corrected (e.g., ...). Figure 8 (As shown). The cross-comparison unit 311 of the orientation correction unit 310 is used to cross-compare the six-DOF orientation OD of the movable object 900 with the six-DOF orientation CD of the movable camera 110. The correction unit 312 is used to correct the six-DOF orientation OD of the movable object 900 with the six-DOF orientation CD of the movable camera 110.
[0108] To reduce the impact of unconscious slight head movements, the six degrees of freedom orientation of the movable camera and movable objects are recalculated, resulting in a virtual screen D2 (shown on...). Figure 1A The shaking causes dizziness. The orientation stabilization unit 320 is used to determine whether the change in the six-degree-of-freedom orientation OD of the movable object 900 or the six-degree-of-freedom orientation CD of the movable camera 110 is less than a preset value, and therefore does not change the six-degree-of-freedom orientation OD of the movable object 900 or the six-degree-of-freedom orientation CD of the movable camera 110.
[0109] The visual axis calculation unit 330 is used to calculate the visual axes of the user's eyes based on the six-degree-of-freedom orientation CD of the movable camera 110.
[0110] The screen orientation calculation unit 340 is used to calculate the six-degree-of-freedom orientation DD of the virtual screen D2 based on the six-degree-of-freedom orientation OD of the movable object 900 and the six-degree-of-freedom orientation CD of the movable camera 110, so that the virtual screen D2 moves together with the movable object 900 (e.g., ...). Figure 1B (as shown), or the viewpoint displayed on the virtual screen D2 changes with the six degrees of freedom orientation of the movable camera 110.
[0111] The image generation unit 351 of the stereoscopic image generation unit 350 is used to generate images based on the six degrees of freedom orientation DD of the virtual screen D2 and the stereoscopic display (e.g., Figure 1A The optical parameters of the MR glasses G1 generate the left-eye and right-eye images of the virtual screen D2. The imaging unit 352 of the stereoscopic image generation unit 350 is used to display the stereoscopic image of the virtual screen D2 on a stereoscopic display (e.g., a stereoscopic display). Figure 1A MR glasses G1).
[0112] The imaging unit 352 of the stereoscopic image generation unit 350 can display the virtual screen D2 at a specific location around the movable object 900 according to user settings.
[0113] In summary, although the present invention has been disclosed above with reference to embodiments, it is not intended to limit the invention. Those skilled in the art will be able to make various modifications and refinements without departing from the spirit and scope of the invention. Therefore, the scope of protection of the present invention shall be determined by the claims.
Claims
1. A method for simultaneously tracking multiple six-degree-of-freedom orientations of a moving object and a moving camera, characterized in that... Methods for simultaneously tracking multiple six-DOF orientations of a moving object and a moving camera include: The movable camera is used to capture multiple images of the movable object. Multiple environmental feature points are extracted from these images. These environmental feature points are matched to calculate multiple camera matrices for the movable camera. The six-degree-of-freedom orientation of the movable camera is then calculated from these camera matrices. The environmental feature points are the feature points of the background outside the movable object in the images. Multiple feature points of the movable object are deduced from the images captured by the movable camera. The coordinates of the feature points of the movable object are corrected by the camera matrices corresponding to the images and by predefined geometric and time constraints. Then, the six degrees of freedom orientation of the movable object is calculated by using the corrected coordinates of the feature points and their corresponding camera matrices.
2. The method for simultaneously tracking the six degrees of freedom orientation of a movable object and a movable camera according to claim 1, wherein the feature points of the movable object are predefined and calculated from the images captured by the movable camera, and the coordinates of the feature points are calculated by comparing them with the images captured by the movable camera.
3. The method for simultaneously tracking the six degrees of freedom orientation of a movable object and a movable camera according to claim 1, wherein the feature points of the movable object are inferred from the images captured by the movable camera, and the coordinates of the feature points are inferred by a neural network inference model, wherein the neural network inference model is pre-trained, the training data consists of manual labeling and automatic expansion, and the geometric constraints and the time constraints are added during the training process.
4. The method for simultaneously tracking the six degrees of freedom orientation of a movable object and a movable camera according to claim 3, wherein when tracking the movable object, the neural network inference model performs incremental learning in the background, the training data for the incremental learning including: The movable camera captures images and automatically augments those images. The manually marked coordinates of the feature points corresponding to the images are replaced with the corrected coordinates of the feature points. The weights in the neural network inference model are adjusted and the neural network inference model is updated to accurately calculate the coordinates of the feature points of the movable object.
5. The method for simultaneously tracking the six degrees of freedom orientation of a moving object and a moving camera according to claim 1, further comprising: Cross-compare the six degrees of freedom orientation of the movable object with the six degrees of freedom orientation of the movable camera to correct the six degrees of freedom orientation of the movable object and the six degrees of freedom orientation of the movable camera. When the changes in the six degrees of freedom orientation of the movable object or the six degrees of freedom orientation of the movable camera are less than the preset value, the six degrees of freedom orientation of the movable object and the six degrees of freedom orientation of the movable camera are not changed. The visual axes of the user's eyes are calculated based on the six degrees of freedom orientation of the movable camera; The six degrees of freedom orientation of the virtual screen is calculated based on the six degrees of freedom orientation of the movable object and the six degrees of freedom orientation of the movable camera; as well as The left and right eye images of the virtual screen are generated based on the six degrees of freedom orientation of the virtual screen and the optical parameters of the stereoscopic display, so as to display the stereoscopic image of the virtual screen on the stereoscopic display. The movable camera is set on the stereoscopic display, the stereoscopic display is worn on the eyes, and the virtual screen is located next to the movable object.
6. The method for simultaneously tracking the six degrees of freedom orientation of a movable object and a movable camera according to claim 5, wherein the virtual screen is set by the user to be displayed around the movable object, and the virtual screen moves with the movable object.
7. The method for simultaneously tracking the six degrees of freedom orientation of a movable object and a movable camera according to claim 1, wherein the step of correcting the coordinates of the feature points of the movable object includes: Using these camera matrices, the two-dimensional coordinates of these feature points are projected onto their corresponding three-dimensional coordinates; Based on this geometric constraint, feature points with 3D coordinate deviations greater than a predetermined value are deleted, or the coordinates of undetected feature points are supplemented using the coordinates of adjacent feature points according to this geometric constraint; and Based on this time constraint, the coordinate changes of these feature points in consecutive images are compared, and then the coordinates of the feature points whose coordinate changes are greater than a set value are corrected using the coordinates of the corresponding feature points in consecutive images.
8. The method for simultaneously tracking the six degrees of freedom orientation of a movable object and a movable camera according to claim 1, wherein in the step of calculating the six degrees of freedom orientation of the movable object, For the movable object in the plane, the fitting plane is calculated using these feature points, and the orientation of the six degrees of freedom of the movable object is defined by the center point and normal vector of the plane; For a non-planar movable object, the orientation of the six degrees of freedom of the movable object is defined by the centroid of the three-dimensional coordinates of the feature points.
9. The method for simultaneously tracking the six degrees of freedom orientation of a movable object and a movable camera according to claim 1, wherein the movable object is a rigid object and the movable camera is mounted on a head-mounted stereoscopic display, a mobile device, a computer, or a robot.
10. A system for simultaneously tracking multiple six-degree-of-freedom orientations of a moving object and a moving camera, characterized in that... Systems that simultaneously track multiple six-degree-of-freedom orientations of a moving object and a moving camera include: The movable camera is used to take pictures of the movable object in order to capture multiple images; A six-DOF orientation calculation unit for a movable camera is used to extract multiple environmental feature points from the images, match these environmental feature points to calculate multiple camera matrices for the movable camera, and then calculate the six-DOF orientation of the movable camera using these camera matrices. The environmental feature points are feature points of the background outside the movable object in the images; and The six-degree-of-freedom orientation calculation unit for movable objects is used to deduce multiple feature points of the movable object from the images captured by the movable camera, correct the coordinates of the feature points of the movable object by using the camera matrices corresponding to the images, as well as predefined geometric constraints and time constraints, and then calculate the six-degree-of-freedom orientation of the movable object using the corrected coordinates of the feature points and the corresponding camera matrices.
11. The system for simultaneously tracking the six degrees of freedom orientation of a movable object and a movable camera according to claim 10, wherein the movable camera six degrees of freedom orientation calculation unit include: An environmental feature extraction unit is used to extract environmental feature points from these images; The camera matrix calculation unit calculates the camera matrices of the movable camera to match the environmental feature points; and The camera orientation calculation unit uses the camera matrices to calculate the six degrees of freedom orientation of the movable camera.
12. The system for simultaneously tracking the six degrees of freedom orientation of a movable object and a movable camera according to claim 10, wherein the six degrees of freedom orientation calculation unit for the movable object comprises: The object feature estimation unit is used to estimate the feature points of the movable object from the images captured by the movable camera. The object feature coordinate correction unit is used to correct the coordinates of the feature points of the movable object by using the camera matrices corresponding to the images, as well as the predefined geometric constraints and time constraints. as well as The object orientation calculation unit calculates the six degrees of freedom orientation of the movable object using the corrected coordinates of the feature points and their corresponding camera matrices.
13. The system for simultaneously tracking the six degrees of freedom orientation of a movable object and a movable camera according to claim 12, wherein the object feature estimation unit estimates the feature points of the movable object from the images captured by the movable camera as predefined, and compares them with the images captured by the movable camera to estimate the coordinates of the feature points.
14. The system for simultaneously tracking the six degrees of freedom orientation of a movable object and a movable camera according to claim 12, wherein the object feature estimation unit estimates the feature points of the movable object from the images captured by the movable camera, and estimates the coordinates of the feature points by a neural network inference model, wherein the neural network inference model is pre-trained, the training data consists of manual labeling and automatic expansion, and the geometric constraints and time constraints are added during the training process.
15. The system for simultaneously tracking the six degrees of freedom orientation of a moving object and a moving camera as claimed in claim 14, wherein, while tracking the moving object, the neural network inference model performs incremental learning in the background, the training data for the incremental learning comprising: The movable camera captures images and automatically augments those images. The manually marked coordinates of the feature points corresponding to the images are replaced with the corrected coordinates of the feature points. The weights in the neural network inference model are adjusted and the neural network inference model is updated to accurately calculate the coordinates of the feature points of the movable object.
16. The system for simultaneously tracking the six degrees of freedom orientation of a moving object and a moving camera according to claim 10, further comprising: The orientation correction unit is used to cross-compare the six degrees of freedom orientation of the movable object with the six degrees of freedom orientation of the movable camera, so as to correct the six degrees of freedom orientation of the movable object and the six degrees of freedom orientation of the movable camera. The orientation stabilization unit does not change the orientation of the six degrees of freedom of the movable object or the orientation of the six degrees of freedom of the movable camera when the change in the orientation of the six degrees of freedom of the movable object or the orientation of the six degrees of freedom of the movable camera is less than a preset value. The visual axis calculation unit is used to calculate the visual axis of the user's eyes based on the six degrees of freedom orientation of the movable camera; The screen orientation calculation unit is used to calculate multiple six-degree-of-freedom orientations of the virtual screen based on the six-degree-of-freedom orientations of the movable object and the six-degree-of-freedom orientations of the movable camera. as well as A stereoscopic image generation unit is used to generate left-eye and right-eye images of the virtual screen based on the six degrees of freedom orientation of the virtual screen and the optical parameters of the stereoscopic display, so as to display the stereoscopic image of the virtual screen on the stereoscopic display. The movable camera is set on the stereoscopic display, the stereoscopic display is worn on the eyes, and the virtual screen is located next to the movable object.
17. The system for simultaneously tracking the six degrees of freedom orientation of a movable object and a movable camera as claimed in claim 16, wherein the virtual screen is set by the user to be displayed around the movable object, and the virtual screen moves with the movable object.
18. The system for simultaneously tracking the six degrees of freedom orientation of a moving object and a moving camera according to claim 12, wherein the object feature coordinate correction unit Using these camera matrices, the two-dimensional coordinates of these feature points are projected onto their corresponding three-dimensional coordinates; and Based on this geometric constraint, feature points with 3D coordinate deviations greater than a predetermined value are deleted, or the coordinates of undetected feature points are supplemented using the coordinates of adjacent feature points according to this geometric constraint; and Based on the time constraint, the coordinate changes of these feature points in consecutive images are compared, and then the coordinates of the feature points whose coordinate changes are greater than a set value are corrected using the coordinates of the corresponding feature points in consecutive images.
19. The system for simultaneously tracking the six degrees of freedom orientation of a moving object and a moving camera according to claim 12, wherein the object orientation calculation unit For the movable object in the plane, the fitting plane is calculated using these feature points, and the orientation of the six degrees of freedom of the movable object is defined by the center point and normal vector of the plane; For a non-planar movable object, the orientation of the six degrees of freedom of the movable object is defined by the centroid of the three-dimensional coordinates of the feature points.
20. The system for simultaneously tracking the six degrees of freedom orientation of a movable object and a movable camera according to claim 10, wherein the movable object is a rigid object and the movable camera is mounted on a head-mounted stereoscopic display, a mobile device, a computer, or a robot.
Citation Information
Patent Citations
Three-dimensional rotation and motion detecting and rotation axis positioning method based on feature points
CN106651942A
A mobile robot tracking control method based on adaptive pose estimation
CN109102525A