Coordinate system alignment method, device, storage medium, and program product
By acquiring image feature data of the display content of the target device through the camera of the head-mounted device, determining the relative pose and adjusting the origin of the coordinate system, the problem of high cost, high efficiency and low efficiency of coordinate system alignment in multi-person head-mounted MR head-mounted display devices is solved, and efficient and low-cost coordinate system alignment is achieved.
Patent Information
- Application Number
- PCT/CN2024/098259
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-06-30
- Filing Date
- 2024-06-07
- Publication Date
- 2026-01-02
AI Technical Summary
In existing technologies, the coordinate system alignment method for multi-person head-mounted MR headsets requires the transmission of a large amount of image or map data, which is costly and inefficient.
The image feature data of the display content of the target device is obtained by the camera of the head-mounted device, the relative pose is determined, and the origin of the coordinate system is adjusted based on the current pose and the relative pose to achieve the alignment of the coordinate systems of multiple head-mounted devices.
It eliminates the need to transmit large amounts of image or map data, reducing costs and improving the efficiency of coordinate system alignment, simplifying algorithms and saving power consumption.
Smart Images

Figure CN2024098259_02012026_PF_FP_ABST
Abstract
Description
Coordinate system alignment method, device, storage medium and program product
[0001] The present application claims priority to the Chinese patent application No. 2023108038379, filed on June 30, 2023, entitled "Coordinate system alignment method, device, storage medium and program product", the whole content of which is incorporated herein by reference. TECHNICAL FIELD
[0002] Embodiments of the present disclosure relate to the technical field of extended reality, and in particular to a coordinate system alignment method, device, storage medium and program product. BACKGROUND
[0003] For MR head-mounted multi-person activities, i.e., games or activities participated by multiple people wearing MR head-mounted devices, such as virtual gun battles or virtual car racing, multiple people need to see the same virtual objects at the same location in the real world to effectively carry out the activities.
[0004] In the related art, the coordinates of the head-mounted devices can be aligned by using spatial anchors.
[0005] However, in the related art, at least the following technical problems exist: the above-mentioned method needs to transmit a large amount of image or map data between devices, which is costly and inefficient.
[0006] SUMMARY
[0007] Embodiments of the present disclosure provide a coordinate system alignment method, device, storage medium and program product to reduce costs and improve the efficiency of coordinate system alignment.
[0008] In a first aspect, embodiments of the present disclosure provide a coordinate system alignment method applied to a first head-mounted device, the head-mounted device being any device in a plurality of head-mounted devices to be aligned, and the method comprising:
[0009] obtaining a current pose of the device in a first coordinate system;
[0010] obtaining image feature data of display content of a target device through a camera of the device at the current pose, and determining a relative pose of the target device in a camera coordinate system of the camera according to the image feature data;
[0011] adjusting an origin of the first coordinate system to a target position according to the current pose and the relative pose, to realize alignment with a coordinate system of another head-mounted device whose coordinate system origin is adjusted to the target position.
[0012] In a second aspect, embodiments of the present disclosure provide a coordinate system alignment device comprising:
[0013] obtain an image feature data of a display content of a target device through a camera of the camera head of the device, and determine a relative pose of the target device in a camera coordinate system of the camera head according to the image feature data;
[0014] obtain an image feature data of a display content of a target device through a camera of the camera head of the device, and determine a relative pose of the target device in a camera coordinate system of the camera head according to the image feature data;
[0015] adjust the origin of the first coordinate system to a target position according to the current pose and the relative pose, so as to align the coordinate system of the device with the coordinate system of the other head-mounted device.
[0016] In a third aspect, an electronic device is provided, including a processor and a memory.
[0017] The memory stores computer-executable instructions.
[0018] The processor executes the computer-executable instructions stored in the memory, so that the at least one processor executes the coordinate system alignment method according to the first aspect and various possible designs of the first aspect.
[0019] In a fourth aspect, a computer-readable storage medium is provided, and the computer-readable storage medium stores computer-executable instructions. When a processor executes the computer-executable instructions, the coordinate system alignment method according to the first aspect and various possible designs of the first aspect is implemented.
[0020] In a fifth aspect, a computer program product is provided, including a computer program. When the computer program is executed by a processor, the coordinate system alignment method according to the first aspect and various possible designs of the first aspect is implemented. BRIEF DESCRIPTION OF DRAWINGS
[0021] In order to more clearly illustrate the technical solutions of the embodiments of the present disclosure or the prior art, the drawings needed in the embodiments or the prior art description will be briefly introduced. Obviously, the drawings in the following description are some embodiments of the present disclosure, and other drawings can be obtained by those skilled in the art without creative labor.
[0022] FIG. 1 is a schematic diagram of a scene application of a coordinate system alignment method according to an embodiment of the present disclosure;
[0023] FIG. 2 is a schematic diagram of a coordinate system alignment method according to an embodiment of the present disclosure;
[0024] FIG. 3 is a schematic diagram of a coordinate system alignment method according to an embodiment of the present disclosure;
[0025] FIG. 4 is a schematic diagram of an MR scene after a virtual object is placed at a reference position of a real reference object according to an embodiment of the present disclosure;
[0026] FIG. 5 is a structural block diagram of a coordinate system alignment device according to an embodiment of the present disclosure;
[0027] FIG. 6 is a schematic diagram of a hardware structure of an electronic device according to an embodiment of the present disclosure. DETAILED DESCRIPTION
[0028] To make the objectives, technical solutions, and advantages of the embodiments of the present disclosure clearer, the following will be combined with the accompanying drawings for the embodiments of the present disclosure to clearly and completely describe the technical solutions of the embodiments of the present disclosure. Obviously, the described embodiments are some but not all of the embodiments of the present disclosure. Based on the embodiments of the present disclosure, all other embodiments obtained by a person of ordinary skill in the art without creative work fall within the protection scope of the present disclosure.
[0029] In the field of virtual reality (VR) / augmented reality (AR) / mixed reality (MR) extended reality (XR) technology, for multi-person activities of an extended reality (XR) head-mounted display device, for example, a virtual gun battle or a virtual racing game in which multiple people wear MR headsets, multiple people need to see the same virtual object at the same location in the real world to effectively carry out the activities. Multiple people wearing MR headsets participate in a virtual gun battle or a virtual racing game by holding virtual guns in a house or placing a racing track on the floor through augmented reality. When multiple people play games with MR headsets, multiple people need to see the same virtual object at the same location in the real world and share the pose and interaction data with each other. For example, A moves and clicks to fire, and B can see the virtual gun in A's hand moving synchronously when A moves the body and raises the hand in real time. Therefore, aligning the coordinate systems of multiple headsets is a key technology to realize multi-person activities.
[0030] In related technologies, the alignment and sharing of the coordinate systems of multiple headsets can be realized by scanning the surrounding environment to construct a map or an anchor point. However, the above method is easily affected by the richness of the environment texture, has high requirements for the environment, and needs to transmit a large amount of image or map data, which is high in cost and low in efficiency.
[0031] To solve the above technical problems, a target device with display content (for example, a handle with multiple light emitting diodes (LEDs), or a display screen displaying display content) can be selected. For each of the multiple head-mounted devices, an image of the display content is captured by a camera of the head-mounted device. Based on the distribution characteristics of the display content in the image, the relative pose of the target device relative to the head-mounted device is obtained, and then based on the relative pose and the current pose of the head-mounted device in its own coordinate system, the origin of the coordinate system of the head-mounted device is adjusted to the target position, that is, the coordinate system origins of the multiple head-mounted devices are adjusted to the same position. The alignment of the coordinate systems of different head-mounted devices is conveniently and quickly achieved, the algorithm is simple, and in this process, a large amount of image or map data does not need to be transmitted between devices, saving power consumption, reducing cost and being efficient. Based on this, the present embodiment provides a coordinate system alignment method.
[0032] FIG. 1 is a schematic diagram of a scene application of a coordinate system alignment method according to an embodiment of the present disclosure. As shown in FIG. 1, the coordinate system alignment system includes multiple head-mounted devices 102 and a target device 101. The target device 101 is provided with display content distributed according to a preset rule, for controlling the display content to emit light. Each head-mounted device 102 is provided with a camera for image capturing of the display content of the target device 101, and based on the captured image, analysis and calculation are performed to determine the relative pose of the target device 101 relative to the head-mounted device 102, and based on the relative pose and the current pose of the head-mounted device 102 in its own coordinate system, the origin of the coordinate system is adjusted to the target position. Optionally, the target device 101 can be any electronic device with display content, such as a tablet device, a target device matched with the head-mounted device 102, etc. The head-mounted device 102 can be an XR device with coordinate system alignment requirements, such as a VR device, an MR device, an AR device, etc.
[0033] In the implementation process, the target device 101 is fixedly placed at a position, and the display content is controlled to be in a light-emitting state at the time of shooting. A plurality of users wear respective head-mounted devices 102. Each head-mounted device 102 performs image shooting on the display content of the target device 101 through the camera of the head-mounted device 102, and determines the relative pose of the target device 101 relative to the head-mounted device 102 based on the image obtained by the shooting, and adjusts the origin of the coordinate system of the head-mounted device 102 to the target position based on the relative pose and the current pose of the head-mounted device 102 in the coordinate system of the head-mounted device 102. The head-mounted device 102 of the present embodiment obtains the image of the display content of the target device 101 through the camera of the plurality of head-mounted devices 102 of the plurality of coordinate systems to be aligned, respectively, and determines the relative pose of the target device 101 relative to each head-mounted device 102 based on the image. Further, based on the current pose of each head-mounted device 102 in the coordinate system of the head-mounted device 102 and the relative pose, the origin of the coordinate system of each head-mounted device 102 is adjusted to the same position, thereby realizing the alignment of the coordinate systems. This way, the algorithm is simple, and there is no need to transmit a large amount of image or map data in this process, which is low in cost and high in efficiency.
[0034] It should be noted that the scene diagram shown in FIG. 1 is only an example, and the coordinate alignment method and the scene described in the embodiments of the present application are used to more clearly illustrate the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided by the embodiments of the present application. It is known to those skilled in the art that as the system evolves and new business scenarios appear, the technical solutions provided by the embodiments of the present application are also applicable to similar technical problems.
[0035] FIG. 2 is a flowchart of a coordinate system alignment method provided by an embodiment of the present disclosure. As shown in FIG. 2, the coordinate system alignment method includes the following steps.
[0036] 201, obtaining the current pose of the device in the first coordinate system.
[0037] The execution subject of the embodiment is any head-mounted device of the plurality of head-mounted devices of the coordinate systems to be aligned. As shown in FIG. 1, the head-mounted device.
[0038] Specifically, after the user wears the head-mounted device, the user can start the tracking and positioning system of the head-mounted device, for example, a head 6DoF tracking and positioning system. The tracking and positioning system can establish a coordinate system, and simultaneously track and obtain the current pose.
[0039] 202, obtaining image feature data of the display content of the target device through the camera of the device, and determining the relative pose of the target device in the camera coordinate system of the camera based on the image feature data.
[0040] Specifically, the user can place the target device at a fixed position in the real space and keep it still, and the target device is provided with display content distributed according to a certain rule. The pixel coordinates of the display content in the images captured by the camera at different poses are different. When the display content emits light, the head-mounted device worn by the user controls the camera provided on the head-mounted device to capture the display content at the current pose, and obtains image feature data. Analysis of the image feature data can obtain the relative pose of the target device relative to the head-mounted device.
[0041] In an embodiment of the present disclosure, considering the power consumption problem, the display content of the target device can be controlled to emit light only when coordinate system alignment is needed, based on which the power consumption can be reduced, and the normal use of the handle of each head-mounted device after alignment will not be disturbed. Specifically, before obtaining the image feature data of the display content of the target device by the camera of the self device, it can further include: establishing a communication connection with the target device;
[0042] Based on the communication connection, a control signal is sent to the target device to make the target device emit light according to the control signal.
[0043] In an embodiment of the present disclosure, obtaining the image feature data of the display content of the target device by the camera of the self device can include: establishing a communication connection with the target device; based on the communication connection, a control signal is sent to the target device to make the target device adjust the light-emitting frequency of the display content of the target device. During the light-emitting process of the display content at the light-emitting frequency, the image feature data of the display content of the target device is obtained by the camera of the self device.
[0044] For example, the target device is a handle. Usually, one target device corresponds to one head-mounted device. The target device can be selected to match the coordinate system to be aligned with multiple head-mounted devices, for example, the target device corresponding to any one of the multiple head-mounted devices. The target device can establish a wireless or wired communication connection with the head-mounted device, for example, a Bluetooth connection. The head-mounted device can obtain the relevant information of the target device based on the above communication connection, including the model of the target device. Different models of target devices correspond to different display contents, and the head-mounted device can further determine which control signal to send to the target device to control the display content of the target device according to the model of the target device. For example, when there are multiple fixed light-emitting electronic elements on the target device, the head-mounted device can determine the relative spatial position relationship between the multiple light-emitting electronic elements; when the target device includes a display screen, the head-mounted device can control or determine the display content such as the displayed pattern, so as to perform subsequent image feature recognition and calculation.
[0045] Since the shooting frame rate of the camera of different head-mounted devices can be different, the shooting time is also usually different. In order to adapt to the shooting frame rate and shooting time of different head-mounted devices, so that the camera of each head-mounted device can capture the image of the display content when it is emitting light, the setting of the lighting frequency of the display content of the target device can be performed through the above communication connection.
[0046] In an embodiment of the present disclosure, the control signal includes the shooting frame frequency of the camera of the head-mounted device, and adjusting the lighting frequency of the display content of the target device according to the control signal can include: adjusting the lighting frequency of the display content of the target device to be consistent with the shooting frame frequency. It can be understood that the shooting frame frequency included in the control signal can include the shooting frame rate and a timestamp representing the shooting frame, so that the target device adjusts the lighting frequency of the display content to the shooting frame rate, and aligns the lighting period with the shooting time according to the timestamp. When the target device is adapted to the head-mounted device and has been set to only light according to the shooting frame rate, the shooting frame frequency included in the control signal can only include the timestamp representing the shooting frame, so that the target device can also be aligned only based on the timestamp.
[0047] Specifically, in the process of capturing the image of the display content by the camera of the plurality of head-mounted devices, the shooting can be completed in various ways. In an implementable manner, the plurality of head-mounted devices can be completed in sequence, for example, the plurality of head-mounted devices include head-mounted devices A, B and C. The target device selects the handle a corresponding to the head-mounted device A. First, the communication connection between A and a can be established, and the lighting frequency of a is set to be the same as the shooting frame frequency of A, so as to complete the shooting of the display content (for example, a plurality of LEDs) of a by the camera of A. Based on the same process, the shooting of the display content of a by the camera of B is completed, and finally, based on the same process, the shooting of the display content of a by the camera of C is completed.
[0048] In an embodiment of the present disclosure, the way of adjusting the lighting frequency of the display content of the target device to be consistent with the shooting frame frequency can be various, for example, the target device can send its lighting frequency to the camera of the head-mounted device, so that the camera adjusts its shooting frame frequency to be consistent with the lighting frequency.
[0049] In another implementation, the plurality of head-mounted devices can be connected to the target device at the same time, and the lighting frequency of the target device is adjusted in sequence to be the same as the shooting frame frequency of the camera of the plurality of head-mounted devices in the process, and the shooting is completed. For example, head-mounted devices A, B and C are all connected to handle a through Bluetooth connection. In the subsequent process, the lighting frequency of a can be first set to be the same as the shooting frame frequency of A, and the shooting of the display content of a by the camera of A is completed. Then, the lighting frequency of a is set to be the same as the shooting frame frequency of B, and the shooting of the display content of a by the camera of B is completed. Finally, the lighting frequency of a is set to be the same as the shooting frame frequency of C, and the shooting of the display content of a by the camera of C is completed. In this way, the user operation is facilitated, and the user does not need to arrange the order.
[0050] In yet another implementation, the lighting frequency of the target device can be set to be the same as the shooting frame frequency of the camera of the plurality of head-mounted devices. For example, the least common multiple of the shooting frame frequencies of head-mounted devices A, B and C can be calculated according to the shooting frame frequencies of A, B and C, and then the lighting frequency of handle a can be set to be the least common multiple. In this way, the efficiency can be improved.
[0051] In still another implementation, adjusting the lighting frequency of the display content of the target device according to the control signal can include: controlling the display content of the target device to be always on during the process of aligning the coordinate systems of the plurality of head-mounted devices. By setting the display content of the target device to be always on, the calculation can be saved. Of course, after the coordinate systems are aligned, the head-mounted device can send an instruction to the target device to close the always-on mode, so that the target device returns to the lighting frequency before the change.
[0052] In an embodiment of the present disclosure, the target device can include a display screen, and the display content is the display content on the display screen, or the target device includes a plurality of fixed light-emitting electronic elements, and the display content is the light spot after the plurality of light-emitting electronic elements are turned on.
[0053] In an embodiment of the present disclosure, the target device can be a handle, which can be a handle corresponding to any of the plurality of head-mounted devices, or the handle can be a handle identifiable by the plurality of head-mounted devices. Specifically, the handle can be a handle matched with the head-mounted device itself, or the handle can be a handle identifiable by the head-mounted device, that is, the head-mounted device can know the distribution of the LED and other electronic light-emitting elements on the handle, so as to facilitate subsequent calculation of image feature data.
[0054] In one embodiment of the present disclosure, determining the relative pose of the target device in the camera coordinate system of the camera according to the image feature data can include: determining the three-dimensional coordinates of the plurality of image feature points corresponding to the display content in the image feature data in the world coordinate system and the pose origin of the target device; and determining the relative pose of the pose origin of the target device in the camera coordinate system of the camera according to the three-dimensional coordinates of the plurality of image feature points and the image feature data based on the PNP pose solving algorithm.
[0055] The pose origin of the target device can be determined according to the display content, for example, one of the plurality of image feature points can be determined as the pose origin of the target device, and the center point of the display content can also be determined as the pose origin of the target device according to the display content.
[0056] Specifically, the images of the display content (for example, a plurality of LEDs) at a plurality of shooting angles can be obtained during movement, and then the plurality of images can be subjected to triangulation processing and projection processing to determine the three-dimensional coordinates of the display content in the world coordinate system. After the three-dimensional coordinates of the display content are determined, the relative pose of the pose origin of the target device in the camera coordinate system of the camera, that is, the relative pose of the target device relative to the head-mounted device, can be determined based on the PNP pose algorithm according to the three-dimensional coordinates of the display content and the image feature data.
[0057] The PNP pose algorithm refers to an algorithm for solving camera extrinsic parameters by minimizing the re-projection error through a plurality of pairs of 3D and 2D matching points, in the case of known or unknown camera intrinsic parameters. The PNP solving algorithm that can be used in the embodiments of the present disclosure selects the Gauss-Newton gradient descent algorithm to perform iterative optimization of 6DoF tracking data.
[0058] In one embodiment of the present disclosure, the target device includes an inertial sensing unit IMU, and determining the relative pose of the pose origin of the target device in the camera coordinate system of the camera according to the three-dimensional coordinates of the plurality of image feature points and the image feature data based on the PNP pose solving algorithm can include: determining the initial pose of the pose origin of the target device in the camera coordinate system of the camera according to the three-dimensional coordinates of the plurality of image feature points and the image feature data based on the PNP pose solving algorithm; and optimizing the initial pose according to the IMU inertial navigation data collected by the IMU to obtain the relative pose of the pose origin of the target device in the camera coordinate system of the camera.
[0059] Specifically, after obtaining the relative pose of the target device relative to the head-mounted device based on the PNP pose solving algorithm, in order to obtain more accurate results, the relative pose obtained by the PNP can be further optimized in combination with the IMU inertial navigation data of the target device to obtain the final relative pose.
[0060] For example, a plurality of head-mounted devices to be aligned with the coordinate system, each device with a handle. The handle a of the head-mounted device A can be placed on the desktop or floor, and the handle A remains stationary until all head-mounted devices are recognized. The communication connection between each head-mounted device and the handle a can be established through Bluetooth, so that the head-mounted device controls the display content of the handle based on the communication connection.
[0061] The cameras of the plurality of head-mounted devices are aligned with the handle a to capture images of the display content when it is lit. Then, based on the PNP pose algorithm and combined with the 6-axis IMU inertial navigation data, the 6DoF positioning data [AR, At] of the handle controller relative to the head-mounted device is estimated in real time. For example, AR represents a 3x3 dimensional relative rotation matrix, and At represents a 3*1 relative displacement matrix.
[0062] 203. Adjust the origin of the first coordinate system to the target position based on the current pose and the relative pose, and realize the alignment of the coordinate system of the other head-mounted device whose coordinate system origin is adjusted to the target position.
[0063] Specifically, after determining the relative pose and the current pose, the current pose and the relative pose can be multiplied to adjust the coordinate origin of the head-mounted device itself, and the coordinate origin is adjusted to the pose origin of the target device.
[0064] For example, the head-mounted device starts the 6DoF tracking positioning system, and the coordinate origin of the 6DoF pose [R, t] tracked by the system is set as the pose origin of the target device. After all the head-mounted devices of the users complete this conversion, the purpose of the unified coordinate system is achieved. Wherein, [R, t]*[AR, At] is the converted 6dof pose. Then, by transmitting the own pose and handle operation to each other through the network, the purpose of sharing the pose and action can be achieved.
[0065] It should be noted that the order of steps 201 and 202 in the embodiments of the present disclosure can be interchanged.
[0066] As can be seen from the above description, the images of the display content of the target device are obtained by the cameras of the plurality of head-mounted devices to be aligned with the coordinate system, and the relative pose of the target device relative to each head-mounted device is determined based on the image, and then the origin of the coordinate system of each head-mounted device is adjusted to the same position based on the current pose of each head-mounted device in the own coordinate system and the relative pose, so as to realize the alignment of the coordinate system. This method is simple in algorithm, and there is no need to transmit a large amount of image or map data between devices in this process, which is low in cost and high in efficiency.
[0067] FIG. 3 is a flowchart of a coordinate system alignment method provided by the embodiments of the present disclosure. In the present embodiment, the determination of the alignment accuracy of the coordinate system alignment is exemplarily illustrated, and the coordinate system alignment method comprises:
[0068] 301、Obtain the current pose of the self device in the first coordinate system.
[0069] 302、In the current pose, obtain the image feature data of the display content of the target device through the camera of the self device, and determine the relative pose of the target device in the camera coordinate system of the camera according to the image feature data.
[0070] 303、According to the current pose and the relative pose, adjust the origin of the first coordinate system to the target position, and realize the alignment of the coordinate system of the other head-mounted device whose coordinate system origin is adjusted to the target position.
[0071] Steps 301 to 303 in the embodiment are similar to steps 201 to 203 in the above embodiment, which will not be described here.
[0072] 304、Place the virtual object in a preset pose to the reference position of the entity reference, and record the relative position relationship between the placed virtual object and the entity reference to obtain the scene record.
[0073] Specifically, after the coordinate system alignment based on the coordinate system alignment algorithm, the effect of the alignment needs to be evaluated in order to improve the algorithm and correct the alignment result. The present inventors have found that the accuracy of the algorithm alignment can be evaluated based on the error between the manual alignment and correction of the pose of the same virtual object in different devices and the calculation algorithm alignment.
[0074] The virtual object can be a single solid or a combination of multiple solids, and the surface can have an asymmetric texture or pattern. The virtual object seen by different head-mounted devices can be the same virtual object generated by any head-mounted device.
[0075] In the specific process, the same virtual object can be shared in the same position in the same scene after the coordinate system alignment between multiple head-mounted devices. Taking a head-mounted device as an MR device as an example, after the coordinate system alignment of multiple MR devices, the same MR scene can be seen through information transmission. MR device A places a virtual object at a certain position in the real scene. At this time, in the MR field of view, the virtual object (which can be a square, a triangle, three lines, etc. any three-dimensional virtual object) can be placed on the corner of a table, the center of a table, the corner of a wall, the floor, etc. through a handle, a gesture or a key. Other users can also see this virtual object in the scene. As shown in FIG. 4, the virtual object is a combination of a cube and a cone, and the virtual object is placed in the corner of a chair. All users can see that due to the existence of alignment error, there can be a certain position or angle difference between the relative position relationship between the virtual object and the chair seen by all users.
[0076] The MR device A walks around the virtual object and looks at the virtual object under the rendering of the virtual object, and records the scene record S (takes pictures or videos) after the placement. The scene record S can include: an animation video stream (one or more) or a taken picture (one or more), wherein the video or picture presents the MR scene after the virtual object is placed. For example, the scene of the virtual block placed on the table, which contains the virtual block, the table, and the real background in the real scene. FIG. 4 is a schematic diagram of an MR scene after a virtual object is placed at a reference position of a physical reference object according to an embodiment of the present disclosure. As shown in FIG. 4, the virtual object is placed in the corner of the chair.
[0077] 305. The scene record is sent to the second head-mounted device, so that the second head-mounted device adjusts the pose of the virtual object according to the scene record, and determines the relative pose between the pose before the adjustment and the pose after the adjustment; and the alignment accuracy of the first coordinate system and the second coordinate system of the second head-mounted device is determined according to the relative pose.
[0078] The second coordinate system is the coordinate system of the second head-mounted device after alignment to the target position.
[0079] Specifically, after the scene record is generated, the scene record can be sent to other head-mounted devices, so that the other head-mounted devices adjust the pose of the virtual object according to the scene record, and the adjustment is the same as the visual effect of the first head-mounted device. After the adjustment, the alignment accuracy can be determined according to the pose of the virtual object before and after the adjustment, that is, the size of the adjustment range.
[0080] For example, A obtains the pose coordinates Ta of the center point of the virtual object in the 6Dof tracking system (SLAM) in the A head-mounted device coordinate system, which includes position and angle information, and is a matrix data [R, t]. For example, R represents a 3x3 relative rotation matrix, and t represents a 3*1 relative displacement matrix. The pose coordinates Ta and the scene record S are uploaded to the server. Subsequently, the head-mounted device B is started in the same environment, and the scene record S is downloaded from the server.
[0081] The head-mounted device B views the video or picture in the scene record S. The same virtual object is also placed in the same way as the virtual object in the scene S, and the orientation and position of the virtual object are consistent with the placement of A. B confirms the correct pose of the placed virtual object by walking around the virtual object. The position and orientation of the virtual object are consistent with what A sees. Specifically, the position and pose of the virtual object before adjustment, i.e., the pose coordinates Tb1 of the 6Dof tracking system (SLAM) under the second coordinate system of the head-mounted device B, contain position and angle information and are a matrix data [R1, t1], for example, R1 represents a 3x3 relative rotation matrix, and t1 represents a 3*1 relative displacement matrix. Adjust the position and orientation of the virtual object under the self picture according to the scene record S, so that the virtual object is consistent with the position and orientation in the picture or video in the scene record S. The position and pose of the virtual object after adjustment, i.e., the pose coordinates Tb2 of the 6Dof tracking system (SLAM) under the coordinate system of the B head-mounted device, contain position and angle information and are a matrix data [R2, t2], for example, R2 represents a 3x3 relative rotation matrix, and t2 represents a 3*1 relative displacement matrix.
[0082] The head-mounted device B calculates the position and angle change amount δTab before and after the adjustment of the virtual object. δTab = Tb1*(Tb2.inv). (Tb2.inv) represents the inverse matrix of the Tb2 matrix. The movement amount δTab of the virtual object position and angle on the B device is obtained by multiplying Tbl by (Tb2.inv), which is the coordinate system alignment error of A and B. δTab contains position and angle information and is a matrix data [δRab, δtab], for example, δRab represents a 3x3 relative rotation matrix, and δtab represents a 3*1 relative displacement matrix. The position and angle change amount δTab [δRab, δtab] is output as the coordinate alignment error value between A and B.
[0083] The other devices repeat the operations of the A and B devices to obtain alignment errors between all devices, δTbc, δTcd, δTac, δTad, etc. Then, all angle errors and displacement errors {δTbc, δTcd, δTac, δTad} can be averaged respectively, i.e., average alignment errors. Each δT contains position and angle information, which is a matrix data [δR, δt], for example, δR represents a 3x3-dimensional relative rotation matrix, and δt represents a 3*1 relative displacement matrix. The unit of δR is converted into degrees, and the average of the accumulations between N device pairs is obtained, i.e., an alignment angle average error Raver=(δRbc+δRcd+δRac+δRad+…) / N. Exemplarily, the unit of δt is converted into meters, and the average of the accumulations between N device pairs is obtained, i.e., an alignment displacement average error taver=(δtbc+δtcd+δtac+δtad+…) / N. The alignment error is Raver degrees and δt meters.
[0084] From the above description, by manually aligning and correcting the poses of the same virtual object in different head-mounted devices, the error between the manual alignment and the algorithm alignment is calculated, and the accuracy of the algorithm alignment is evaluated based on the error, so that the alignment algorithm can be improved based on the accuracy, and the alignment result can be corrected.
[0085] The coordinate system alignment method, device, storage medium, and program product provided in this embodiment obtain a current pose of the device in a first coordinate system, and in the current pose, image feature data of display content of a target device is obtained through a camera of the device, and a relative pose of the target device in a camera coordinate system of the camera is determined according to the image feature data. The origin of the first coordinate system is adjusted to a target position according to the current pose and the relative pose, so as to realize alignment of the coordinate system with other head-mounted devices whose coordinate system origins are adjusted to the target position. The method provided in this embodiment obtains images of display content of a target device through cameras of multiple head-mounted devices of multiple coordinate systems to be aligned, respectively, and determines relative poses of the target device with respect to the respective head-mounted devices based on the images, respectively. Then, the origins of the respective head-mounted devices in the respective coordinate systems are adjusted to the same position based on the current poses of the respective head-mounted devices in the respective coordinate systems and the relative poses, so as to realize alignment of the coordinate systems. This method is simple in algorithm, and a large amount of image or map data does not need to be transmitted between devices in this process, which is low in cost and high in efficiency.
[0086] Corresponding to the coordinate system alignment method in the above embodiment, FIG. 5 is a structural block diagram of a coordinate system alignment device provided in an embodiment of the present disclosure. For ease of illustration, only parts related to the embodiment of the present disclosure are shown. Referring to FIG. 5, the device 500 includes an obtaining module 501, a determining module 502, and an adjusting module 503.
[0087] The acquisition module 501 is configured to acquire a current pose of the device in a first coordinate system.
[0088] The determination module 502 is configured to acquire image feature data of display content of a target device through a camera of the device, and determine a relative pose of the target device in a camera coordinate system of the camera according to the image feature data.
[0089] The adjustment module 503 is configured to adjust an origin of the first coordinate system to a target position according to the current pose and the relative pose, so as to align the coordinate system of the device with a coordinate system of another head-mounted device whose origin is adjusted to the target position.
[0090] In an embodiment of the present disclosure, the determination module 502 is further configured to: establish a communication connection with the target device; and send a control signal to the target device based on the communication connection, so that the target device is controlled to light up according to the control signal.
[0091] In an embodiment of the present disclosure, the determination module 502 is specifically configured to: establish a communication connection with the target device; send a control signal to the target device based on the communication connection, so that the target device is controlled to adjust a light-up frequency of the display content according to the control signal; and acquire image feature data of the display content of the target device through the camera of the device in a process in which the display content emits light at the light-up frequency.
[0092] In an embodiment of the present disclosure, the determination module 502 is specifically configured to: adjust the light-up frequency of the display content of the target device to be consistent with a shooting frame frequency.
[0093] In an embodiment of the present disclosure, the determination module 502 is specifically configured to: send the light-up frequency of the display content of the target device to the camera of the head-mounted device, so that the camera of the head-mounted device adjusts a shooting frame frequency to be consistent with the light-up frequency.
[0094] In an embodiment of the present disclosure, the target device is a handle, the handle is a handle corresponding to any device of a plurality of head-mounted devices, or the handle is a handle identifiable by the plurality of head-mounted devices.
[0095] In an embodiment of the present disclosure, the target device includes a display screen, and correspondingly, the display content is display content on the display screen, or the target device includes a plurality of fixed light-emitting electronic elements, and correspondingly, the display content is light spots after the plurality of light-emitting electronic elements light up.
[0096] In an embodiment of the present disclosure, the adjusting module 503 is specifically configured to: determine the three-dimensional coordinates of the plurality of image feature points corresponding to the display content in the image feature data in the world coordinate system and the pose origin of the target device; and determine the relative pose of the pose origin of the target device in the camera coordinate system of the camera based on the PNP pose solving algorithm and according to the three-dimensional coordinates of the plurality of image feature points and the image feature data.
[0097] In an embodiment of the present disclosure, the adjusting module 503 is specifically configured to: determine an image feature point in the plurality of image feature points as the pose origin of the target device.
[0098] In an embodiment of the present disclosure, the target device includes an inertial sensing unit IMU, and the adjusting module 503 is specifically configured to: determine the initial pose of the pose origin of the target device in the camera coordinate system of the camera based on the PNP pose solving algorithm and according to the three-dimensional coordinates of the plurality of image feature points and the image feature data; and optimize the initial pose according to the IMU inertial navigation data collected by the IMU to obtain the relative pose of the pose origin of the target device in the camera coordinate system of the camera.
[0099] In an embodiment of the present disclosure, the device 500 further includes a detecting module (not shown) configured to:
[0100] place the virtual object at the reference position of the physical reference object in the preset pose;
[0101] record the relative position relationship between the placed virtual object and the physical reference object to obtain a scene record;
[0102] send the scene record to a second head-mounted device, so that the second head-mounted device adjusts the pose of the virtual object according to the scene record and determines the relative pose between the adjusted pose and the adjusted pose; and determine the alignment accuracy of the first coordinate system and a second coordinate system of the second head-mounted device according to the relative pose.
[0103] The device provided in the embodiment can be used to execute the technical solutions of the method embodiments, and has similar implementation principles and technical effects, which will not be described here in detail.
[0104] In order to implement the above-mentioned embodiments, the present disclosure further provides an electronic device.
[0105] Referring to FIG. 6, a structural diagram of an electronic device 600 suitable for implementing embodiments of the disclosure is shown, which can be a terminal device or a server. The terminal device can include, but is not limited to, a mobile terminal such as a mobile phone, a notebook computer, a digital broadcast receiver, a Personal Digital Assistant (PDA), a Portable Android Device (PAD), a Portable Media Player (PMP), a car terminal (e.g., a car navigation terminal), and the like, and a stationary terminal such as a digital TV, a desktop computer, and the like. The electronic device shown in FIG. 6 is merely an example and should not impose any limitation on the functions and use range of embodiments of the disclosure.
[0106] As shown in FIG. 6, the electronic device 600 can include a processing device (e.g., a central processor, a graphic processor, etc.) 601 that can perform various appropriate actions and processes according to a program stored in a Read Only Memory (ROM) 602 or a program loaded into a Random Access Memory (RAM) 603 from a storage device 608. Various programs and data required for the operation of the electronic device 600 are also stored in the RAM 603. The processing device 601, the ROM 602, and the RAM 603 are connected to each other through a bus 604. An Input / Output (I / O) interface 605 is also connected to the bus 604.
[0107] In general, the following devices can be connected to the I / O interface 605: an input device 606 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, and the like; an output device 607 including, for example, a Liquid Crystal Display (LCD), a speaker, a vibrator, and the like; a storage device 608 including, for example, a magnetic tape, a hard disk, and the like; and a communication device 609. The communication device 609 can allow the electronic device 600 to communicate with other devices wirelessly or wired to exchange data. Although FIG. 6 shows the electronic device 600 having various devices, it should be understood that all of the shown devices are not required to be implemented or provided. More or less devices can be alternatively implemented or provided.
[0108] In particular, according to embodiments of the present disclosure, the processes described above with reference to the flowcharts can be implemented as a computer software program. For example, embodiments of the present disclosure include a computer program product comprising a computer program carried on a computer readable medium, the computer program comprising program code for performing the methods illustrated by the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via the communication device 609, or installed from the storage device 608, or installed from the ROM 602. When the computer program is executed by the processing device 601, the above-mentioned functions defined in the methods of the embodiments of the present disclosure are performed.
[0109] Note that the computer readable medium described above in the present disclosure can be a computer readable signal medium or a computer readable storage medium or any combination thereof. The computer readable storage medium may, for example, be, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or apparatus, or any suitable combination of the foregoing. More specific examples of the computer readable storage medium can include, but are not limited to, an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing. In the present disclosure, the computer readable storage medium can be any tangible medium that contains or stores a program used or used in conjunction with an instruction execution system, device, or apparatus. In the present disclosure, the computer readable signal medium can include a data signal carried in a baseband or as a part of a carrier wave, which carries computer readable program code. Such a propagated data signal can take on many forms, including but not limited to an electromagnetic signal, an optical signal, or any suitable combination of the foregoing. The computer readable signal medium can also be any computer readable medium that is not a computer readable storage medium and that can transmit, propagate, or transport program for use by or in connection with an instruction execution system, device, or apparatus. The program code contained on the computer readable medium can be transmitted by any suitable medium, including but not limited to a wire, cable, optical fiber, RF (radio frequency), or any suitable combination of the foregoing.
[0110] The computer readable medium described above can be included in the electronic device described above; or can exist separately from the electronic device and not be assembled into the electronic device.
[0111] The computer readable medium described above carries one or more programs, which, when executed by the electronic device, cause the electronic device to perform the methods shown in the embodiments described above.
[0112] Computer program code for carrying out operations of the present disclosure can be written in any combination of one or more programming languages, including an object oriented programming language such as Java, Smalltalk, C++ or the like and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider).
[0113] The flow diagrams and the block diagrams in the drawings are illustrations of architectures, functionalities, and operations of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flow diagrams or block diagrams can represent a module, a segment, or a portion of code, which comprises one or more executable instructions for implementing the specified logical function(s). It should also be noted that in some alternative implementations, the functions noted in the block can occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently or the blocks can sometimes be executed in the reverse order, depending upon the functionality involved. It will also be noted that each block of the block diagrams and / or flow diagrams, and combinations thereof, can be implemented by special purpose hardware-based systems that perform the specified functions or operations, or combinations of special purpose hardware and computer instructions.
[0114] The units described in the embodiments of the present disclosure can be implemented by software, or by hardware. In some cases, the name of the unit does not constitute a limitation on the unit itself. For example, the first obtaining unit can also be described as a unit for obtaining at least two Internet protocol addresses.
[0115] The functions described in this specification can be implemented in part or in whole by one or more hardware logic components. For example, and without limitation, illustrative types of hardware logic components that can be used include Field-programmable Gate Arrays (FPGAs), Program-specific Integrated Circuits (ASICs), Program-specific Standard Products (ASSPs), System-on-a-chip systems (SOCs), Complex Programmable Logic Devices (CPLDs), etc.
[0116] In the context of this disclosure, a machine-readable medium can be a tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include but is not limited to an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium would include a lined- up electrical connection, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0117] The above description is only preferred embodiments of the present disclosure and the explanation of the applied technical principles. It should be understood by those skilled in the art that the disclosure range involved in the present disclosure is not limited to the technical solutions formed by the specific combinations of the above technical features, and also covers other technical solutions formed by any combination of the above technical features or equivalent features without departing from the above disclosed concept. For example, the technical solutions formed by replacing the above features with the technical features disclosed in the present disclosure (but not limited to) having similar functions.
[0118] In addition, although each operation is described in a particular order, this should not be understood as requiring the operations to be performed in the specific order shown or in a sequential order. In certain circumstances, multitasking and parallel processing can be advantageous. Similarly, although several implementation details are included in the above discussion, these should not be interpreted as limiting the scope of the present disclosure. Certain features described in the context of separate embodiments can also be combined in a single embodiment. Conversely, various features described in the context of a single embodiment can also be separated and implemented in multiple embodiments.
[0119] Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are disclosed as example forms of implementing the claims.
Claims
1. A coordinate system alignment method, applied to a first head-mounted device, wherein the first head-mounted device is any one of a plurality of head-mounted devices whose coordinate systems are to be aligned, the method comprising: Obtain the current pose of the device in the first coordinate system; In the current pose, image feature data of the display content of the target device is acquired through the camera of the device itself, and the relative pose of the target device in the camera coordinate system of the camera is determined based on the image feature data. Based on the current pose and the relative pose, the origin of the first coordinate system is adjusted to the target position to achieve alignment with the coordinate systems of other head-mounted devices that have adjusted their coordinate system origins to the target position.
2. The method according to claim 1, wherein the method further comprises: Establish a communication connection with the target device; Based on the communication connection, a control signal is sent to the target device so that the target device lights up according to the control signal.
3. The method according to claim 2, wherein sending a control signal to the target device based on the communication connection, so that the target device lights up according to the control signal, comprises: Based on the communication connection, a control signal is sent to the target device so that the target device adjusts the lighting frequency of the display content according to the control signal; Accordingly, acquiring image feature data of the target device's display content through the device's own camera includes: During the process of the displayed content emitting light at the specified lighting frequency, image feature data of the displayed content of the target device is acquired through the camera of the device itself.
4. The method according to claim 3, wherein the control signal includes the shooting frame rate of the camera of the head-mounted device, and the step of adjusting the illumination frequency of the display content of the target device according to the control signal includes: Adjust the illumination frequency of the display content on the target device to match the shooting frame rate.
5. The method according to claim 4, wherein adjusting the illumination frequency of the display content of the target device to be consistent with the shooting frame rate includes: The lighting frequency of the display content of the target device is sent to the camera of the head-mounted device so that the camera of the head-mounted device adjusts its shooting frame rate to match the lighting frequency.
6. The method according to any one of claims 1-5, wherein The target device is a handle, which can be the handle corresponding to any one of a plurality of head-mounted devices, or... The handle is a handle that can be recognized by multiple head-mounted devices.
7. The method according to any one of claims 1-5, wherein The target device includes a display screen, and correspondingly, the display content is the content displayed on the display screen, or... The target device includes multiple fixed light-emitting electronic components, and the corresponding display content is the light spot of the multiple light-emitting electronic components after they are lit.
8. The method according to any one of claims 1-5, wherein determining the relative pose of the target device in the camera coordinate system of the camera based on the image feature data comprises: Determine the three-dimensional coordinates of multiple image feature points corresponding to the displayed content in the image feature data in the world coordinate system and the pose origin of the target device; Based on the PNP pose calculation algorithm, the relative pose of the target device's pose origin in the camera coordinate system of the camera is determined according to the three-dimensional coordinates of multiple image feature points and the image feature data.
9. The method according to claim 8, wherein determining the pose origin of the target device comprises: One of the multiple image feature points is determined as the pose origin of the target device.
10. The method according to claim 8, wherein the target device includes an inertial sensing unit (IMU); the step of determining the relative pose of the target device's pose origin in the camera coordinate system of the camera based on the PNP pose calculation algorithm, according to the three-dimensional coordinates of multiple image feature points and the image feature data, includes: Based on the PNP pose calculation algorithm, the initial pose of the target device's pose origin in the camera coordinate system of the camera is determined according to the three-dimensional coordinates of multiple image feature points and the image feature data. The initial pose is optimized based on the IMU inertial navigation data collected by the IMU to obtain the relative pose of the target device's pose origin in the camera coordinate system of the camera.
11. The method according to any one of claims 1-5, wherein after adjusting the origin of the first coordinate system to the target position according to the current pose and the relative pose, the method further comprises: Place the virtual object at the reference position of the physical reference object with a preset posture; Record the relative positional relationship between the placed virtual object and the physical reference object to obtain a scene record; The scene record is sent to the second head-mounted device, so that the second head-mounted device adjusts the pose of the virtual object according to the scene record and determines the relative pose between the pose before adjustment and the pose after adjustment; and determines the alignment accuracy between the first coordinate system and the second coordinate system of the second head-mounted device according to the relative pose.
12. A coordinate system alignment device, comprising: The acquisition module is used to acquire the current pose of the device in the first coordinate system; The determination module is used to acquire image feature data of the display content of the target device through its own device's camera, and determine the relative pose of the target device in the camera coordinate system of the camera based on the image feature data; The adjustment module is used to adjust the origin of the first coordinate system to the target position according to the current pose and the relative pose, so as to achieve alignment with the coordinate systems of other head-mounted devices that have adjusted their coordinate system origins to the target position.
13. An electronic device, comprising: Processor and memory; The memory stores computer-executed instructions; The processor executes computer execution instructions stored in the memory, causing the processor to perform the coordinate system alignment method as described in any one of claims 1 to 11.
14. A computer-readable storage medium storing computer-executable instructions that, when executed by a processor, implement the coordinate system alignment method as described in any one of claims 1 to 11.
15. A computer program product comprising a computer program, wherein the computer program, when executed by a processor, implements the coordinate system alignment method as described in any one of claims 1 to 11.