Map alignment method, multi-device system and non-transitory computer readable storage medium

The map alignment method for multi-device systems addresses inefficiencies in immersive experience systems by aligning keyframes to establish map consistency, reducing processing resource burden and enhancing user experience.

TWI932361BActive Publication Date: 2026-07-11HTC CORP
0 Cites 0 Cited by

Patent Information

Application Number
TW114129958
Authority / Receiving Office
TW · TW
Patent Type
Patents
Current Assignee / Owner
Priority Date
2024-12-29
Filing Date
2025-08-06
Publication Date
2026-07-11
Estimated Expiration
2045-08-05

AI Technical Summary

Technical Problem

Existing immersive experience systems face inefficiencies in map alignment between head-mounted devices and peripheral devices, leading to a significant burden on processing resources.

Method used

A map alignment method for multi-device systems, involving a host device and a client device, where keyframes are acquired and aligned to establish map consistency using feature extraction-based localization techniques, reducing processing resource burden.

Benefits of technology

Efficient map alignment across multiple devices, achieving lower processing resource usage and improved user experience in immersive environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IMG-2_DRAW_04_A0101_DRAWINGS_1
    Figure IMG-2_DRAW_04_A0101_DRAWINGS_1
  • Figure IMG-2_DRAW_04_A0101_DRAWINGS_2
    Figure IMG-2_DRAW_04_A0101_DRAWINGS_2
  • Figure IMG-2_DRAW_04_A0101_DRAWINGS_3
    Figure IMG-2_DRAW_04_A0101_DRAWINGS_3
Patent Text Reader

Abstract

This disclosure provides a map alignment method, a multi-device system, and a non-transitory computer-readable storage medium. The multi-device system is used to operate in a physical environment and includes a host device and a client device. The map alignment method includes: acquiring host-side keyframes and client-side keyframes at preset time points by the host device and client device, respectively; generating a first client pose by the host device based on the host-side keyframes; and aligning the client map established by the client device detecting the physical environment with the host-side map established by the host device detecting the physical environment by the client device based on the first client pose.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to a method and system, and in particular to a map alignment method and a multi-device system. Prior Technology

[0002] In the field of immersive experience systems (e.g., Virtual Reality (VR), Augmented Reality (AR), Mixed Reality (MR), etc.), some related technologies allow a head-mounted device and at least one peripheral device (e.g., controllers, trackers, etc.) to exchange map data to align the map of the peripheral device with that of the head-mounted device. However, this method is inefficient and places a significant burden on processing resources. Summary of the Invention

[0003] One aspect of this disclosure is a map alignment method. This map alignment method is applicable to a multi-device system, wherein the multi-device system is used to operate in a physical environment and includes a host device and a client device. The map alignment method includes: acquiring a host-side keyframe and a client-side keyframe at a preset time point, respectively, using the host device; generating a first client pose based on the host-side keyframe using the host device; and aligning a client map established by the client device based on the first client pose with a host-side map established by the host device based on the physical environment detected by the client device.

[0004] Another embodiment of this disclosure is a multi-device system. This multi-device system is designed to operate in a physical environment and includes a host device and a client device. The host device detects the physical environment to build a host-side map. The client device detects the physical environment to build a client-side map. The host device and the client device acquire a host-side keyframe and a client-side keyframe respectively at a preset time point. The host device generates a first client pose based on the host-side keyframe, and the client device aligns the client-side map with the host-side map based on the first client pose.

[0005] Another embodiment of this disclosure is a non-transitory computer-readable storage medium having a computer program for executing a map alignment method, wherein the map alignment method is applicable to a multi-device system, wherein the multi-device system is used to operate in a physical environment and includes a host device and a client device, and the map alignment method includes: acquiring a host-side keyframe and a client-side keyframe at a preset time point by the host device and the client device respectively; generating a first client pose by the host device based on the host-side keyframe; and aligning a client map established by the client device in detecting the physical environment with a host-side map established by the host device in detecting the physical environment by the client device based on the first client pose.

[0006] In summary, by generating keyframe pairs (i.e., host-side keyframes and client-side keyframes) between the host device and the client device, the client device can continuously align the client map with the host map using the poses corresponding to the host keyframes (e.g., first client pose, third client pose, etc.). In this way, multi-device systems can efficiently achieve map consistency across multiple devices and have advantages such as lower processing resource burden. Simple Explanation of the Diagram

[0007] Figure 1 is a block diagram illustrating a multi-device system according to some embodiments of the present disclosure. Figure 2 is a flowchart illustrating a map alignment method according to some embodiments of the present disclosure. Figure 3 is a flowchart illustrating an operation of a map alignment method according to some embodiments of the present disclosure. Figure 4 is a schematic diagram illustrating a usage scenario of a multi-device system according to some embodiments of the present disclosure. Figure 5 is a flowchart illustrating a map alignment method according to some embodiments of the present disclosure. Figure 6 is a schematic diagram of a multi-device system illustrated according to some embodiments of the present disclosure. Implementation

[0008] The following detailed description uses examples and accompanying drawings. However, the specific embodiments described are only for explaining this case and are not intended to limit this case. The description of the structural operation is not intended to limit the order of its execution. Any structure that is recombined from the components and produces a device with equivalent function is within the scope of this disclosure.

[0009] The terms "coupled" or "connected" as used in this article can refer to two or more components making direct physical or electrical contact with each other, or making indirect physical or electrical contact with each other, or to two or more components operating or moving together.

[0010] Please refer to Figure 1, which is a block diagram illustrating a multi-device system 100 according to some embodiments of the present disclosure. In some embodiments, the multi-device system 100 can be operated by a user U1 in a physical environment E1 (e.g., a game venue, workplace, residence, etc.) and can provide the user U1 with an immersive experience.

[0011] In some embodiments, as shown in Figure 1, the multi-device system 100 includes a host device 11 and at least one client device 13. In some practical applications, the host device 11 may be implemented by a wearable display device (e.g., a head-mounted device) of an immersive system, while the client device 13 may be implemented by a controller device (e.g., a handheld controller, a wearable controller, etc.) of an immersive system.

[0012] In some embodiments, the host device 11 is used to locate itself and the client device 13 in the physical environment E1, and to provide visual feedback to the user U1 based on the location of the host device 11 and the client device 13. Accordingly, as shown in Figure 1, the host device 11 includes a processor 110, a camera 112, and a display panel 114. The processor 110 is electrically and / or communicatively coupled to the camera 112 and the display panel 114.

[0013] In the above embodiment of the host device 11, the camera 112 is used to capture multiple host-based images in the physical environment E1. It should be understood that these host-based images may include at least one of multiple images of the entire or partial physical environment E1, multiple images of the client device 13, and multiple images of the user U1. By applying some feature extraction-based positioning techniques (e.g., simultaneous localization and mapping) to the host-based images captured by the camera 112, the processor 110 can be used to build a host-based map MH of the physical environment E1, and can also be used to calculate the position and / or orientation of the host device 11 in the host-based map MH. The processor 110 is used to use some interactive tracking techniques (e.g., optical tracking) to calculate the position and / or orientation of the client device 13 relative to the host device 11. Furthermore, the processor 110 is used to generate at least one visual content based on the position and / or orientation of the host device 11 and the client device 13. Display panel 114 is used to display at least one visual content generated by processor 110, thereby providing user U1 with an immersive content CI (i.e., visual feedback).

[0014] In some embodiments, the host device 11 may obstruct the direct visibility of the user U1 to the physical environment E1. In this case, the immersive content CI may be a virtual reality (VR) environment or a mixed reality (MR) environment. Specifically, a virtual reality environment may include at least one virtual reality object that cannot be directly seen by the user U1 in the physical environment E1. The mixed reality environment simulates the physical environment E1 and enables at least one virtual reality object to interact with the simulated physical environment. However, this disclosure is not limited to this. For example, the immersive content CI may be a simulated physical environment that does not contain virtual reality objects, also known as a pass-through view.

[0015] In some embodiments, the host device 11 does not obstruct the user U1's direct visibility of the physical environment E1. In this case, the immersive content CI can be an augmented reality (AR) environment. Specifically, the augmented reality environment uses at least one virtual reality object to enhance the physical environment E1 that the user U1 directly sees.

[0016] In some embodiments, the client device 13 is used to locate itself in the physical environment E1 and to interact with the host device 11, so that the host device 11 can locate the client device 13. Accordingly, as shown in Figure 1, the client device 13 includes a processor 130, a camera 132, and at least one trackable object 134. The processor 130 is electrically and / or communicatively coupled to the camera 132 and the trackable object 134. Specifically, the trackable object 134 is disposed on the outer surface of the client device 13 so that the user U1 can directly see it or the camera 112 of the host device 11 can directly capture it. Furthermore, following the above embodiments where the immersive content CI is a virtual reality environment, a mixed reality environment, or an augmented reality environment, the user U1 can use the client device 13 to control at least one virtual reality object in the immersive content CI.

[0017] In the above embodiment of client device 13, camera 132 is used to capture multiple client-based images in the physical environment E1. It should be understood that these client-based images may include at least one of multiple images of the entire or partial physical environment E1, multiple images of the host device 11, and multiple images of the user U1. By applying some feature extraction-based positioning techniques (e.g., simultaneous localization and mapping) to the client-based images captured by camera 132, processor 130 can be used to build a client map MC of the physical environment E1, and can also be used to calculate the position and / or orientation of client device 13 in the client map MC. Furthermore, processor 130 is used to actuate trackable object 134 so that client device 13 can interact with host device 11. For example, when trackable object 134 is actuated, processor 110 of host device 11 can identify the image of trackable object 134 located on client device 13 from host-based images captured by camera 112 of host device 11.

[0018] In the above embodiments, processor 110 and processor 130 may each be implemented by a central processing unit (CPU), graphics processing unit (GPU), application-specific integrated circuit (ASIC), microprocessor, system-on-a-chip (SoC), or other suitable processing circuitry. Display panel 114 may be implemented by an active-matrix organic light-emitting diode (AMOLED) display, organic light-emitting diode (OLED) display, or other suitable display. Trackable object 134 may be implemented by an infrared light-emitting diode (LED), but is not limited thereto. For example, in some embodiments, trackable object 134 may be the physical shape of the entire or part of the client device 13, which may be pre-stored in the host device 11 and recognizable by the host device 11.

[0019] In addition, the host device 11 and the client device 13 may each include motion sensors (e.g., an inertial measurement unit (IMU) including an accelerometer, gyroscope, and / or magnetometer), storage (e.g., volatile memory, non-volatile memory, etc.), and / or communicators (e.g., Wi-Fi modules, Bluetooth Low Energy (BLE) modules, Bluetooth modules, etc.). The motion sensors are used to sense the movement of the host device 11 or the client device 13 to generate motion data accordingly. Mathematical operations can be used to calculate the position and / or orientation of the host device 11 or the client device 13 based on the motion data. The storage can be used to store signals, data, and / or information, such as motion data, the aforementioned images, host-side map MH, client-side map MC, the physical shape of the client device 13 (used as a trackable object 134), and the position and / or orientation of the host device 11 or the client device 13. The host device 11 and the client device 13 can communicate with each other or other devices via the communicator (e.g., transmit signals, data, and / or information).

[0020] In the above embodiments, the host device 11 and at least one client device 13 in the multi-device system 100 must achieve map consistency to improve the user experience of user U1 and the immersion of user U1 in immersive content CI. It is worth noting that the multi-device system 100 can achieve map consistency between the host device 11 and at least one client device 13 by implementing a map alignment method 200, which will be described in detail in the following paragraphs with reference to Figure 2.

[0021] Please also refer to Figure 2, which is a flowchart illustrating a map alignment method 200 according to some embodiments of the present disclosure. In some embodiments, as shown in Figure 2, the map alignment method 200 includes multiple operations S201-S205. However, the present disclosure is not limited thereto.

[0022] In operation S201, client device 13 determines whether it can acquire a client keyframe IKFC. In some embodiments, the processor 130 of client device 13 calculates its processor utilization to determine whether client device 13 is available to acquire the client keyframe IKFC. For example, when the processor utilization is below an execution threshold (e.g., 90%), processor 130 determines that client device 13 is available to acquire the client keyframe IKFC, thereby executing operation S202. When the processor utilization exceeds the execution threshold, processor 130 determines that client device 13 is not available to acquire the client keyframe IKFC, thereby executing operation S201 again.

[0023] In operation S202, the host device 11 determines whether it can acquire a host-side keyframe IKFH and whether it can see at least one trackable object 134 on the client device 13. In some embodiments, the processor 110 of the host device 11 calculates its processor utilization to determine whether the host device 11 can acquire the host-side keyframe IKFH, which is similar to the operation of the processor 130 of the client device 13. Simultaneously, the processor 110 searches for or identifies the image of the trackable object 134 from the host-side image captured by the camera 112 to determine whether the host device 11 can see the trackable object 134. When the image of the trackable object 134 can be found or identified, the processor 110 determines that the host device 11 can see the trackable object 134. When the image of the trackable object 134 cannot be found or identified, the processor 110 determines that the host device 11 cannot see the trackable object 134.

[0024] In some embodiments of operation S202, the processor 110 determines that the host device 11 can obtain the host-side keyframe IKFH and determines that the host device 11 can see the trackable object 134, thereby executing operation S203. Furthermore, in some embodiments of operation S202, the processor 110 determines that the host device 11 cannot obtain the host-side keyframe IKFH or determines that the host device 11 cannot see the trackable object 134, thereby executing operation S201 again.

[0025] In operation S203, the host device 11 and the client device 13 respectively acquire the host-side keyframe IKFH and the client-side keyframe IKFC at a preset time point. In some embodiments of operation S203, at the preset time point, the processor 110 selects at least one image from the host-side images as the host-side keyframe IKFH using a feature extraction-based localization technique, while the processor 130 selects at least one image from the client-side images as the client-side keyframe IKFC using a feature extraction-based localization technique. As can be seen from the description of operations S201 to S203, when the host device 11 can see the trackable object 134 on the client device 13, the host device 11 and the client device 13 will acquire their respective keyframes (i.e., the host-side keyframe IKFH and the client-side keyframe IKFC) at the preset time point agreed upon by both parties.

[0026] In some embodiments, before operation S203 is executed, host device 11 and client device 13 synchronize their time. For example, host device 11 may exchange timestamps with client device 13, thereby allowing the time difference between two clock signals followed by processors 110 and 130 respectively and / or the data transmission delay between host device 11 and client device 13 to be calculated. One of host device 11 and client device 13 may adjust its clock signal based on the time difference and / or the data transmission delay, so that the two clock signals have the same phase and frequency.

[0027] Following the above explanation, processor 110 determines to acquire the host-side keyframe IKFH at a preset time point. As shown in Figure 1, host device 11 can transmit a timestamp T1 indicating the preset time point to client device 13 through processor 110. By receiving timestamp T1, processor 130 knows that it should acquire the client-side keyframe IKFC at the preset time point. In short, after time synchronization, host device 11 notifies client device 13 of the preset time point for acquiring both host-side keyframe IKFH and client-side keyframe IKFC.

[0028] In operation S204, the host device 11 generates a first client pose PSC1 based on the host keyframe IKFH. In some embodiments, the host keyframe IKFH includes an image of the trackable object 134. Accordingly, the processor 110 can perform, for example, triangulation on the host keyframe IKFH to calculate the position and / or orientation of the trackable object 134 relative to the camera 112. It should be understood that since the origin of the host map MH can be the position and / or orientation of the camera 112, the position and / or orientation of the trackable object 134 relative to the camera 112 can be used to represent the position and / or orientation of the client device 13 in the host map MH of the physical environment E1. In the embodiment of Figure 1, the processor 110 directly uses the attitude data of the trackable object 134 (i.e., the position and / or orientation of the trackable object 134 relative to the camera 112) as the first client attitude PSC1, wherein the first client attitude PSC1 can indicate the position and / or orientation of the client device 13 in the host map MH.

[0029] In some embodiments, as shown in Figure 1, while the host device 11 generates a first client pose PSC1, the client device 13 generates a second client pose PSC2 corresponding to a client keyframe IKFC using a feature extraction-based localization technique. Specifically, the processor 130 extracts multiple feature points from the client keyframe IKFC and matches these feature points with multiple map points PM (shown in Figure 4) of the client map MC to determine the position and / or orientation of the client device 13 in the client map MC of the physical environment E1. In some embodiments, the processor 130 directly uses the pose data of the client device 13 (i.e., the position and / or orientation of the client device 13 in the client map MC) as the second client pose PSC2.

[0030] In some embodiments, after operation S204, the host device 11 transmits the first client orientation PSC1 and the host keyframe IKFH to the client device 13, thereby executing operation S205. In operation S205, the client device 13, based on the first client orientation PSC1, aligns the client map MC established by the client device 13 detecting the entity environment E1 with the host map MH established by the host device 11 detecting the entity environment E1. This will be described in detail in the following paragraphs with reference to Figure 3. Figure 3 is a flowchart illustrating operation S205 according to some embodiments of this disclosure. In some embodiments, as shown in Figure 3, operation S205 includes multiple sub-operations S301-S302.

[0031] In sub-operation S301, client device 13 replaces the second client pose PSC2 corresponding to client keyframe IKFC with the first client pose PSC1. In other words, after sub-operation S301, client keyframe IKFC corresponds to the first client pose PSC1.

[0032] In sub-operation S302, client device 13 transforms multiple map points PM in client map MC according to transformation data (not shown), wherein the transformation data is used to transform second client pose PSC2 into first client pose PSC1. In some embodiments, processor 130 calculates the transformation data by performing a matrix calculation between first client pose PSC1 and second client pose PSC2. For example, processor 130 multiplies the inverse of first client pose PSC1 and second client pose PSC2 to obtain data that allows second client pose PSC2 to become first client pose PSC1 as transformation data. In some embodiments of sub-operation S302, processor 130 multiplies each map point PM in client map MC by the transformation data. After sub-operation S302, the coordinates of each map point PM in client map MC become relative to the origin of host map MH, rather than the origin of client map MC.

[0033] The map alignment method 200 disclosed herein is not limited to the embodiment in Figure 2, which will be described in detail below in conjunction with Figures 4 and 5. Figure 4 is a schematic diagram illustrating a usage scenario of a multi-device system 100 according to some embodiments of this disclosure. Figure 5 is another flowchart illustrating the map alignment method 200 according to some embodiments of this disclosure.

[0034] In some embodiments, after the host-side keyframe IKFH and client-side keyframe IKFC are acquired and the first client pose PSC1 and the second client pose PSC2 are calculated, the user U1 operating the multi-device system 100 may move in the physical environment E1. As the user U1 moves in the physical environment E1, as shown in Figure 4, the pose of the host device 11 and the field of view 411 of the camera 112 change as the host device 11 moves along arrow L11, while the pose of the client device 13 and the field of view 413 of the camera 132 change as the client device 13 moves along arrow L13. During the change in the pose of the host device 11, the processor 110 of the host device 11 can acquire a new keyframe (not shown) using a feature extraction-based localization technique. In the embodiment of Figure 4, this new keyframe may include an image of a trackable object 134, because the trackable object 134 on the client device 13 is located within the field of view 411 of the camera 112 of the host device 11.

[0035] Following the description of the above embodiments, the processor 110 can use new keyframes to update the host-side map MH of the entity environment E1. For example, since a new pose of the host device 11 corresponding to the new keyframe is used as the new origin of the host-side map MH, as shown in Figure 5, the first client pose PSC1 corresponding to the host-side keyframe IKFH of the host-side map MH is updated to a third client pose. In the embodiment of Figure 5, after the first client pose PSC1 becomes the third client pose, operation S501 is executed.

[0036] In operation S501, the host device 11 transmits the third client attitude to the client device 13. In some embodiments, operation S502 is executed when the client device 13 receives the third client attitude. In operation S502, the client device 13 updates the client map MC based on the third client attitude.

[0037] In some embodiments of operation S502, the processor 130 of the client device 13 replaces the first client pose PSC1 corresponding to the client keyframe IKFC with a third client pose. Further, the processor 130 can update transformation data or regenerate another transformation data based on the second client pose PSC2 and the third client pose, as described in sub-operation S302. For example, the processor 130 can obtain another transformation data by multiplying the third client pose by the inverse of the second client pose PSC2. Accordingly, the processor 130 can transform multiple map points PM in the client map MC using this other transformation data, where the other transformation data is used to convert the second client pose PSC2 into the third client pose. With this configuration, the client map MC established by the client device 13 will be updated and aligned with the host map MH updated by the host device 11.

[0038] As can be seen from the above embodiments disclosed herein, by generating keyframe pairs (i.e., host-side keyframe IKFH and client-side keyframe IKFC) between host device 11 and client device 13, client device 13 can continuously align the client map MC with the host map MH using the pose corresponding to the host-side keyframe IKFH (e.g., first client pose PSC1, third client pose, etc.). In this way, the multi-device system 100 can efficiently achieve map consistency among multiple devices and has the advantage of lower processing resource burden.

[0039] It should be understood that the map alignment method 200 is not limited to application to the multi-device system 100 shown in Figure 1, which will be described in detail in the following paragraphs with reference to Figure 6. Figure 6 is a schematic diagram of a multi-device system 100 according to some embodiments of the present disclosure. In some embodiments, as shown in Figure 6, the multi-device system 100 also includes another client device 15, that is, the multi-device system 100 includes a host device 11, a client device 13, and a client device 15. The client device 15 may be a tracker for an immersive system, wherein the tracker of the immersive system is used to track the movement of the user U1 in the physical environment E1, and has a configuration similar to that of the client device 13. For example, in Figure 6, the client device 15 is provided with four trackable objects 154, and the camera of the client device 15 has a field of view 615.

[0040] In the embodiment of Figure 6, client device 13 and client device 15 can execute map alignment method 200. For example, client device 13 determines whether it can see the trackable object 154 on client device 15 within its field of view 413. When client device 13 can see the trackable object 154 on client device 15, client device 13 and client device 15 can respectively obtain their keyframes (hereinafter referred to as the first keyframe and the second keyframe) at a mutually agreed time point. Then, client device 13 calculates a first pose of client device 15 corresponding to the first keyframe, as described in operation S204. Client device 15 can align its map with the client map MC of client device 13 by replacing the second pose of client device 15 corresponding to the second keyframe with the first pose, as described in operation S205.

[0041] The methods disclosed herein can exist in the form of program code. The program code can be contained in physical media, such as floppy disks, optical discs, hard disks, or any other transient or non-transitory computer-readable storage media, wherein when the program code is loaded and executed by a computer, the computer becomes an apparatus for implementing the methods. The program code can also be transmitted via some transmission medium, such as wires or cables, via optical fibers, or via any other transmission method, wherein when the program code is received, loaded, and executed by a computer, the computer becomes an apparatus for implementing the methods. When implemented on a general-purpose processor, the program code, in conjunction with the processor, provides a unique device that operates similarly to an application-specific logic circuit.

[0042] Although the present disclosure has been described above with reference to embodiments, it is not intended to limit the present disclosure. Those skilled in the art can make various modifications and refinements without departing from the spirit and scope of the present disclosure. Therefore, the scope of protection of the present disclosure shall be determined by the appended claims.

[0043] 11: Main unit 13, 15: Customer Units 100: Multi-device system 110, 130: Processor 112,132: Camera 114: Display panel 134,154: Trackable objects 200: Map Alignment Methods 411, 413, 615: Field of View CI: Immersive Content E1: Physical Environment IKFC: Client Keyframe IKFH: Host-side keyframes L11, L13: Arrows MC: Client Map MH: Host-side Map PM: Map Points PSC1: First Client Stance PSC2: Second Client Stance S201~S205, S501~S502: Operation S301~S302: sub-operation T1: Timestamp U1: User

[0044] Domestic storage information (please note in order of storage institution, date, and number) none Overseas storage information (please note in the order of storage country, institution, date, and number) none

Claims

1. A map alignment method applicable to a multi-device system, wherein the multi-device system is operated in a physical environment and includes a host device and a client device, the map alignment method comprising: using the host device to capture a plurality of host-based images and obtaining a host-based keyframe from the host-based images at a preset time point using a positioning technique; using the client device to capture a plurality of client-based images and obtaining a client-based keyframe from the client-based images at the preset time point using the positioning technique; using the host device to generate a first client pose based on the host-based keyframe; and using the client device to align a client map established by the client device detecting the physical environment with a host-based map established by the host device detecting the physical environment based on the first client pose.

2. The map alignment method as described in claim 1, wherein before the host device and the client device respectively obtain the host-side keyframe and the client-side keyframe at the preset time point, the map alignment method further includes: synchronizing time between the host device and the client device; and notifying the client device of the preset time point at which the host-side keyframe and the client-side keyframe are obtained by the host device.

3. The map alignment method as described in claim 1, wherein generating the first client pose based on the host keyframe by means of the host device comprises: calculating pose data of at least one trackable object on the client device from the host keyframe by means of the host device to generate the first client pose.

4. The map alignment method as described in claim 1, wherein aligning the client map with the host map by means of the client device according to the first client pose comprises: replacing a second client pose corresponding to the client keyframe with the first client pose by means of the client device; and transforming a plurality of map points in the client map according to a transformation data by means of the client device, wherein the transformation data is used to transform the second client pose into the first client pose.

5. The map alignment method as described in claim 1, further comprising: generating a second client pose corresponding to the client keyframe by means of the client device and the positioning technology based on feature extraction.

6. The map alignment method as described in Request 1, further comprising: determining, by means of the client device, whether the client device is able to obtain the client keyframe.

7. The map alignment method as described in claim 6 further comprises: when the client device is able to obtain the client keyframe, the host device determines whether the host device is able to obtain the host keyframe and whether it can see at least one trackable object on the client device, wherein when the host device is able to obtain the host keyframe and can see the at least one trackable object on the client device, the host device and the client device respectively obtain the host keyframe and the client keyframe at a preset time point.

8. The map alignment method as described in claim 1, wherein after the first client pose is updated to a third client pose, the map alignment method further comprises: transmitting the third client pose to the client device via the host device.

9. The map alignment method as described in claim 8, further comprising: updating the client map based on the attitude of the third client using the client device.

10. The map alignment method as described in claim 9, wherein updating the client map based on the third client pose via the client device comprises: replacing the first client pose corresponding to the client keyframe with the third client pose via the client device.

11. A multi-device system for operation in a physical environment, comprising: a host device for detecting the physical environment to establish a host-side map; and a client device for detecting the physical environment to establish a client-side map, wherein the host device is configured to capture a plurality of host-side based images and obtain a host-side keyframe from the host-side based images at a preset time point using a positioning technique, the client device is configured to capture a plurality of client-side based images and obtain a client-side keyframe from the client-side based images at the preset time point using the positioning technique, and wherein the host device is configured to generate a first client pose based on the host-side keyframe, and the client device is configured to align the client-side map with the host-side map based on the first client pose.

12. A non-transitory computer-readable storage medium having a computer program for executing a map alignment method, wherein the map alignment method is applicable to a multi-device system for operation in a physical environment and includes a host device and a client device, and the map alignment method includes: using the host device to capture a plurality of host-based images and obtain a host-based keyframe from the host-based images at a preset time point using a positioning technique; using the client device to capture a plurality of client-based images and obtain a client-based keyframe from the client-based images at the preset time point using the positioning technique; using the host device to generate a first client pose based on the host-based keyframe; and using the client device to align a client map established by the client device detecting the physical environment with a host-based map established by the host device detecting the physical environment based on the first client pose.