AR glasses collaborative system combining large-space visual positioning
By combining a large-space visual positioning AR glasses collaborative system, the problems of virtual content fusion error and multi-user collaborative interaction on the optical see-through AR glasses end have been solved, realizing a high-precision AR experience and multi-user collaborative display.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- HANGZHOU YIXIAN XIANJIN TECH CO LTD
- Filing Date
- 2023-02-14
- Publication Date
- 2026-05-26
AI Technical Summary
The lack of mature AR glasses systems based on visual large-space positioning technology in the current technology leads to poor user experience, especially in optical see-through AR glasses, where the fusion error between virtual content and the real environment is large, and the latency and positioning error of human-to-human collaborative interaction are serious.
An AR glasses collaborative system combining large-space visual positioning is adopted, including AR glasses, a visual positioning server, and a coordinate synchronization server. Through real-time image acquisition and matching point cloud maps, relative pose is calculated to achieve high-precision overlay of virtual content and multi-user collaborative display.
It achieves high-precision integration of virtual content with the real environment, enhancing the user's AR experience in large-space scenarios, especially in optical perspective AR glasses, enabling collaborative interaction and content sharing among multiple users.
Smart Images

Figure CN116228862B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of augmented reality, and in particular to an AR glasses collaborative system that combines large spatial visual positioning. Background Technology
[0002] Vision-based large-space positioning technology uses images captured by a camera to perform real-time spatial positioning of a device. By matching the images with a positioning map, the real-time pose (position and orientation) of the camera is returned.
[0003] In related technologies, vision-based large-space positioning technology has been widely applied in terminal devices such as mobile phones, such as Apple's geolocation anchoring and Google's VPS (Visual Positioning System).
[0004] However, in AR glasses, especially optical see-through AR glasses, there are few systems based on visual large-space positioning technology. Furthermore, there are no mature and reliable technical solutions in this field for human-environment collaboration and multi-person collaboration based on visual large-space positioning technology. Summary of the Invention
[0005] This application provides an AR glasses collaborative system, method, electronic device, and storage medium that combines large-space visual positioning, in order to at least solve the problem of poor user experience caused by the lack of large-space AR glasses system solutions in related technologies.
[0006] In a first aspect, embodiments of this application provide an AR glasses collaborative system that combines large-space visual positioning, applied in large scene spaces. The system includes: AR glasses, a visual positioning server, and a coordinate synchronization server.
[0007] The AR glasses are used to acquire real-time images of the target scene and track and obtain the first pose, and send them to the visual positioning server;
[0008] The visual positioning server is used to obtain and return the second pose of the AR glasses in the map coordinate system by matching the real-time image in the point cloud map, and to calculate the first relative pose of the AR glasses' local coordinate system relative to the map coordinate system based on the first pose and the second pose.
[0009] The coordinate system synchronization server is used to obtain the second relative pose between different AR glasses based on the second pose of multiple AR glasses, and return it to the corresponding AR glasses;
[0010] The AR glasses are also used to locally display AR content in the target scene based on the first relative pose.
[0011] Furthermore, the AR content is collaboratively displayed in the target scene based on the second relative pose.
[0012] In some embodiments, the AR glasses include a tracking camera for acquiring a first pose in real time using SALM technology, the first pose being the position and orientation of the AR glasses in the local coordinate system.
[0013] In some embodiments, the point cloud map is a large-scale 3D point cloud map with real-scale information of the target scene, and the second pose is the position and posture of the AR glasses in the coordinates of the point cloud map.
[0014] In some embodiments, the visual positioning server includes: an image judgment module, a pose judgment module, and a visual positioning module, wherein,
[0015] The image judgment module is used to obtain the motion blur state of the real-time image, determine the optimal image from multiple real-time images based on the motion blur state, and input the optimal image to the visual positioning module.
[0016] The pose determination module is used to identify the current viewing angle of the AR glasses based on the first pose, and determine the target image corresponding to the best viewing angle as the optimal image based on preset rules, and input the optimal image into the visual positioning module.
[0017] The visual positioning module is used to match the global features of the optimal image in the point cloud map using a visual positioning algorithm to obtain similar map frames.
[0018] Furthermore, a 2D-3D observation is established between the local features of the optimal image and the local features of the similar map frame, and the second pose of the AR glasses in the point cloud map coordinate system is obtained based on the 2D-3D observation.
[0019] In some embodiments, the visual positioning server calculates a first relative pose of the local coordinate system relative to the point cloud map based on the first pose and the second pose using the following formula:
[0020] T m_w =T m_cam *T world_cam ·inverse
[0021] Among them, T m_w It is the first relative pose, T m_camIt is the second pose in the point cloud map coordinates, T world_cam It is the first pose in the local coordinate system obtained by tracking the camera.
[0022] In some embodiments, the AR glasses include a user-environment interaction module and a user-user interaction module, wherein,
[0023] The user-environment interaction module is used to, when the AR glasses' current first relative pose triggers a preset condition, obtain a corresponding preset virtual effect from the point cloud map based on the first relative pose.
[0024] And, instruct the AR glasses to overlay the virtual effects with the real image in the current field of view of the AR glasses to generate the AR content, and display the AR content at the user's eye observation position;
[0025] The user-user interaction module is used to, after displaying the AR content at the user's eye viewing position:
[0026] When a sharing request is received from the second AR glasses, the second relative pose between the AR glasses and the coordinate system synchronization server is obtained, and the AR content is shared to the second AR glasses according to the second relative pose.
[0027] In some embodiments, the AR glasses further include an AR interaction module, wherein the AR interaction module is used for:
[0028] The system receives user input commands and edits virtual effects in the AR content according to these commands. When the AR content is displayed locally, the input commands are input locally on the device.
[0029] When the AR content is displayed in a collaborative manner, the operation instructions can be input locally on the device or by the second AR glasses.
[0030] Secondly, embodiments of this application provide a collaborative method for AR glasses that combines large-space visual positioning, the method comprising:
[0031] The AR glasses capture real-time images of the target scene and track the first pose to obtain the first pose, which is then sent to the visual positioning server.
[0032] The real-time image is matched in the point cloud map through the visual positioning server to obtain the second pose of the AR glasses in the map coordinate system and return it. Based on the first pose and the second pose, the first relative pose of the AR glasses in the local coordinate system relative to the map coordinate system is calculated.
[0033] By using a coordinate system synchronization server, the second relative pose between different AR glasses is obtained based on the second pose of multiple AR glasses, and then returned to the corresponding AR glasses;
[0034] The AR glasses display AR content locally in the target scene according to the first relative pose, and collaboratively display the AR content in the target scene according to the second relative pose.
[0035] Thirdly, embodiments of this application provide a computer device including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the method described in the second aspect above.
[0036] Fourthly, embodiments of this application provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method described in the second aspect above.
[0037] Compared to related technologies, this application provides an AR glasses collaboration method combining large-space visual positioning. The method involves AR glasses acquiring real-time images of a target scene and tracking them to obtain a first pose, which is then sent to a visual positioning server. The visual positioning server matches the real-time images in a point cloud map to obtain the second pose of the AR glasses in the map coordinate system and returns it. Based on the first and second poses, a first relative pose of the AR glasses' local coordinate system relative to the map coordinate system is calculated. A coordinate system synchronization server obtains the second relative poses between different AR glasses based on the second poses of multiple AR glasses and returns them to the corresponding AR glasses. The AR glasses display AR content locally in the target scene based on the first relative pose and collaboratively display AR content in the target scene based on the second relative poses. This solves the problem of poor user experience caused by the lack of large-space AR glasses system solutions in traditional technologies, achieving a virtual overlay of large-space AR experiences and collaborative AR experiences among multiple users. Attached Figure Description
[0038] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:
[0039] Figure 1 This is a schematic diagram of an application environment for an AR glasses collaborative system that combines visual large spatial positioning according to an embodiment of this application.
[0040] Figure 2 This is a structural block diagram of an AR glasses collaborative system combining visual large spatial positioning according to an embodiment of this application;
[0041] Figure 3 This is a schematic diagram of a coordinate transformation relationship according to an embodiment of this application;
[0042] Figure 4 This is a flowchart of an AR glasses collaborative method combining large spatial visual positioning according to an embodiment of this application;
[0043] Figure 5 This is a schematic diagram of the internal structure of an electronic device according to an embodiment of this application. Detailed Implementation
[0044] To make the objectives, technical solutions, and advantages of this application clearer, the application is described and illustrated below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application. All other embodiments obtained by those skilled in the art based on the embodiments provided in this application without inventive effort are within the scope of protection of this application.
[0045] Obviously, the accompanying drawings described below are merely some examples or embodiments of this application. Those skilled in the art can apply this application to other similar scenarios based on these drawings without any inventive effort. Furthermore, it is understood that although the efforts made in this development process may be complex and lengthy, for those skilled in the art related to the content disclosed in this application, any changes to design, manufacturing, or production based on the technical content disclosed in this application are merely conventional technical means and should not be construed as insufficient disclosure of the content of this application.
[0046] In this application, the reference to "embodiment" means that a specific feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment that is mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described in this application may be combined with other embodiments without conflict.
[0047] Unless otherwise defined, the technical or scientific terms used in this application shall have the ordinary meaning understood by one of ordinary skill in the art to which this application pertains. The terms “a,” “an,” “an,” “the,” and similar words used in this application do not indicate quantity limitation and may indicate singular or plural. The terms “comprising,” “including,” “having,” and any variations thereof used in this application are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or device that includes a series of steps or modules (units) is not limited to the listed steps or units, but may also include steps or units not listed, or may include other steps or units inherent to these processes, methods, products, or devices. The terms “connected,” “linked,” “coupled,” and similar words used in this application are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. “Multiple” used in this application refers to two or more. “And / or” describes the relationship between related objects, indicating that three relationships may exist; for example, “A and / or B” can represent: A alone, A and B simultaneously, and B alone. The character " / " generally indicates that the preceding and following objects are in an "or" relationship. The terms "first," "second," and "third" used in this application are merely to distinguish similar objects and do not represent a specific ordering of the objects.
[0048] In this document, it should be understood that the terms used may be technical means used to implement part of the present invention or other summary technical terms. For example, the terms may include:
[0049] AR (Argument Reality): Augmented Reality is a technology that cleverly integrates virtual information with the real world. It simulates and applies computer-generated text, images, 3D models, videos, and other virtual information to the real world. The two types of information complement each other, thus "enhancing" the real world.
[0050] Pose: Position and orientation (facing), for example, in 2D it is generally (x, y, yaw), and in 3D it is generally (x, y, z, yaw, pitch, roll), including 6 degrees of freedom, 6 of (6 Degrees Of Freedom). The last three elements describe the object's orientation, where yaw is the heading angle, rotating around the Z-axis, pitch is the pitch angle, rotating around the Y-axis, and roll is the roll angle, rotating around the X-axis.
[0051] Large-space visual positioning system (VisualPositionSystem): Uses real-time images for positioning in a visual map coordinate system in large scenes such as scenic spots and parks, with a position uncertainty of generally 0.2m.
[0052] The AR glasses collaborative system combining visual large-space positioning provided in this application can be applied to, for example... Figure 1 In the application environment shown, Figure 1 This is a schematic diagram of an application environment for an AR glasses collaborative system combining visual large-space positioning according to an embodiment of this application, such as... Figure 1 As shown, users operate AR glasses 10 to experience AR in offline scenarios, such as scenic spots and parks. Furthermore, the AR glasses can acquire and track poses in real time, and obtain the pose of the AR glasses relative to the point cloud map coordinate system through a visual positioning server. Based on this pose, corresponding virtual overlaid AR content can be obtained for interactive experiences. The AR glasses can obtain relative poses with other devices through a coordinate transformation server, enabling AR content sharing and collaborative operations among multiple users. Through this application, AR experiences involving people and their environment, as well as collaborative AR experiences among multiple users, are realized in large-space scenarios using optically perceptible AR glasses.
[0053] Figure 2 This is a structural block diagram of an AR glasses collaborative system combining visual large-space positioning according to an embodiment of this application, such as... Figure 2 As shown, the process includes: AR glasses 20, visual positioning server 21, and coordinate synchronization server 22;
[0054] AR glasses 20 are used to collect real-time images of the target scene and track and obtain the first pose, and send it to the visual positioning server 21;
[0055] In this embodiment, the target scene can be a large-space offline scene such as a scenic spot or park, and the AR glasses mentioned above are optical see-through glasses; the AR glasses obtain the first pose by running a SLAM tracking system. Specifically, in this embodiment, any common SLAM scheme can be applied, as long as the real-time camera pose can be obtained through such a scheme.
[0056] The visual positioning server 21 is used to obtain the second pose of the AR glasses in the map coordinate system by matching the real-time image in the point cloud map and return it to the AR glasses 20, and to calculate the first relative pose of the local coordinate system relative to the map coordinate system based on the first pose and the second pose.
[0057] Furthermore, the visual positioning server 21 pre-stores a point cloud map of the target scene, and obtains the pose of the AR glasses 20 in the map coordinate system by matching the image information uploaded by the AR glasses in the map.
[0058] The coordinate system synchronization server 22 is used to obtain the second relative pose between different AR glasses based on the second pose of multiple AR glasses 20, and return it to the corresponding AR glasses;
[0059] The coordinate synchronization server and visual positioning server can be deployed on a public cloud server or in hardware in the offline scene where the AR glasses are located. The visual positioning server and coordinate synchronization server are respectively connected to the AR glasses network and realize data transmission and reception through the network connection channel.
[0060] AR glasses are also used to locally display specific AR content associated with the real scene in a target scene based on a first relative pose, and to collaboratively display specific AR content in a target scene based on a second relative pose.
[0061] Through visual positioning servers and coordinate transformation servers, AR glasses can know their pose relative to the real physical environment in their local coordinate system, as well as their pose relative to the local coordinate systems of other AR glasses. Furthermore, based on the above two pose data, users can experience AR content superimposed on the real environment within the current camera's field of view, share the same AR content with other users, and collaborate with other users to edit AR content. This enables human-environment interaction and multi-user interaction in large-space scenarios.
[0062] One specific implementation method of this system is as follows:
[0063] 1. Perform 3D reconstruction on a large-space physical scene to generate a point cloud map of the scene. Specifically, the 3D reconstruction method can be any image-based reconstruction method. However, it should be noted that in this embodiment, the mapping method must be able to restore the true scale information of the physical environment, that is, 1 meter in the map corresponds to 1 meter in the real physical environment.
[0064] 2. Store the point cloud map on a visual large-space positioning server, which can be deployed on a cloud platform or in hardware form in the target scene.
[0065] 3. While capturing real-time footage, the AR glasses acquire their pose T relative to the local SLAM coordinate system via a SLAM tracking system. world_cam At the same time, the pose and camera images are sent to the aforementioned positioning server at preset intervals.
[0066] 4. Upon receiving a location request from the AR glasses, the location server uses a visual positioning algorithm to match the real-time image sent by the AR glasses against a point cloud map, thereby obtaining the pose information T of the image in the point cloud map. m_camAnd return it to the AR glasses;
[0067] 5. Combination T world_cam and T m_cam Two sets of 6DoF poses can be used to determine the relative pose of the SLAM local coordinate system with respect to the point cloud map coordinate system, i.e., T. m_w =T m_cam *T world_cam • inverse;
[0068] 6. Coordinate system synchronization server: This server maintains the 6DoF poses of different AR glasses relative to the same point cloud map coordinate system. At the same time, it calculates the relative 6DoF poses between any two AR glasses and sends broadcast information to all AR glasses.
[0069] Figure 3 This is a schematic diagram of a coordinate transformation relationship according to an embodiment of this application. For example... Figure 3 As shown, after the above steps, all AR glasses can know their pose T in their local coordinate system relative to the point cloud map. m_wi (Wi represents the i-th local coordinate system), and we can also know the pose T of its local coordinate system relative to the local coordinate systems of other AR glasses. wj_wi (wj represents the local coordinate system of the j-th eyepiece).
[0070] 7. Based on the pose data obtained from the above process, multi-faceted collaborative interaction is achieved, specifically including a user-environment interaction module and a user-user interaction module, wherein:
[0071] (1) User-Environment Interaction Module: The AR glasses can identify the pose relative to the map coordinate system by recognizing the point cloud map, and then know the pose of the virtual content placed in the point cloud map. Thus, when the wearer of the glasses arrives at the physical environment corresponding to the point cloud map, through successful visual recognition, they can observe the virtual content bound to the physical environment and obtain an immersive experience of virtual and real superposition.
[0072] (2) User-to-user interaction module: When different glasses wearers are in the same physical space at the same physical moment, they can learn about each other's relative poses by recognizing the same point cloud map, so that they can share the interaction process and AR effect of other glasses wearers in real time, and multiple people can cooperate to operate the same AR content in the environment.
[0073] For example, different users can experience the AR effect from different directions of real-world objects in the environment. For instance, an AR effect of a "phoenix" flying out of a pavilion in the middle of a lake can be displayed from different angles based on their real-time poses. Users can also perform interactive operations such as rotating and moving the "phoenix".
[0074] The AR glasses collaborative system provided in this embodiment designs a 6DoF visual large-space positioning system for optical see-through AR glasses. On the optical see-through AR glasses based on the 6DoF visual large-space positioning system, a system-level solution for human-environment collaboration and human-human interaction collaboration is realized, filling a gap in the field and greatly improving the user experience when using AR devices.
[0075] It should be noted that on optical see-through AR glasses, users directly observe the real environment through the lenses, so virtual content will be superimposed on the real environment observed by the user; while on smartphones, users experience AR by observing the display screen, which superimposes virtual content on a video stream to generate AR effects.
[0076] It's understandable that on AR glasses, because the human eye uses the real environment as a reference, users can easily perceive the error in the blending of virtual and real elements. However, the method of overlaying AR effects onto video streams on mobile devices masks a significant portion of this fusion error. These two AR interaction solutions are fundamentally different; conventional smartphone-based AR interaction solutions cannot be directly applied to optical see-through AR glasses.
[0077] The inventors in this case, through comparison and summarization, discovered that:
[0078] For interactive experience solutions on optical perspective AR glasses, during human-environment collaboration, errors arise from the local SLAM tracking errors of each AR device, the positioning errors of each device in the map (errors in the second pose), and optical display calibration errors. During human-to-human interactive collaboration, the local SLAM tracking errors and the positioning errors in the map will be further amplified as the number of people increases. At the same time, human-to-human interaction involves the transmission of relative poses over the network, and the latency of pose data will also affect the spatiotemporal errors of multi-person interactive AR content.
[0079] This embodiment aims to solve the above-mentioned problems:
[0080] For optical perspective AR devices, a high-precision display calibration method is adopted to achieve high-precision virtual-real anchoring. Furthermore, an image judgment module and a pose judgment module are introduced to determine the motion blur state of the image and select the best viewing angle, thereby minimizing the positioning error in the map. Furthermore, in human-human collaboration scenarios with high requirements for spatiotemporal accuracy, pose prediction and interpolation methods are introduced to reduce errors caused by factors such as network latency.
[0081] In some embodiments, the AR glasses include a tracking camera used to acquire a first pose in real time via SLAM technology. The first pose is the position and orientation of the AR glasses in the local coordinate system. The point cloud map is a large-scale 3D point cloud map of the target scene with real-scale information. The second pose is the position and orientation of the AR glasses in the point cloud map coordinate system.
[0082] In some embodiments, the visual positioning server includes: an image judgment module, a pose judgment module, and a visual positioning module, wherein,
[0083] The image judgment module is used to acquire the motion blur state of the real-time image and determine the optimal image based on the motion blur state. The optimal image is then input into the visual positioning module. Through this module, the motion blur state of the positioning image can be judged, thereby selecting high-quality images for positioning and obtaining more reliable positioning results.
[0084] The pose determination module is used to identify the current viewing angle of the AR glasses based on the first pose, determine the target image corresponding to the best viewing angle, and input the target image into the visual positioning module. Through this module, the current viewing angle of the camera can be determined, and then the best positioning angle can be selected for visual positioning to improve the success rate of visual positioning.
[0085] The visual positioning module is used to match the global features of the target image in the point cloud map to obtain similar map frames through a visual positioning algorithm, and to establish 2D-3D observation between the local features of the target image and the local features of the similar map frames, and to obtain the second pose of the AR glasses in the point cloud map coordinate system based on the 2D-3D observation.
[0086] It should be noted that this application does not specify which positioning algorithm is used in the visual positioning module. It should be understood that the application of this algorithm can achieve efficient and accurate visual positioning in large space scenarios.
[0087] In some embodiments, the visual positioning server calculates a first relative pose of the local coordinate system relative to the point cloud map based on the first pose and the second pose using the following formula:
[0088] T m_w =T m_cam *T world_cam ·inverse
[0089] Among them, T m_w It is the first relative pose, T m_cam It is the second pose in point cloud map coordinates, T world_cam It is the first pose in the local coordinate system obtained by tracking the camera.
[0090] In some embodiments, the AR glasses include a user-environment interaction module and a user-user interaction module, wherein,
[0091] The user-environment interaction module is used to obtain the corresponding preset virtual effects in the point cloud map based on the first relative pose of the AR glasses when the preset conditions are triggered. It also instructs the AR glasses to overlay the virtual effects with the real images in the current field of view of the AR glasses to generate AR content, and displays the AR content at the user's eye observation position. Specifically, when the user walks to a specific location in an offline scene, he / she can observe the virtual content bound to the physical environment through successful visual recognition and obtain an immersive AR experience of virtual and real overlay.
[0092] The user-user interaction module is used to, after the user-environment interaction module instructs the display of AR content at the observation position of the user's glasses: if a second AR glasses is detected in the current scene, request the coordinate system synchronization server to obtain the second relative pose between itself and the second AR glasses, and share the AR content with the second AR glasses according to the second relative pose.
[0093] AR glasses also include an AR interaction module, which is used to: receive operation instructions from each user after sharing AR content to the second AR glasses, and edit the AR content of this device and the AR content shared from other AR glasses according to the operation instructions. In other words, the operation instructions can be input locally on the device or input by the second AR glasses.
[0094] Through the various modules mentioned above in the AR glasses, when different glasses wearers are in the same physical space at the same physical moment, they can know each other's relative poses by recognizing the same point cloud map. Thus, they can share the interaction process and AR effects of other glasses wearers in real time, and multiple people can also cooperate to operate the same AR content in the environment. In the above-mentioned application scenario of "Lake Pavilion", four users A, B, C and D, located in the four directions of east, west, south and north around the lake pavilion, can experience AR effects that interact with the lake pavilion from different angles. Furthermore, each user can not only edit the AR effects themselves, but also see the AR effects and added edits of other users.
[0095] This application also provides a collaborative method for AR glasses that combines large-space visual positioning. Figure 4 This is a flowchart of an AR glasses collaborative method combining large-space visual positioning according to an embodiment of this application, such as... Figure 4 As shown, the process includes the following steps:
[0096] S401 acquires real-time images and tracks the target scene using AR glasses to obtain the first pose and sends it to the visual positioning server;
[0097] S402, through the visual positioning server, by matching the real-time image in the point cloud map, the second pose of the AR glasses is obtained and returned to the AR glasses, and the first relative pose of the local coordinate system relative to the point cloud map is calculated based on the first pose and the second pose.
[0098] S403: Through the coordinate system synchronization server, the second relative pose between different AR glasses is obtained based on the second pose of multiple AR glasses, and then returned to the corresponding AR glasses;
[0099] S404, the AR glasses locally display specific AR content associated with the real scene in the target scene according to the first relative pose, and collaboratively display the specific AR content in the target scene according to the second relative pose.
[0100] Through the above steps S301 to S304, the problem of poor user experience caused by the lack of large-space AR glasses system solutions in traditional technologies is solved, realizing a virtual overlay large-space AR experience, as well as a collaborative AR experience between multiple users.
[0101] In one embodiment, Figure 5 This is a schematic diagram of the internal structure of an electronic device according to an embodiment of this application, such as... Figure 5 As shown, an electronic device is provided, which can be a server, and its internal structure diagram can be as follows. Figure 5 As shown, the electronic device includes a processor, a network interface, internal memory, and non-volatile memory connected via an internal bus. The non-volatile memory stores an operating system, computer programs, and a database. The processor provides computing and control capabilities, the network interface communicates with external terminals via a network connection, the internal memory provides an environment for the operation of the operating system and computer programs, the computer programs are executed by the processor to implement an AR glasses collaborative method combining large-space visual positioning, and the database stores data.
[0102] Those skilled in the art will understand that Figure 5 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the electronic device to which the present application is applied. The specific electronic device may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements.
[0103] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. This computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and RAMbus dynamic RAM (RDRAM), etc.
[0104] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0105] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this patent application should be determined by the appended claims.
Claims
1. An AR glasses collaborative system combining large-space visual positioning, characterized in that, For applications in large-scale environments, the system includes: AR glasses, a visual positioning server, and a coordinate synchronization server; The AR glasses are used to acquire real-time images of the target scene and track and obtain the first pose, and send them to the visual positioning server; The visual positioning server is used to obtain and return the second pose of the AR glasses in the map coordinate system by matching the real-time image in the point cloud map of the target scene, and to calculate the first relative pose of the AR glasses' local coordinate system relative to the map coordinate system based on the first pose and the second pose. The coordinate system synchronization server is used to obtain the second relative pose between different AR glasses based on the second pose of multiple AR glasses, and return it to the corresponding AR glasses. The second relative position is the pose of any AR glasses in the local coordinate system relative to the local coordinate system of other AR glasses. The AR glasses are also used to locally display specific AR content associated with the real scene in the target scene, based on the first relative pose. Furthermore, based on the second relative pose, AR content shared with other users is collaboratively displayed in the target scene, and AR content is collaboratively edited with other users.
2. The system according to claim 1, characterized in that, The AR glasses include a tracking camera, which is used to acquire a first pose in real time through the SALM system. The first pose is the position and orientation of the AR glasses in the local coordinate system.
3. The system according to claim 1, characterized in that, The point cloud map is a large-scale 3D point cloud map with real-scale information in the target scene, and the second pose is the position and posture of the AR glasses in the coordinates of the point cloud map.
4. The system according to claim 1, characterized in that, The visual positioning server includes: an image judgment module, a pose judgment module, and a visual positioning module, wherein... The image judgment module is used to obtain the motion blur state of the real-time image, determine the optimal image from multiple real-time images based on the motion blur state, and input the optimal image to the visual positioning module. The pose determination module is used to identify the current viewing angle of the AR glasses based on the first pose, and determine the target image corresponding to the best viewing angle as the optimal image based on preset rules, and input the optimal image into the visual positioning module. The visual positioning module is used to match the global features of the optimal image in the point cloud map using a visual positioning algorithm to obtain similar map frames. Furthermore, a 2D model is established between the local features of the optimal image and the local features of the similar map frames. 3D observation, based on the aforementioned 2D 3D observation yields the second pose of the AR glasses in the point cloud map coordinate system.
5. The system according to any one of claims 1 to 4, characterized in that, The visual positioning server calculates the first relative pose of the local coordinate system relative to the point cloud map based on the first pose and the second pose using the following formula: Tm_w = Tm_cam * Tworld_cam * inverse, where Tm_w is the first relative pose, Tm_cam is the second pose in the point cloud map coordinates, and Tworld_cam is the first pose in the local coordinate system obtained by the tracking camera.
6. The system according to claim 1, characterized in that, The AR glasses include the user Environment interaction module and user The user interaction module, in which... The user The environmental interaction module is used to, when the AR glasses trigger a preset condition in the current first relative pose, obtain the corresponding preset virtual effect in the point cloud map according to the first relative pose, and instruct the AR glasses to overlay the virtual effect with the real image in the current field of view of the AR glasses to generate the AR content, and display the AR content at the user's eye observation position. The user-user interaction module is used to, after displaying the AR content at the user's eye viewing position: When a sharing request is received from the second AR glasses, the second relative pose between the AR glasses and the coordinate system synchronization server is obtained, and the AR content is shared to the second AR glasses according to the second relative pose.
7. The system according to claim 6, characterized in that, The AR glasses also include an AR interaction module, which is used to: receive user operation commands and edit virtual effects in the AR content according to the operation commands; wherein, when the AR content is displayed locally, the operation commands are input locally by the device. When the AR content is displayed in a collaborative manner, the operation instructions can be input locally on the device or by the second AR glasses.
8. A collaborative method for AR glasses combining large-space visual positioning, characterized in that, The method includes: The AR glasses capture real-time images of the target scene and track the first pose to obtain the first pose, which is then sent to the visual positioning server. The real-time image is matched with the point cloud map of the target scene through the visual positioning server to obtain the second pose of the AR glasses in the map coordinate system and return it. Based on the first pose and the second pose, the first relative pose of the AR glasses' local coordinate system relative to the map coordinate system is calculated. By using a coordinate system synchronization server, the second relative pose between different AR glasses is obtained based on the second pose of multiple AR glasses, and then returned to the corresponding AR glasses. Here, the second relative position is the pose of any AR glass in the local coordinate system relative to the local coordinate system of other AR glasses. The AR glasses, based on the first relative pose, locally display specific AR content associated with the real scene in the target scene, and based on the second relative pose, collaboratively display AR content shared with other users in the target scene, and collaboratively edit the AR content with other users.
9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, The processor executes the computer program to implement the method as described in claim 8.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the method as described in claim 8.