Method and device for presenting ar label information

By storing the sparse point cloud map, texture map and AR tag information of the scene in the spatial positioning database, and using the camera positioning information to overlay the AR tag information in the scene positioning image, the problem of insufficient correlation between the large spatial positioning system and the object detection system data is solved, and the unified management of AR tag information and the efficient utilization of the three-dimensional map by the object detection module is realized.

WO2025130086A1PCT designated stage expired Publication Date: 2025-06-26HISCENE INFORMATION TECH CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/111879
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-20
Filing Date
2024-08-13
Publication Date
2025-06-26

AI Technical Summary

Technical Problem

In the prior art, there is a lack of correlation between the data of the large-space positioning system and the target detection system, which leads to the large-space positioning system being unable to effectively manage dynamic targets in the scene, and the target detection system is unable to effectively utilize the texture map and sparse map in the large-space positioning, thereby affecting the efficiency and accuracy of target detection.

Method used

By establishing or updating the spatial positioning database, the sparse point cloud map, texture map and AR tag information of the scene are stored, and the camera position information of the current scene is obtained, and the AR tag information is superimposed and presented in the scene positioning image based on this information, so as to realize the unified management of AR tag information and the utilization of the three-dimensional map by the target detection module.

Benefits of technology

It realizes unified management of AR tag information in augmented reality scenarios, improves the utilization efficiency of the object detection module on three-dimensional maps in large spaces, can efficiently mark target training samples, and quickly and accurately present AR tag information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024111879_26062025_PF_FP_ABST
    Figure CN2024111879_26062025_PF_FP_ABST
Patent Text Reader

Abstract

The present application is aimed at providing a method and device for presenting AR label information. The method specifically comprises: establishing or updating a spatial positioning database; acquiring a scene positioning image of a current scene, and determining photographic pose information corresponding to the scene positioning image; on the basis of the scene positioning image, acquiring label position information of one or more pieces of scene AR label information; and on the basis of the photographic pose information and the label position information of the one or more pieces of scene AR label information, superimposing and presenting label content information of the one or more pieces of scene AR label information in the scene positioning image. The present application can quickly and accurately superimpose and present corresponding AR label information in the current scene.
Need to check novelty before this filing date? Find Prior Art

Description

A method and device for presenting AR tag information

[0001] This application is based on and claims priority to an application with CN application number 202311771011.5 and application date 2023.12.20. The disclosed content of the CN application is hereby incorporated into this application as a whole. Technical Field

[0002] The present application relates to the field of image processing, and in particular to a technology for presenting AR tag information. Background Art

[0003] In the existing technology, the large-scale spatial positioning system is responsible for generating sparse point cloud maps and texture maps. The sparse point cloud map is a map that describes the point features or line features in the scene. The point features and line features stored based on the sparse point cloud map can be used for positioning and target recognition; the texture map is used to describe the dense texture structure in the scene. It is a digital three-dimensional model of the scene. Based on the texture map, target editing and AR tag editing of the scene can be realized.

[0004] Summary of the Invention

[0005] One objective of the present application is to provide a method and device for presenting AR tag information.

[0006] According to one aspect of the present application, a method for presenting AR tag information is provided, wherein the method includes:

[0007] Establishing or updating a spatial positioning database, wherein the spatial positioning database includes scene records for one or more scenes, each scene record including a sparse point cloud map and a texture map of the corresponding scene, and one or more AR tag information of the scene, wherein the AR tag information includes tag content information;

[0008] Obtaining a scene positioning image of the current scene and determining camera pose information corresponding to the scene positioning image;

[0009] Acquire tag position information of one or more scene AR tag information based on the scene positioning image, wherein the one or more scene AR tag information is included in the AR tag information of the spatial positioning database;

[0010] According to the camera posture information and the tag position information of the one or more scene AR tag information, tag content information of the one or more scene AR tag information is superimposed and presented in the scene positioning image.

[0011] According to another aspect of the present application, a device for identifying AR tag information is provided, wherein the device includes:

[0012] A module for establishing or updating a spatial positioning database, wherein the spatial positioning database includes scene records of one or more scenes, each scene record includes a sparse point cloud map and a texture map of the corresponding scene, and one or more AR tag information of the scene, wherein the AR tag information includes tag content information;

[0013] Module 12 is used to obtain a scene positioning image of the current scene and determine the camera posture information corresponding to the scene positioning image;

[0014] A module 13 is configured to obtain tag position information of one or more scene AR tag information based on the scene positioning image, wherein the one or more scene AR tag information is included in the AR tag information of the spatial positioning database;

[0015] A fourth module is used to overlay and present the tag content information of the one or more scene AR tag information in the scene positioning image based on the camera posture information and the tag position information of the one or more scene AR tag information.

[0016] According to one aspect of the present application, a computer device is provided, wherein the device includes:

[0017] processor; and

[0018] A memory arranged to store computer executable instructions, which when executed cause the processor to perform the steps of any of the methods described above.

[0019] According to one aspect of the present application, a computer-readable storage medium is provided, on which a computer program / instruction is stored, characterized in that when the computer program / instruction is executed, the system performs the steps of any of the methods described above.

[0020] According to one aspect of the present application, a computer program product is provided, comprising a computer program / instruction, wherein the computer program / instruction implements the steps of any of the above methods when executed by a processor.

[0021] Compared with the existing technology, this application realizes the mutual utilization and association between map data in large space positioning and target detection module data through the spatial positioning database. On the one hand, it realizes the unified management of AR tag information in augmented reality scenes, and on the other hand, it realizes the use of three-dimensional maps in large spaces by target detection modules. It can efficiently and accurately mark target training samples, and use them to quickly and accurately overlay and present corresponding AR tag information in the current scene. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Other features, objects and advantages of the present application will become more apparent upon reading the detailed description of non-limiting embodiments made with reference to the following drawings:

[0023] FIG1 shows a flow chart of a method for presenting AR tag information according to one embodiment of the present application;

[0024] FIG2 shows a device structure diagram of a computer device according to another embodiment of the present application;

[0025] FIG3 illustrates an exemplary system that may be used to implement various embodiments described herein.

[0026] The same or similar reference numerals in the drawings represent the same or similar components. DETAILED DESCRIPTION

[0027] The present application is described in further detail below with reference to the accompanying drawings.

[0028] In a typical configuration of the present application, the terminal, the device of the service network, and the trusted party each include one or more processors (eg, a central processing unit (CPU), an input / output interface, a network interface, and a memory).

[0029] Memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash memory. Memory is an example of a computer-readable medium.

[0030] Computer-readable media include permanent and non-permanent, removable and non-removable media that can be used to store information by any method or technology. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PCM), programmable random access memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, tape disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information that can be accessed by a computing device.

[0031] The devices referred to in this application include but are not limited to user devices, network devices, or devices formed by integrating user devices and network devices through a network. The user devices include but are not limited to any mobile electronic product that can interact with a user (for example, through a touchpad), such as a smartphone, a tablet computer, smart glasses, etc. The mobile electronic product can use any operating system, such as the Android operating system, the iOS operating system, etc. Among them, the network device includes an electronic device that can automatically perform numerical calculations and information processing according to pre-set or stored instructions, and its hardware includes but is not limited to a microprocessor, an application specific integrated circuit (ASIC), a programmable logic device (PLD), a field programmable gate array (FPGA), a digital signal processor (DSP), an embedded device, etc. The network device includes but is not limited to a computer, a network host, a single network server, a set of multiple network servers, or a cloud composed of multiple servers; here, the cloud is composed of a large number of computers or network servers based on cloud computing, wherein cloud computing is a type of distributed computing, a virtual supercomputer composed of a group of loosely coupled computers. The network includes but is not limited to the Internet, a wide area network, a metropolitan area network, a local area network, a VPN network, a wireless self-organizing network (Ad Hoc network), etc. Preferably, the device may also be a program running on the user device, the network device, or a device formed by integrating the user device and the network device, the network device and the touch terminal, or the network device and the touch terminal via a network.

[0032] Of course, those skilled in the art should understand that the above-mentioned devices are only examples, and other existing or future devices that are applicable to this application should also be included in the scope of protection of this application and are included here by reference.

[0033] In the description of the present application, “plurality” means two or more, unless otherwise clearly defined.

[0034] FIG1 illustrates a method for presenting AR tag information according to one aspect of the present application, wherein the method is applied to a computer device and includes steps S101, S102, S103, and S104. In step S101, a spatial positioning database is established or updated, wherein the spatial positioning database includes scene records for one or more scenes, each scene record including a sparse point cloud map and a texture map of the corresponding scene, and one or more AR tag information of the scene, wherein the AR tag information includes tag content information; in step S102, a scene positioning image of the current scene is obtained, and the camera pose information corresponding to the scene positioning image is determined; in step S103, tag position information of one or more scene AR tag information is obtained based on the scene positioning image, wherein the one or more scene AR tag information is included in the AR tag information in the spatial positioning database; in step S104, tag content information of the one or more scene AR tag information is superimposed and presented in the scene positioning image based on the camera pose information and the tag position information of the one or more scene AR tag information. Wherein, the computer device includes but is not limited to a user device, a network device, or a collection of a user device and a network device; wherein, the user device includes but is not limited to any electronic product that can interact with a user, such as a smart phone, a tablet computer, smart glasses, a drone, a surveillance camera, etc.; the network device includes but is not limited to a computer, a network host, a single network server, a set of multiple network servers, or a cloud composed of multiple servers, such as a ground control center server, a business platform, etc., wherein the computer device can use different devices in different steps, which is not limited here. Here, the smart glasses of this application are used as an example to illustrate the following embodiments. Those skilled in the art should understand that the following embodiments are also applicable to other computer devices, etc. Generally, large-space positioning technology and target detection technology are executed independently. The main defect of this independent execution implementation scheme is that there is no association between the data of the large-space positioning system and the target detection system, resulting in the large-space positioning system being unable to manage dynamic targets in the scene, and the target detection system cannot effectively utilize the texture map in large-space positioning on the one hand, and on the other hand, it cannot utilize the sparse map in large-space positioning to achieve more efficient target detection. This solution integrates map data and target detection modules in large-scale spatial positioning to achieve interoperability between the two. On the one hand, it realizes the unified management of static and dynamic AR (augmented reality) tags in augmented reality scenes. On the other hand, it enables the target detection module to utilize three-dimensional maps in large spaces. For example, the target detection module can use three-dimensional texture maps to efficiently label target training samples and use three-dimensional sparse maps to quickly detect and identify targets.

[0035] In step S101, a spatial positioning database is established or updated. The spatial positioning database includes scene records for one or more scenes, each scene record including a sparse point cloud map and texture map for the corresponding scene, and one or more AR tag information for the scene, wherein the AR tag information includes tag content information. For example, a computer device establishes or updates a corresponding spatial positioning database to achieve data interconnection between large-scale spatial positioning map data and a target detection module. Here, the spatial positioning database can be established on the computer device or on another device for data access by the computer device, and can be updated based on map data uploaded by the computer device. The spatial positioning database is used to store and update scene records corresponding to one or more mapping scenes. The scene records include a sparse point cloud map and texture map for each mapping scene. The sparse point cloud map includes a map for describing point features or line features in the scene. The point features and line features stored in the sparse point cloud map can be used for positioning and target detection. The texture map includes a digitized three-dimensional model of the scene for describing dense texture structures in the scene. The texture map can be used to implement target editing and AR tag editing for the scene. In some cases, the scene record of each mapping scene also includes one or more AR tag information contained in the mapping scene, wherein the AR tag information includes annotation information for indicating a point, line, surface, object or area in the corresponding scene, and each AR tag information includes tag content information of the tag, for example, a picture, video, 3D model, PDF file, office document, form information, audio, hyperlink, application call information (used to execute relevant instructions of the application, such as opening the application, calling specific functions of the application - such as making a call, etc.), real-time sensor information (used to connect to the sensor and obtain sensor data of the target object), graffiti and other annotation information added to a certain point, line, surface, object or area. The spatial positioning database can be created by a computer device by importing map data corresponding to an initial mapping scene, or it can be created by subsequently superimposing and calculating other scenes and updating the spatial positioning database with relevant data of other scenes as mapping scene data.

[0036] In some embodiments, the method further includes step S105 (not shown), in which a scene image corresponding to the mapping scene is obtained, and a sparse point cloud map of the mapping scene corresponding to the scene image and a texture map of the mapping scene are obtained; and a scene record of the mapping scene is established based on the sparse point cloud map and texture map of the mapping scene. For example, the scene image of the mapping scene includes multiple images of the mapping scene, such as an image sequence or video of the mapping scene, etc. For another example, in addition to the image of the mapping scene, the scene image of the mapping scene also includes corresponding depth information (such as obtaining the corresponding image and depth information through a depth camera, etc.). The scene image of the mapping scene can be collected by a corresponding camera device (for example, a local camera or an external camera, etc.), and can also be an image sequence / video of the mapping scene obtained from other devices based on a communication connection with other devices. The computer device can perform data processing on the scene image, for example, three-dimensional reconstruction, etc., to obtain the corresponding sparse point cloud map and texture map. For example, the sparse map is constructed first, and then the texture map is constructed. The construction of the sparse map can be based on the existing incremental motion structure recovery technology (Srtucture From Motion, SFM) and the construction of the texture map can be based on the traditional multi-view stereo vision technology (Multi-view stereo, MVS). The construction of the sparse map and the texture map is only for example and not limited here. After the computer device obtains the sparse point cloud map and texture map of the mapping scene, on the one hand, it can establish a scene record of the mapping scene based on the sparse point cloud map and texture map, so as to facilitate the subsequent establishment or update of the spatial positioning database based on the scene record. On the other hand, it can manage the scene AR tag information of the mapping scene based on the sparse point cloud map and / or texture map (such as adding, modifying, deleting, etc.), and then establish a scene record of the mapping scene based on the sparse point cloud map, texture map and scene AR tag information, so as to facilitate the subsequent establishment or update of the spatial positioning database based on the scene record, for example, generating corresponding mapping AR tag information based on the current user's editing operation on the sparse point cloud map and / or texture map.In some embodiments, the method further includes step S106 (not shown), in which, based on the sparse point cloud map and / or the texture map of the mapping scene, mapping AR tag information in the mapping scene is obtained, wherein the mapping AR tag information includes tag content information; wherein, in step S105, the scene record of the mapping scene is established based on the sparse point cloud map and texture map of the mapping scene, including: establishing a scene record of the mapping scene based on the sparse point cloud map, texture map and the mapping AR tag information of the mapping scene; wherein, in step S101, a spatial positioning database is established or updated based on the scene record of the mapping scene, wherein the spatial positioning database includes scene records of one or more scenes, each scene record includes a sparse point cloud map and a texture map of the corresponding scene and one or more AR tag information of the scene, the AR tag information includes tag content information, and the mapping scene is included in the one or more scenes. In some embodiments, after obtaining the sparse point cloud map and the texture map of the mapping scene, the user (for example, the user corresponding to the mapping scene, etc.) can perform editing operations on the sparse point cloud map and / or texture map of the mapping scene, etc. to generate mapping AR tag information about the mapping scene, thereby establishing a scene record of the mapping scene based on the sparse point cloud map, texture map and mapping AR tag information of the mapping scene, and further establishing or updating the spatial positioning database based on the scene record. In other embodiments, the method further includes step S107 (not shown), in which, based on the sparse point cloud map and / or the texture map of the mapping scene, the mapping AR tag information in the mapping scene is obtained, wherein the mapping AR tag information includes tag content information; and the scene record corresponding to the mapping scene in the spatial positioning database is updated according to the mapping AR tag information. In some cases, the computer device may also first upload the sparse point cloud map and texture map of the mapping scene to the spatial positioning database, and generate mapping AR tag information about the mapping scene based on subsequent user (e.g., the user corresponding to the mapping scene or other users accessing the spatial positioning database) editing operations on the sparse point cloud map and / or texture map of the mapping scene. In other words, the mapping AR tag information of the mapping scene may include the information obtained by the corresponding user performing an editing operation on the sparse point cloud map and / or texture map before uploading the sparse point cloud map and texture map obtained for mapping to the spatial positioning database, and may also include the information obtained by the user calling the sparse point cloud map and / or texture map and performing an editing operation after uploading the sparse point cloud map and texture map obtained for mapping to the spatial positioning database.The computer device can establish a scene record of the mapping scene based on the sparse point cloud map and texture map of the mapping scene. In some cases, the process of establishing the scene record may also include the mapping AR tag information obtained before uploading to the spatial positioning database. In some cases, after the scene record is uploaded to the spatial positioning database, it can also be updated based on the mapping AR tag information obtained by subsequent users calling and editing the scene data. Accordingly, the spatial positioning database can be established or updated based on the creation of the scene record of the mapping scene, and updated as subsequent scene records are updated. Here, the texture map can be open to users for visual editing, and users can complete the editing of AR tags and target areas on the texture map; the sparse point cloud map can be used for positioning and target detection.

[0037] In some embodiments, the mapping AR tag information includes, but is not limited to: static AR tag information, wherein the static AR tag information also includes corresponding tag location information; and dynamic AR tag information, wherein the dynamic AR tag information also includes target area identification information corresponding to a dynamic target. For example, when editing AR tag information on a sparse point cloud map and / or texture map, multiple tag types may be set based on different editing object types. These multiple tag types may be determined based on a user's selection of a tag type, or may be determined based on target detection to classify the AR tag information. For example, if target detection determines that the AR tag information editing object is a static object whose spatial position and physical structure remain unchanged (e.g., a wall, the ground, a street lamp, etc.), the corresponding AR tag information is determined to be static AR tag information. For another example, if target detection determines that the AR tag information editing object is a dynamic rigid object whose spatial position may change (e.g., a chair, a pedestrian on the road, a vehicle, etc.), the corresponding AR tag information is determined to be dynamic AR tag information. Specifically, for static AR tag information, when the AR tag information is stored or updated, the tag location information of the AR tag information is usually also stored. The tag location information is used to indicate the location information of the point, line, surface, area, or object edited by the tag, such as the world coordinates of the absolute position, the image coordinates of the relative position, or the annotation position on the sparse point cloud map and / or texture map. In some cases, the static AR tag information includes the storage of text descriptions, AR models, and the 6-dimensional pose data + 1-dimensional scale data of the AR model in the map. For dynamic AR tag information, when the AR tag information is stored or updated, the target area identification information of the dynamic target edited by the AR tag information is usually also stored. The target area identification information is used to indicate the unique identifier of the dynamic target edited by the dynamic AR tag information, such as the name, serial number, or identification feature. In some cases, the dynamic AR tag information includes the storage of target IDs, text descriptions, and AR models.

[0038] In some embodiments, the method further includes step S108 (not shown), in which the sparse point cloud map of the mapping scene is subjected to mask optimization processing to obtain an optimized sparse point cloud map after optimization; wherein, in step S105, a scene record of the mapping scene is established based on the optimized sparse point cloud map and texture map of the mapping scene. For example, in order to save computing resources and improve computing efficiency, the computer device can perform mask optimization processing on the sparse point cloud map. The mask technology is a technology for selecting, filtering or hiding specific areas in an image. The optimized sparse point cloud map after mask optimization processing is a matrix with the same dimension as the original sparse point cloud map, wherein the optimized sparse point cloud map has an optimized area, which refers to the image pixels of the corresponding area that are modified or blocked so that no calculation is performed when adding labels or feature calculations in the subsequent process. After the mask processing, the scene record of the mapping scene is established based on the optimized sparse point cloud map and texture map of the mapping scene.

[0039] In step S102, a scene positioning image of the current scene is obtained, and the corresponding camera pose information of the scene positioning image is determined. For example, the computer device may capture the scene positioning image of the current scene using a camera device, or receive the scene positioning image of the current scene transmitted by another device based on a communication connection with another device. Here, the scene positioning image may be one or more images of the current scene. The following embodiments are described using a single image as an example. If the corresponding scene positioning image is multiple images, the processing process for each image is similar to the processing process for the single image, and the embodiments are also applicable to the processing process for multiple images. The computer device may perform scene positioning on the current scene based on the scene positioning image to determine the corresponding camera pose information. For example, the computer device may perform feature extraction and matching on the scene positioning image to determine the corresponding camera pose information; or, the computer device may obtain a scene feature map of the scene positioning image and use feature point matching and PNP calculation on the scene feature map to obtain the camera pose information of the scene positioning image. The camera pose information includes the camera position information and camera pose information of the camera device when the corresponding scene positioning image was captured.

[0040] In step S103, tag location information of one or more scene AR tag information is obtained based on the scene positioning image, wherein the one or more scene AR tag information is included in the AR tag information in the spatial positioning database. For example, after the computer device obtains the camera pose information of the scene positioning image, it can overlay and present some / all AR tag information in the spatial positioning database on the scene positioning image based on the camera pose information. For example, the computer device can directly access all static AR tag information in the spatial positioning database. Since static AR tag information includes tag location information, all static AR tag information can be overlaid and presented based on the tag location information and camera pose information of the static AR tag information. Alternatively, the computer device can first perform scene matching to determine scene AR tag information stored in the spatial positioning database and included in the current scene positioning image. For example, the computer device can determine static AR tag information within the current scene and target area identification information and target area location information in the scene positioning image identified by target detection. Based on the target area identification information, the computer device can obtain corresponding dynamic AR tag information, and determine the target area location information as the tag location information of the dynamic AR tag information, thereby overlaying and presenting the scene AR tag information based on the camera pose information and tag location information. Here, the scene AR tag information determined in the scene positioning image is included in the AR tag information stored in the spatial positioning database. In other words, the one or more scene AR tag information includes part / all of the AR tag information in the spatial positioning database.

[0041] In step S104, based on the camera pose information and the tag position information of the one or more scene AR tag information, the tag content information of the one or more scene AR tag information is superimposed and presented in the scene positioning image. For example, after the computer device obtains the camera pose information and the one or more scene AR tag information of the scene positioning image of the current scene, the scene AR tag information can be superimposed and presented. For example, for static AR tag information, the computer device calculates the image position of the static AR tag information in the scene positioning image based on the camera pose information and the tag position information of the static AR tag information, thereby superimposing and presenting the corresponding AR tag information at the image position. For dynamic AR tag information, the target area identification information and the target area position information in the scene positioning image identified by target detection are used to obtain the corresponding dynamic AR tag information through the target area identification information. The target area position information is determined as the tag position information of the dynamic AR tag information, and the corresponding image position is calculated based on the camera pose information and the tag position information, thereby achieving the superimposed presentation of the AR tag information.

[0042] In some embodiments, the method further includes step S109 (not shown), in which an editing operation of the user in the texture map of the scene is obtained, and one or more AR tag information of the scene is determined based on the editing operation, wherein the AR tag information includes tag content information. For example, after obtaining the texture map of the mapping scene, the computer device can perform an editing operation on the AR tag information based on the texture map. The editing operation can be before the texture map is stored in the spatial positioning database, or after the texture map is stored in the spatial positioning database for the user to call and perform the editing operation of the corresponding AR tag. The editing operation includes but is not limited to selecting a point, line, surface, object or three-dimensional area in the texture map and adding, modifying, deleting a mark / annotation, etc. The computer device can generate AR tag information corresponding to the editing operation based on the point, line, surface, object or three-dimensional area selected by the editing operation and the mark / annotation content added, modified, or deleted.

[0043] In some embodiments, the AR tag information includes static AR tag information; wherein the method further comprises step S110 (not shown), in which, based on the edit location information of the editing operation in the texture map of the scene, tag location information of the static AR tag information is determined and stored; wherein, based on the scene positioning image, in step S103, one or more corresponding scene AR tag information is determined from the one or more static AR tag information in the spatial positioning database, and the tag location information of the one or more scene AR tag information is queried. For example, after a computer device obtains a user's editing operation in a three-dimensional texture map, if the tag information corresponding to the editing operation is static AR tag information, the computer device may directly determine the tag location information of the static AR tag information based on the location of the point, line, surface, or three-dimensional region corresponding to the editing operation in the three-dimensional texture map, such as directly determining the location information of the static AR tag based on the location in the three-dimensional texture map, or further calculating the corresponding spatial location based on the location in the three-dimensional texture map to determine the location information of the static AR tag. The editing position information is used to indicate the indicated position of the point, line, surface, or three-dimensional area corresponding to the editing operation in the three-dimensional texture map, and can be the positions of all points of the point, line, surface, or three-dimensional area corresponding to the editing operation, or the positions of some points (e.g., endpoints and / or center points, etc.) of the point, line, surface, or three-dimensional area corresponding to the editing operation. After the computer device obtains the scene positioning image, it can directly call one or more static AR tag information in the spatial positioning database, such as directly determining all static AR tag information in the spatial positioning data as scene AR tag information, or, for example, determining the static AR tag information contained in the current scene positioning image as the corresponding scene AR tag information through position matching.

[0044] In some embodiments, the AR tag information includes dynamic AR tag information; wherein, the method further includes step S111 (not shown), in which, according to the editing operation, an area is edited in the texture map of the scene, and target area identification information of the target area is determined; wherein, the step S103 includes sub-step S1031 (not shown) and sub-step S1032 (not shown), in which, in step S1031, target detection is performed based on the scene positioning image to determine at least one target area identification information in the scene positioning image, and position information of the at least one target area identification information; in step S1032, at least one corresponding dynamic AR tag information is determined based on the at least one target area identification information, the at least one dynamic AR tag information is determined as the corresponding scene AR tag information, and the position information of the at least one target area identification information is determined as the tag position information of the at least one scene AR tag information, wherein the at least one scene AR tag information is included in the AR tag information of the spatial positioning database. For example, a computer device presents a visual texture map of a scene. A user can select an area of ​​a specific dynamic object in the texture map and annotate the area with area identification information (thereby completing the determination of the target area and the target area identification information of the target area). Furthermore, the AR tag information of the target area can be edited to generate corresponding dynamic AR tag information, wherein the dynamic AR tag information includes target area identification information for indicating the target area of ​​the dynamic object, such as directly using the name of the dynamic object as the target area identification information or obtaining the corresponding serial number as the target area identification information of the target area. After the computer device determines the target area identification information of the target area, it stores the target area identification information and the corresponding dynamic AR tag information in a spatial positioning database for subsequent target detection and determination of the target area identification information on the scene positioning image, and then matching the dynamic AR tag. After the computer device obtains the scene positioning image, it can perform target detection on the scene positioning image, identify at least one target area identification information contained in the scene positioning image, and simultaneously detect and determine the location information of the at least one identified target area identification information, such as the image location information / spatial location information in the scene positioning image. The computer device can query and match at least one target area identification information in the spatial positioning database, determine the dynamic AR tag information that matches the at least one target area identification information, determine the at least one dynamic AR tag information as the corresponding scene AR tag information, and determine the location information of the at least one target area identification information as the tag location information of the at least one scene AR tag information.Specifically, the spatial positioning database stores each dynamic AR tag information and the corresponding target area identification information. If the target area identification information determined by target detection matches the target area identification information stored in the spatial positioning database, the dynamic AR tag information corresponding to the target area identification information stored in the spatial positioning database is determined as the matching dynamic AR tag information, and is used as the scene AR tag information of the scene positioning image, etc.

[0045] In some embodiments, the method further includes step S112 (not shown), in which sparse point cloud feature information of the target area of ​​the editing operation in the sparse point cloud map of the scene is determined based on the target area, and a corresponding target detection database is established or updated based on the sparse point cloud feature information of the target area; wherein, in step S1031, the scene sparse point cloud feature information corresponding to the scene positioning image is extracted, and the scene sparse point cloud feature information is matched with the target area sparse point cloud feature information in the target detection database to determine at least one target area identification information contained in the scene positioning image, and the location information of the at least one target area identification information. For example, in order to speed up the detection process, the target detection of the target area can be performed based on a specific target detection database, which can be another database independent of the spatial positioning database, or a part / sub-data of the spatial positioning database, etc. For example, there are corresponding sparse point cloud maps and texture maps for corresponding scenes. When the user edits the target area in the texture map, the target sparse area of ​​the target area in the sparse point cloud map can be determined based on the target area through the correspondence between the sparse point cloud map and the texture map, and the point features, line features, etc. contained in the target sparse area are determined as the target area sparse point cloud feature information, etc. That is, the computer device can complete the annotation of the target area in the scene (such as the determination of the target area and the target area identification information of the target area, etc.) and the editing of the AR tag information on the texture map, and automatically obtain the target area sparse point cloud feature information of the target area on the sparse point cloud map. The computer device can store multiple target area sparse point cloud feature information corresponding to multiple editing operations in the target detection database for subsequent scene positioning images to perform feature matching, etc., wherein each target area sparse point cloud feature information is mapped to the corresponding target area identification information. For example, after the computer device obtains the scene positioning image, it extracts the scene sparse point cloud feature information of the scene positioning image, and performs similarity matching on the scene sparse point cloud feature information with the target area sparse point cloud feature information stored in the target detection database, for example, through nearest neighbor, deep learning, etc., thereby determining at least one target area identification information contained in the scene positioning image, and determining the at least one target area identification information as the target area identification information contained in the scene positioning image, and determining the identification position information of the target area identification information in the scene positioning image as the position information corresponding to the target area identification information, etc.

[0046] In some embodiments, the method further includes step S113 (not shown), in which, based on the target area, regional sample information of the edit operation in the texture map of the scene is determined, and a corresponding target detection network model is trained based on the regional sample information. In step S1031, the scene localization image is input into the target detection network model to determine at least one target area identification information contained in the scene localization image, as well as the location information of the at least one target area identification information. For example, target detection and recognition through feature matching is fast, but has poor detection robustness, and due to the variability of the posture or background environment of dynamic targets, the corresponding recognition rate is low. Here, the computer device can also achieve accurate target detection and recognition by obtaining regional sample information of the target area and training the corresponding target detection network model. For example, after the target area is marked in the three-dimensional texture map, the three-dimensional target area can be quickly and batch-mapped onto the image through an image projection process (such as perspective projection), generating an image set containing the target area as training samples for target detection. In some embodiments, the image set containing the target area can also be stored in a target detection database for training the target detection network model. The target detection network model may be included in the target detection database, or may be independent of the target detection database, etc. The target detection network model includes a deep learning model for inputting an image and outputting an object category and position coordinates, such as YOLO (You Only Look Once), Faster R-CNN, SSD (Single Shot MultiBox Detector) model, etc. Compared with the traditional method of obtaining an atlas containing the target area by photographing the target area (such as the area where the target object is located) with a camera and marking the target area of ​​each image on the atlas, the training sample generation scheme in the present invention is fast and efficient. After the computer device obtains the scene positioning image, the scene positioning image is input into the target detection network model, and the target area identification information contained in the model output result is determined as at least one target area identification information contained in the scene positioning image, etc., and the recognition position information contained in the output result is determined as the position information corresponding to the target area identification information, etc.

[0047] In some embodiments, the method further includes step S114 (not shown), in which sparse point cloud feature information of the target area of ​​the editing operation in the sparse point cloud map of the scene is determined based on the target area, and a corresponding target detection database is established or updated based on the sparse point cloud feature information of the target area; regional sample information of the editing operation in the texture map of the scene is determined based on the target area, and a corresponding target detection network model is trained based on the regional sample information; wherein, in step S1031, the scene sparse point cloud feature information corresponding to the scene positioning image is extracted, and the scene sparse point cloud feature information is matched with the target area sparse point cloud feature information in the target detection database. If the match is successful, at least one target area identification information contained in the scene positioning image and the identification position information of the at least one target area identification information are determined; if the match fails, the scene positioning image is input into the target detection network model to determine at least one target area identification information contained in the scene positioning image and the identification position information of the at least one target area identification information. For example, on the one hand, the advantage of the target detection method based on feature point clouds is that it can be created and used immediately, and the recognition speed is fast. However, the disadvantage is that it has poor robustness. When the target posture or background environment changes, the recognition rate of this method will decrease. Due to the poor robustness of feature point clouds, the target detection module needs to use feature point clouds for recognition on the one hand, and deep learning models for recognition on the other hand. Therefore, the computer device can first perform target detection on the scene positioning image through feature point matching. For example, the target detection task can be quickly completed when the target position and scene do not occur through sparse point cloud features. If the position or scene changes significantly, the feature matching fails, and the target detection based on the deep learning model with high robustness is continued. The target area identification information contained in the model output result is determined as at least one target area identification information contained in the scene positioning image, and the recognition position information contained in the output result is determined as the position information corresponding to the target area identification information.

[0048] In some embodiments, in step S102, a scene positioning image of the current scene is obtained, and the camera pose information corresponding to the scene positioning image is determined using feature point matching and a PNP algorithm based on the scene positioning image and the scene sparse point cloud map of one or more scene records in the spatial positioning database. Feature point matching is an image matching method that uses points with certain local special properties extracted from the image (called feature points) as conjugate entities, and attribute parameters of the feature points, i.e., feature descriptions, as matching entities, to achieve conjugate entity registration by calculating a similarity measure. The PNP algorithm is a method for solving the correspondence between three-dimensional space and two-dimensional points, and generally solves the camera pose, etc., given the coordinates of a 3D point and the coordinates of the corresponding 2D point and an internal parameter matrix. The computer device first extracts features from the scene positioning image, then matches it with the scene sparse point cloud map in the spatial positioning database, and then obtains the camera pose information, etc., through PNP calculation.

[0049] In some embodiments, in step S104, based on the camera pose information and the tag position information of the one or more scene AR tag information, tag image position information of the one or more scene AR tag information is determined, so that the tag image position information of the one or more scene AR tag information is superimposed on and presented in the scene positioning image with the tag content information of the one or more scene AR tag information. For example, after obtaining the tag position information of the scene AR tag information, the computer device may perform coordinate conversion based on the camera pose information and the tag position information to determine the tag image position information of the AR tag information in the image coordinate system of the scene positioning image, thereby superimposing and presenting the tag content of the AR tag position at the corresponding position in the scene positioning image.

[0050] The above mainly introduces the various embodiments of a method for presenting AR tag information in one aspect of the present application. In addition, the present application also provides specific devices that can implement the above embodiments, which are introduced below in conjunction with Figure 2.

[0051] FIG2 shows a device for presenting AR tag information according to one aspect of the present application, wherein the device includes a first module 101, a second module 102, a third module 103, and a fourth module 104. Module 101 is configured to establish or update a spatial positioning database, wherein the spatial positioning database includes scene records for one or more scenes, each scene record including a sparse point cloud map and a texture map of the corresponding scene, and one or more AR tag information for the scene, wherein the AR tag information includes tag content information; module 102 is configured to obtain a scene positioning image of the current scene and determine the camera pose information corresponding to the scene positioning image; module 103 is configured to obtain tag position information of one or more scene AR tag information based on the scene positioning image, wherein the one or more scene AR tag information is included in the AR tag information in the spatial positioning database; and module 104 is configured to overlay tag content information of the one or more scene AR tag information in the scene positioning image based on the camera pose information and the tag position information of the one or more scene AR tag information. The computer device includes, but is not limited to, a user device, a network device, or a combination of a user device and a network device; the user device includes, but is not limited to, any electronic product capable of human-computer interaction with a user, such as a smartphone, tablet computer, smart glasses, drones, surveillance cameras, etc.; the network device includes, but is not limited to, a computer, a network host, a single network server, a collection of multiple network servers, or a cloud consisting of multiple servers, such as a ground control center server, a business platform, etc. The following embodiments are described using smart glasses as an example, and those skilled in the art will appreciate that the following embodiments are equally applicable to other computer devices, etc.

[0052] Here, the specific implementations corresponding to module 101, module 102, module 103 and module 104 shown in Figure 2 are the same or similar to the embodiments of step S101, step S102, step S103 and step S104 shown in Figure 1 above, and are therefore not repeated here and are included herein by reference.

[0053] In some embodiments, the device further includes a module (not shown) for acquiring a scene image corresponding to a mapping scene, and acquiring a sparse point cloud map of the mapping scene and a texture map of the mapping scene corresponding to the scene image; and establishing a scene record of the mapping scene based on the sparse point cloud map and texture map of the mapping scene. In some embodiments, the device further includes a module (not shown) for obtaining mapping AR tag information in the mapping scene based on the sparse point cloud map and / or the texture map of the mapping scene, wherein the mapping AR tag information includes tag content information; wherein the establishing a scene record of the mapping scene based on the sparse point cloud map and texture map of the mapping scene includes: establishing a scene record of the mapping scene based on the sparse point cloud map, texture map and the mapping AR tag information of the mapping scene; wherein, module 101 is used to establish or update a spatial positioning database based on the scene record of the mapping scene, wherein the spatial positioning database includes scene records of one or more scenes, each scene record includes a sparse point cloud map and a texture map of the corresponding scene and one or more AR tag information of the scene, the AR tag information includes tag content information, and the mapping scene is included in the one or more scenes. In some embodiments, the device further includes a module (not shown) for obtaining mapping AR tag information in the mapping scene based on the sparse point cloud map and / or the texture map of the mapping scene, wherein the mapping AR tag information includes tag content information; and updating the scene record corresponding to the mapping scene in the spatial positioning database according to the mapping AR tag information.

[0054] In some embodiments, the mapping AR tag information includes but is not limited to: static AR tag information, wherein the static AR tag information also includes corresponding tag location information; dynamic AR tag information, wherein the dynamic AR tag information also includes target identification information of corresponding dynamic targets.

[0055] In some embodiments, the device further includes an eight-module (not shown) for performing mask optimization processing on the sparse point cloud map of the mapping scene to obtain an optimized sparse point cloud map; wherein, establishing a scene record of the mapping scene based on the sparse point cloud map and texture map of the mapping scene includes: establishing a scene record of the mapping scene based on the optimized sparse point cloud map and texture map of the mapping scene.

[0056] In some embodiments, the device further includes a module (not shown) for obtaining a user's editing operations in the texture map of the scene, and determining one or more AR tag information of the scene based on the editing operations, wherein the AR tag information includes tag content information.

[0057] In some embodiments, the AR tag information includes static AR tag information; wherein, the device further includes a module (not shown) for determining and storing tag location information of the static AR tag information based on the editing location information of the editing operation in the texture map of the scene; wherein, the module 103 is used to determine, based on the scene positioning image, one or more corresponding scene AR tag information in one or more static AR tag information in the spatial positioning database, and query the tag location information of the one or more scene AR tag information.

[0058] In some embodiments, the AR tag information includes dynamic AR tag information; wherein, the device further includes an eleventh module (not shown) for editing an area in the texture map of the scene according to the editing operation, and determining the target area identification information of the target area; wherein, the three-module 103 includes a three-one unit (not shown) and a three-two unit (not shown); the three-one unit is used to perform target detection based on the scene positioning image, determine at least one target area identification information in the scene positioning image, and the position information of the at least one target area identification information; the three-two unit is used to query and determine the corresponding at least one dynamic AR tag information based on the at least one target area identification information, determine the at least one dynamic AR tag information as the corresponding scene AR tag information, and determine the position information of the at least one target area identification information as the tag position information of the at least one scene AR tag information, wherein the at least one scene AR tag information is included in the AR tag information of the spatial positioning database.

[0059] In some embodiments, the device also includes a twelve-module (not shown), which is used to determine the target area sparse point cloud feature information corresponding to the target area in the sparse point cloud map of the scene according to the target area determined by the editing operation, and establish or update the corresponding target detection database based on the target area sparse point cloud feature information; wherein, the one three one unit is used to extract the scene sparse point cloud feature information corresponding to the scene positioning image, and match the scene sparse point cloud feature information with the target area sparse point cloud feature information in the target detection database to determine at least one target area identification information contained in the scene positioning image, and the location information of the at least one target area identification information.

[0060] In some embodiments, the device also includes a thirteenth module (not shown), which is used to determine the regional sample information of the target area in the texture map of the scene according to the target area determined by the editing operation, and train the corresponding target detection network model based on the regional sample information; wherein, the one-three-one unit is used to input the scene positioning image into the target detection network model, determine at least one target area identification information contained in the scene positioning image, and the location information of the at least one target area identification information.

[0061] In some embodiments, the device also includes a fourteenth unit (not shown), which is used to determine the sparse point cloud feature information of the target area of ​​the editing operation in the sparse point cloud map of the scene based on the target area, and establish or update the corresponding target detection database based on the sparse point cloud feature information of the target area; determine the regional sample information of the editing operation in the texture map of the scene based on the target area, and train the corresponding target detection network model based on the regional sample information; wherein, the one three one unit is used to extract the scene sparse point cloud feature information corresponding to the scene positioning image, and match the scene sparse point cloud feature information with the target area sparse point cloud feature information in the target detection database. If the match is successful, at least one target area identification information contained in the scene positioning image and the identification position information of the at least one target area identification information are determined; if the match fails, the scene positioning image is input into the target detection network model to determine the at least one target area identification information contained in the scene positioning image and the identification position information of the at least one target area identification information.

[0062] In some embodiments, module 102 is used to obtain a scene positioning image of the current scene, and determine the camera posture information corresponding to the scene positioning image by using feature point matching and PNP algorithm based on the scene positioning image and the scene sparse point cloud map recorded in one or more scenes in the spatial positioning database.

[0063] In some embodiments, module 104 is used to determine the label image position information of the one or more scene AR label information based on the camera posture information and the label position information of the one or more scene AR label information, so as to superimpose the label image position information of the one or more scene AR label information on the scene positioning image to present the label content information of the one or more scene AR label information.

[0064] Here, the specific implementations corresponding to the modules 15 to 14 are the same as or similar to the embodiments of the aforementioned steps S105 to S114, and thus are not described in detail and are included herein by reference.

[0065] In addition to the methods and devices described in the above embodiments, the present application also provides a computer-readable storage medium, which stores computer code. When the computer code is executed, the method described in any of the above items is executed.

[0066] The present application also provides a computer program product. When the computer program product is executed by a computer device, the method described in any one of the preceding items is executed.

[0067] The present application also provides a computer device, comprising:

[0068] one or more processors;

[0069] a memory for storing one or more computer programs;

[0070] When the one or more computer programs are executed by the one or more processors, the one or more processors are caused to implement the method as described in any one of the preceding items.

[0071] FIG3 illustrates an exemplary system that may be used to implement various embodiments described herein;

[0072] As shown in FIG3 , in some embodiments, system 300 can function as any of the aforementioned devices in the various embodiments. In some embodiments, system 300 may include one or more computer-readable media (e.g., system memory or non-volatile memory (NVM) / storage device 320) having instructions and one or more processors (e.g., processor(s) 305) coupled to the one or more computer-readable media and configured to execute the instructions to implement the modules and thereby perform the actions described herein.

[0073] For one embodiment, system control module 310 may include any suitable interface controller to provide any suitable interface to at least one of processor(s) 305 and / or any suitable device or component in communication with system control module 310 .

[0074] The system control module 310 may include a memory controller module 330 to provide an interface to the system memory 315. The memory controller module 330 may be a hardware module, a software module, and / or a firmware module.

[0075] System memory 315 can be used, for example, to load and store data and / or instructions for system 300. For one embodiment, system memory 315 can include any suitable volatile memory, such as a suitable DRAM. In some embodiments, system memory 315 can include Double Data Rate 4 SDRAM (DDR4 SDRAM).

[0076] For one embodiment, system control module 310 may include one or more input / output (I / O) controllers to provide interfaces to NVM / storage device 320 and communication interface(s) 325 .

[0077] For example, NVM / storage 320 may be used to store data and / or instructions. NVM / storage 320 may include any suitable non-volatile memory (e.g., flash memory) and / or may include any suitable non-volatile storage device(s) (e.g., one or more hard disk drives (HDDs), one or more compact disc (CD) drives, and / or one or more digital versatile disc (DVD) drives).

[0078] NVM / storage device 320 may include storage resources that are physically part of the device on which system 300 is installed, or it may be accessible to the device without being part of the device. For example, NVM / storage device 320 may be accessed over a network via communication interface(s) 325.

[0079] Communication interface(s) 325 may provide an interface for system 300 to communicate over one or more networks and / or with any other suitable devices. System 300 may wirelessly communicate with one or more components of a wireless network in accordance with any of one or more wireless network standards and / or protocols.

[0080] For one embodiment, at least one of the processor(s) 305 may be packaged together with the logic of one or more controllers of the system control module 310 (e.g., the memory controller module 330). For one embodiment, at least one of the processor(s) 305 may be packaged together with the logic of one or more controllers of the system control module 310 to form a system in a package (SiP). For one embodiment, at least one of the processor(s) 305 may be integrated on the same die with the logic of one or more controllers of the system control module 310. For one embodiment, at least one of the processor(s) 305 may be integrated on the same die with the logic of one or more controllers of the system control module 310 to form a system on chip (SoC).

[0081] In various embodiments, system 300 may be, but is not limited to, a server, a workstation, a desktop computing device, or a mobile computing device (e.g., a laptop computing device, a handheld computing device, a tablet computer, a netbook, etc.). In various embodiments, system 300 may have more or fewer components and / or a different architecture. For example, in some embodiments, system 300 includes one or more cameras, a keyboard, a liquid crystal display (LCD) screen (including a touchscreen display), a non-volatile memory port, multiple antennas, a graphics chip, an application-specific integrated circuit (ASIC), and a speaker.

[0082] It should be noted that the application can be implemented in software and / or a combination of software and hardware, for example, can be implemented using an application specific integrated circuit (ASIC), a general purpose computer or any other similar hardware device. In one embodiment, the software program of the application can be executed by a processor to realize the steps or functions described above. Similarly, the software program of the application (including relevant data structures) can be stored in a computer-readable recording medium, for example, a RAM memory, a magnetic or optical drive or a floppy disk and similar devices. In addition, some steps or functions of the application can be implemented using hardware, for example, as a circuit that cooperates with a processor to perform each step or function.

[0083] In addition, a part of the present application may be applied as a computer program product, such as a computer program instruction, which, when executed by a computer, can call or provide the method and / or technical solution according to the present application through the operation of the computer. Those skilled in the art should understand that the form in which the computer program instruction exists in a computer-readable medium includes but is not limited to a source file, an executable file, an installation package file, etc. Accordingly, the way in which the computer program instruction is executed by the computer includes but is not limited to: the computer directly executes the instruction, or the computer compiles the instruction and then executes the corresponding compiled program, or the computer reads and executes the instruction, or the computer reads and installs the instruction and then executes the corresponding installed program. Here, the computer-readable medium can be any available computer-readable storage medium or communication medium that can be accessed by the computer.

[0084] Communication media include media by which communication signals containing, for example, computer-readable instructions, data structures, program modules, or other data are transmitted from one system to another. Communication media may include guided transmission media such as cables and wires (e.g., fiber optic, coaxial, etc.) and wireless (unguided transmission) media that can propagate energy waves, such as acoustic, electromagnetic, radio frequency (RF), microwave, and infrared. Computer-readable instructions, data structures, program modules, or other data may be embodied as, for example, a modulated data signal in a wireless medium such as a carrier wave or similar mechanism such as that embodied as part of spread spectrum technology. The term "modulated data signal" refers to a signal that has one or more characteristics changed or set in such a manner as to encode information in the signal. Modulation may be analog, digital, or a hybrid modulation technique.

[0085] By way of example and not limitation, computer-readable storage media may include volatile and non-volatile, removable and non-removable media implemented in any method or technology for storage of information such as computer-readable instructions, data structures, program modules or other data. For example, computer-readable storage media include, but are not limited to, volatile memory, such as random access memory (RAM, DRAM, SRAM); and non-volatile memory, such as flash memory, various read-only memories (ROM, PROM, EPROM, EEPROM), magnetic and ferromagnetic / ferroelectric memories (MRAM, FeRAM); and magnetic and optical storage devices (hard disks, magnetic tapes, CDs, DVDs); or other media now known or later developed that can store computer-readable information / data for use by a computer system.

[0086] Here, according to one embodiment of the present application, a device is included, which includes a memory for storing computer program instructions and a processor for executing the program instructions, wherein, when the computer program instructions are executed by the processor, the device is triggered to run the methods and / or technical solutions based on the aforementioned multiple embodiments of the present application.

[0087] It is obvious to those skilled in the art that the present application is not limited to the details of the above-mentioned exemplary embodiments, and that the present application can be implemented in other specific forms without departing from the spirit or basic characteristics of the present application. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-restrictive, and the scope of the present application is defined by the appended claims rather than the above description, and it is intended that all changes that fall within the meaning and scope of the equivalent elements of the claims are included in the present application. Any figure mark in the claims should not be regarded as limiting the claims involved. In addition, it is obvious that the word "comprising" does not exclude other units or steps, and the singular does not exclude the plural. Multiple units or devices stated in the device claim can also be implemented by one unit or device through software or hardware. Words such as first and second are used to indicate names and do not indicate any particular order.

Claims

1. A method for presenting AR tag information, wherein: The method includes: Establishing or updating a spatial positioning database, wherein the spatial positioning database includes scene records of one or more scenes, each scene record includes a sparse point cloud map and a texture map of the corresponding scene and one or more AR tag information of the scene, and the AR tag information includes tag content information; Acquire a scene positioning image of the current scene, and determine the camera position information corresponding to the scene positioning image; Acquire tag position information of one or more scene AR tag information based on the scene positioning image, wherein the one or more scene AR tag information is included in the AR tag information of the spatial positioning database; According to the camera posture information and the tag position information of the one or more scene AR tag information, tag content information of the one or more scene AR tag information is superimposed and presented in the scene positioning image.

2. The method according to claim 1, wherein: The method further comprises: Acquire a scene image corresponding to a mapping scene, and acquire a sparse point cloud map of the mapping scene corresponding to the scene image and a texture map of the mapping scene; A scene record of the mapping scene is established based on the sparse point cloud map and the texture map of the mapping scene.

3. The method according to claim 2, wherein: The method further comprises: Based on the sparse point cloud map and / or the texture map of the mapping scene, obtaining mapping AR tag information in the mapping scene, wherein the mapping AR tag information includes tag content information; Wherein, the step of establishing a scene record of the mapping scene based on the sparse point cloud map and the texture map of the mapping scene includes: Establishing a scene record of the mapping scene based on the sparse point cloud map, the texture map and the mapping AR tag information of the mapping scene; Wherein, the establishing or updating of a spatial positioning database includes scene records of one or more scenes, each scene record includes a sparse point cloud map and a texture map of the corresponding scene and one or more AR tag information of the scene, and the AR tag information includes tag content information, including: A spatial positioning database is established or updated based on the scene record of the mapping scene, wherein the spatial positioning database includes scene records of one or more scenes, each scene record includes a sparse point cloud map and a texture map of the corresponding scene, and one or more AR tag information of the scene, the AR tag information includes tag content information, and the mapping scene is included in the one or more scenes.

4. The method according to claim 2, wherein: The method further comprises: Based on the sparse point cloud map and / or the texture map of the mapping scene, obtaining mapping AR tag information in the mapping scene, wherein the mapping AR tag information includes tag content information; The scene record corresponding to the mapping scene in the spatial positioning database is updated according to the mapping AR tag information.

5. The method according to claim 3 or 4, wherein: The mapping AR tag information includes at least one of the following: Static AR tag information, wherein the static AR tag information also includes corresponding tag location information; Dynamic AR tag information, wherein the dynamic AR tag information also includes target area identification information corresponding to the dynamic target.

6. The method according to claim 2, wherein: The method further comprises: Performing Mask optimization processing on the sparse point cloud map of the mapping scene to obtain an optimized sparse point cloud map; Wherein, the step of establishing a scene record of the mapping scene based on the sparse point cloud map and the texture map of the mapping scene includes: A scene record of the mapping scene is established based on the optimized sparse point cloud map and the texture map of the mapping scene.

7. The method according to claim 1, wherein: The method further comprises: An editing operation of a user in the texture map of a scene is obtained, and one or more AR tag information of the scene is determined based on the editing operation, where the AR tag information includes tag content information.

8. The method according to claim 7, wherein: The AR tag information includes static AR tag information; wherein the method further includes: Determining and storing tag position information of the static AR tag information according to the editing position information of the editing operation in the texture map of the scene; The step of acquiring tag location information of one or more scene AR tag information based on the scene positioning image includes: According to the scene positioning image, one or more corresponding scene AR tag information is determined from one or more static AR tag information in the spatial positioning database, and tag position information of the one or more scene AR tag information is queried.

9. The method according to claim 7, wherein: The AR tag information includes dynamic AR tag information; wherein the method further includes: Editing an area in a texture map of the scene according to the editing operation, and determining target area identification information of the target area; The step of acquiring tag location information of one or more scene AR tag information based on the scene positioning image includes: Performing target detection according to the scene positioning image, determining at least one target area identification information in the scene positioning image, and position information of the at least one target area identification information; Determine at least one corresponding dynamic AR tag information based on the at least one target area identification information query, determine the at least one dynamic AR tag information as the corresponding scene AR tag information, and determine the location information of the at least one target area identification information as the tag location information of the at least one scene AR tag information.

10. The method according to claim 9, wherein: The method further comprises: Determine the editing operation according to the target area, determine the target area sparse point cloud feature information corresponding to the target area in the sparse point cloud map of the scene, and establish or update the corresponding target detection database according to the target area sparse point cloud feature information; The performing of target detection according to the scene positioning image to determine at least one target area identification information in the scene positioning image and the position information of the at least one target area identification information includes: Extract scene sparse point cloud feature information corresponding to the scene positioning image, and match the scene sparse point cloud feature information with the target area sparse point cloud feature information in the target detection database to determine at least one target area identification information contained in the scene positioning image and the position information of the at least one target area identification information.

11. The method according to claim 9, wherein: The method further comprises: Determine the target area in the texture map of the scene according to the target area for the editing operation regional sample information, and training a corresponding target detection network model according to the regional sample information; The performing of target detection according to the scene positioning image to determine at least one target area identification information in the scene positioning image and the position information of the at least one target area identification information includes: The scene localization image is input into the target detection network model to determine at least one target area identification information contained in the scene localization image and the position information of the at least one target area identification information.

12. The method according to claim 9, wherein: The method further comprises: Determine, according to the target area, sparse point cloud feature information of the target area of ​​the editing operation in the sparse point cloud map of the scene, and establish or update a corresponding target detection database according to the sparse point cloud feature information of the target area; Determine, according to the target area, regional sample information of the editing operation in the texture map of the scene, and train a corresponding target detection network model according to the regional sample information; The performing of target detection according to the scene positioning image to determine at least one target area identification information in the scene positioning image and the position information of the at least one target area identification information includes: Extracting scene sparse point cloud feature information corresponding to the scene positioning image, and matching the scene sparse point cloud feature information with the target area sparse point cloud feature information in the target detection database, and if the match is successful, determining at least one target area identification information contained in the scene positioning image, and identification position information of the at least one target area identification information; If the matching fails, the scene localization image is input into the target detection network model to determine at least one target area identification information contained in the scene localization image and the identification position information of the at least one target area identification information.

13. The method according to claim 1, wherein: The step of acquiring a scene positioning image of the current scene and determining the camera position information corresponding to the scene positioning image includes: Obtain a scene positioning image of the current scene, and determine the camera posture information corresponding to the scene positioning image by using feature point matching and PNP algorithm based on the scene positioning image and the scene sparse point cloud map recorded in one or more scenes in the spatial positioning database.

14. The method according to claim 1, wherein: The step of superimposing and presenting the tag content information of the one or more scene AR tag information in the scene positioning image according to the camera posture information and the tag position information of the one or more scene AR tag information includes: According to the camera posture information and the tag position information of the one or more scene AR tag information, the tag image position information of the one or more scene AR tag information is determined, so that the tag image position information of the one or more scene AR tag information is superimposed on the tag content information of the one or more scene AR tag information in the scene positioning image.

15. A device for presenting AR tag information, wherein: The equipment includes: A module for establishing or updating a spatial positioning database, wherein the spatial positioning database includes scene records of one or more scenes, each scene record includes a sparse point cloud map and a texture map of the corresponding scene and one or more AR tag information of the scene, and the AR tag information includes tag content information; Module one and module two are used to obtain a scene positioning image of the current scene and determine the camera position information corresponding to the scene positioning image; A module is used to obtain tag position information of one or more scene AR tag information based on the scene positioning image, wherein the one or more scene AR tag information is included in the AR tag of the spatial positioning database. information; A module 4 is used to overlay and present the tag content information of the one or more scene AR tag information in the scene positioning image according to the camera posture information and the tag position information of the one or more scene AR tag information.

16. A computer device, wherein: The equipment includes: Processor; and A memory arranged to store computer executable instructions which, when executed, cause the processor to perform the steps of the method as claimed in any one of claims 1 to 14.

17. A computer-readable storage medium having a computer program / instruction stored thereon, characterized in that: The computer program / instructions, when executed, cause the system to perform the steps of the method as claimed in any one of claims 1 to 12.

18. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instructions are executed by a processor, the steps of the method according to any one of claims 1 to 14 are implemented.

Citation Information

Patent Citations

  • Air-ground integrated city ecological civilization managing system and method based on Beidou positioning

    CN104021586A

  • Augmented reality data display method and device, electronic equipment and storage medium

    CN113345108A

  • Method for determining and presenting target mark information and equipment

    CN113741698A

  • Method and device for presenting AR label information

    CN117745988A

  • Method and device for presenting AR virtual information

    CN118250447A