Augmented reality positioning method and apparatus, electronic device, and storage medium

CN118196198BActive Publication Date: 2026-09-22MIGU COMIC CO LTD +2
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410396761.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-04-02
Publication Date
2026-09-22
Estimated Expiration
2044-04-02

AI Technical Summary

Technical Problem

[0003]本申请实施例提供一种增强现实的定位方法、装置、电子设备及存储介质,以解决现有技术中对于增强现实的定位效果较差的问题

Benefits of technology

[0044]本申请提供一种增强现实的定位方法、装置、电子设备及存储介质,该方法包括:获取场景数据,所述场景数据包括目标建筑内的多个座位所对应的多个座位号和多张场景地图;根据所述多个座位号在三维坐标系中确定多个三维坐标,所述多个三维坐标与所述多个座位号一一对应,所述三维坐标系根据所述目标建筑的三维模型生成;根据目标用户输入的目标座位号,在所述多张场景地图中确定至少一张目标场景地图以及确定所述目标座位号对应的三维坐标;基于所述至少一张目标场景地图和所述目标座位号所对应的三维坐标,在所述目标建筑内对所述目标用户输入的待定位图像进行增强现实定位,得到定位结果。本申请通过对目标建筑内的多个座位号确定多个三维坐标,并通过目标需要进行定位的座位号确定在预先拍摄的多张场景地图中确定出至少一张目标场景地图,从而根据目标场景地图和进行定位的座位号对用户输入的图像进行增强现实定位,从而提高了增强现实的定位效果。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118196198B_ABST
    Figure CN118196198B_ABST
Patent Text Reader

Abstract

The application provides an augmented reality positioning method and device, electronic equipment and storage medium, the method comprising: acquiring scene data; determining a plurality of three-dimensional coordinates in a three-dimensional coordinate system according to a plurality of seat numbers; determining at least one target scene map and determining the three-dimensional coordinates corresponding to the target seat number according to the target seat number input by the target user in a plurality of scene maps; based on at least one target scene map and the three-dimensional coordinates corresponding to the target seat number, performing augmented reality positioning on the image to be positioned input by the target user in the target building to obtain the positioning result. The application determines a plurality of three-dimensional coordinates according to a plurality of seat numbers in the target building, and determines at least one target scene map in a plurality of scene maps pre-shot by the seat number to be positioned, so as to perform augmented reality positioning on the image input by the user according to the target scene map and the seat number to be positioned, thereby improving the positioning effect of augmented reality.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer vision technology, and in particular to an augmented reality positioning method, apparatus, electronic device, and storage medium. Background Technology

[0002] With the development of Augmented Reality (AR) positioning technology, it is being applied in an increasing number of fields, such as using AR for positioning in buildings like stadiums. Current solutions, however, have high requirements for the visual features of the scene. The scene to be modeled must have rich texture features, no repeating similar areas, and the actual positioning must maintain the same lighting and environmental conditions as during modeling. For stadium-type scenes, there are generally many similar areas in the architectural structure, resulting in a high degree of similarity. These stadiums are often used for concerts or opening ceremonies, leading to poor positioning results with AR. Summary of the Invention

[0003] This application provides an augmented reality positioning method, apparatus, electronic device, and storage medium to solve the problem of poor positioning performance in the prior art.

[0004] To solve the above problems, this application is implemented as follows:

[0005] In a first aspect, embodiments of this application provide an augmented reality positioning method, the method comprising:

[0006] Acquire scene data, which includes multiple seat numbers corresponding to multiple seats within the target building and multiple scene maps;

[0007] Multiple three-dimensional coordinates are determined in a three-dimensional coordinate system based on the multiple seat numbers, and the multiple three-dimensional coordinates correspond one-to-one with the multiple seat numbers. The three-dimensional coordinate system is generated based on the three-dimensional model of the target building.

[0008] Based on the target seat number input by the target user, determine at least one target scene map from the multiple scene maps and determine the three-dimensional coordinates corresponding to the target seat number;

[0009] Based on the at least one target scene map and the three-dimensional coordinates corresponding to the target seat number, augmented reality positioning is performed on the image to be located input by the target user within the target building to obtain the positioning result.

[0010] Optionally, before determining the multiple three-dimensional coordinates in the three-dimensional coordinate system based on the multiple seat numbers, the method further includes:

[0011] The target building is modeled in three dimensions to obtain a three-dimensional model of the target building;

[0012] The target building is divided into multiple areas by dividing the multiple seats into areas, and each target area includes at least one of the seats;

[0013] Each key point corresponding to the target region is determined to obtain multiple key points, which are the vertices of the target region.

[0014] The three-dimensional coordinate system is generated based on the three-dimensional model and the multiple key points.

[0015] Optionally, determining multiple three-dimensional coordinates in a three-dimensional coordinate system based on the multiple seat numbers includes:

[0016] Based on the multiple seat numbers, determine the multiple target two-dimensional coordinates corresponding to the multiple key points;

[0017] Based on the three-dimensional coordinate system, the two-dimensional coordinates of the multiple targets are transformed to obtain the three-dimensional coordinates of the multiple targets, and the three-dimensional coordinates of the multiple targets correspond one-to-one with the multiple key points;

[0018] Based on the multiple target three-dimensional coordinates and the linear interpolation algorithm, the multiple two-dimensional coordinates corresponding to the multiple seat numbers are transformed in the three-dimensional coordinate system to obtain the multiple three-dimensional coordinates.

[0019] Optionally, determining at least one target scene map from the plurality of scene maps based on the target seat number input by the target user and determining the three-dimensional coordinates corresponding to the target seat number includes:

[0020] Determine the first three-dimensional coordinates of the target seat number in the three-dimensional coordinate system, and the second three-dimensional coordinates of the center point of the target building in the three-dimensional coordinate system;

[0021] The direction vector is determined based on the first three-dimensional coordinates and the second three-dimensional coordinates;

[0022] Based on the first three-dimensional coordinates and the direction vector, the multiple scene maps are filtered to determine at least one target scene map.

[0023] Optionally, the step of filtering the multiple scene maps based on the first three-dimensional coordinates and the direction vector to determine at least one target scene map includes:

[0024] A first video frame set is determined from the multiple scene maps. The first video frame set includes multiple first video frames, and the shooting position of the multiple first video frames is less than a preset distance from the first three-dimensional coordinate.

[0025] A second video frame set is determined based on the first video frame set. The second video frame set includes multiple second video frames, and the angle between the multiple second video frames and the direction vector is less than a preset angle.

[0026] In the second set of video frames, at least one target scene map is determined, wherein the target scene map is captured at a location within the observation range of the target seat number.

[0027] Optionally, the step of performing augmented reality localization on the target user's input image within the target building based on the at least one target scene map and the three-dimensional coordinates corresponding to the target seat number, to obtain a localization result, includes:

[0028] Feature extraction is performed on the image to be located input by the target user to obtain multiple location features;

[0029] The matching results are obtained by matching the multiple positioning features and the three-dimensional coordinates corresponding to the target seat number with the at least one target scene map;

[0030] Based on the matching result, augmented reality localization is performed on the image to be located within the target building to obtain the localization result.

[0031] Optionally, after performing augmented reality localization on the image to be located input by the target user within the target building based on the at least one target scene map and the three-dimensional coordinates corresponding to the target seat number, and obtaining the localization result, the method further includes:

[0032] If the positioning result indicates that positioning has failed, keyframe matching is performed on the at least one target scene map to obtain the target keyframe;

[0033] Perform feature matching between the image to be located and the target keyframe to obtain the feature matching result;

[0034] The transformation relationship is obtained based on the feature matching results;

[0035] Based on the transformation relationship, augmented reality localization is performed on the target image input by the target user within the target building to obtain the target localization result.

[0036] Secondly, embodiments of this application also provide an augmented reality positioning device, comprising:

[0037] The acquisition module is used to acquire scene data, which includes multiple seat numbers corresponding to multiple seats in the target building and multiple scene maps;

[0038] The first determining module is used to determine multiple three-dimensional coordinates in a three-dimensional coordinate system based on the multiple seat numbers, wherein the multiple three-dimensional coordinates correspond one-to-one with the multiple seat numbers, and the three-dimensional coordinate system is generated based on the three-dimensional model of the target building;

[0039] The second determining module is used to determine at least one target scene map from the multiple scene maps and determine the three-dimensional coordinates corresponding to the target seat number based on the target seat number input by the target user.

[0040] The positioning module is used to perform augmented reality positioning on the target user's input image within the target building based on the at least one target scene map and the three-dimensional coordinates corresponding to the target seat number, and to obtain the positioning result.

[0041] Thirdly, embodiments of this application also provide an electronic device, including: a transceiver, a memory, a processor, and a program stored in the memory and executable on the processor; the processor is configured to read the program in the memory to implement the steps in the method described in the first aspect above.

[0042] Fourthly, embodiments of this application also provide a readable storage medium for storing a program, which, when executed by a processor, implements the steps of the method described in the first aspect above.

[0043] Fifthly, embodiments of this application also provide a computer program product, which is stored in a storage medium and executed by at least one processor to implement the steps in the method as described in the first aspect.

[0044] This application provides an augmented reality positioning method, apparatus, electronic device, and storage medium. The method includes: acquiring scene data, the scene data including multiple seat numbers corresponding to multiple seats within a target building and multiple scene maps; determining multiple three-dimensional coordinates in a three-dimensional coordinate system based on the multiple seat numbers, the multiple three-dimensional coordinates corresponding one-to-one with the multiple seat numbers, the three-dimensional coordinate system being generated based on a three-dimensional model of the target building; determining at least one target scene map and the corresponding three-dimensional coordinates based on a target seat number input by a target user from the multiple scene maps; and performing augmented reality positioning on the image input by the target user within the target building based on the at least one target scene map and the corresponding three-dimensional coordinates, to obtain a positioning result. This application improves the augmented reality positioning effect by determining multiple three-dimensional coordinates for multiple seat numbers within a target building and determining at least one target scene map from multiple pre-captured scene maps based on the target seat number to be positioned, thereby performing augmented reality positioning on the user-input image based on the target scene map and the seat number to be positioned. Attached Figure Description

[0045] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments of this application will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0046] Figure 1 This is a flowchart illustrating the augmented reality positioning method provided in an embodiment of this application;

[0047] Figure 2 This is a schematic diagram of the augmented reality positioning device provided in the embodiments of this application;

[0048] Figure 3 This is a schematic diagram of the structure of the communication device provided in the embodiments of this application. Detailed Implementation

[0049] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0050] The terms "first," "second," etc., used in the embodiments of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to these processes, methods, products, or devices. Additionally, the use of "and / or" in this application indicates at least one of the connected objects, such as A and / or B and / or C, representing seven possibilities: including A alone, B alone, C alone, and the presence of both A and B, both B and C, both A and C, and the presence of A, B, and C.

[0051] See Figure 1 , Figure 1 This is a flowchart illustrating the augmented reality positioning method provided in an embodiment of this application. Figure 1 As shown, augmented reality positioning methods may include the following steps:

[0052] Step 101: Obtain scene data, which includes multiple seat numbers corresponding to multiple seats in the target building and multiple scene maps.

[0053] In this embodiment, the scene data includes the location information of multiple seats within the target building and multiple scene maps taken. The target building can be a large venue such as a stadium or concert hall. The multiple seats are for users to watch performances, and each seat has a unique seat number, which accurately determines its location within the venue. The multiple scene maps are obtained by collecting scene data using a drone or a camera equipped with an IMU and GPS module, resulting in multiple scene maps of the target building. These scene maps clearly display images of the target building taken from different directions.

[0054] Step 102: Determine multiple three-dimensional coordinates in a three-dimensional coordinate system based on the multiple seat numbers. The multiple three-dimensional coordinates correspond one-to-one with the multiple seat numbers. The three-dimensional coordinate system is generated based on the three-dimensional model of the target building.

[0055] In this embodiment, a three-dimensional coordinate system is generated based on the three-dimensional model of the target building. Specifically, a drone or a camera equipped with an IMU and GPS module is used to collect scene data, perform three-dimensional modeling, and use the drone's RTK or the camera's IMU module data to restore the scale of the three-dimensional model to the true scale. The three-dimensional coordinate system can consist of an X-axis, a Y-axis, and a Z-axis, where the X-axis is based on the length of the target building, the Y-axis on the width, and the Z-axis on the height. Multiple three-dimensional coordinates are determined in the three-dimensional coordinate system based on the positions of multiple seat numbers, and these coordinates are in the form of (X, Y, Z). Therefore, different seat numbers correspond to different three-dimensional coordinates.

[0056] Step 103: Based on the target seat number input by the target user, determine at least one target scene map from the multiple scene maps and determine the three-dimensional coordinates corresponding to the target seat number.

[0057] In this embodiment, when a target user is located, they need to input the target seat number and the captured image on their terminal. Depending on the seat number, the scene map associated with that seat number will also differ. For example, the target seat may be captured in a scene map, thus allowing the scene map containing the target seat to be determined. Additionally, the target seat number input by the user needs to be converted into three-dimensional coordinates.

[0058] Step 104: Based on the at least one target scene map and the three-dimensional coordinates corresponding to the target seat number, perform augmented reality positioning on the image to be located input by the target user within the target building to obtain the positioning result.

[0059] In this embodiment, augmented reality positioning is performed on the user-input image based on at least one determined target scene map and the three-dimensional coordinates corresponding to the target seat number, thereby generating a positioning result that indicates the location of the image taken by the user within the target building.

[0060] It's important to note that Augmented Reality (AR) positioning enhances a user's perception by overlaying computer-generated images, videos, or other data onto their real-world location. AR positioning typically relies on the camera, sensors, and software of mobile devices (such as smartphones or AR glasses) to identify the user's location and orientation in the real world and render virtual information on the screen in real time.

[0061] For example, after entering the AR stage, the user's camera is pointed at the target scene, and structured AR positioning is performed. Following the conventional structured process, feature extraction, feature matching, keyframe retrieval, map point matching, and PNP pose estimation are performed.

[0062] This application provides an augmented reality (AR) positioning method, comprising: acquiring scene data, the scene data including multiple seat numbers corresponding to multiple seats within a target building and multiple scene maps; determining multiple three-dimensional coordinates in a three-dimensional coordinate system based on the multiple seat numbers, wherein the multiple three-dimensional coordinates correspond one-to-one with the multiple seat numbers, and the three-dimensional coordinate system is generated based on a three-dimensional model of the target building; determining at least one target scene map and the corresponding three-dimensional coordinates based on a target seat number input by a target user; and performing AR positioning on the image input by the target user within the target building based on the at least one target scene map and the corresponding three-dimensional coordinates, thereby obtaining a positioning result. This application improves the AR positioning effect by determining multiple three-dimensional coordinates for multiple seat numbers within a target building and determining at least one target scene map from multiple pre-captured scene maps based on the target seat number to be positioned, thus performing AR positioning on the user-input image based on the target scene map and the seat number to be positioned.

[0063] In some feasible implementations, optionally, before determining the plurality of three-dimensional coordinates in the three-dimensional coordinate system based on the plurality of seat numbers, the method further includes:

[0064] The target building is modeled in three dimensions to obtain a three-dimensional model of the target building;

[0065] The target building is divided into multiple areas by dividing the multiple seats into areas, and each target area includes at least one of the seats;

[0066] Each key point corresponding to the target region is determined to obtain multiple key points, which are the vertices of the target region.

[0067] The three-dimensional coordinate system is generated based on the three-dimensional model and the multiple key points.

[0068] In this embodiment, the target building is first modeled in 3D. Specifically, scene data is collected using a drone or a camera equipped with an IMU and GPS module to create a 3D model. The 3D model is then reconstructed to its true scale using drone RTK or camera IMU module data. The coordinate system used in this step is the world coordinate system used for subsequent modeling and positioning. This 3D model stores information such as keyframes and map points. The keyframe information includes the camera pose when the photo was taken, the map points observable in that pose, and the keyframe feature points corresponding to those map points. This keyframe constitutes the scene map, which is used for subsequent matching with the target seat number.

[0069] Subsequently, by dividing the target building into multiple seating areas, multiple target regions were determined. Specifically, on the venue drawings, the seating areas were divided into regions according to regular shapes (such as rectangles). For each region, key points on its corners were selected (taking a rectangle as an example, the four corner points are the key points), and their two-dimensional coordinates on the drawing were extracted, denoted as {x}. Gi y Gi}

[0070] Therefore, the X-axis, Y-axis, and Z-axis of the three-dimensional coordinate system are determined by the three-dimensional model and multiple key points, thereby generating the three-dimensional coordinate system.

[0071] Optionally, determining multiple three-dimensional coordinates in a three-dimensional coordinate system based on the multiple seat numbers includes:

[0072] Based on the multiple seat numbers, determine the multiple target two-dimensional coordinates corresponding to the multiple key points;

[0073] Based on the three-dimensional coordinate system, the two-dimensional coordinates of the multiple targets are transformed to obtain the three-dimensional coordinates of the multiple targets, and the three-dimensional coordinates of the multiple targets correspond one-to-one with the multiple key points;

[0074] Based on the multiple target three-dimensional coordinates and the linear interpolation algorithm, the multiple two-dimensional coordinates corresponding to the multiple seat numbers are transformed in the three-dimensional coordinate system to obtain the multiple three-dimensional coordinates.

[0075] In this embodiment, the two-dimensional coordinates of multiple targets corresponding to the multiple key points are first determined by multiple seat numbers. Specifically, for example, {x Gi y Gi Based on the aforementioned three-dimensional coordinate system, the two-dimensional coordinates are transformed; specifically, they are transformed into a three-dimensional model, and a three-dimensional coordinate system is generated. Specifically, this is denoted by {x}. Wi y Wi , z Wi Since the world coordinate system is already restored to the actual scale of reality, height can be applied as the third dimension in the drawing coordinate system, thus converting it into three-dimensional coordinates, i.e., {x}. Gi y Gi , z Wi For the same key points, apply the sim3 transformation to calculate the transformation relationship between the two coordinate systems:

[0076] x Wi =s*R*x Gi +t

[0077] y Wi =s*R*y Gi +t

[0078] Where s represents the scale factor, R is the rotation matrix, and t is the translation vector. These three parameters can be obtained by the above sim3 transformation.

[0079] Linear interpolation is used to estimate the value of an unknown point between two known points. It assumes that the change between the two known points is linear, meaning that any point on the line segment formed by the two points can be calculated using the equation of the line.

[0080] Specifically, by exporting the two-dimensional coordinates of all seats from the venue drawings, and using the aforementioned similarity transformation relationship, these coordinates are converted to the world coordinate system of the 3D model. Since the height difference between seats on each floor is the same, the corresponding height z-values ​​can be calculated based on the differences between key points in the area, and then filled in sequentially. Using this method, even when the 3D model lacks the coordinates of all seats, the coordinate annotation of seats on a very large scale can be quickly completed using the two-dimensional coordinates of the venue drawings, along with area division and linear interpolation.

[0081] Optionally, determining at least one target scene map from the plurality of scene maps based on the target seat number input by the target user and determining the three-dimensional coordinates corresponding to the target seat number includes:

[0082] Determine the first three-dimensional coordinates of the target seat number in the three-dimensional coordinate system, and the second three-dimensional coordinates of the center point of the target building in the three-dimensional coordinate system;

[0083] The direction vector is determined based on the first three-dimensional coordinates and the second three-dimensional coordinates;

[0084] Based on the first three-dimensional coordinates and the direction vector, the multiple scene maps are filtered to determine at least one target scene map.

[0085] In this embodiment, due to the nature of sports venue performances, there is generally a stage core point, which is the center point of the target building. No matter where the audience sits, they will look towards the stage. Furthermore, the AR special effects that blend reality and virtuality are generally integrated and interact with the stage.

[0086] By determining the first three-dimensional coordinates of the target seat number and the second three-dimensional coordinates of the center point, a direction vector V for the user's seat coordinates and the stage center point coordinates is calculated based on these first and second three-dimensional coordinates. Then, by matching and filtering multiple scene maps using the values ​​of the direction vector and the first three-dimensional coordinates, the scene map that is closest to or contains the target seat number is determined.

[0087] Once the user's seat coordinates and the approximate direction of the stage are known, a customized personal map can be created for each user (or each small area of ​​users) and distributed to the client for location tracking. Because the computational load is significantly reduced, the cloud only needs to handle map cropping and distribution based on each user's location; the subsequent complete location tracking process only needs to be performed on the user's end. This effectively solves the pain points of mislocation caused by a large number of similar textures in sports venues and the inability to handle the concurrency issues caused by tens of thousands of spectators simultaneously requesting cloud location tracking.

[0088] Optionally, the step of filtering the multiple scene maps based on the first three-dimensional coordinates and the direction vector to determine at least one target scene map includes:

[0089] A first video frame set is determined from the multiple scene maps. The first video frame set includes multiple first video frames, and the shooting position of the multiple first video frames is less than a preset distance from the first three-dimensional coordinate.

[0090] A second video frame set is determined based on the first video frame set. The second video frame set includes multiple second video frames, and the angle between the multiple second video frames and the direction vector is less than a preset angle.

[0091] In the second set of video frames, at least one target scene map is determined, wherein the target scene map is captured at a location within the observation range of the target seat number.

[0092] In this embodiment, multiple scene maps are filtered multiple times. Specifically, a first video frame set is first determined from the multiple scene maps. This first video frame set includes multiple first video frames. It should be noted that the shooting position of the multiple first video frames is less than a preset distance φ from the first three-dimensional coordinates.

[0093] Then, a second video frame set is selected from the first video frame set. This second video frame set contains multiple second video frames, where the angle between these frames and the direction V (facing the current user towards the stage center) is less than a preset angle θ. It should be noted that the preset distance φ and the preset angle θ can be customized based on the scene size and the number of seats; however, they are not specifically limited in this embodiment.

[0094] In the second set of video frames, observable map points are selected and combined to form the user's personal map. Finally, the above method is used to traverse all seats, generate personal maps for all seats in turn, and store them in the cloud server.

[0095] Optionally, the step of performing augmented reality localization on the target user's input image within the target building based on the at least one target scene map and the three-dimensional coordinates corresponding to the target seat number, to obtain a localization result, includes:

[0096] Feature extraction is performed on the image to be located input by the target user to obtain multiple location features;

[0097] The matching results are obtained by matching the multiple positioning features and the three-dimensional coordinates corresponding to the target seat number with the at least one target scene map;

[0098] Based on the matching result, augmented reality localization is performed on the image to be located within the target building to obtain the localization result.

[0099] Optionally, after performing augmented reality localization on the image to be located input by the target user within the target building based on the at least one target scene map and the three-dimensional coordinates corresponding to the target seat number, and obtaining the localization result, the method further includes:

[0100] If the positioning result indicates that positioning has failed, keyframe matching is performed on the at least one target scene map to obtain the target keyframe;

[0101] Perform feature matching between the image to be located and the target keyframe to obtain the feature matching result;

[0102] The transformation relationship is obtained based on the feature matching results;

[0103] Based on the transformation relationship, augmented reality localization is performed on the target image input by the target user within the target building to obtain the target localization result.

[0104] In this embodiment, after post-processing optimization, the structured positioning method described above still has the potential for positioning failure due to the existence of corner locations, viewing angle differences, and long distances within the stadium. To ensure the overall algorithm's compatibility, a rotation positioning method based on the essential matrix is ​​executed after the above method fails to locate the position, thus compensating for positioning failures at certain locations. The specific implementation is as follows:

[0105] Keyframe matching is performed, the best matching keyframe is selected, and feature matching between the image to be located and the best keyframe is performed. Then, the transformation relationship between the two is estimated by calculating the essential matrix. The rotation part in the transformation relationship is extracted and combined with the user's seat coordinates as the translation part. The two are combined as the final camera pose to achieve a simple AR positioning with a high success rate.

[0106] It should be noted that this embodiment can also verify the positioning results. Existing structured positioning verification schemes generally estimate the accuracy of camera pose by checking the number of interior points after PNP projection (e.g., >200). However, sports stadium performances often have issues such as lighting changes and stage setups that cannot be modeled in advance, resulting in a very low number of successfully matched interior points, leading to positioning failures in traditional methods. In this proposal, by using a smaller number of interior points (e.g., >10) and the positional deviation (i.e., the difference between the final estimated camera pose and the user's seat coordinates) as post-verification criteria, the positioning success rate is effectively improved.

[0107] This application determines multiple three-dimensional coordinates by identifying multiple seat numbers within a target building, and determines at least one target scene map from multiple pre-captured scene maps by identifying the seat number to be located. Based on the target scene map and the seat number to be located, augmented reality positioning is performed on the image input by the user, thereby improving the augmented reality positioning effect.

[0108] See Figure 2 , Figure 2 This is a structural diagram of the augmented reality positioning device provided in an embodiment of this application. Figure 2 As shown, the augmented reality positioning device 200 includes:

[0109] The acquisition module 210 is used to acquire scene data, which includes multiple seat numbers corresponding to multiple seats in the target building and multiple scene maps;

[0110] The first determining module 220 is used to determine multiple three-dimensional coordinates in a three-dimensional coordinate system based on the multiple seat numbers, wherein the multiple three-dimensional coordinates correspond one-to-one with the multiple seat numbers, and the three-dimensional coordinate system is generated based on the three-dimensional model of the target building;

[0111] The second determining module 230 is used to determine at least one target scene map from the plurality of scene maps and determine the three-dimensional coordinates corresponding to the target seat number based on the target seat number input by the target user.

[0112] The positioning module 240 is used to perform augmented reality positioning on the image to be positioned input by the target user within the target building based on the at least one target scene map and the three-dimensional coordinates corresponding to the target seat number, and obtain the positioning result.

[0113] Optional, also includes:

[0114] The model building module is used to perform three-dimensional modeling of the target building to obtain a three-dimensional model of the target building;

[0115] The area division module is used to divide the multiple seats included in the target building into multiple target areas, and each target area includes at least one of the seats;

[0116] The key point determination module is used to determine the key points corresponding to each target region, thereby obtaining multiple key points, where the key points are the vertices of the target regions;

[0117] The coordinate system generation module is used to generate the three-dimensional coordinate system based on the three-dimensional model and the multiple key points.

[0118] Optionally, the first determining module 220 includes:

[0119] The first determining submodule is used to determine the two-dimensional coordinates of multiple targets corresponding to the multiple key points based on the multiple seat numbers;

[0120] The first transformation submodule is used to transform the two-dimensional coordinates of the plurality of targets based on the three-dimensional coordinate system to obtain the three-dimensional coordinates of the plurality of targets, wherein the three-dimensional coordinates of the plurality of targets correspond one-to-one with the plurality of key points;

[0121] The second transformation submodule is used to transform the multiple two-dimensional coordinates corresponding to the multiple seat numbers in the three-dimensional coordinate system based on the multiple target three-dimensional coordinates and a linear interpolation algorithm to obtain the multiple three-dimensional coordinates.

[0122] Optionally, the second determining module 230 includes:

[0123] The second determining submodule is used to determine the first three-dimensional coordinates of the target seat number in the three-dimensional coordinate system, and the second three-dimensional coordinates of the center point of the target building in the three-dimensional coordinate system;

[0124] The third determining submodule is used to determine the direction vector based on the first three-dimensional coordinates and the second three-dimensional coordinates;

[0125] The filtering submodule is used to filter the multiple scene maps based on the first three-dimensional coordinates and the direction vector to determine at least one target scene map.

[0126] Optionally, the filtering submodule includes:

[0127] The first determining unit is used to determine a first video frame set in the plurality of scene maps, the first video frame set including a plurality of first video frames, wherein the shooting position of the plurality of first video frames is less than a preset distance from the first three-dimensional coordinates;

[0128] The second determining unit is configured to determine a second video frame set based on the first video frame set, wherein the second video frame set includes multiple second video frames, and the angle between the multiple second video frames and the direction vector is less than a preset angle.

[0129] The third determining unit is used to determine the at least one target scene map in the second video frame set, wherein the shooting location of the target scene map is within the observation range of the target seat number.

[0130] Optionally, the positioning module 240 includes:

[0131] The feature extraction submodule is used to extract features from the image to be located input by the target user to obtain multiple location features;

[0132] The feature matching submodule is used to match the multiple positioning features and the three-dimensional coordinates corresponding to the target seat number with the at least one target scene map to obtain the matching result;

[0133] The positioning submodule is used to perform augmented reality positioning of the image to be positioned within the target building based on the matching result, and obtain the positioning result.

[0134] Optional, also includes:

[0135] The keyframe matching module is used to perform keyframe matching on the at least one target scene map to obtain target keyframes when the positioning result indicates that the positioning has failed.

[0136] The feature matching module is used to perform feature matching between the image to be located and the target keyframe to obtain the feature matching result.

[0137] The relationship determination module is used to obtain the transformation relationship based on the feature matching results;

[0138] The image localization module is used to perform augmented reality localization on the target image input by the target user within the target building according to the transformation relationship, and obtain the target localization result.

[0139] This application determines multiple three-dimensional coordinates by identifying multiple seat numbers within a target building, and determines at least one target scene map from multiple pre-captured scene maps by identifying the seat number to be located. Based on the target scene map and the seat number to be located, augmented reality positioning is performed on the image input by the user, thereby improving the augmented reality positioning effect.

[0140] This application also provides a communication device. Please refer to [link to relevant documentation]. Figure 3The communication device may include a processor 301, a memory 302, and a program 3021 stored in the memory 302 and capable of running on the processor 301.

[0141] When the communication device is an electronic device, program 3021 can be executed by processor 301. Figure 1 Any step in the corresponding method embodiment:

[0142] Acquire scene data, which includes multiple seat numbers corresponding to multiple seats within the target building and multiple scene maps;

[0143] Multiple three-dimensional coordinates are determined in a three-dimensional coordinate system based on the multiple seat numbers, and the multiple three-dimensional coordinates correspond one-to-one with the multiple seat numbers. The three-dimensional coordinate system is generated based on the three-dimensional model of the target building.

[0144] Based on the target seat number input by the target user, determine at least one target scene map from the multiple scene maps and determine the three-dimensional coordinates corresponding to the target seat number;

[0145] Based on the at least one target scene map and the three-dimensional coordinates corresponding to the target seat number, augmented reality positioning is performed on the image to be located input by the target user within the target building to obtain the positioning result.

[0146] Optionally, before determining the multiple three-dimensional coordinates in the three-dimensional coordinate system based on the multiple seat numbers, the method further includes:

[0147] The target building is modeled in three dimensions to obtain a three-dimensional model of the target building;

[0148] The target building is divided into multiple areas by dividing the multiple seats into areas, and each target area includes at least one of the seats;

[0149] Each key point corresponding to the target region is determined to obtain multiple key points, which are the vertices of the target region.

[0150] The three-dimensional coordinate system is generated based on the three-dimensional model and the multiple key points.

[0151] Optionally, determining multiple three-dimensional coordinates in a three-dimensional coordinate system based on the multiple seat numbers includes:

[0152] Based on the multiple seat numbers, determine the multiple target two-dimensional coordinates corresponding to the multiple key points;

[0153] Based on the three-dimensional coordinate system, the two-dimensional coordinates of the multiple targets are transformed to obtain the three-dimensional coordinates of the multiple targets, and the three-dimensional coordinates of the multiple targets correspond one-to-one with the multiple key points;

[0154] Based on the multiple target three-dimensional coordinates and the linear interpolation algorithm, the multiple two-dimensional coordinates corresponding to the multiple seat numbers are transformed in the three-dimensional coordinate system to obtain the multiple three-dimensional coordinates.

[0155] Optionally, determining at least one target scene map from the plurality of scene maps based on the target seat number input by the target user and determining the three-dimensional coordinates corresponding to the target seat number includes:

[0156] Determine the first three-dimensional coordinates of the target seat number in the three-dimensional coordinate system, and the second three-dimensional coordinates of the center point of the target building in the three-dimensional coordinate system;

[0157] The direction vector is determined based on the first three-dimensional coordinates and the second three-dimensional coordinates;

[0158] Based on the first three-dimensional coordinates and the direction vector, the multiple scene maps are filtered to determine at least one target scene map.

[0159] Optionally, the step of filtering the multiple scene maps based on the first three-dimensional coordinates and the direction vector to determine at least one target scene map includes:

[0160] A first video frame set is determined from the multiple scene maps. The first video frame set includes multiple first video frames, and the shooting position of the multiple first video frames is less than a preset distance from the first three-dimensional coordinate.

[0161] A second video frame set is determined based on the first video frame set. The second video frame set includes multiple second video frames, and the angle between the multiple second video frames and the direction vector is less than a preset angle.

[0162] In the second set of video frames, at least one target scene map is determined, wherein the target scene map is captured at a location within the observation range of the target seat number.

[0163] Optionally, the step of performing augmented reality localization on the target user's input image within the target building based on the at least one target scene map and the three-dimensional coordinates corresponding to the target seat number, to obtain a localization result, includes:

[0164] Feature extraction is performed on the image to be located input by the target user to obtain multiple location features;

[0165] The matching results are obtained by matching the multiple positioning features and the three-dimensional coordinates corresponding to the target seat number with the at least one target scene map;

[0166] Based on the matching result, augmented reality localization is performed on the image to be located within the target building to obtain the localization result.

[0167] Optionally, after performing augmented reality localization on the image to be located input by the target user within the target building based on the at least one target scene map and the three-dimensional coordinates corresponding to the target seat number, and obtaining the localization result, the method further includes:

[0168] If the positioning result indicates that positioning has failed, keyframe matching is performed on the at least one target scene map to obtain the target keyframe;

[0169] Perform feature matching between the image to be located and the target keyframe to obtain the feature matching result;

[0170] The transformation relationship is obtained based on the feature matching results;

[0171] Based on the transformation relationship, augmented reality localization is performed on the target image input by the target user within the target building to obtain the target localization result.

[0172] This application determines multiple three-dimensional coordinates by identifying multiple seat numbers within a target building, and determines at least one target scene map from multiple pre-captured scene maps by identifying the seat number to be located. Based on the target scene map and the seat number to be located, augmented reality positioning is performed on the image input by the user, thereby improving the augmented reality positioning effect.

[0173] This application also provides a computer-readable storage medium storing a computer program. When executed by a processor, this computer program implements the various processes of the above-described augmented reality positioning method embodiments and achieves the same technical effects. To avoid repetition, it will not be described again here. The computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0174] This application also provides a computer program product, which is stored in a storage medium and executed by at least one processor to implement the various processes of the above-described augmented reality positioning method embodiments, and can achieve the same technical effect. To avoid repetition, it will not be described again here.

[0175] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.

[0176] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0177] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.

Claims

1. An augmented reality positioning method, characterized in that, The method includes: Acquire scene data, which includes multiple seat numbers corresponding to multiple seats within the target building and multiple scene maps; Multiple three-dimensional coordinates are determined in a three-dimensional coordinate system based on the multiple seat numbers, and the multiple three-dimensional coordinates correspond one-to-one with the multiple seat numbers. The three-dimensional coordinate system is generated based on the three-dimensional model of the target building. Based on the target seat number input by the target user, determine at least one target scene map from the multiple scene maps and determine the three-dimensional coordinates corresponding to the target seat number; Based on the at least one target scene map and the three-dimensional coordinates corresponding to the target seat number, augmented reality positioning is performed on the image to be located input by the target user within the target building to obtain the positioning result.

2. The method according to claim 1, characterized in that, Before determining multiple three-dimensional coordinates in a three-dimensional coordinate system based on the multiple seat numbers, the method further includes: The target building is modeled in three dimensions to obtain a three-dimensional model of the target building; The target building is divided into multiple areas by dividing the multiple seats into areas, and each target area includes at least one of the seats; Each key point corresponding to the target region is determined to obtain multiple key points, which are the vertices of the target region; The three-dimensional coordinate system is generated based on the three-dimensional model and the multiple key points.

3. The method according to claim 2, characterized in that, The step of determining multiple three-dimensional coordinates in a three-dimensional coordinate system based on the multiple seat numbers includes: Based on the multiple seat numbers, determine the multiple target two-dimensional coordinates corresponding to the multiple key points; Based on the three-dimensional coordinate system, the two-dimensional coordinates of the multiple targets are transformed to obtain the three-dimensional coordinates of the multiple targets, and the three-dimensional coordinates of the multiple targets correspond one-to-one with the multiple key points; Based on the multiple target three-dimensional coordinates and the linear interpolation algorithm, the multiple two-dimensional coordinates corresponding to the multiple seat numbers are transformed in the three-dimensional coordinate system to obtain the multiple three-dimensional coordinates.

4. The method according to claim 1, characterized in that, The step of determining at least one target scene map from multiple scene maps based on the target seat number input by the target user and determining the three-dimensional coordinates corresponding to the target seat number includes: Determine the first three-dimensional coordinates of the target seat number in the three-dimensional coordinate system, and the second three-dimensional coordinates of the center point of the target building in the three-dimensional coordinate system; The direction vector is determined based on the first three-dimensional coordinates and the second three-dimensional coordinates; Based on the first three-dimensional coordinates and the direction vector, the multiple scene maps are filtered to determine at least one target scene map.

5. The method according to claim 4, characterized in that, The step of filtering the multiple scene maps based on the first three-dimensional coordinates and the direction vector to determine at least one target scene map includes: A first video frame set is determined in the multiple scene maps. The first video frame set includes multiple first video frames, and the shooting position of the multiple first video frames is less than a preset distance from the first three-dimensional coordinate. A second video frame set is determined based on the first video frame set. The second video frame set includes multiple second video frames, and the angle between the multiple second video frames and the direction vector is less than a preset angle. In the second set of video frames, at least one target scene map is determined, wherein the target scene map is captured at a location within the observation range of the target seat number.

6. The method according to claim 1, characterized in that, The step of performing augmented reality localization on the target user's input image within the target building based on the at least one target scene map and the three-dimensional coordinates corresponding to the target seat number, to obtain a localization result, includes: Feature extraction is performed on the image to be located input by the target user to obtain multiple location features; The matching results are obtained by matching the multiple positioning features and the three-dimensional coordinates corresponding to the target seat number with the at least one target scene map; Based on the matching result, augmented reality localization is performed on the image to be located within the target building to obtain the localization result.

7. The method according to claim 1, characterized in that, After performing augmented reality localization on the image to be located input by the target user within the target building based on the at least one target scene map and the three-dimensional coordinates corresponding to the target seat number, and obtaining the localization result, the method further includes: If the positioning result indicates that positioning has failed, keyframe matching is performed on the at least one target scene map to obtain the target keyframe; Perform feature matching between the image to be located and the target keyframe to obtain the feature matching result; The transformation relationship is obtained based on the feature matching results; Based on the transformation relationship, augmented reality localization is performed on the target image input by the target user within the target building to obtain the target localization result.

8. An augmented reality positioning device, characterized in that, The device includes: The acquisition module is used to acquire scene data, which includes multiple seat numbers corresponding to multiple seats in the target building and multiple scene maps; The first determining module is used to determine multiple three-dimensional coordinates in a three-dimensional coordinate system based on the multiple seat numbers, wherein the multiple three-dimensional coordinates correspond one-to-one with the multiple seat numbers, and the three-dimensional coordinate system is generated based on the three-dimensional model of the target building; The second determining module is used to determine at least one target scene map from the multiple scene maps and determine the three-dimensional coordinates corresponding to the target seat number based on the target seat number input by the target user. The positioning module is used to perform augmented reality positioning on the target user's input image within the target building based on the at least one target scene map and the three-dimensional coordinates corresponding to the target seat number, and to obtain the positioning result.

9. An electronic device, comprising: A memory, a processor, and a program stored in the memory and executable on the processor; characterized in that the processor is configured to read the program from the memory to implement the steps of the augmented reality positioning method as claimed in any one of claims 1 to 7.

10. A readable storage medium for storing a program, characterized in that, When the program is executed by the processor, it implements the steps of the augmented reality positioning method as described in any one of claims 1 to 7.

11. A computer program product, characterized in that, The computer program product is stored in a storage medium, and the computer program product is executed by at least one processor to implement the steps in the augmented reality positioning method as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Vehicle positioning method and device, computer equipment and storage medium

    CN115326084A

  • Cross-reality system with fast positioning

    CN115461787A