AR content display method and device, computer equipment and storage medium
By acquiring and processing 3D surround view images and point cloud data, segmenting and verifying feature points, and generating AR maps, the problem of inaccurate AR content display is solved, high-precision alignment of virtual reality and the real world is achieved, and the robustness of AR display and user experience are improved.
Patent Information
- Application Number
- CN202410347292.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-26
- Publication Date
- 2025-10-03
AI Technical Summary
In the existing technology, AR content cannot be accurately displayed on objects in the real world, resulting in inaccurate position matching between virtual content and reality.
By obtaining the 3D surround view image and point cloud data of the target space, dividing it into multiple sub-images and ensuring the overlapping area, the sub-image is converted into a plane image, and the initial feature points are verified and optimized using the point cloud data to generate the target AR map, and the AR content is displayed according to the camera posture.
It improves the display accuracy and robustness of AR content, ensures the alignment accuracy of virtual content with the real world, optimizes AR visual effects, and enhances user experience.
Smart Images

Figure CN120747422A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of virtual reality technology, and in particular to a method, apparatus, computer device, and storage medium for displaying AR content. Background Art
[0002] With the rapid development of digital and virtual technologies, augmented reality (AR) has become a crucial interactive platform, particularly in museums, education, and retail. A core feature of AR technology is AR tags, which anchor virtual content to real-world objects, providing users with an informative and interactive experience. This virtual content can be displayed via a camera's video stream or dedicated AR glasses.
[0003] In existing technologies, after acquiring images of the real environment, it is impossible to accurately obtain the feature points of objects in the environment image, and thus it is impossible to provide high-precision position matching between reality and virtuality. Therefore, when anchoring virtual content to objects in the real world, AR content cannot be displayed accurately. Summary of the Invention
[0004] One purpose of the embodiments of the present application is to provide a method, apparatus, device, and storage medium for displaying AR content to solve the technical problem of inaccurate AR content display in related technologies.
[0005] In a first aspect, a method for displaying AR content is provided, comprising:
[0006] Obtain 3D surround image and point cloud data of the target space;
[0007] Splitting the 3D surround view image into a plurality of sub-images, wherein two adjacent sub-images have overlapping image areas;
[0008] Converting each sub-image into a planar image, wherein the planar image includes a plurality of initial feature points;
[0009] Performing verification and optimization processing on the initial feature points of the plane image according to the point cloud data to obtain target feature points;
[0010] Generate a target AR map according to the target feature points;
[0011] Acquire a target image captured by a camera in the target space;
[0012] Determining a target pose of the target image captured by the camera according to the AR map;
[0013] Determining target AR content corresponding to the target pose;
[0014] The target AR content is displayed at a position corresponding to the target posture.
[0015] In combination with the first aspect, in a possible implementation method, the verification optimization processing includes verification processing and position optimization processing, and the verification optimization processing is performed on the initial feature points of the plane image according to the point cloud data to obtain target feature points, including: verifying the initial feature points of the plane image according to the point cloud data to obtain candidate feature points that meet the verification conditions; obtaining the candidate spatial positions of the candidate feature points in the world coordinate system; performing position optimization processing on the candidate spatial positions to obtain the target spatial positions; and generating a target AR map based on the target spatial positions of the candidate feature points.
[0016] As can be seen, the spatial positions of the feature points obtained through verification and position optimization in this method will be considered the final target spatial positions. Based on these optimized feature points, a target AR map can be generated. This map records the positions of these feature points in the physical world in detail, providing the necessary spatial reference, accurately positioning virtual objects, and dynamically adjusting the display of virtual content based on the user's perspective and position.
[0017] In combination with the first aspect, in a possible implementation, the initial feature points of the plane image are verified based on the point cloud data to obtain candidate feature points that meet the verification conditions, including: projecting the point cloud data onto the image coordinate system of the plane image to obtain a projected point cloud, wherein the projected point cloud includes multiple projection points; searching for multiple projection points that meet the interpolation conditions with the initial feature points among the multiple projection points as multiple target projection points, and the initial feature points are located in a closed area formed by the multiple target projection points; determining the effectiveness attributes of the closed area; and verifying the initial feature points of the plane image according to the effectiveness attributes of the closed area to obtain candidate feature points that meet the verification conditions.
[0018] It can be seen that this method can effectively screen out those candidate feature points that are geometrically reliable and match the point cloud data from a large number of initial feature points. The verified feature points provide a solid foundation for the precise positioning and stable display of subsequent AR content.
[0019] In combination with the first aspect, in a possible implementation method, the validity attribute includes a valid attribute and an invalid attribute, and the initial feature points of the plane image are verified according to the validity attribute of the closed area to obtain candidate feature points that meet the verification conditions, including: if the validity attribute of the closed area is a valid attribute, then according to the positions of multiple target projection points of the closed area and the position of the initial feature point, the depth of the initial feature point is determined; according to a preset world transformation matrix, the position of the initial feature point and the depth, the spatial position of the initial feature point in the world coordinate system is determined; according to the spatial position of the initial feature point of each of the plane images, whether the initial feature point meets the verification conditions is judged; if the verification conditions are met, the initial feature point is determined as a candidate feature point.
[0020] It can be seen that this method can effectively screen out candidate feature points with precise spatial positions from a large number of initial feature points, thereby improving the accuracy of feature point positioning.
[0021] In combination with the first aspect, in a possible implementation method, judging whether the initial feature points meet the verification conditions based on the spatial positions of the initial feature points of each of the planar images includes: calculating the variance value of the spatial positions of the initial feature points of all the planar images; if the variance value is greater than the preset variance value, judging that the initial feature points do not meet the verification conditions, and eliminating the initial feature points; if the variance value is less than or equal to the preset variance value, judging that the initial feature points meet the verification conditions, and retaining the initial feature points.
[0022] It can be seen that the variance-based verification process of this method provides an effective method to screen out feature points that are stable and consistent in spatial position. By eliminating feature points with inconsistent positions, the risk of mismatching can be reduced, ensuring the correct rendering and display of AR content.
[0023] In combination with the first aspect, in a possible implementation method, determining the effectiveness attribute of the closed area includes: obtaining the depth value of each target projection point in the closed area; calculating the difference in depth values between each two adjacent target projection points to obtain a depth difference; if all the depth differences are less than a preset difference, determining that the closed area has a valid attribute.
[0024] It can be seen that this method can effectively screen out geometrically stable and consistent feature points, providing a solid foundation for building reliable 3D models and AR scenes.
[0025] In combination with the first aspect, in a possible implementation method, the position optimization processing of the candidate spatial positions to obtain the target spatial position includes: obtaining the candidate spatial position of the candidate feature point on each of the sub-images; calculating the average value of all the candidate spatial positions of the candidate feature points to obtain the average spatial position of the candidate feature points; calculating the position deviation based on the average spatial position of the candidate feature point and each candidate spatial position; and calculating the target spatial position of the candidate feature point based on the average spatial position and the position deviation.
[0026] It can be seen that this method can determine an optimized spatial position for each feature point. This position is the most likely coordinate that best represents the actual spatial position of the feature point in a statistical sense.
[0027] In combination with the first aspect, in a possible implementation, in the AR map, each of the plane images is configured with a reference global feature vector, and determining the target pose of the target image captured by the camera based on the AR map includes: obtaining the target global feature vector of the target image; determining a reference global feature vector that satisfies a preset similarity condition with the target global feature vector as a candidate global feature vector, and the plane image corresponding to the candidate global feature vector as a candidate plane image; performing feature matching on each of the candidate plane images with the target image to obtain feature matching points; and determining the target pose based on a preset pose estimation algorithm and the spatial position of each feature matching point in the world coordinate system.
[0028] It can be seen that this method can understand the position and orientation of the scene captured by the camera in the physical world, so as to accurately superimpose the computer-generated image or 3D model on the user's actual environment.
[0029] In a second aspect, a device for displaying AR content is provided, the device comprising:
[0030] An acquisition unit, used to acquire 3D surround view images and point cloud data of the target space;
[0031] a segmentation unit, configured to segment the 3D surround view image into a plurality of sub-images, wherein two adjacent sub-images have overlapping image areas;
[0032] a conversion unit, configured to convert each sub-image into a planar image, wherein the planar image includes a plurality of initial feature points;
[0033] an optimization unit, configured to perform verification and optimization processing on the initial feature points of the planar image according to the point cloud data to obtain target feature points;
[0034] A generating unit, configured to generate a target AR map according to the target feature points;
[0035] The acquisition unit is further configured to acquire a target image captured by the camera in the target space;
[0036] An acquisition unit, configured to determine a target pose for the camera to acquire the target image according to the AR map;
[0037] a determination unit, configured to determine a target AR content corresponding to the target posture;
[0038] The display unit is further configured to display the target AR content at a position corresponding to the target posture.
[0039] In a third aspect, an embodiment of the present invention provides a computer device, including:
[0040] at least one processor; and,
[0041] a memory communicatively connected to the at least one processor; wherein,
[0042] The memory stores instructions that can be executed by the at least one processor. The instructions are executed by the at least one processor to enable the at least one processor to perform the method according to the first aspect.
[0043] In a fourth aspect, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores a computer program, wherein the computer program includes program instructions, and when the program instructions are executed by a processor, the processor is caused to perform the method according to the first aspect.
[0044] In the scheme implemented by the above-mentioned display method, device, computer equipment and storage medium of AR content, first, a 3D surround view image and point cloud data of the target space are obtained, and the 3D surround view image is divided into multiple sub-images, and there is an overlapping image area between two adjacent sub-images. Secondly, each sub-image is converted into a plane image, and the plane image includes multiple initial feature points. The initial feature points of the plane image are verified and optimized according to the point cloud data to obtain target feature points. Then, a target AR map is generated according to the target feature points, and the target image captured by the camera in the target space is obtained. The target posture of the target image captured by the camera is determined according to the AR map, and the target AR content corresponding to the target posture is determined. Finally, the target AR content is displayed at the position corresponding to the target posture. This method enhances the continuity and integrity of the scene by dividing the 3D surround view image into multiple sub-images and ensuring that there are overlapping areas between them, thereby improving the accuracy of subsequent feature point extraction and matching. It further converts the sub-images into planar images and includes initial feature points, simplifying the image processing process and providing preliminary data for the accurate verification and optimization of feature points. It also uses point cloud data to verify and optimize the initial feature points of the planar image, significantly improving the reliability and accuracy of the feature points, laying the foundation for building high-quality AR maps. The target AR map generated based on precise feature points improves the alignment accuracy of virtual content and the real world in augmented reality, optimizes AR visual effects, improves the robustness and user experience of the AR display method, and ensures the spatial consistency and real-time performance of AR display content. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments of the present application. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0046] Figure 1 1 is a schematic diagram of a scene of a method for displaying AR content in one embodiment of the present invention;
[0047] Figure 2A 1 is a flow chart of a method for displaying AR content in one embodiment of the present invention;
[0048] Figure 2B This is a schematic diagram of an equi-cylindrical projection model in one embodiment of the present invention;
[0049] Figure 2C 1 is a schematic diagram of a barycentric coordinate interpolation method in one embodiment of the present invention;
[0050] Figure 2D is a schematic diagram of a triangular facet in one embodiment of the present invention;
[0051] Figure 3 is a schematic structural diagram of a display device for AR content in one embodiment of the present invention;
[0052] Figure 4 It is a structural diagram of a computer device in one embodiment of the present invention. DETAILED DESCRIPTION
[0053] In order to make the purpose, technical solutions and advantages of this application more clearly understood, the present application is further described in detail below with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0054] It should be noted that, if there is no conflict, the various features in the embodiments of the present application can be combined with each other and are all within the scope of protection of the present application. In addition, although the functional modules are divided in the device schematic and the logical order is shown in the flow chart, in some cases, the steps shown or described can be performed in a different order than the module division in the device or the order in the flow chart. Furthermore, the words "first", "second", "third", etc. used in this application do not limit the data and execution order, but only distinguish between the same items or similar items with basically the same functions and effects.
[0055] The present invention is described in detail below through specific examples.
[0056] The technical solution of this application can be applied to various types of AR devices, etc.
[0057] In the existing technology, building AR tags for indoor scenes usually requires a detailed digital map so that the corresponding virtual content can be displayed at a specific location. Figure 1 Generally, it relies on high-precision indoor 3D scanning technology, which can generate high-quality point cloud data and surround view images. This data is then used to assist visual positioning algorithms to ensure that virtual content can accurately correspond to physical locations in the real world.
[0058] However, existing technologies face a series of challenges and limitations. First, the feature point extraction and matching process is easily disturbed in complex indoor environments, resulting in inaccurate positioning of AR content. In addition, traditional scanning and matching technologies are often difficult to adapt to dynamically changing environments, which affects the practicality and user experience of AR tags. In addition, incorrect matches may be caused by repeated textures or light changes in the environment. Although existing filtering algorithms (such as RANSAC) can solve this problem to a certain extent, their computational cost is high and they are still not robust enough in some cases.
[0059] In addition to the matching issue, existing technologies for aligning 3D scan data with 2D images typically rely on complex mathematical models and algorithms to calculate projection and inverse projection. This process is computationally demanding and highly dependent on the accuracy of the camera's intrinsic and extrinsic parameters. Without accurate camera parameters, existing methods struggle to achieve precise visual positioning.
[0060] Furthermore, existing systems require high mobility and flexibility in scanning equipment during the digitization of indoor spaces, requiring scanning at multiple locations to obtain complete indoor data. This not only increases the time required to build digital maps but also places high demands on operators.
[0061] In summary, although the existing technology provides a certain basis for the implementation of AR tags, there are still obvious deficiencies in the accuracy of feature point matching, algorithm robustness, computational cost, and user experience. Therefore, a method, device, computer equipment, and storage medium for displaying AR content are proposed.
[0062] See also Figure 1 , Figure 1 This is a scene diagram of a method for displaying AR content provided in an embodiment of the present application. In this diagram, there is a 3D scanner 100 involved in this solution, and a camera 200 and a laser radar 300 are built into the 3D scanner 100. Specifically, the 3D scanner 100 is an indoor surveying and mapping device that integrates a camera 200 and a laser radar 300 to accurately capture and build a three-dimensional digital model of the indoor space. The camera 200 is used to capture high-resolution images of the indoor environment, while the laser radar 300 emits a laser beam and receives the reflected laser, and calculates the distance to the object by measuring the round-trip time of the laser, thereby obtaining accurate point cloud data of the space.
[0063] exist Figure 1The figure describes an indoor environment equipped with various furniture and electronic devices. The 3D scanner 100 is placed in a corner of the indoor space and performs a 360-degree scan of the entire room, collecting surround-view images and point cloud data for building a high-precision digital map. The above data is used to create a digital twin environment and accurately place AR tags in the environment. During the scanning process, the laser or infrared light emitted by the scanner is reflected by the surface of objects in the room and then captured by the scanner's sensor. In this way, it can measure the size, shape and exact position of objects in space. In addition, an AR display device is shown, which may be AR glasses or other wearable devices. The OLED screen of the device may display the AR map and related virtual content generated by the 3D scanner. These devices can use surround-view images and point cloud data to accurately track the user's position and line of sight, and then present corresponding virtual information or images in the user's field of view, thereby creating an interactive augmented reality experience.
[0064] It should be noted that in the above process, the following four coordinate systems are involved: lidar coordinate system, image coordinate system and world coordinate system.
[0065] The world coordinate system is used to define the position and orientation of all objects in an indoor environment. In this solution, the point cloud data and surround view images generated by the 3D scanner 100 are mapped to this world coordinate system, ensuring consistency and comparability between different data.
[0066] The camera 200 and the LiDAR 300 each have their own internal coordinate systems, which are used to describe the relative positions of the components within the device. During data processing, the data in these coordinate systems needs to be converted to the world coordinate system for integration and analysis.
[0067] Two-dimensional coordinate system: When the image captured by the camera 200 is processed, it will be mapped into a two-dimensional image coordinate system, which is directly related to the pixel position and is used for image processing and feature extraction.
[0068] Three-dimensional coordinate system: The point cloud data obtained by the laser radar 300 scan is first located in the laser radar's own coordinate system, which is a three-dimensional coordinate system. It then needs to be mapped to the world coordinate system through a conversion algorithm to achieve fusion with other data.
[0069] Through the conversion between the above coordinate systems, this solution ensures the accuracy and efficiency of the coherent process from the actual environment to the digital model and then to the AR display.
[0070] In view of this, this application proposes a method for displaying AR content to solve the above problems, which is described in detail below.
[0071] See also Figure 2A , Figure 2A A flowchart of a method for displaying AR content provided by an embodiment of the present invention is provided, wherein the method includes the following steps:
[0072] S10: Acquire 3D surround view images and point cloud data of the target space.
[0073] The 3D surround view image is captured by a camera in a 3D scanner. The 3D surround view image can record visual information in the target space. The 3D surround view image can be understood as a panoramic image that captures the target space in all directions.
[0074] Specifically, the process of obtaining a 3D surround view image usually involves fixing the camera at different viewing angles at any point in the target space to capture a panoramic view of the interior. This allows for detailed images of every corner of the interior to be obtained. The detailed images are equirectangular images, which are then used to generate a 3D surround view image, providing users with a continuous and seamless visual environment.
[0075] In augmented reality (AR) applications, 3D surround images are associated with physical locations in the real world. For example, during indoor 3D scanning, converting panoramic images into equirectangular projections makes it easier to locate feature points in the 2D image and match these points with physical locations in 3D space. This is the basis for creating indoor digital maps and accurately placing virtual content in augmented reality applications.
[0076] Among them, 3D surround view images and point cloud data can be synchronized and matched based on time series. For example, capture devices (such as 3D cameras and laser scanners) can record a timestamp at the same time when capturing each data point (whether it is an image pixel or a point in the point cloud), ensuring that all data are given accurate timestamps during the capture of 3D surround view images and point cloud data. This means that each sub-image and each point cloud data point has a clear time stamp, indicating the exact moment they were captured. It can be seen that through time series-based matching, the spatial correspondence between the surround view image or sub-image and the point cloud data can be ensured to be both accurate and consistent.
[0077] like Figure 2B As shown, Figure 2B This is a schematic diagram of the equirectangular projection model. The position of a point P in the spherical coordinate system is defined by two angles, θ and φ, where θ is a horizontal angle ranging from -π to +π, and φ is a vertical angle. This model converts the viewing direction represented by the azimuth coordinates (φ, θ) into planar coordinates P′. The correspondence between the azimuth coordinates and the circumferential image plane is obtained by mapping the horizontal and vertical coordinates of P′ from (-π, π) and (0, π), respectively, to (0, W) and (0, H), where W and H represent the width and height of the image, respectively.
[0078] The point cloud data is generated by the 3D scanner's LiDAR (LiDAR). This data measures the position of objects by emitting laser light and receiving reflected light. As the LiDAR scans the entire indoor space, the position information of each reflected point is compiled into a point cloud, representing the three-dimensional structure of the indoor space. This point cloud data accurately reflects the shape, size, and relative position of objects in space.
[0079] Among them, 3D surround view images provide rich texture and color information for point cloud data, while point cloud data provides accurate three-dimensional spatial information, such as the exact position and size of objects in space.
[0080] It can be seen that this method can accurately reconstruct the three-dimensional model of the target space by combining surround view images and point cloud data.
[0081] S20 , dividing the 3D surround view image into a plurality of sub-images, where two adjacent sub-images have overlapping image areas.
[0082] In particular, the 3D surround view image and point cloud data mentioned in S10 can be synchronized and matched based on a time series. Furthermore, when the 3D surround view image is segmented into sub-images, each sub-image inherits the timestamp range or specific time stamp of the original image, indicating the moment or time period represented by the sub-image. Simultaneously, the point cloud data is organized or segmented based on timestamps to match the timestamps of the image data. By comparing the timestamps of the sub-images with the timestamps of the point cloud data points, it is possible to identify which point cloud data is contemporaneous with a particular sub-image, i.e., they represent spatial information from the same moment or time period. Once the point cloud data corresponding to a particular sub-image is determined, it can be assigned to the corresponding sub-image, forming a complete spatial dataset that contains not only the visual information of the sub-image but also the spatial geometry information at the corresponding moment. Finally, through this time series-based correspondence, the visual feature points in each sub-image can be matched and fused with the corresponding point cloud data for further analysis and processing, such as feature point verification and augmented reality content positioning.
[0083] Splitting the full 3D surround view image into smaller sub-images and ensuring a certain amount of overlap between them increases the chances of feature points appearing from different perspectives. This way, even if a feature point is unclear in one sub-image or difficult to identify due to perspective, it may be more visible in another overlapping sub-image.
[0084] The surround view image is a panoramic image within the target space. Panoramic images are usually large and difficult to process completely at one time. Therefore, ensuring that there is an overlapping area between adjacent sub-images can help match and align the above sub-images in subsequent steps.
[0085] The overlapping image regions provide a continuous visual context, helping to identify and associate features that appear in different sub-images. By segmenting and setting overlapping regions, the continuity of spatial information can be better maintained, making the resulting model more complete and coherent. This is critical for building accurate AR scenes because it ensures the accurate positioning and natural integration of virtual content.
[0086] It can be seen that in practical applications, due to factors such as occlusion and illumination changes, some parts of the image may not provide sufficient feature information. Through segmentation and overlap, this method ensures that even if part of the image is affected, the remaining images can still provide sufficient information for feature point matching, thereby enhancing the adaptability to complex environments and overall robustness.
[0087] S30 , converting each sub-image into a planar image, where the planar image includes a plurality of initial feature points.
[0088] The process of converting a sub-image into a two-dimensional image is essentially a projection process, typically using perspective or orthographic projection. The goal is to transform a 3D view into a two-dimensional image. Perspective projection preserves the sense of depth, making distant objects appear smaller and more realistic to the human eye. Orthographic projection, on the other hand, ignores depth of field, rendering all objects the same size regardless of distance, and is often used in situations requiring precise measurement.
[0089] Feature points are points in an image that have significant visual characteristics, such as corners, edges, or specific texture areas. Feature points are also matchable across multiple images, making it easier to identify the same features in the same scene captured in different images.
[0090] Specifically, these initial feature points are key to aligning virtual content with the real world. By identifying and tracking these feature points, the system understands the scene structure from the user's perspective and accurately positions and renders virtual objects accordingly. Especially in dynamic environments, high-quality feature points enable the system to update the position of virtual content in real time, maintaining consistency between virtual content and the real world.
[0091] Specifically, feature point extraction usually utilizes algorithms in computer vision, such as SIFT (Scale Invariant Feature Transform), SURF (Speeded Up Robust Features), ORB (Oriented FAST and Rotated BRIEF), etc., to automatically detect and extract these salient points.
[0092] Specifically, when converting each sub-image in the three-dimensional coordinate system into a plane image in the two-dimensional coordinate system, the intrinsic parameter matrix K of the sub-image and the orientation (rotation matrix) R in the three-dimensional coordinate system are given. The intrinsic parameter matrix K includes the focal length (f x , f y ) and optical center (c x , v y ) parameters, the parameters in the intrinsic parameter matrix describe the optical properties of the camera lens and the geometric center of the image sensor. The rotation matrix R is used to transform the 3D coordinates of the surround image to a new 3D coordinate system relative to a specific observation point.
[0093] The mapping relationship between the sub-image coordinates and the surround image coordinates in any direction is established according to the following formula. Sub-image coordinates (u0 v0) T Coordinates of the surround image (u1 v1) T The mapping relationship is:
[0094]
[0095]
[0096] θ=arcsin(z) Formula 3
[0097]
[0098]
[0099] Among them, (xyz) T is the unit direction vector, s is the modulus of vector v,
[0100] In the above transformation process, 3D points are represented by their positions in the spherical coordinate system (defined by angles θ and φ), so the 3D points are first transformed by the rotation matrix R and then mapped to the two-dimensional image plane using the intrinsic parameter matrix K. Among them, Formula 2 describes that v is the inverse matrix RK through K and R ―1 The resulting 2D image coordinates. Formulas 3 and 4 are used to calculate coordinates in the equidistant cylindrical projection. These coordinates define the position of a 3D point on the 2D image plane, where θ is the latitude angle of the point and φ is the longitude angle. Using the above formulas, spherical coordinates are converted to 2D image coordinates, transforming visual information originally captured in a 3D environment into a 2D image format that can be used by feature point detection algorithms.
[0101] Among them, the plane image can use a feature point detection algorithm (such as SIFT or D2-Net) to identify the feature points in the image. Please refer to the previous step for details and will not be repeated here.
[0102] It can be seen that this method can detect and match feature points more accurately and optimize the efficiency of feature point matching by dividing the surround view image into multiple sub-images and ensuring that there is overlapping area between the sub-images. It can further convert the sub-images into two-dimensional plane images in a three-dimensional coordinate system, which can maintain the visual information and structural layout in the space and improve the quality of AR content rendering.
[0103] S40 , performing verification and optimization processing on the initial feature points of the planar image according to the point cloud data to obtain target feature points.
[0104] Specifically, the corresponding position of each initial feature point in the two-dimensional image in the point cloud data needs to be determined. This can be achieved by using pre-determined camera parameters and scanner positions, as well as pose information recorded during the capture process. The two-dimensional feature points in the image are then mapped back to three-dimensional space and compared with the points in the point cloud data.
[0105] Among them, the verification optimization processing can be divided into verification processing and position optimization processing.
[0106] Specifically, the verification process involves comparing the expected positions of image feature points (based on the 3D model reconstructed from the point cloud data) with their actual positions. If the position of a feature point in 3D space matches the geometric characteristics of the corresponding position in the point cloud data (such as edges, corners, or other significant structures), the feature point is considered valid. Conversely, if there is a mismatch, the feature point may need to be adjusted or removed.
[0107] Specifically, for those feature points that are verified to be valid, further optimization processing may include adjusting their position in the image to better match the corresponding geometric features in the point cloud data, or adjusting their weight (i.e., their importance in subsequent processing) to reflect the different degrees of their reliability. Optimization may also involve merging feature points that are approximately overlapping, or extracting new feature points from the point cloud data to supplement the image data.
[0108] It can be seen that through verification and optimization processing, the target feature points finally obtained by this method are feature points that have been carefully screened and adjusted. The target feature points are not only visually significant but also spatially geometrically accurate, providing a reliable foundation for subsequent AR content positioning and rendering.
[0109] S50: Generate a target AR map according to the target feature points.
[0110] Among them, the target feature points are accurate feature points obtained after verification and optimization processing in S40. According to the positions of the target feature points in three-dimensional space, a database containing the spatial coordinates of the target feature points is constructed. The coordinates reflect the position of each target feature point relative to a known reference point (such as the starting position of the scan). The spatial information of the target feature points is further used to generate a detailed AR environment map.
[0111] An AR map is a digital model of an environment that features and is used to identify and locate points in real physical space. The AR map not only marks the locations of all feature points but also includes information about other aspects of the environment, such as object edges, surface textures, or other recognizable landmarks, though these are not intended to be exclusive.
[0112] Among them, this method uses the verified and optimized feature points to construct a detailed AR environment map, which not only reflects the spatial layout of the physical environment, but also contains sufficient feature information to accurately locate and render virtual objects.
[0113] S60: Acquire a target image captured by a camera in the target space.
[0114] The target image is a real-time image captured by the camera in the target space. The target image can be a continuous image stream of the target space, providing visual information about objects and surfaces in the target space, including color, shape, size and their relative positions.
[0115] The camera could be part of a smartphone, tablet, AR glasses, or any device with photography capabilities.
[0116] S70: Determine the target pose of the target image captured by the camera according to the AR map.
[0117] The target pose includes the position (point in three-dimensional space) and direction (i.e., orientation) of the camera, including its coordinates in three-dimensional space and the direction about its visual axis.
[0118] Specifically, in determining the target pose of the target image captured by the camera based on the AR map, feature points can be first identified in the target image captured by the camera, and then the feature points can be matched with corresponding points in the AR map. Feature points can be significant corners, edges or other unique visual marks in the image. Further, based on the known feature points and their positions on the AR map, the camera's intrinsic parameter matrix and extrinsic parameter information are used to calculate the coordinates of the feature points in the current camera perspective. Geometric calibration is also required to correct any perspective deviation between the camera and the target space, which can be performed through a transformation matrix. The transformation matrix describes the relationship between the camera and the global reference frame (i.e., the coordinate system in the AR map).
[0119] For example, the edge of a door or the corner of furniture is detected in the target image, and the features of the door edge or the corner of the furniture are matched with the corresponding physical features recorded in the AR map. Through matching, the target pose of the camera relative to the object, that is, the precise position and orientation, can be calculated. The target pose is then used to render virtual objects or information at the appropriate position and scale in augmented reality applications.
[0120] It can be seen that this method can ensure that virtual objects and information are correctly rendered on the user's device, and their position and orientation are accurately aligned with the user's position and perspective in the real environment, providing a coherent and responsive augmented reality experience.
[0121] S80: Determine target AR content corresponding to the target posture.
[0122] S50 is a process of selecting appropriate virtual content for display based on the exact position (posture) of the user or the camera in the three-dimensional space.
[0123] Specifically, it is necessary to first determine the user or camera's perspective, which includes position and orientation. The target pose can indicate which direction the user is facing and their relative position to the surrounding objects. Secondly, based on the camera's pose, select content that matches the pose from the preset AR content library. For example, if the user is facing an exhibit, select information or images related to the exhibit. In some cases, the content may be customized in real time based on the user's specific position and perspective, further indicating that the same object may display different information at different viewing angles.
[0124] It can be seen that this method ensures that the provided content not only matches the user's physical location, but also can be dynamically adjusted according to the user's sight and behavior.
[0125] S90: Display the target AR content at a position corresponding to the target posture.
[0126] The display process involves rendering the selected AR content onto the user's device, ensuring that the virtual content appears at the correct position and scale within the user's field of view. This process requires rendering the virtual content so that it can be displayed on the user's AR display device. This may include text, images, videos, or 3D models.
[0127] For example, when a user views an exhibit in a museum through AR glasses, the system analyzes the real-time image provided by the glasses' camera and matches it with reference points on the AR map. Once the user's position and perspective are determined, information labels or 3D models of the exhibits can be correctly placed in the user's field of view, providing a rich and interactive visitor experience.
[0128] It can be seen that the AR content displayed by this method corresponds to the user's position, providing an immersive experience, blurring the boundaries between virtual elements and the real-world environment, and greatly enhancing the user's interactivity and sense of participation.
[0129] This method enhances the continuity and integrity of the scene by dividing the 3D surround view image into multiple sub-images and ensuring that there are overlapping areas between them, thereby improving the accuracy of subsequent feature point extraction and matching. It further converts the sub-images into planar images and includes initial feature points, simplifying the image processing process and providing preliminary data for the accurate verification and optimization of feature points. It also uses point cloud data to verify and optimize the initial feature points of the planar image, significantly improving the reliability and accuracy of the feature points, laying the foundation for building high-quality AR maps. The target AR map generated based on precise feature points improves the alignment accuracy of virtual content and the real world in augmented reality, optimizes AR visual effects, improves the robustness and user experience of the AR display method, and ensures the spatial consistency and real-time performance of AR display content.
[0130] In one possible example, the verification optimization processing includes verification processing and position optimization processing, and the verification optimization processing is performed on the initial feature points of the plane image according to the point cloud data to obtain target feature points, including: verifying the initial feature points of the plane image according to the point cloud data to obtain candidate feature points that meet the verification conditions; obtaining the candidate spatial positions of the candidate feature points in the world coordinate system; performing position optimization processing on the candidate spatial positions to obtain the target spatial positions; and generating a target AR map based on the target spatial positions of the candidate feature points.
[0131] The verification optimization process includes verification processing and position optimization processing.
[0132] Among them, the candidate feature points are initial feature points that meet the verification conditions.
[0133] Specifically, the verification process can be to check whether each initial feature point corresponds to an actual geometric feature in the point cloud, such as an edge, corner, or specific surface texture. This verification can be based on the point cloud density, shape, or other geometric properties near the feature point. Only when a feature point can be matched with a corresponding geometric feature in the point cloud data is it considered valid and retained as a candidate feature point.
[0134] Among them, the acquisition of candidate spatial positions can be based on the process of converting the two-dimensional image coordinates of the candidate feature points into three-dimensional spatial coordinates, based on the calibration parameters of the camera, the position information of the scanner, and any known scene geometric constraints. Specifically, after determining the feature points in the surround view image, the point cloud data is used to determine the exact three-dimensional position of the feature points in the real world. The point cloud data is captured by a 3D scanner using lidar technology and contains the X, Y, and Z coordinates of a large number of points in space. By comparing the feature points in the surround view image with the corresponding points in the point cloud, the exact spatial position of the feature points can be determined in a preset world coordinate system (a reference coordinate system used to define the position of an object in space).
[0135] Among them, position optimization involves minimizing the projection error between the feature points in the image domain and the point cloud data, or adjusting the feature point position to better conform to the geometric layout of the surrounding point cloud. This can be achieved through various optimization algorithms, such as the Iterative Closest Point (ICP) algorithm or the least squares method.
[0136] The AR map is a digital model that contains the spatial locations of all feature points. It is used to locate and track the user's environment. The AR map allows the user's perspective and position to be identified, and then virtual content can be accurately overlaid on the user's device.
[0137] Among them, after the AR map is established, when the environment is updated or adjusted, only the relevant feature points and point cloud data need to be updated, without having to rebuild the entire AR map from scratch.
[0138] As can be seen, the spatial positions of the feature points obtained through verification and position optimization in this method will be considered the final target spatial positions. Based on these optimized feature points, a target AR map can be generated. This map records the positions of these feature points in the physical world in detail, providing the necessary spatial reference, accurately positioning virtual objects, and dynamically adjusting the display of virtual content based on the user's perspective and position.
[0139] In one possible example, the initial feature points of the plane image are verified based on the point cloud data to obtain candidate feature points that meet the verification conditions, including: projecting the point cloud data onto the image coordinate system of the plane image to obtain a projected point cloud, wherein the projected point cloud includes multiple projection points; searching for multiple projection points that meet the interpolation conditions with the initial feature points among the multiple projection points as multiple target projection points, and the initial feature points are located in a closed area formed by the multiple target projection points; determining the effectiveness attributes of the closed area; and verifying the initial feature points of the plane image according to the effectiveness attributes of the closed area to obtain candidate feature points that meet the verification conditions.
[0140] Among them, point cloud data is a data set consisting of a series of points in three-dimensional space, and each point contains X, Y, and Z coordinates.
[0141] The projection process is a dimensionality reduction operation that converts the three-dimensional coordinates of each point into a two-dimensional coordinate system, which is the image coordinate system of the plane image. Depth information is usually lost, but the position of each point on the two-dimensional plane (for example, the position seen by the camera) can be obtained. Therefore, the projected point cloud is position information on the two-dimensional plane without depth information.
[0142] The target projection point is a point cloud projection point that is geometrically adjacent to the feature point in the plane image, such as Figure 2D In P1, P2, and P3, the target projection point is close to the feature point in the physical space.
[0143] The interpolation conditions are based on the adjacency of points in the point cloud, which can be determined by projecting the point cloud onto a 2D image. During the search process, a nearest neighbor search algorithm can be used to find the point cloud projection point closest to the feature point. For example, if the feature point lies within the triangle formed by the three target projection points, barycentric coordinate interpolation can be used to estimate the depth.
[0144] The validity attributes include valid and invalid attributes. The validity attributes determine whether the geometric structure formed by the target projection points within the closed area is consistent with the expected physical reality. Valid attributes mean that the distribution of projection points within the closed area reflects the consistent geometric features of the actual object surface, while invalid attributes mean that the distribution of projection points within the area does not conform to any reasonable physical structure, which may be due to errors caused by noise, occlusion, or other factors.
[0145] Specifically, if an initial feature point is located within a valid closed area, it means that the corresponding position of the feature point in 3D space has reliable geometric support. Therefore, the feature point is considered to meet the verification conditions and can be used as a candidate feature point. Conversely, if the initial feature point is located within a closed area with invalid validity attributes, or is not located in any closed area at all, then the point may be eliminated.
[0146] It can be seen that this method can effectively screen out those candidate feature points that are geometrically reliable and match the point cloud data from a large number of initial feature points. The verified feature points provide a solid foundation for the precise positioning and stable display of subsequent AR content.
[0147] In one possible example, the validity attribute includes a valid attribute and an invalid attribute, and the initial feature points of the plane image are verified according to the validity attribute of the closed area to obtain candidate feature points that meet the verification conditions, including: if the validity attribute of the closed area is a valid attribute, then the depth of the initial feature point is determined according to the positions of multiple target projection points of the closed area and the position of the initial feature point; the spatial position of the initial feature point in the world coordinate system is determined according to a preset world transformation matrix, the position of the initial feature point and the depth; based on the spatial position of the initial feature point of each of the plane images, whether the initial feature point meets the verification conditions; if the verification conditions are met, the initial feature point is determined as a candidate feature point.
[0148] The validity attribute includes a valid attribute and an invalid attribute.
[0149] Among them, the depth of the initial feature points is usually obtained by combining with point cloud data. The point cloud data contains the original three-dimensional spatial information, which involves comparing the original 3D point cloud and two-dimensional image feature points. Stereo vision methods can be used to estimate the depth of the feature points.
[0150] The above-mentioned preset world transformation matrix can be manually set or conventional data. The world transformation matrix is usually composed of the rotation R pin and pan t pin The points converted from the two-dimensional image coordinate system to the world coordinate system can be correctly positioned in space.
[0151] Specifically, such as Figure 2C As shown, Figure 2C This is a schematic diagram of the barycentric coordinate interpolation method. For further information, please refer to Figure 2D , the specific process of finding the spatial position of the unknown point P4 is calculated by weighted interpolation of the three known points P1, P2, P3 (whose 3D positions are known to us) and the unknown point P4. The specific interpolation method is as follows:
[0152]
[0153]
[0154] w3=1.0―w1―w2
[0155] d4=w1d1+w2d2+w3d3
[0156] Where d1, d2, d3, and d4 are the depths of P1, P2, P3, and P4 respectively.
[0157] After obtaining the depth of P4, calculate the spatial coordinates of the interpolation point in the world coordinate system:
[0158]
[0159] where R pin , t pin are the rotation matrix and position vector of the sub-image in the world coordinate system, with dimensions of 3x3 and 3x1 respectively.
[0160] That is, through the above process, the spatial position in the world coordinate system can be obtained.
[0161] Among them, the depth information of the known target projection point is used to estimate the depth of the unknown feature point. The interpolation can be linear or based on a more complex model, such as polynomial interpolation or barycentric coordinate interpolation. Figure 2D The barycentric coordinate interpolation process is shown in Figure 1. Depth d4 is calculated using the depth information of three known points, P1, P2, and P3, and the geometric relationship between these points and the unknown point, P4. Weights w1, w2, and w3 are calculated based on the relationship between the position of P4 in the 2D coordinate system and the positions of P1, P2, and P3 in the 2D coordinate system. These weights are then multiplied by the corresponding depths d1, d2, and d3, and the sum is used to estimate the depth of P4.
[0162] Among them, after mapping the initial feature points to the three-dimensional space, it is necessary to determine whether these feature points meet further verification conditions. The verification conditions may include the spatial relationship between the feature points and the surrounding environment, the distribution density of the feature points in the three-dimensional space, etc., and further screen out those feature points that are accurately positioned in the three-dimensional space and consistent with the geometric characteristics of the environment.
[0163] It can be seen that this method can effectively screen out candidate feature points with precise spatial positions from a large number of initial feature points, thereby improving the accuracy of feature point positioning.
[0164] In one possible example, judging whether the initial feature points meet the verification conditions based on the spatial positions of the initial feature points of each of the planar images includes: calculating the variance value of the spatial positions of the initial feature points of all the planar images; if the variance value is greater than a preset variance value, judging that the initial feature points do not meet the verification conditions, and eliminating the initial feature points; or, if the variance value is less than or equal to the preset variance value, judging that the initial feature points meet the verification conditions, and retaining the initial feature points.
[0165] Among them, the variance value P of the spatial position of each initial feature point under all viewing angles is calculated s The variance value is a statistic that measures the degree of data dispersion. Here, it measures the consistency of the spatial position of the same feature point under different viewing angles. The smaller the variance value, the more consistent the position of the feature point in each image, while the smaller the variance value, the greater the difference in position.
[0166] Specifically, the variance is calculated by comparing the spatial position of each feature point with its average position P a For each initial feature point, the three-dimensional position P obtained in all plane images is obtained. i The average value P a This average is calculated by summing the positions of the same feature points in each view and dividing by the number of views, where n is the number of views.
[0167] The preset variance value can be set manually or based on historical data and is not a single limit here. If the calculated variance value is greater than the preset variance threshold, it indicates that the position of the feature point varies significantly under different viewing angles. This may be due to mismatching, noise, or other disturbance factors. Therefore, this feature point is not considered stable and will be rejected by the system.
[0168] Among them, if the variance value is less than or equal to the preset threshold, it means that the position consistency of the feature point under different viewing angles is good and reliable, and it is retained as a valid feature point for subsequent processing.
[0169] Specifically, let 3d point P w It is a 3D point in the world coordinate system, which is observed by images taken from different points. These images are sub-images of the pinhole camera model, denoted by I0, I1, ..., I n , their corresponding poses in the world coordinate system are R0, t0, R1, t1..., R n ,t n , where R is a 3x3 rotation matrix and t is a 3x1 position vector. Let d i is sub-image I i For 3d point Pw The result of depth interpolation.
[0170] According to the following formula, P w Coordinates in the world coordinate system:
[0171]
[0172] Where (u i ,v i ) is image I i P w Observation, a total of n+1 three-dimensional coordinates P can be obtained 0~n ; Calculate P 0~n The mean and variance of :
[0173]
[0174]
[0175] If the variance P s If one of the dimensional components is greater than the threshold, the point is removed.
[0176] It can be seen that the variance-based verification process of this method provides an effective method to screen out feature points that are stable and consistent in spatial position. By eliminating feature points with inconsistent positions, the risk of mismatching can be reduced, ensuring the correct rendering and display of AR content.
[0177] In one possible example, determining the validity attribute of the closed area includes: obtaining the depth value of each target projection point in the closed area; calculating the difference in depth values between each two adjacent target projection points to obtain a depth difference; if all the depth differences are less than a preset difference, determining that the closed area has a valid attribute.
[0178] The depth value of the target projection point can be directly obtained by a depth sensor or calculated using a stereo vision algorithm using images taken from multiple angles. This is not a limitation. The depth value reflects the distance of each projection point relative to the observation point (usually the position of a camera or scanner).
[0179] Among them, the depth difference represents the relative height change of the target projection point in three-dimensional space, which can be used to determine the geometric continuity of the surface in the closed area.
[0180] The preset difference value may be set manually or by other settings, and is not limited here.
[0181] Specifically, if all calculated depth differences are less than a preset difference, then it can be considered that the surface changes in this closed area are smooth and there is no mutation, which means that the area corresponds to a part of the surface of the actual object and its spatial geometric information is coherent, so the area is marked as having valid attributes.
[0182] Specifically, if all depth differences exceed or are equal to the preset difference, it indicates that the surface within the closed area may be discontinuous, which may be caused by mismatched feature points, image noise or object edges. Therefore, this closed area may be deemed to have invalid attributes, and the feature points in the area may be eliminated by subsequent processes.
[0183] For example, Figure 2D The three vertices of the triangle in the image may be near the edge of a structure with a large depth difference. In this case, the depth inside the triangle is very inconsistent with the depth of the actual structure. To avoid this, calculate the depth difference between the three vertices, denoted as d 12 ,d 23 ,d 13 , set the threshold t d , a valid triangle must satisfy any d ij <t d , otherwise the triangle is considered invalid.
[0184] It can be seen that this method can effectively screen out geometrically stable and consistent feature points, providing a solid foundation for building reliable 3D models and AR scenes.
[0185] In one possible example, the position optimization processing of the candidate spatial positions to obtain the target spatial position includes: obtaining the candidate spatial positions of the candidate feature points on each of the sub-images; calculating the average value of all the candidate spatial positions of the candidate feature points to obtain the average spatial position of the candidate feature points; calculating the position deviation based on the average spatial position of the candidate feature points and each candidate spatial position; and calculating the target spatial position of the candidate feature points based on the average spatial position and the position deviation.
[0186] The candidate spatial positions may be obtained by converting the two-dimensional image coordinates of the candidate feature points into three-dimensional spatial coordinates, based on the camera calibration parameters, the scanner position information, and any known scene geometry constraints.
[0187] After obtaining the spatial position of each candidate feature point from each perspective, the arithmetic mean of these positions is calculated to obtain the average spatial position of each feature point. The average value helps reduce measurement errors or accidental deviations that may be caused by a single perspective, thereby obtaining a more stable and reliable spatial position estimate.
[0188] Among them, the position deviation is calculated by comparing the difference between each candidate spatial position and the average position. This deviation reflects the range of change in the spatial position of the feature point under different viewing angles.
[0189] Among them, the final position optimization processing is performed based on the average spatial position and position deviation. This process may involve using statistical models, filtering algorithms or mathematical methods to minimize errors to fine-tune the position of each feature point to ensure that the obtained target spatial position is as close as possible to the actual position of the feature point in the real world.
[0190] Specifically, the candidate feature points after verification are 3D points, and their mean value P is taken a As the initial value, the 3D point coordinates are optimized in combination with visual observation and the error is statistically analyzed.
[0191] 3d point P a To sub-image I i The projection equation is:
[0192]
[0193] Where (u a ,v a ) is the 3D coordinate P a In image I i The projection point on the , further summarize the above process into the projection function f:
[0194]
[0195] Compute the optimal 3D coordinates P by minimizing the reprojection error * :
[0196]
[0197] Based on the least squares method, the optimal coordinate P can be obtained. * , then calculate P * In each image I i If the reprojection errors on are all less than the threshold, then P * As the final position representation of the 3D point, otherwise it will be eliminated.
[0198] It can be seen that this method can determine an optimized spatial position for each feature point. This position is the most likely coordinate that best represents the actual spatial position of the feature point in a statistical sense.
[0199] In one possible example, in the AR map, each of the plane images is configured with a reference global feature vector, and determining the target pose of the target image captured by the camera based on the AR map includes: obtaining the target global feature vector of the target image; determining a plurality of reference global feature vectors that satisfy a preset similarity condition with the target global feature vector as a plurality of candidate global feature vectors, and the plane images corresponding to the candidate global feature vectors as candidate plane images; performing feature matching on each of the candidate plane images with the target image to obtain feature matching points; and determining the target pose based on the spatial position of each feature matching point in the world coordinate system.
[0200] The global feature vector is a high-dimensional representation of the image, typically containing key information extracted from the image, such as corners, edges, or other salient features. The global feature vector can be analyzed using feature extraction algorithms (such as SIFT, SURF, and ORB) to extract a global feature vector representing the image content.
[0201] The setting of the preset similarity condition may be based on the Euclidean distance, cosine similarity or other similarity metrics between feature vectors, which is not limited here.
[0202] The preset similarity condition can be satisfied when the global feature vector of the target image is compared with the reference global feature vector stored in the AR map. That is, the global feature vector of the target image is compared with the reference global feature vector stored in the AR map, and the reference feature vector with the highest similarity is selected as the candidate feature vector.
[0203] Feature matching can be achieved using algorithms such as brute force matching and FLANN matching. The goal is to find the best pair of corresponding points, i.e., feature matching points. Feature matching is performed on each candidate plane image and the target image to find the common feature points in the two images and establish matching pairs between them.
[0204] The preset pose estimation algorithm may be RANSAC (Random Sample Consensus) or PnP (Perspective-n-Point), etc., to determine the position and orientation of the camera. The preset pose estimation algorithm can estimate the 3D position and orientation of the camera relative to the reference plane image.
[0205] Specifically, the pose estimation algorithm requires the position of the feature points in the image and the known position of the feature points in the world coordinate system. Based on the feature matching points and their spatial positions, the algorithm can calculate the pose of the camera relative to the world coordinate system.
[0206] It can be seen that this method can understand the position and orientation of the scene captured by the camera in the physical world, so as to accurately superimpose the computer-generated image or 3D model on the user's actual environment.
[0207] It should be noted that, in each of the above-mentioned embodiments, there is not necessarily a certain order between the above-mentioned steps. A person skilled in the art can understand, based on the description of the embodiments of this application, that in different embodiments, the above-mentioned steps may have different execution orders, that is, they may be executed in parallel, or may be executed interchangeably, etc.
[0208] As another aspect of the embodiments of the present application, an embodiment of the present application provides a device for displaying AR content. The device for displaying AR content may be a software module, which includes several instructions stored in a memory. A processor may access the memory and call the instructions for execution to implement the AR content display method described in each of the above embodiments.
[0209] See also Figure 3 , Figure 3 Schematic diagram of the structure of an AR content display device provided in an embodiment of the present application. Figure 3 As shown, it is characterized in that the display device of the AR content includes:
[0210] An acquisition unit 301 is used to acquire 3D surround view images and point cloud data of the target space;
[0211] A segmentation unit 302 is configured to segment the 3D surround view image into a plurality of sub-images, wherein two adjacent sub-images have overlapping image areas;
[0212] A conversion unit 303 is configured to convert each sub-image into a planar image, wherein the planar image includes a plurality of initial feature points;
[0213] An optimization unit 304 is configured to perform verification and optimization processing on the initial feature points of the planar image according to the point cloud data to obtain target feature points;
[0214] A generating unit 305 is configured to generate a target AR map according to the target feature points;
[0215] The acquisition unit 301 is further configured to acquire a target image captured by a camera in the target space;
[0216] An acquisition unit 306 is configured to determine a target pose for the camera to acquire the target image according to the AR map;
[0217] A determination unit 307 is configured to determine a target AR content corresponding to the target posture;
[0218] The display unit 308 is further configured to display the target AR content at a position corresponding to the target posture.
[0219] This method enhances the continuity and integrity of the scene by dividing the 3D surround view image into multiple sub-images and ensuring that there are overlapping areas between them, thereby improving the accuracy of subsequent feature point extraction and matching. It further converts the sub-images into planar images and includes initial feature points, simplifying the image processing process and providing preliminary data for the accurate verification and optimization of feature points. It also uses point cloud data to verify and optimize the initial feature points of the planar image, significantly improving the reliability and accuracy of the feature points, laying the foundation for building high-quality AR maps. The target AR map generated based on precise feature points improves the alignment accuracy of virtual content and the real world in augmented reality, optimizes AR visual effects, improves the robustness and user experience of the AR display method, and ensures the spatial consistency and real-time performance of AR display content.
[0220] In one embodiment, the verification optimization processing includes verification processing and position optimization processing. In the verification optimization processing of the initial feature points of the plane image according to the point cloud data to obtain the target feature points, the optimization unit 304 is also used to: verify the initial feature points of the plane image according to the point cloud data to obtain candidate feature points that meet the verification conditions; obtain the candidate spatial positions of the candidate feature points in the world coordinate system; perform position optimization processing on the candidate spatial positions to obtain the target spatial positions; and generate a target AR map based on the target spatial positions of the candidate feature points.
[0221] In one embodiment, in the verification processing of the initial feature points of the plane image based on the point cloud data to obtain candidate feature points that meet the verification conditions, the optimization unit 304 is further used to: project the point cloud data onto the image coordinate system of the plane image to obtain a projected point cloud, wherein the projected point cloud includes a plurality of projection points; search for a plurality of projection points that meet the interpolation conditions with the initial feature points among the plurality of projection points as a plurality of target projection points, wherein the initial feature points are located in a closed area formed by the plurality of target projection points; determine the effectiveness attributes of the closed area; and verify the initial feature points of the plane image according to the effectiveness attributes of the closed area to obtain candidate feature points that meet the verification conditions.
[0222] In one embodiment, the validity attribute includes a valid attribute and an invalid attribute. In the verification processing of the initial feature points of the plane image according to the validity attribute of the closed area to obtain candidate feature points that meet the verification conditions, the optimization unit 304 is also used to: if the validity attribute of the closed area is a valid attribute, determine the depth of the initial feature point according to the positions of multiple target projection points of the closed area and the position of the initial feature point; determine the spatial position of the initial feature point in the world coordinate system according to a preset world transformation matrix, the position of the initial feature point and the depth; judge whether the initial feature point meets the verification conditions according to the spatial position of the initial feature point of each of the plane images; if the verification conditions are met, determine the initial feature point as a candidate feature point.
[0223] In one embodiment, in the step of determining whether the initial feature points satisfy the verification conditions based on the spatial positions of the initial feature points of each of the planar images, the optimization unit 304 is further used to: calculate the variance value of the spatial positions of the initial feature points of all the planar images; if the variance value is greater than a preset variance value, determine that the initial feature points do not satisfy the verification conditions, and eliminate the initial feature points; if the variance value is less than or equal to the preset variance value, determine that the initial feature points satisfy the verification conditions, and retain the initial feature points.
[0224] In one embodiment, in determining the effectiveness attribute of the closed area, the optimization unit 304 is further used to: obtain the depth value of each target projection point in the closed area; calculate the difference in depth values between each two adjacent target projection points to obtain a depth difference; if all the depth differences are less than a preset difference, determine that the closed area has a valid attribute.
[0225] In one embodiment, in the step of performing position optimization processing on the candidate spatial positions to obtain the target spatial positions, the optimization unit 304 is further used to: obtain the candidate spatial positions of the candidate feature points on each of the sub-images; calculate the average value of all candidate spatial positions of the candidate feature points to obtain the average spatial position of the candidate feature points; calculate the position deviation based on the average spatial position of the candidate feature points and each candidate spatial position; and calculate the target spatial position of the candidate feature points based on the average spatial position and the position deviation.
[0226] In one embodiment, in the AR map, each of the plane images is configured with a reference global feature vector. In determining the target pose of the target image captured by the camera based on the AR map, the generation unit 305 is further used to: obtain the target global feature vector of the target image; determine a plurality of reference global feature vectors that satisfy a preset similarity condition with the target global feature vector as a plurality of candidate global feature vectors, and the plane images corresponding to the candidate global feature vectors are candidate plane images; perform feature matching on each of the candidate plane images with the target image to obtain feature matching points; and determine the target pose based on a preset pose estimation algorithm and the spatial position of each feature matching point in the world coordinate system.
[0227] The display device for AR content can also be built from hardware devices. For example, the display device for AR content can be built from one or more chips, and the chips can work in coordination with each other to complete the display methods of AR content described in the above embodiments. For another example, the display device for AR content can also be built from various logic devices, such as a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a single-chip microcomputer, an ARM (Acorn RISC Machine) or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination of these components.
[0228] It should be noted that the above-mentioned AR content display device can execute the AR content display method provided in the embodiments of this application, and has the corresponding functional modules and beneficial effects of the execution method. For technical details not fully described in the embodiments of the AR content display device, please refer to the AR content display method provided in the embodiments of this application.
[0229] See also Figure 4 , Figure 4 4 is a schematic diagram of the structure of a computer device provided in an embodiment of the present application. The computer device includes one or more processors 41 and a memory 42. The memory 42 is connected to the one or more processors 41, for example, via a bus.
[0230] The processor 41 is configured to support the computer device in executing the corresponding functions of the method in the above method embodiment. The processor 41 can be a central processing unit (CPU), a network processor (NP), a hardware chip, or any combination thereof. The above hardware chip can be an application specific integrated circuit (ASIC), a programmable logic device (PLD), or a combination thereof. The above PLD can be a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.
[0231] Memory 42 is used to store program code, etc. Memory 42 may include volatile memory (VM), such as random access memory (RAM); non-volatile memory (NVM), such as read-only memory (ROM), flash memory, hard disk drive (HDD), or solid-state drive (SSD); or a combination of the aforementioned types of memory.
[0232] Memory 42 can be used to store non-volatile software programs, non-volatile computer executable programs, and modules, such as the program instructions / modules corresponding to the vehicle diagnostic method in the embodiments of the present application. Processor 41 executes the non-volatile software programs, instructions, and modules stored in memory to execute various functional applications and data processing of the vehicle diagnostic method and computer device, thereby implementing the functions of the vehicle diagnostic method and various modules or units of the computer device provided in the above-mentioned method embodiments.
[0233] Memory 42 may include a program storage area and a data storage area. The program storage area may store an operating system and application programs required for at least one function. The data storage area may store data generated based on the use of the vehicle diagnostic device. In some embodiments, memory 42 may optionally include a remote memory device located relative to the processor. Such remote memory device may be connected to the vehicle diagnostic device via a network. Examples of such networks include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0234] The one or more modules are stored in the memory, and when executed by the one or more processors, the display method of AR content in any of the above method embodiments is executed, for example, the method steps described in the above method embodiments are executed to realize the functions of the modules described in the above device embodiments.
[0235] An embodiment of the present application further provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, wherein the computer program includes program instructions, and when the program instructions are executed by a computer, the computer executes the method as described in the above embodiment.
[0236] Those skilled in the art will appreciate that all or part of the processes in the above-described method embodiments can be implemented by instructing related hardware through a computer program. The program can be stored in a computer-readable storage medium, and when executed, the program can include the processes in the above-described method embodiments. The storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM).
[0237] The above disclosure is only a preferred embodiment of the present application, and certainly cannot be used to limit the scope of rights of the present application. Therefore, equivalent changes made according to the claims of the present application are still within the scope covered by the present application.
Claims
1. A method for displaying AR content, characterized in that: include: Obtain 3D surround image and point cloud data of the target space; Splitting the 3D surround view image into a plurality of sub-images, wherein two adjacent sub-images have overlapping image areas; Converting each sub-image into a planar image, wherein the planar image includes a plurality of initial feature points; Performing verification and optimization processing on the initial feature points of the plane image according to the point cloud data to obtain target feature points; Generate a target AR map according to the target feature points; Acquire a target image captured by a camera in the target space; Determining a target pose of the target image captured by the camera according to the AR map; Determining target AR content corresponding to the target pose; The target AR content is displayed at a position corresponding to the target posture.
2. The display method according to claim 1, wherein: The verification optimization process includes verification processing and position optimization processing, and the verification optimization process is performed on the initial feature points of the plane image according to the point cloud data to obtain target feature points, including: Performing verification processing on the initial feature points of the planar image according to the point cloud data to obtain candidate feature points that meet verification conditions; Obtaining the candidate spatial position of the candidate feature point in the world coordinate system; Performing position optimization processing on the candidate spatial positions to obtain a target spatial position; A target AR map is generated according to the target spatial positions of the candidate feature points.
3. The display method according to claim 2, wherein: The verifying process of the initial feature points of the planar image according to the point cloud data to obtain candidate feature points that meet verification conditions includes: Projecting the point cloud data onto the image coordinate system of the plane image to obtain a projected point cloud, wherein the projected point cloud includes a plurality of projection points; Searching for multiple projection points that satisfy the interpolation condition together with the initial feature point among the multiple projection points as multiple target projection points, wherein the initial feature point is located in a closed area formed by the multiple target projection points; determining a validity attribute of the enclosed region; The initial feature points of the planar image are verified according to the validity attributes of the closed area to obtain candidate feature points that meet the verification conditions.
4. The display method according to claim 3, wherein: The validity attributes include valid attributes and invalid attributes. The initial feature points of the planar image are verified according to the validity attributes of the closed area to obtain candidate feature points that meet the verification conditions, including: If the effectiveness attribute of the closed area is a valid attribute, determining the depth of the initial feature point according to the positions of the plurality of target projection points of the closed area and the position of the initial feature point; Determining the spatial position of the initial feature point in a world coordinate system according to a preset world transformation matrix, the position of the initial feature point, and the depth; determining, based on the spatial positions of the initial feature points of each of the planar images, whether the initial feature points meet a verification condition; If the verification condition is met, the initial feature point is determined to be a candidate feature point.
5. The display method according to claim 4, wherein: The determining, based on the spatial positions of the initial feature points of each of the planar images, whether the initial feature points meet a verification condition includes: Calculating the variance values of the spatial positions of the initial feature points of all the planar images; If the variance value is greater than the preset variance value, it is determined that the initial feature point does not meet the verification condition and the initial feature point is removed; If the variance value is less than or equal to the preset variance value, it is determined that the initial feature point meets the verification condition, and the initial feature point is retained.
6. The display method according to claim 3, wherein: Determining the effectiveness attribute of the closed area includes: Obtaining the depth value of each target projection point in the closed area; Calculate the difference in depth between each two adjacent target projection points to obtain the depth difference; If all the depth differences are smaller than a preset difference, the closed area is determined to be a valid attribute.
7. The display method according to claim 2, wherein: The performing position optimization processing on the candidate spatial position to obtain the target spatial position includes: Obtaining candidate spatial positions of the candidate feature points on each of the sub-images; Calculating the average of all candidate spatial positions of the candidate feature points to obtain the average spatial position of the candidate feature points; Calculating a position deviation based on the average spatial position of the candidate feature points and each candidate spatial position; The target spatial position of the candidate feature point is calculated according to the average spatial position and the position deviation.
8. The display method according to claim 1, wherein: In the AR map, each of the planar images is configured with a reference global feature vector, and determining the target pose of the target image captured by the camera according to the AR map includes: Obtaining a target global feature vector of the target image; Determine a reference global feature vector that satisfies a preset similarity condition with the target global feature vector as a candidate global feature vector, and a plane image corresponding to the candidate global feature vector as a candidate plane image; Perform feature matching on each candidate plane image and the target image to obtain feature matching points; The target pose is determined based on a preset pose estimation algorithm and the spatial position of each feature matching point in the world coordinate system.
9. A display device for AR content, characterized in that: The device comprises: An acquisition unit, used to acquire 3D surround view images and point cloud data of the target space; a segmentation unit, configured to segment the 3D surround view image into a plurality of sub-images, wherein two adjacent sub-images have overlapping image areas; a conversion unit, configured to convert each sub-image into a planar image, wherein the planar image includes a plurality of initial feature points; an optimization unit, configured to perform verification and optimization processing on the initial feature points of the planar image according to the point cloud data to obtain target feature points; A generating unit, configured to generate a target AR map according to the target feature points; The acquisition unit is further configured to acquire a target image captured by the camera in the target space; An acquisition unit, configured to determine a target pose for the camera to acquire the target image according to the AR map; a determination unit, configured to determine a target AR content corresponding to the target posture; The display unit is further configured to display the target AR content at a position corresponding to the target posture.
10. A computer device, characterized in that: The computer device comprises a memory and a processor, wherein the memory is connected to the processor, and the processor is used to execute one or more computer programs stored in the memory. When the processor executes the one or more computer programs, the computer device implements the method according to any one of claims 1 to 8.
11. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, wherein the computer program includes program instructions. When the program instructions are executed by a processor, the processor is caused to perform the method according to any one of claims 1 to 8.