Planar detection method for augmented reality devices, augmented reality devices and media

By acquiring and processing environmental image data and sensor data from AR devices, generating map points and performing semantic segmentation, the accuracy and speed issues of AR devices in detecting multiple adjacent planes are solved, achieving more efficient plane detection.

CN122313342APending Publication Date: 2026-06-30HANGZHOU LINGBAN TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HANGZHOU LINGBAN TECH CO LTD
Filing Date
2026-03-19
Publication Date
2026-06-30

AI Technical Summary

Technical Problem

Existing AR devices have low accuracy and long processing time when detecting multiple adjacent planes. Monocular SLAM algorithms require a long time to move and generate map points in order to update the plane information.

Method used

By acquiring environmental image data, historical environmental image data, and sensor data, map points are generated and semantic segmentation is performed. Planar information is generated using a pre-set database, and plane detection is performed by combining map points and image mask data.

Benefits of technology

It improves the accuracy and speed of plane detection and reduces detection time, especially when there are multiple adjacent planes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122313342A_ABST
    Figure CN122313342A_ABST
Patent Text Reader

Abstract

This disclosure provides embodiments of a planar detection method, augmented reality device, and medium for augmented reality devices. One specific implementation of the method includes: acquiring environmental image data, historical environmental image data, and sensor data; generating map points based on the environmental image data, historical environmental image data, and sensor data; performing semantic segmentation processing on the environmental image data based on the map points to obtain image mask data; and generating planar information based on a preset database, the map points, and the image mask data. This implementation can improve the speed and accuracy of planar detection and reduce the time spent on planar detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of this disclosure relate to the field of computer technology, and more specifically to a planar detection method for augmented reality devices, augmented reality devices, and media. Background Technology

[0002] Plane detection is a method used by AR devices to detect planar objects (walls, desktops, etc.) in the surrounding environment, in order to adjust or interact with virtual objects associated with the planar structure. Currently, the common approach to plane detection is for AR devices to directly generate map points using a monocular SLAM algorithm to detect or update planes in the surrounding environment.

[0003] However, when using the above method for planar detection, the following technical problems often arise: When multiple adjacent planes exist, the accuracy of plane detection can be low. Furthermore, when using monocular SLAM algorithms to detect or update planes, it is necessary to wait for the AR device to move for a long time to generate a large number of map points before the plane information can be detected or updated based on the generated map points, which can easily lead to a long time consumption during plane detection.

[0004] The information disclosed in this background section is only intended to enhance the understanding of the background of the present disclosure concept, and therefore may contain information that does not constitute prior art known to those skilled in the art. Summary of the Invention

[0005] The summary portion of this disclosure is intended to provide a brief overview of the concepts, which will be described in detail in the detailed description portion. This summary portion is not intended to identify key or essential features of the claimed technical solutions, nor is it intended to limit the scope of the claimed technical solutions.

[0006] Some embodiments of this disclosure provide a planar detection method for augmented reality devices, augmented reality devices, and computer-readable media to address one or more of the technical problems mentioned in the background section above.

[0007] In a first aspect, some embodiments of this disclosure provide a planar detection method for an augmented reality device. The method includes: acquiring environmental image data, historical environmental image data, and sensor data; generating map points based on the environmental image data, the historical environmental image data, and the sensor data; performing semantic segmentation processing on the environmental image data based on the map points to obtain image mask data; and generating planar information based on a preset database, the map points, and the image mask data.

[0008] Secondly, some embodiments of this disclosure provide an augmented reality device, including: one or more display screens for imaging in front of a user; one or more processors; and a storage device storing one or more programs that, when executed by one or more processors, cause the one or more processors to implement the method described in any implementation of the first aspect above.

[0009] Thirdly, some embodiments of this disclosure provide a computer-readable medium having a computer program stored thereon, wherein the program, when executed by a processor, implements the method described in any of the implementations of the first aspect above.

[0010] The various embodiments of this disclosure have the following beneficial effects: the plane detection method for augmented reality devices according to some embodiments of this disclosure can improve the accuracy of plane detection, increase the speed of plane detection, and reduce the time spent on plane detection. Specifically, the reason for the long time spent on plane detection is that when there are multiple adjacent planes, the accuracy of plane detection is easily reduced. Moreover, when using a monocular SLAM algorithm to detect or update planes, it is necessary to wait for the AR device to move for a long time to generate a large number of map points, and then detect or update the plane information based on the generated map points, which easily leads to a long time spent on plane detection. Based on this, the plane detection method for augmented reality devices according to some embodiments of this disclosure first acquires environmental image data, historical environmental image data, and sensor data. Thus, the raw data to be processed can be obtained. Second, based on the above-mentioned environmental image data, the above-mentioned historical environmental image data, and the above-mentioned sensor data, various map points are generated. Thus, device pose data and various map points can be obtained. Then, based on the above-mentioned map points, semantic segmentation processing is performed on the above-mentioned environmental image data to obtain various image mask data. Thus, various image mask data can be obtained. Finally, based on the preset database, the aforementioned map points, and the aforementioned image mask data, various planar information is generated. Thus, the planar information can be obtained. Because plans can be detected jointly using the generated map points and image mask data, when adjacent plans exist, the semantic labels of different plans can be further determined using the image mask data, thereby improving the accuracy of planar detection. Furthermore, because planar information can be detected or updated by combining map points and image mask data during planar detection, the number of map points required for detection can be reduced, thus reducing the movement time of the AR device and consequently reducing the time spent on planar detection. Attached Figure Description

[0011] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and elements are not necessarily drawn to scale.

[0012] Figure 1 This is a flowchart of some embodiments of a planar detection method for augmented reality devices according to the present disclosure; Figure 2 These are flowcharts of other embodiments of the planar detection method for augmented reality devices according to this disclosure; Figure 3 This is a schematic diagram of the structure of an augmented reality device suitable for implementing some embodiments of the present disclosure. Detailed Implementation

[0013] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0014] It should also be noted that, for ease of description, only the parts relevant to the invention are shown in the accompanying drawings. Unless otherwise specified, the embodiments and features described in this disclosure can be combined with each other.

[0015] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.

[0016] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".

[0017] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.

[0018] This disclosure will now be described in detail with reference to the accompanying drawings and embodiments.

[0019] Figure 1A flow 100 of some embodiments of a plane detection method for augmented reality devices according to the present disclosure is shown. The plane detection method for augmented reality devices includes the following steps: Step 101: Acquire environmental image data, historical environmental image data, and sensor data.

[0020] In some embodiments, the execution entity (e.g., a computing device) of the planar detection method for an augmented reality device can acquire environmental image data, historical environmental image data, and sensor data. The augmented reality device can be a device capable of combining a computer-generated virtual environment with a real environment. For example, the augmented reality device can be AR glasses. The environmental image data can be images obtained after photographing a target environmental area. The environmental image data corresponds to a time point. The time point corresponding to the environmental image data can be the time point at which the environmental image data was generated. The target environmental area can be a pre-defined, real environment. The historical environmental image data can be images obtained when photographing the target environmental area before photographing the environmental image data. The historical environmental image data corresponds to a time point. The time point corresponding to the historical environmental image data can be the time point at which the historical environmental image data was generated. The sensor data can be data collected by the inertial measurement unit in the augmented reality device when the augmented reality device moves. The sensor data can include angular velocity data and acceleration data. The angular velocity data can be vectors of the angular velocities of the augmented reality device in the X, Y, and Z directions of the inertial unit coordinate system, respectively, when the device moves. The aforementioned acceleration data can be vectors of acceleration of the augmented reality device in the X, Y, and Z directions of the inertial unit coordinate system during movement. The aforementioned sensor data corresponds to a time point. The time point corresponding to the aforementioned sensor data can characterize the time when the aforementioned sensor data was generated. The aforementioned inertial unit coordinate system can be the reference coordinate system corresponding to the inertial measurement unit used to acquire the aforementioned sensor data. The aforementioned inertial measurement unit can be mounted on the aforementioned augmented reality device.

[0021] For example, the aforementioned inertial unit coordinate system can be a three-dimensional coordinate system with the center of gravity of the augmented reality device as the origin, the horizontal direction to the right as the x-axis (i.e., the X-axis), the forward direction as the y-axis (i.e., the Y-axis), and the vertical direction upwards perpendicular to the ground as the z-axis (i.e., the Z-axis). The aforementioned execution entity can be a server.

[0022] As an example, the aforementioned executing entity can send a preset AR data acquisition instruction to the aforementioned augmented reality device. This AR data acquisition instruction can be an instruction for acquiring the aforementioned environmental image data, the aforementioned historical environmental image data, and the aforementioned sensor data. Then, it can receive the environmental image data, historical environmental image data, and sensor data sent by the aforementioned augmented reality device.

[0023] Step 102: Generate map points based on environmental image data, historical environmental image data, and sensor data.

[0024] In some embodiments, the executing entity may generate various map points based on the environmental image data, the historical environmental image data, and the sensor data. Each map point can be the three-dimensional coordinates of a pixel in the environmental image data within a camera coordinate system. The camera coordinate system can be the camera coordinate system corresponding to the camera in the augmented reality device at the current moment.

[0025] In practice, the aforementioned environmental image data and sensor data can first be input into a visual inertial odometry system (VIOS) to obtain a 3x4 matrix as the environmental pose matrix. This environmental pose matrix can be a matrix that maps the world coordinate system to the camera coordinate system corresponding to the environmental image data. The VINSS can be either VINS-Mono or VINS-Fusion.

[0026] Next, the preset acquisition information and the aforementioned historical environmental image data are sent to the target terminal. The target terminal can be a terminal used by a technician. The acquisition information can be sensor data used to acquire the historical environmental image data. Then, the sensor data sent by the target terminal can be received as historical sensor data.

[0027] Next, the aforementioned historical environment image data and historical sensor data can be input into the aforementioned visual inertial odometry to obtain a 3x4 matrix as the historical environment pose matrix. This historical environment pose matrix can be a matrix capable of mapping the world coordinate system to the camera coordinate system corresponding to the aforementioned historical environment image data.

[0028] Then, the submatrices corresponding to the first three rows and first three columns of the above environment pose matrix can be determined as the environment rotation matrix. The column vector of the fourth column of the above environment pose matrix can be determined as the environment vector. Then, the submatrices corresponding to the first three rows and first three columns of the above historical environment pose matrix can be determined as the historical environment rotation matrix. The column vector of the fourth column of the above historical environment pose matrix can be determined as the historical environment vector.

[0029] Then, the transpose of the aforementioned environment rotation matrix can be determined as the environment transpose matrix. Then, the product of the aforementioned environment transpose matrix and the aforementioned historical environment rotation matrix can be determined as the camera rotation matrix.

[0030] Then, the difference between the historical environment vector and the environment vector can be used to determine the vector to be processed. The product of the transpose of the environment matrix and the vector to be processed can be used to determine the camera translation vector.

[0031] Then, the ORB algorithm (Oriented Fast and Rotated BRIEF) can be used to extract feature points from the aforementioned environmental image data and historical environmental image data, respectively. This yields individual feature points from the environmental image data as environmental feature points and individual feature points from the historical environmental image data as historical environmental feature points. Each of these feature points corresponds to a descriptor and pixel coordinates.

[0032] Next, for each of the aforementioned environmental feature points, the corresponding descriptor can be determined as the environmental descriptor. Then, a matching technique can be used to find the historical environmental feature point whose descriptor has the highest similarity to the aforementioned environmental descriptor, as the target historical environmental feature point. Finally, the aforementioned environmental feature point and the target historical environmental feature point can be combined into an environmental feature point pair. The matching technique described above can be any technique capable of matching two feature points. For example, the matching technique could be a brute-force matcher.

[0033] Then, for each of the obtained environmental feature point pairs, the pixel coordinates corresponding to the environmental feature point and the target historical environmental feature point in the above environmental feature point pair can be determined as environmental coordinates and historical environmental coordinates, respectively. The above environmental coordinates and the above historical environmental coordinates can be combined into environmental coordinate data.

[0034] Then, the inverse of the preset camera intrinsic parameter matrix can be determined as the intrinsic parameter inverse matrix. Here, the aforementioned camera intrinsic parameter matrix can be a third-order intrinsic parameter matrix corresponding to the camera in the augmented reality device.

[0035] Then, for each environmental coordinate data point obtained, a preset coordinate value can be added to both the environmental coordinates and historical environmental coordinates included in the environmental coordinate data. This results in environmental coordinates with the added preset coordinate value used as environmental 3D coordinates, and historical environmental coordinates with the added preset coordinate value used as historical environmental 3D coordinates. The preset coordinate value can be 1. For example, when the environmental coordinates are (2,3), the environmental 3D coordinates can be (2,3,1). Then, the product of the intrinsic parameter inverse matrix and the transpose of the environmental 3D coordinates can be used to determine the normalized coordinates. The product of the intrinsic parameter inverse matrix and the transpose of the historical environmental 3D coordinates can be used to determine the historical normalized coordinates. Finally, the normalized coordinates and the historical normalized coordinates can be combined to form normalized coordinate data.

[0036] Then, for each normalized coordinate data in the obtained normalized coordinate data, the first preset depth symbol, the aforementioned camera rotation matrix, and the product of the normalized coordinates included in the aforementioned normalized coordinate data can be determined as the first depth term number. The aforementioned first preset depth symbol can be the string "z1". Then, the sum of the aforementioned first depth term number and the aforementioned camera translation vector can be determined as the depth polynomial.

[0037] Then, the product of the second preset depth symbol and the historical normalized coordinates included in the above normalized coordinate data can be determined as the second depth term. The second preset depth symbol can be the string "z2".

[0038] Then, the number of the second depth term can be equal to the depth polynomial to construct an equation as the image depth equation. A fitting algorithm can then be used to fit this image depth equation to obtain the numerical value corresponding to the first preset depth symbol, which is taken as the coordinate depth value. The fitting algorithm can be any algorithm capable of fitting the equation. For example, the fitting algorithm can be the least squares method. Finally, the product of the coordinate depth value and the normalized coordinates included in the normalized coordinate data can be used to determine the map point.

[0039] Step 103: Based on each map point, perform semantic segmentation on the environmental image data to obtain each image mask data.

[0040] In some embodiments, the execution entity can perform semantic segmentation processing on the environmental image data based on the map points to obtain various image mask data. Each image mask data can be a binary matrix of the same size as the environmental image data, where each element value represents the category of the corresponding pixel. The element values ​​in each image mask data can be 0 or 1. An element value of "0" in the image mask data can represent the category of the pixel corresponding to that element value as "background." An element value of "1" in the image mask data can represent the category of the pixel corresponding to that element value as the semantic label of the image mask data. The semantic label can represent the category of the pixel corresponding to the element value of "1" in the image mask data. For example, the semantic label can be "table," "wall," or "ground." Each image mask data corresponds to bounding box data, a semantic label, and a confidence score. The bounding box data can be data used to describe the bounding box of the object corresponding to the semantic label in the environmental image data. The bounding box data can include the coordinates of the upper left corner and the lower right corner. The top-left corner coordinates mentioned above can be the pixel coordinates of the top-left vertex of the bounding box in the environment image data. The bottom-right corner coordinates mentioned above can be the pixel coordinates of the bottom-right vertex of the bounding box in the environment image data.

[0041] In practice, firstly, in response to the determination that the number of each of the aforementioned map points exceeds a preset map point threshold, the aforementioned environmental image data can be input into a pre-trained semantic segmentation model to obtain image mask data for each point. The aforementioned semantic segmentation model can be a pre-trained Mask R-CNN. The aforementioned pre-training can be a process of fine-tuning Mask R-CNN using a sample dataset and a cross-entropy loss function. Each sample data in the aforementioned sample dataset can be used to train Mask R-CNN. Each sample data in the aforementioned sample dataset may include sample environmental image data, semantic labels for each sample, image mask data for each sample, and bounding box data for each sample.

[0042] The aforementioned sample environment image data can be the environment image data used to train the aforementioned semantic segmentation model. Each sample semantic label in the aforementioned sample semantic labels can be a label representing the category corresponding to an object in the sample environment image data. Each sample image mask data in the aforementioned sample image mask data can be the image mask data corresponding to an object in the sample environment image data. Each sample bounding box data in the aforementioned sample bounding box data can be the bounding box data corresponding to the sample environment image data.

[0043] The aforementioned preset map point threshold can be a pre-defined value. Here, the specific setting of the aforementioned preset map point threshold is not limited.

[0044] Optionally, after step 103, the aforementioned executing entity may also perform the following steps: The first step involves performing semantic segmentation on the environmental image data in response to determining that the number of each map point is less than or equal to a preset map point threshold, thereby obtaining image mask data. In practice, in response to determining that the number of each map point is less than or equal to the preset map point threshold, the environmental image data can first be input into the semantic segmentation model to obtain image mask data.

[0045] The second step is to perform the following steps for each image mask data in the above image mask data: The first sub-step generates target historical plane information, image mask coordinates, and mask depth values ​​based on a preset camera intrinsic parameter matrix, a preset database, and the aforementioned image mask data. The aforementioned camera intrinsic parameter matrix can be a third-order intrinsic parameter matrix corresponding to the camera in the augmented reality device.

[0046] The aforementioned preset database can be a database used to store previously generated planar information. Each planar information stored in the aforementioned preset database can be a planar equation obtained by fitting various historical map points. Each planar information stored in the aforementioned preset database corresponds to various historical map points, semantic tags, and historical environmental image data. Each historical map point can be a previously generated map point used to fit and generate the corresponding planar information.

[0047] The aforementioned target historical planar information can be planar information stored in the aforementioned preset database whose corresponding semantic tags are the same as the semantic tags of the aforementioned image mask data.

[0048] Each of the above image mask coordinates can be the pixel coordinate of the element value "1" in the image mask data in the above environmental image data. The above mask depth value can be a numerical value used to characterize the depth corresponding to each of the above image mask coordinates.

[0049] In practice, firstly, the planar information of the semantic tags corresponding to the preset database and the semantic tags of the above image mask data can be determined as the target historical planar information.

[0050] Secondly, the pixel coordinates corresponding to each element value of "1" in the above image mask data in the above environmental image data can be determined as the image mask coordinates. For example, when the element value in the third row and fourth column is "1", the image mask coordinates are (3,4).

[0051] Then, the element values ​​in the first row and first column, and the second row and second column of the camera intrinsic parameter matrix can be determined as the first focal length and the second focal length, respectively. Similarly, the element values ​​in the first row and third column, and the second row and third column of the camera intrinsic parameter matrix can be determined as the first principal point value and the second principal point value, respectively.

[0052] Then, for each of the aforementioned image mask coordinates, firstly, the historical map points corresponding to the aforementioned target historical plane information can be determined as the respective target historical map points. Secondly, for each of the aforementioned target historical map points, the third value among the aforementioned target historical map points can be determined as the target historical depth value. For example, when the target historical map point is (1,2,3), the target historical depth value is "3". Finally, the average of the determined target historical depth values ​​can be determined as the mask depth value corresponding to the aforementioned image mask coordinates.

[0053] The second sub-step involves performing back-projection processing on the image mask coordinates based on the aforementioned camera intrinsic parameter matrix and mask depth values ​​to obtain the three-dimensional coordinates of each mask. Each of these three-dimensional mask coordinates can be a three-dimensional coordinate obtained by projecting the image mask coordinates onto the camera coordinate system.

[0054] In practice, firstly, based on the aforementioned camera intrinsic parameter matrix, the first principal point values ​​and second principal point values ​​of the first focal length and second focal length corresponding to the aforementioned camera intrinsic parameter matrix can be determined. The method for determining the first principal point values ​​and second principal point values ​​of the first focal length and second focal length corresponding to the aforementioned camera intrinsic parameter matrix can be referred to the specific implementation method in step 103, and will not be repeated here.

[0055] Then, for each of the aforementioned image mask coordinates, the difference between the first value and the first principal point value can be determined as a first difference. The ratio of the first difference to the first focal length can be determined as a first ratio. The product of the first ratio and the mask depth value can be determined as the mask's horizontal coordinate value. Then, the difference between the second value and the second principal point value can be determined as a second difference. The ratio of the second difference to the second focal length can be determined as a second ratio. The product of the second ratio and the mask depth value can be determined as the mask's vertical coordinate value. Finally, the mask's horizontal coordinate value, vertical coordinate value, and depth value can be combined to form a three-dimensional coordinate system as the mask's three-dimensional coordinate system.

[0056] The third sub-step involves semantically updating the historical plane information of the target based on the three-dimensional coordinates of each of the aforementioned masks.

[0057] In practice, firstly, the obtained 3D coordinates of each mask and the aforementioned historical map points of each target can be combined into a set as the 3D coordinate set to be updated. Then, for each 3D coordinate in the aforementioned set to be updated, the coordinate can be substituted into the general equation of the plane to obtain the equation to be updated. The general equation of the plane can be a general form equation of the plane, i.e., "Ax + By + Cz + D = 0". "A", "B", "C", and "D" are constants, and "A", "B", and "C" are not all 0 simultaneously. "x", "y", and "z" can be the unknowns that need to be substituted into the solution.

[0058] Then, the obtained equations to be updated can be fitted using an equation fitting algorithm to obtain the values ​​corresponding to "A", "B", "C", and "D" in the aforementioned general plane equation. These values ​​can then be substituted into the general plane equation to obtain the resulting general plane equation, which serves as the target historical plane information for updating the target historical plane information in the preset database. The equation fitting algorithm can be one capable of fitting polynomials. For example, the least squares method could be used.

[0059] Step 104: Generate various planar information based on the preset database, various map points, and various image mask data.

[0060] In some embodiments, the execution entity may generate various planar information based on a preset database, the various map points, and the various image mask data.

[0061] Each of the aforementioned planar information can be a three-dimensional planar equation obtained by fitting various map points.

[0062] In some optional implementations of certain embodiments, the execution entity may generate various planar information based on a preset database, the aforementioned map points, and the aforementioned image mask data through the following steps: The first step is to determine that each of the above image mask data meets the preset valid mask conditions. For each of the above map points, the map point is matched with the plane information stored in the preset database to obtain a matching result group.

[0063] The aforementioned valid mask condition requires that at least one valid image mask data exists among the aforementioned image mask data. The aforementioned valid image mask data can be image mask data with a confidence level greater than a preset confidence threshold. The aforementioned preset confidence threshold can be a pre-defined value. Here, the specific setting of the aforementioned preset confidence threshold is not limited. Each matching result in the aforementioned matching result group can be a label used to characterize whether a map point matches various planar information stored in a preset database. For example, the matching result can be "match" or "not match". Each matching result in the aforementioned matching result group corresponds to one map point and one planar information.

[0064] In practice, the aforementioned map points and a preset camera intrinsic parameter matrix can be input into a projection function to obtain the corresponding two-dimensional map points, which are then used as the output coordinates of the projection function. The projection function can be any function capable of projecting three-dimensional coordinates onto a two-dimensional plane. For example, the projection function could be `cv2.projectPoints()`.

[0065] Secondly, for each of the aforementioned two-dimensional map points and for each of the aforementioned image mask data, the category corresponding to the two-dimensional map point in the aforementioned image mask data can be determined as the pixel category. The confidence level corresponding to the aforementioned image mask data can be determined as the confidence level of the aforementioned pixel category.

[0066] As an example, when the 2D map point is (1,2), the element value of the first row and second column of the image mask data can be determined as the mask element value. When the mask element value is 0, the pixel category is "background"; when the mask element value is 1 and the semantic label corresponding to the image mask data is "desktop", the pixel category is "desktop". Then, the pixel category with the highest confidence among the determined pixel categories can be determined as the target pixel category. Then, for each piece of planar information stored in the preset database, in response to determining that the semantic label corresponding to the above planar information is the same as the target pixel category, "match" can be determined as the matching result between the above map point and the above planar information. In response to determining that the semantic label corresponding to the above planar information is not the same as the target pixel category, "no match" can be determined as the matching result between the above map point and the above planar information.

[0067] The second step involves generating planar information based on the aforementioned image mask data and map points, in response to the determination that the obtained matching result groups do not meet the preset planar expansion conditions. The planar expansion conditions can be that any one of the matching results included in the aforementioned matching result groups is a "match".

[0068] In some optional implementations of certain embodiments, the aforementioned execution entity may generate various planar information based on the aforementioned image mask data and map points through the following steps: The first step is to perform coordinate projection processing on each of the aforementioned map points to obtain camera map points. These camera map points can be two-dimensional coordinates obtained by projecting the map points onto the aforementioned environmental image data.

[0069] In practice, for each of the above map points, the first, second, and third values ​​can be determined as the horizontal coordinate, vertical coordinate, and depth of the map point, respectively.

[0070] Secondly, the element values ​​in the first row and first column, and the second row and second column of the preset camera intrinsic parameter matrix can be determined as the horizontal axis focal length and the vertical axis focal length, respectively. The element values ​​in the first row and third column, and the second row and third column of the same camera intrinsic parameter matrix can be determined as the principal point's horizontal coordinate value and the principal point's vertical coordinate value, respectively. The aforementioned camera intrinsic parameter matrix can be a third-order intrinsic parameter matrix corresponding to the camera in the augmented reality device.

[0071] Then, the ratio of the above map point's x-coordinate value to the above map point's depth value can be determined as the map point's x-axis ratio. Then, the product of the above map point's x-axis ratio and the above x-axis focal length can be determined as the map point's x-axis product. Then, the sum of the above map point's x-axis product and the above principal point's x-coordinate value can be determined as the camera's x-coordinate value.

[0072] Next, the ratio of the map point's ordinate value to its depth value can be determined as the map point's ordinate ratio. Then, the product of the map point's ordinate ratio and its focal length can be determined as the map point's ordinate product. Finally, the sum of the map point's ordinate product and the principal point's ordinate value can be determined as the camera's ordinate value.

[0073] Finally, the above camera x-coordinate values ​​and y-coordinate values ​​can be combined into two-dimensional coordinates to serve as camera map points.

[0074] The second step is to determine the image mask data that meets the preset confidence level condition from the aforementioned image mask data as each valid mask data. The aforementioned confidence level condition can be that the confidence level corresponding to the image mask data is greater than the preset confidence threshold.

[0075] Third, for each valid mask data in the above valid mask data, perform the following steps: The first sub-step involves clustering the obtained camera map points based on the aforementioned effective mask data to obtain individual target camera map points. Each target camera map point can be a camera map point whose corresponding category is a semantic label associated with the aforementioned effective mask data.

[0076] In practice, firstly, the element values ​​among the various element values ​​included in the aforementioned effective mask data that satisfy the preset target element conditions can be determined as the target element values. The preset target element conditions can be that the category corresponding to the element value is the semantic label corresponding to the aforementioned effective mask data, i.e., the element value is "1". Secondly, the pixel coordinates corresponding to the aforementioned target element values ​​in the aforementioned environmental image data can be determined as the target pixel coordinates. Then, the camera map points among the aforementioned camera map points that are identical to any one of the aforementioned target pixel coordinates can be determined as the target camera map points.

[0077] The second sub-step involves fitting the map points of each target camera to obtain planar information. This planar information can be the plane equation corresponding to a plane in the real environment within the camera coordinate system.

[0078] In practice, for each of the aforementioned target camera map points, the map point corresponding to that target camera map point can be determined as a 3D map point. Then, the first value of the aforementioned 3D map point can be substituted into the preset plane general equation "x", the second value of the aforementioned 3D map point can be substituted into the aforementioned plane general equation "y", and the third value of the aforementioned 3D map point can be substituted into the aforementioned plane general equation "z", to obtain the equation after substitution as the plane equation to be fitted.

[0079] The general equation of the plane mentioned above can be a general form equation of the plane, namely "Ax + By + Cz + D = 0". "A", "B", "C", and "D" are constants, and "A", "B", and "C" are not all 0 at the same time. "x", "y", and "z" can be unknowns that need to be substituted into the solution.

[0080] Then, the obtained equations for each plane to be fitted can be fitted using the above equation fitting algorithm to obtain the values ​​corresponding to "A", "B", "C", and "D". These values ​​can then be substituted into the general plane equation to obtain the plane information. Simultaneously, the confidence level and semantic labels corresponding to the effective mask data can be determined as the confidence level and semantic labels corresponding to the plane information.

[0081] In some optional implementations of certain embodiments, the execution entity may generate various planar information based on a preset database, the aforementioned map points, and the aforementioned image mask data through the following steps: The first step is to determine that each of the above image mask data meets the preset valid mask conditions. For each of the above map points, the map point is matched with the plane information stored in the preset database to obtain a matching result group.

[0082] In practice, in response to determining that the above image mask data satisfies the above valid mask conditions, the specific steps for generating a matching result group corresponding to each of the above map points can be referred to the specific implementation method described in step 104, and will not be repeated here.

[0083] The second step involves determining, in response to the determination that each obtained set of matching results satisfies the preset planar expansion conditions, each planar information to be expanded is determined based on the aforementioned sets of matching results and the planar information stored in the preset database. Each of the planar information to be expanded can be the planar information corresponding to the target matching result stored in the preset database. The target matching result can be the matching result marked as "matched" in the aforementioned sets of matching results. The aforementioned planar expansion conditions are the planar expansion conditions themselves.

[0084] In practice, firstly, the matching results marked "matched" in each of the above matching result groups can be identified as the target matching results. Secondly, the planar information corresponding to each of the above target matching results can be identified as the planar information to be expanded.

[0085] The third step involves performing planar expansion processing on the aforementioned map points to obtain various expanded planar information. Each of these expanded planar information pieces can be the expanded planar information itself.

[0086] In practice, for each of the aforementioned planar information to be expanded, firstly, the image mask data whose corresponding semantic labels are the same as the semantic labels to be expanded can be determined as the image mask data to be processed. Here, the semantic labels to be expanded can be the semantic labels corresponding to the aforementioned planar information to be expanded.

[0087] Secondly, the matching results with a value of "match" among the matching results corresponding to the above-mentioned planar information to be expanded can be determined as the set of matching results to be expanded. Then, in response to determining that the above-mentioned set of matching results to be expanded is not empty, the map points corresponding to the above-mentioned set of matching results to be expanded can be determined as the set of map points to be expanded.

[0088] Then, the historical map points corresponding to the plane information to be expanded can be determined as the plane map points to be expanded.

[0089] Next, the aforementioned set of map points to be expanded and the aforementioned individual planar map points to be expanded can be combined into a single set as a planar 3D coordinate set. Then, for each planar 3D coordinate in the aforementioned planar 3D coordinate set, the planar 3D coordinate can be substituted into the aforementioned general planar equation to obtain the substituted general planar equation as the planar 3D equation. Then, the obtained planar 3D equations can be fitted using the aforementioned equation fitting algorithm to obtain the values ​​corresponding to "A", "B", "C", and "D" in the aforementioned general planar equation. The values ​​corresponding to "A", "B", "C", and "D" can then be substituted into the aforementioned general planar equation to obtain the substituted general planar equation as the planar expansion information.

[0090] The fourth step is to determine the above-mentioned plane extension information as the plane information, and update the plane information corresponding to the above-mentioned plane extension information.

[0091] The various embodiments of this disclosure have the following beneficial effects: the plane detection method for augmented reality devices according to some embodiments of this disclosure can improve the accuracy of plane detection, increase the speed of plane detection, and reduce the time spent on plane detection. Specifically, the reason for the long time spent on plane detection is that when there are multiple adjacent planes, the accuracy of plane detection is easily reduced. Moreover, when using a monocular SLAM algorithm to detect or update planes, it is necessary to wait for the AR device to move for a long time to generate a large number of map points, and then detect or update the plane information based on the generated map points, which easily leads to a long time spent on plane detection. Based on this, the plane detection method for augmented reality devices according to some embodiments of this disclosure first acquires environmental image data, historical environmental image data, and sensor data. Thus, the raw data to be processed can be obtained. Second, based on the above-mentioned environmental image data, the above-mentioned historical environmental image data, and the above-mentioned sensor data, various map points are generated. Thus, device pose data and various map points can be obtained. Then, based on the above-mentioned map points, semantic segmentation processing is performed on the above-mentioned environmental image data to obtain various image mask data. Thus, various image mask data can be obtained. Finally, based on the preset database, the aforementioned map points, and the aforementioned image mask data, various planar information is generated. Thus, the planar information can be obtained. Because plans can be detected jointly using the generated map points and image mask data, when adjacent plans exist, the semantic labels of different plans can be further determined using the image mask data, thereby improving the accuracy of planar detection. Furthermore, because planar information can be detected or updated by combining map points and image mask data during planar detection, the number of map points required for detection can be reduced, thus reducing the movement time of the AR device and consequently reducing the time spent on planar detection.

[0092] Further reference Figure 2 This illustrates a flow 200 of another embodiment of a plane detection method for augmented reality devices. The flow 200 of this plane detection method for augmented reality devices includes the following steps: Step 201: Acquire environmental image data, historical environmental image data, and sensor data.

[0093] In some embodiments, the aforementioned implementing entity may acquire environmental image data, historical environmental image data, and sensor data.

[0094] Step 202: Generate device pose data based on environmental image data, historical environmental image data, and sensor data.

[0095] In some embodiments, the executing entity may generate device pose data based on the environmental image data, the historical environmental image data, and the sensor data. The device pose data may be data characterizing the pose of the augmented reality device at the current moment. The device pose data may include a rotation matrix and a translation vector. The rotation matrix may be a matrix capable of mapping the camera coordinate system corresponding to the environmental image data to the camera coordinate system of the historical environmental image data. The translation vector may be a vector characterizing the position of the optical center of the camera coordinate system corresponding to the environmental image data within the camera coordinate system of the historical environmental image data.

[0096] In addressing the aforementioned technical problems in the application scenario of using AR glasses for cargo recognition and navigation, the following technical issues often arise: relying solely on a single visual SFM to generate rotation matrices and translation vectors requires extracting a large number of feature points from the image and performing complex optimizations, which can easily lead to high computational resource consumption in generating rotation matrices and translation vectors. Considering the following requirements of this application scenario: warehouse staff need to wear AR glasses for extended periods for cargo recognition and navigation, thus placing high demands on the battery life of the AR glasses. Furthermore, given the limited computational resources of AR glasses, it is necessary to minimize computational resource consumption to extend the battery life of the AR glasses. Therefore, we have decided to adopt the following solution: In some optional implementations of certain embodiments, the execution entity may generate device pose data based on the aforementioned environmental image data, the aforementioned historical environmental image data, and the aforementioned sensor data through the following steps: The first step involves integrating the angular velocity and acceleration data from the aforementioned sensors to obtain an integral angular velocity vector and a velocity vector. The integral angular velocity vector can be obtained by integrating the angular velocity data, and the velocity vector can be obtained by integrating the acceleration data.

[0097] In practice, the aforementioned executing entity can use an integral method to integrate the angular velocity data to obtain an integral angular velocity vector. Then, the aforementioned integral method can be used to integrate the acceleration data to obtain a velocity vector. Specifically, the aforementioned integral method can be the median integral method.

[0098] As an example, when integrating, the lower limit of integration can be the time point corresponding to the aforementioned historical environmental image data, and the upper limit of integration can be the time point corresponding to the aforementioned environmental image data.

[0099] The second step is to perform an exponential mapping on the aforementioned angular velocity integral vector to obtain a relative rotation matrix. This relative rotation matrix can be a third-order matrix obtained after performing an exponential mapping on the angular velocity integral vector.

[0100] In practice, the aforementioned executing entity can input the angular velocity integral vector into the Rodriguez formula to obtain a third-order matrix as the relative rotation matrix.

[0101] The third step is to integrate the velocity vector to obtain the position vector. The position vector can be the velocity vector after integration.

[0102] In practice, the aforementioned execution entity can use the aforementioned integration method to integrate the aforementioned velocity vector to obtain the position vector.

[0103] The fourth step involves performing feature point extraction and matching processing on the aforementioned environmental image data and historical environmental image data to obtain an image feature coordinate array and a historical image feature coordinate array. Each image feature coordinate in the image feature coordinate array can be the pixel coordinate corresponding to a feature point in the aforementioned environmental image data. Similarly, each historical image feature coordinate in the historical image feature coordinate array can be the pixel coordinate corresponding to a feature point in the aforementioned historical environmental image data.

[0104] In practice, firstly, feature points can be extracted from the aforementioned environmental image data and historical environmental image data using a feature point extraction algorithm. Each feature point in the environmental image data is then used as a separate image feature point, and the same feature points in the historical environmental image data are used as separate historical image feature points. The ORB algorithm can be used for this feature point extraction. Each extracted feature point corresponds to a descriptor.

[0105] Secondly, for each of the aforementioned image feature points, firstly, the pixel coordinates of the image feature point in the aforementioned environmental image data can be determined as image feature coordinates. Then, the descriptor corresponding to the image feature point can be determined as an image descriptor. Next, using the aforementioned matching technique, the historical image feature point with the highest similarity to the aforementioned image descriptor can be found from among the descriptors corresponding to the aforementioned historical image feature points as the target historical image feature point. Then, the pixel coordinates of the target historical image feature point in the aforementioned historical environmental image data can be determined as historical image feature coordinates. At this point, it can be determined that the aforementioned image feature coordinates correspond to the aforementioned historical image feature coordinates.

[0106] Then, the determined image pixel coordinates can be arranged into a single-row array as the image feature coordinate array. Following the order in the image feature coordinate array, the historical image feature coordinates corresponding to the aforementioned image pixel coordinates can be arranged into a single-row array as the historical image feature coordinate array.

[0107] Fifth, based on the aforementioned image feature coordinate array, the aforementioned historical image feature coordinate array, and the preset camera intrinsic parameter matrix, an essential matrix is ​​generated. This essential matrix can be a matrix containing information about the relative rotation and translation between the environmental image data and the historical environmental image data.

[0108] In practice, the aforementioned image feature coordinate array, the aforementioned historical image feature coordinate array, and the aforementioned camera intrinsic parameter matrix can be input into the essential matrix generation function to obtain the third-order matrix output by the essential matrix generation function as the essential matrix. The essential matrix generation function can be the cv.findEssentialMat() function.

[0109] Step 6: Perform singular value decomposition on the aforementioned essential matrix to obtain a direction matrix and a direction vector. The direction matrix can be a third-order matrix obtained after singular value decomposition of the essential matrix. The direction vector can be a column vector of length 3 obtained after singular value decomposition of the essential matrix.

[0110] In practice, the essential matrix described above can be input into a decomposition function. The resulting 3x3 matrix can be used as the direction matrix, and the resulting 3-column vector can be used as the direction vector. The decomposition function can be any function capable of decomposing the essential matrix. For example, the decomposition function could be `cv.recoverPose()`.

[0111] Step 7: Based on the preset historical projection matrix, the aforementioned camera intrinsic parameter matrix, the aforementioned direction matrix, the aforementioned direction vector, the aforementioned image feature coordinate array, and the aforementioned historical image feature coordinate array, generate the three-dimensional coordinates of each feature point.

[0112] The 3D coordinates of each feature point can be non-homogeneous 3D coordinates generated based on the camera intrinsic parameter matrix, the aforementioned direction matrix, the aforementioned direction vector, the aforementioned image feature coordinate array, and the aforementioned historical image feature coordinate array. The aforementioned historical projection matrix can be a matrix capable of mapping the world coordinate system to the camera coordinate system corresponding to the aforementioned historical environmental image data.

[0113] Then, the aforementioned direction matrix and direction vector can be combined into a 3x4 matrix as the image matrix. For example, the aforementioned direction vector can be added to the fourth column of the aforementioned direction matrix to obtain a 3x4 matrix as the image matrix. Then, the product of the aforementioned camera intrinsic parameter matrix and the aforementioned image matrix can be used to determine the projection matrix.

[0114] Then, the aforementioned projection matrix, the aforementioned historical projection matrix, the aforementioned image feature coordinate array, and the aforementioned historical image feature coordinate array can be input into the homogeneous coordinate generation function to obtain the arrays of length 4 output by the homogeneous coordinate generation function as the respective homogeneous coordinates. The aforementioned homogeneous coordinate generation function can be the cv.triangulatePoints() function.

[0115] Then, for each of the above homogeneous coordinates, each homogeneous coordinate can be converted into a three-dimensional coordinate as the three-dimensional coordinate of each feature point. For example, when the homogeneous coordinates are (4,6,8,2), the first three values ​​of the homogeneous coordinates, 4, 6, and 8, can be divided by the last value, 2, to obtain 2, 3, and 4 respectively. Then, 2, 3, and 4 can be combined into the three-dimensional coordinates (2,3,4) as the three-dimensional coordinates of the feature point.

[0116] Step 8: Perform data alignment processing on the above-mentioned direction matrix, direction vector, three-dimensional coordinates of each feature point, velocity vector, relative rotation matrix and position vector to obtain gravity vector and scale factor.

[0117] The gravity vector can be a vector of length 3, used to characterize the gravitational force acting on the object in different directions. For example, the gravity vector could be (0.1, 0.2, -9.81).

[0118] The aforementioned scale factor can be a value used to restore the fuzzy scale of the motion inference structure to the true physical scale.

[0119] In practice, the aforementioned execution entity can input the aforementioned environmental image data, the aforementioned orientation matrix, the aforementioned orientation vector, the aforementioned velocity vector, the aforementioned relative rotation matrix, and the aforementioned position vector into the alignment function to obtain the gravity vector and scale factor output by the alignment function. The aforementioned alignment function can be the VisualIMUAlignment() function.

[0120] Step 9: Based on the aforementioned orientation matrix, orientation vector, gravity vector, and scale factor, generate the pose matrix. This pose matrix can be a 3x4 matrix capable of projecting the world coordinate system onto the camera coordinate system corresponding to the aforementioned environmental image data.

[0121] In practice, firstly, the executing entity can determine the true direction vector by multiplying the scale factor and the direction vector. Secondly, the ratio of the gravity vector to the gravity vector magnitude data can be determined as the normalized gravity vector. Here, the gravity vector magnitude data can be the magnitude length of the gravity vector. Then, the ratio of the preset target gravity vector to the target magnitude data can be determined as the normalized target vector. Here, the target magnitude data can be the magnitude length of the target gravity vector. Here, the target gravity vector can be (0, 0, -9.81).

[0122] Then, the cross product of the above-mentioned normalized gravity vector and the above-mentioned normalized target vector can be determined as the rotation axis vector.

[0123] Then, the dot product of the aforementioned normalized gravity vector and the aforementioned normalized target vector can be determined as the vector cosine value. Then, the aforementioned vector cosine value can be input into the inverse cosine function to obtain the angle output by the inverse cosine function as the cosine angle. Then, the sine value corresponding to the aforementioned cosine angle can be determined as the vector sine value. The square of the aforementioned vector sine value can be determined as the vector square sine value. Then, the difference between the preset minuend and the aforementioned vector cosine value can be determined as the vector cosine difference value. The preset minuend can be 1. Then, the ratio of the aforementioned vector cosine difference value to the aforementioned vector square sine value can be determined as the vector ratio.

[0124] Then, the rotation axis vector can be input into an antisymmetric function to obtain the rotation axis matrix as the output matrix of the antisymmetric function. The antisymmetric function can be any function that converts a vector of length 3 into its corresponding antisymmetric matrix. For example, the antisymmetric function could be the `skew()` function. Then, the square of the rotation axis matrix can be used to determine the rotation axis square matrix. Finally, the product of the third-order identity matrix, the rotation axis matrix, the rotation axis square matrix, and the ratio of the vectors can be used to determine the gravity transformation matrix.

[0125] Then, the product of the gravity transformation matrix and the true direction vector can be used to determine the pose translation vector. Then, the product of the gravity transformation matrix and the direction matrix can be used to determine the pose rotation matrix.

[0126] Then, the pose rotation matrix and the pose translation vector can be combined into a 3x4 matrix as the pose matrix. For example, the pose translation vector can be added to the fourth column of the pose rotation matrix to obtain a 3x4 matrix as the pose matrix.

[0127] Step 10: Based on the pose matrix and the historical projection matrix mentioned above, generate the rotation matrix and translation vector.

[0128] In practice, firstly, the submatrices corresponding to the first three rows and first three columns of the aforementioned pose matrix and the aforementioned historical projection matrix can be respectively determined as the pose rotation matrix and the historical pose rotation matrix. Secondly, the column vectors of the fourth column of the aforementioned pose matrix and the aforementioned historical projection matrix can be respectively determined as the pose translation vector and the historical pose translation vector. Then, the product of the historical pose rotation matrix and the transpose of the aforementioned pose rotation matrix can be determined as the rotation matrix.

[0129] Then, the product of the rotation matrix and the pose translation vector can be used to determine the translation vector to be processed. Then, the difference between the historical pose translation vector and the translation vector to be processed can be used to determine the translation vector.

[0130] Step 11: Combine the above rotation matrix and translation vector into device pose data.

[0131] The above-described technical solution and its related content, as an inventive point of this disclosure, solve the problem of "excessive consumption of computational resources." Factors leading to excessive computational resource consumption often include: when relying solely on a single visual SFM to generate rotation matrices and translation vectors, a large number of feature points need to be extracted from the image and complex optimizations are required, easily resulting in high computational resource consumption during generation. Solving these factors can reduce computational resource consumption. To achieve this effect, this disclosure first integrates the angular velocity and acceleration data included in the sensor data to obtain an angular velocity integral vector and a velocity vector. Secondly, it performs exponential mapping on the angular velocity integral vector to obtain a relative rotation matrix. Then, it integrates the velocity vector to obtain a position vector. Finally, it performs feature point extraction and matching processing on the environmental image data and the historical environmental image data to obtain an image feature coordinate array and a historical image feature coordinate array. Then, based on the aforementioned image feature coordinate array, the aforementioned historical image feature coordinate array, and the preset camera intrinsic parameter matrix, an essential matrix is ​​generated. This yields the essential matrix. Next, singular value decomposition (SVD) is performed on the essential matrix to obtain a direction matrix and a direction vector. This allows the essential matrix to be decomposed, yielding a direction matrix and a direction vector that only represent direction. Then, based on the preset historical projection matrix, the aforementioned camera intrinsic parameter matrix, the aforementioned direction matrix, the aforementioned direction vector, the aforementioned image feature coordinate array, and the aforementioned historical image feature coordinate array, the 3D coordinates of each feature point are generated. This yields the 3D coordinates of each feature point. Then, data alignment processing is performed on the aforementioned direction matrix, the aforementioned direction vector, the aforementioned 3D coordinates of each feature point, the aforementioned velocity vector, the aforementioned relative rotation matrix, and the aforementioned position vector to obtain a gravity vector and a scale factor. This allows the visual SFM data to be aligned with the IMU data after pre-integration processing. Then, based on the aforementioned direction matrix, the aforementioned direction vector, the aforementioned gravity vector, and the aforementioned scale factor, a pose matrix is ​​generated. This allows the direction matrix and direction vector to be adjusted using the obtained gravity vector and scale factor. Then, based on the aforementioned pose matrix and historical projection matrix, a rotation matrix and translation vector are generated. Thus, the rotation matrix and translation vector are obtained. Finally, the rotation matrix and translation vector are combined to form the device pose data. Because the rotation matrix and translation vector can be generated by combining visual SFM and pre-integrated IMU data, highly reliable rotation matrix and translation vectors can be generated without complex optimization of feature points. Therefore, the computational resources consumed by complex optimization processes can be reduced when generating the rotation matrix and translation vectors.

[0132] Step 203: Generate various map points based on device pose data, environmental image data, and historical environmental image data.

[0133] In some embodiments, the execution entity may generate various map points based on the device pose data, the environmental image data, and the historical environmental image data.

[0134] In addressing the aforementioned technical problems in the application scenario of using AR glasses for immersive guided tours, the following technical issues often arise: Generating pixel coordinate depth values ​​using a dedicated deep learning model requires parallel computation using a high-performance GPU, which can lead to significant computational resource consumption during runtime. Considering the specific requirements of this application scenario—that visitors often need to wear AR glasses for extended periods, placing high demands on battery life—and given the limited computational resources of AR glasses, we have decided to adopt the following solution: In some optional implementations of certain embodiments, the aforementioned execution entity may generate various map points based on the aforementioned device pose data, the aforementioned environmental image data, and the aforementioned historical environmental image data through the following steps: The first step involves extracting feature points from the aforementioned environmental image data to obtain individual feature points. Each feature point can be a feature point extracted from the environmental image data. Each feature point corresponds to a descriptor and pixel coordinates. The descriptor can be a feature vector used to quantify the local texture and gradient information of the feature point in the image. The pixel coordinates can be the pixel coordinates of the feature point in the corresponding image data.

[0135] In practice, the aforementioned execution entity can extract each feature point and its corresponding descriptor from the aforementioned environmental image data using a feature point extraction algorithm. This feature point extraction algorithm can be any algorithm capable of extracting feature points from an image. For example, it could be a scale-invariant feature transformation algorithm. Then, the pixel coordinates of each feature point in the aforementioned environmental image data can be determined as the pixel coordinates corresponding to each feature point.

[0136] The second step involves extracting feature points from the aforementioned historical environmental image data to obtain various historical feature points. Each of these historical feature points can be a feature point extracted from the historical environmental image data. Each historical feature point corresponds to a descriptor and pixel coordinates.

[0137] In practice, the aforementioned execution entity can use the aforementioned feature point extraction algorithm to extract various feature points from the aforementioned historical environmental image data as various historical feature points, and extract the descriptors corresponding to the aforementioned historical feature points. Then, the pixel coordinates of the aforementioned historical feature points in the aforementioned historical environmental image data can be determined as the pixel coordinates corresponding to the aforementioned historical feature points.

[0138] The third step involves matching each of the aforementioned feature points with each of the aforementioned historical feature points to obtain feature point matching pairs. Each feature point matching pair can be a data group composed of highly similar feature points and historical feature points. Each feature point matching pair includes both the feature point and historical feature points.

[0139] In practice, the above-mentioned historical feature points can first be combined into a set of historical feature points.

[0140] Secondly, for each of the aforementioned feature points, the following feature matching steps are performed: the cosine similarity between the aforementioned feature point and each historical feature point in the historical feature point set is determined as the feature point similarity. Then, the feature point with the largest value among the determined feature point similarities is determined as the target feature point similarity. Next, the aforementioned feature point and the historical feature points corresponding to the target feature point similarity are combined into feature point matching pairs. Then, the historical feature points corresponding to the target feature point similarity are removed from the historical feature point set using a deletion function. The feature matching steps are then performed again on the historical feature point set and the updated historical feature point set to obtain each feature point matching pair. The deletion function can be a function capable of deleting data from a set. For example, the deletion function can be the remove() function.

[0141] Fourth, for each of the above feature point matching pairs, based on a preset camera intrinsic parameter matrix, back-projection processing is performed on the feature points and historical feature points included in the feature point matching pair to obtain a projection matching pair. The projection matching pair can be the data obtained by projecting the pixel coordinates corresponding to the feature points and historical feature points. The projection matching pair includes projection points and historical projection points. The projection points can be the three-dimensional coordinates obtained by projecting the pixel coordinates corresponding to the feature points onto a preset normalized plane. The historical projection points can be the three-dimensional coordinates obtained by projecting the pixel coordinates corresponding to historical feature points onto the normalized plane. The camera intrinsic parameter matrix is ​​the aforementioned camera intrinsic parameter matrix. The normalized plane can be a virtual plane located one unit length in front of the camera.

[0142] In practice, firstly, the inverse of the aforementioned camera intrinsic parameter matrix can be determined as the camera intrinsic parameter inverse matrix. Secondly, for each feature point matching pair, the pixel coordinates corresponding to the feature points included in the matching pair can be determined as the first pixel coordinates. The pixel coordinates corresponding to the historical feature points included in the matching pair can be determined as the second pixel coordinates. Then, a preset element value can be added after the last value in the first pixel coordinates to obtain the added first pixel coordinates as the first projection coordinates. For example, when the first pixel coordinates are "(1,2)" and the preset element value is 1, the obtained first projection coordinates can be "(1,2,1)". The preset element value can be a pre-set value. For example, the preset element value can be 1. Then, the preset element value can be added after the last value in the second pixel coordinates to obtain the added second pixel coordinates as the second projection coordinates. Then, the transposes of the first and second projection coordinates can be determined as the first transpose vector and the second transpose vector, respectively. Then, the product of the aforementioned camera intrinsic parameter inverse matrix and the aforementioned first transpose vector can be determined as the projection point. Next, the product of the aforementioned camera intrinsic parameter inverse matrix and the aforementioned second transpose vector can be used to determine the historical projection points. Finally, the aforementioned projection points and the aforementioned historical projection points can be combined into projection matching pairs.

[0143] Fifth step: For each of the obtained projection matching pairs, construct a map point depth equation based on the preset depth character, the preset historical depth character, the above device pose data, and the above projection matching pair.

[0144] The aforementioned depth character can be a character used to characterize the depth value of a pixel in the aforementioned environmental image data. For example, the aforementioned depth character can be "s2". The aforementioned historical depth character can be a character used to characterize the depth value of a pixel in the aforementioned historical environmental image data. For example, the aforementioned historical depth character can be "s1". The aforementioned map point depth equation can be an equation composed of the aforementioned depth character, the aforementioned historical depth character, the aforementioned device pose data, and the aforementioned projection matching pair.

[0145] In practice, for each of the obtained projection matching pairs, firstly, the product of the aforementioned historical depth character and the historical projection points included in the projection matching pair can be determined as the historical depth data. Then, the product of the aforementioned depth character, the rotation matrix included in the aforementioned device pose data, and the projection points included in the aforementioned projection matching pair can be determined as the initial depth term number. Next, the sum of the aforementioned initial depth term number and the translation vector included in the aforementioned device pose data can be determined as the depth data. Finally, the aforementioned historical depth data can be made equal to the aforementioned depth data, thereby constructing an equation as the map point depth equation.

[0146] Step 6: For each of the constructed map point depth equations, perform a fitting process to obtain the depth value. This depth value can be the numerical value corresponding to the aforementioned depth character.

[0147] In practice, for each of the above map point depth equations, the above fitting algorithm can be used to fit the map point depth equation to obtain the numerical value corresponding to the depth character in the map point depth equation as the depth value.

[0148] Step 7: Based on the aforementioned feature points and the obtained depth values, generate each map point. In practice, for each feature point, the first and second values ​​of the corresponding pixel coordinates can be determined as the pixel's horizontal and vertical coordinates, respectively.

[0149] Secondly, the element values ​​in the first row and first column, and the second row and second column of the aforementioned camera intrinsic parameter matrix can be determined as the horizontal axis focal length and the vertical axis focal length, respectively. Similarly, the element values ​​in the first row and third column, and the second row and third column of the aforementioned camera intrinsic parameter matrix can be determined as the horizontal coordinate value and the vertical coordinate value of the principal point, respectively.

[0150] Then, the difference between the aforementioned pixel x-coordinate value and the aforementioned principal point x-coordinate value can be determined as the x-coordinate difference. Next, the ratio of the aforementioned x-coordinate difference to the aforementioned horizontal axis focal length can be determined as the x-coordinate ratio. Finally, the product of the aforementioned x-coordinate ratio and the depth value corresponding to the aforementioned feature point can be determined as the map point x-coordinate value.

[0151] Then, the difference between the aforementioned pixel ordinate value and the aforementioned principal point ordinate value can be determined as the ordinate difference value. Then, the ratio of the aforementioned ordinate difference value to the aforementioned vertical axis focal length can be determined as the ordinate ratio value. Finally, the product of the aforementioned ordinate ratio value and the depth value corresponding to the aforementioned feature point can be determined as the map point ordinate value.

[0152] Finally, the x-coordinate value of the map point, the y-coordinate value of the map point, and the depth value corresponding to the feature point can be combined into a three-dimensional coordinate system as the map point. As an example, when the x-coordinate value of the map point is 1, the y-coordinate value of the map point is 2, and the depth value corresponding to the feature point is 3, the resulting map point can be (1,2,3).

[0153] The above-described technical solution and its related content, as an inventive point of this disclosure, solve the problem of "high computational resource consumption." Factors leading to high computational resource consumption often include: when using a dedicated deep learning model to generate depth values ​​for pixel coordinates, high-performance GPUs are required for parallel computation, easily resulting in high computational resource consumption during runtime. Solving these factors can reduce computational resource consumption. To achieve this effect, this disclosure first performs feature point extraction processing on the aforementioned environmental image data to obtain various feature points. Secondly, feature point extraction processing is performed on the aforementioned historical environmental image data to obtain various historical feature points. Then, the aforementioned feature points and the aforementioned historical feature points are matched to obtain various feature point matching pairs, wherein each feature point matching pair includes a feature point and a historical feature point. Thus, the various feature points and the aforementioned historical feature points can be matched. Then, for each of the aforementioned feature point matching pairs, based on a preset camera intrinsic parameter matrix, back-projection processing is performed on the feature points and historical feature points included in the feature point matching pair to obtain a projection matching pair, wherein the projection matching pair includes a projection point and a historical projection point. Thus, projection matching pairs are obtained. Then, for each of the obtained projection matching pairs, a map point depth equation is constructed based on a preset depth character, a preset historical depth character, the aforementioned device pose data, and the aforementioned projection matching pair. Thus, a map point depth equation is constructed. Then, for each of the constructed map point depth equations, the map point depth equation is fitted to obtain a depth value. Thus, a depth value can be obtained through a fitting algorithm. Finally, based on the aforementioned feature points and the obtained depth values, each map point is generated. Thus, each map point is obtained. Because the depth values ​​of pixels can be generated in a non-iterative, geometric operation mode using triangulation, without running a deep learning model or relying on a high-performance GPU, the computational resources consumed in generating pixel depth values ​​can be reduced.

[0154] Optionally, after generating each map point based on the device pose data, the environmental image data, and the historical environmental image data, the execution entity may further perform the following steps: For each piece of information about a plane stored in the preset database, perform the following steps: The first step involves generating an occlusion detection result based on the aforementioned environmental image data, device pose data, historical map points corresponding to the aforementioned planar information, and historical environmental image data. This occlusion detection result can be information used to characterize whether the plane represented by the aforementioned planar information has been occluded or moved. For example, the occlusion detection result can be "occluded" or "no occlusion".

[0155] In practice, firstly, the historical environmental image data corresponding to the aforementioned planar information can be determined as historical planar image data. Secondly, for each historical map point, coordinate projection processing can be performed on the historical map point to obtain the corresponding camera map point as the historical camera map point. The method for generating the camera map point corresponding to the historical map point can be referred to the specific implementation of step 104, and will not be repeated here. Then, the image region of a preset size corresponding to the historical camera map point in the aforementioned historical planar image data can be determined as the historical image region to be compared. The preset size can be a pre-defined image size. For example, the image region of 9*9 centered on the aforementioned historical camera map point in the aforementioned historical planar image data can be determined as the historical image region to be compared. Then, the product of the aforementioned historical map point and the rotation matrix included in the aforementioned device pose data can be determined as the transformed map point.

[0156] Then, coordinate projection processing can be performed on the above-mentioned transformed map points to obtain the camera map points corresponding to the transformed map points as transformed camera map points. The method for generating the camera map points corresponding to the transformed map points can be referred to the specific implementation method of step 104, and will not be repeated here.

[0157] Next, for each of the obtained converted camera map points, firstly, the image region of the preset size corresponding to the converted camera map point in the environmental image data can be determined as the image region to be compared. Secondly, the historical image region to be compared corresponding to the converted camera map point can be determined as the target historical image region. Then, the target historical image region and the image region to be compared can be compared using an image comparison algorithm to obtain the image similarity. The image comparison algorithm can be any algorithm capable of comparing the similarity between two images. For example, the image comparison algorithm can be AHash (Average Hash). The image similarity can be a numerical value used to characterize the degree of similarity between the target historical image region and the image region to be compared.

[0158] Then, in response to determining that the obtained image similarities meet a preset similarity condition, "occlusion" can be identified as an occlusion detection result. The aforementioned similarity condition can be that the similarity of each image is greater than a preset similarity threshold. The aforementioned similarity threshold can be a pre-set value. Here, the specific setting of the aforementioned similarity threshold is not limited. In response to determining that the obtained image similarities do not meet the aforementioned similarity condition, "no occlusion" can be identified as an occlusion detection result.

[0159] The second step involves generating image mask data based on the environmental image data in response to the determination that the occlusion detection result meets a preset occlusion condition. The occlusion condition can be that the occlusion detection result is "occluded". In practice, the method for generating image mask data based on the environmental image data can be found in the specific implementation described in step 103, and will not be repeated here.

[0160] The third step involves filtering the planar information stored in the preset database based on the aforementioned image mask data to obtain the planar information to be updated. Each of the planar information to be updated can be a planar information stored in the preset database that has been filtered.

[0161] In practice, in response to determining that the aforementioned image mask data satisfies the aforementioned valid mask conditions, for each piece of planar information stored in the preset database, firstly, the semantic label corresponding to the aforementioned planar information can be determined as a planar semantic label. Secondly, the image mask data whose corresponding semantic label is the aforementioned planar semantic label can be determined as planar image mask data.

[0162] Then, in response to determining that the confidence level corresponding to the aforementioned planar image mask data is greater than the aforementioned preset confidence threshold, each converted camera map point corresponding to each of the aforementioned historical map points can be generated as each image camera map point based on each of the aforementioned historical map points corresponding to the planar information. The method for generating each of the aforementioned converted camera map points corresponding to each of the aforementioned historical map points can be referred to the specific implementation of step 203, and will not be repeated here.

[0163] Then, for each of the aforementioned image camera map points, the element value in the aforementioned planar image mask data corresponding to the aforementioned image camera map point can be determined as the planar image element value. For example, when the image camera map point is (1,2), the planar image element value can be the element value in the first row and second column of the aforementioned planar image mask data. In response to determining that the aforementioned planar image element value satisfies a preset label element condition, the aforementioned planar image element value can be determined as the intersection element value. The label element condition can be that the aforementioned planar image element value is "1". Then, the number of the determined intersection element values ​​can be determined as the intersection quantity. The number of the aforementioned image camera map points can be determined as the map point projection quantity. Furthermore, the ratio of the aforementioned intersection quantity to the aforementioned map point projection quantity can be determined as the overlap degree. In response to determining that the aforementioned overlap degree is greater than a preset overlap degree threshold, the aforementioned planar information can be determined as the planar information to be updated. The aforementioned overlap degree threshold can be a preset value. Here, the specific setting of the aforementioned overlap degree threshold is not limited. Thus, various planar information to be updated can be generated.

[0164] The fourth step is to perform semantic updates on the above image mask data to obtain the updated plane information.

[0165] In practice, for each plane information to be updated, firstly, the image mask data whose semantic labels are the same as those of the plane information to be updated can be determined as the target image mask data. Then, the camera map points corresponding to the map points can be generated as the camera projection map points. The method for generating the camera map points corresponding to the map points can be referred to the specific implementation of step 104, and will not be repeated here.

[0166] Next, the camera projection map points that meet the preset label conditions among the aforementioned camera projection map points can be identified as target camera projection map points. The preset label conditions can be that the element value corresponding to the camera projection map point in the target image mask data is "1". Then, the map points corresponding to the target camera projection map points can be identified as map points to be supplemented. Finally, the historical map points corresponding to the planar information to be updated and the map points to be supplemented can be combined into a set of map points to be updated.

[0167] Next, for each map point in the aforementioned set of map points to be updated, the map point to be updated can be substituted into the aforementioned general equation of the plane to obtain the general equation of the plane after substitution, which can be used as the equation of the plane to be updated. The method of substituting the map point to be updated into the general equation of the plane can be referred to the specific implementation of step 104, and will not be repeated here.

[0168] Then, the above equation fitting algorithm can be used to fit each of the obtained plane equations to be updated, to obtain the values ​​corresponding to "A", "B", "C" and "D" in each plane equation to be updated, and the values ​​corresponding to "A", "B", "C" and "D" can be substituted into the above general plane equation to obtain the plane equation as the plane information to be updated, so as to update the plane information to be updated.

[0169] In practice, when the detected plane moves or is obstructed, the plane information can be updated using the steps described above.

[0170] Step 204: Based on each map point, perform semantic segmentation on the environmental image data to obtain each image mask data.

[0171] In some embodiments, the execution entity may perform semantic segmentation processing on the environmental image data based on the map points to obtain image mask data.

[0172] In practice, in response to the determination that the number of each of the aforementioned map points exceeds the aforementioned preset map point threshold, the aforementioned environmental image data can be input into a pre-trained image segmentation model to obtain image mask data for each point. The aforementioned image segmentation model can be a pre-trained U-Net model. The aforementioned pre-training can be a process of fine-tuning the U-Net model using the aforementioned sample dataset and the cross-entropy loss function.

[0173] Step 205: Generate various planar information based on the preset database, various map points, and various image mask data.

[0174] In some embodiments, the execution entity may generate various planar information based on a preset database, the various map points, and the various image mask data.

[0175] The various embodiments of this disclosure have the following beneficial effects: the plane detection method for augmented reality devices according to some embodiments of this disclosure can reduce the time spent on plane detection. Specifically, the reason for the long time spent on plane detection is that when using monocular SLAM algorithm to detect weak texture regions, the plane detection is easily incomplete due to insufficient extracted feature points, which requires re-detection of the incomplete plane, resulting in a long re-detection time. Based on this, the plane detection method for augmented reality devices according to some embodiments of this disclosure first acquires environmental image data, historical environmental image data, and sensor data. This yields the raw data to be processed. Second, based on the environmental image data, historical environmental image data, and sensor data, device pose data is generated. This yields the device pose data. Then, based on the device pose data, environmental image data, and historical environmental image data, map points are generated. This yields the map points. Then, based on the map points, semantic segmentation processing is performed on the environmental image data to obtain image mask data. This yields the image mask data. Finally, based on the preset database, the aforementioned map points, and the aforementioned image mask data, various planar information is generated. Thus, various planar information can be obtained. Because planar information can be generated by combining map points and image mask data during planar detection, the number of map points needed to be generated when detecting weak texture areas can be reduced, as can the number of feature points needed to be generated. This reduces the probability of incomplete planar detection due to insufficient extracted feature points, thereby reducing the number of times incomplete plans need to be re-detected and reducing the time spent on re-detection.

[0176] The following is for reference. Figure 3 It shows a schematic diagram of the structure of an augmented reality device 300 suitable for implementing some embodiments of the present disclosure. Figure 3 The augmented reality device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.

[0177] like Figure 3 As shown, the augmented reality device 300 may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 301, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 302 or a program loaded from a storage device 308 into a random access memory (RAM) 303. The RAM 303 also stores various programs and data required for the operation of the augmented reality device 300. The processing unit 301, ROM 302, and RAM 303 are interconnected via a bus 304. An input / output (I / O) interface 305 is also connected to the bus 304.

[0178] Typically, the following devices can be connected to I / O interface 305: input devices 306 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 307 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 308 including, for example, magnetic tapes, hard disks, etc.; and communication devices 309. Communication device 309 allows augmented reality device 300 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 3 An augmented reality device 300 with various devices is shown, but it should be understood that implementation or possession of all the devices shown is not required. More or fewer devices may be implemented or possessed alternatively. Figure 3 Each box shown can represent a device or multiple devices as needed.

[0179] In particular, according to some embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, some embodiments of this disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication device 309, or installed from storage device 308, or installed from ROM 302. When the computer program is executed by processing device 301, it performs the functions defined in the methods of some embodiments of this disclosure.

[0180] It should be noted that, in some embodiments of this disclosure, the computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium may be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In some embodiments of this disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In some embodiments of this disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.

[0181] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol) and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.

[0182] The aforementioned computer-readable medium may be included in the aforementioned augmented reality device; or it may exist independently and not assembled into the augmented reality device. The aforementioned computer-readable medium carries one or more programs that, when executed by the augmented reality device, cause the augmented reality device to: acquire environmental image data, historical environmental image data, and sensor data; generate various map points based on the aforementioned environmental image data, historical environmental image data, and sensor data; perform semantic segmentation processing on the aforementioned environmental image data based on the aforementioned map points to obtain various image mask data; and generate various planar information based on a preset database, the aforementioned map points, and the aforementioned image mask data.

[0183] Computer program code for performing operations of some embodiments of this disclosure can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, and C++, and conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0184] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0185] The functions described above in this document can be performed at least in part by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip (SoCs), complex programmable logic devices (CPLDs), and so on.

[0186] The above description is merely a selection of preferred embodiments of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of the invention involved in the embodiments of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described inventive concept. For example, technical solutions formed by substituting the above-described features with (but not limited to) technical features with similar functions disclosed in the embodiments of this disclosure.

Claims

1. A method for plane detection in augmented reality devices, comprising: Acquire environmental image data, historical environmental image data, and sensor data; Based on the environmental image data, the historical environmental image data, and the sensor data, various map points are generated; Based on the map points, semantic segmentation processing is performed on the environmental image data to obtain image mask data. Based on the preset database, the map points, and the image mask data, various planar information is generated.

2. The method according to claim 1, wherein, The process of generating map points based on the environmental image data, the historical environmental image data, and the sensor data includes: Based on the environmental image data, the historical environmental image data, and the sensor data, device pose data is generated; Based on the device pose data, the environmental image data, and the historical environmental image data, various map points are generated.

3. The method according to claim 1, wherein, The default database stores information about each plane. And the generation of various planar information based on the preset database, the various map points, and the various image mask data includes: In response to determining that each image mask data satisfies a preset valid mask condition, for each map point among the map points, the map point is matched with each plane information stored in a preset database to obtain a matching result group; In response to the determination that the obtained matching result groups do not meet the preset planar expansion conditions, each planar information is generated based on the image mask data and the map points.

4. The method according to claim 3, wherein, The generation of planar information based on the image mask data and the map points includes: For each of the map points, perform coordinate projection processing to obtain the camera map point; Each image mask data that meets the preset confidence level condition is determined as a valid mask data. For each valid mask data in the aforementioned valid mask data, perform the following steps: Based on the effective mask data, clustering is performed on each obtained camera map point to obtain each target camera map point; The map points of each target camera are fitted to obtain planar information.

5. The method according to claim 1, wherein, The default database stores information about each plane, and each plane in the default database corresponds to a semantic tag and a historical map point. And the generation of various planar information based on the preset database, the various map points, and the various image mask data includes: In response to determining that each image mask data satisfies a preset valid mask condition, for each map point among the map points, the map point is matched with each plane information stored in a preset database to obtain a matching result group; In response to determining that each obtained matching result group satisfies the preset plane expansion condition, each plane information to be expanded is determined based on each matching result group and each plane information stored in the preset database. Based on the map points, the planar expansion information to be expanded is processed to obtain the planar expansion information. Each plane extension information is determined as a plane information, and the plane information corresponding to each plane extension information is updated.

6. The method according to claim 2, wherein, The preset database stores information about each plane, and each plane in the preset database corresponds to a historical map point and historical environmental image data. And after generating each map point based on the device pose data, the environmental image data, and the historical environmental image data, the method further includes: For each piece of information about a plane stored in the preset database, perform the following steps: Based on the environmental image data, the device pose data, the historical map points corresponding to the planar information, and the historical environmental image data, an occlusion detection result is generated. In response to determining that the occlusion detection result meets the preset occlusion conditions, each image mask data is generated based on the environmental image data; Based on the image mask data, the plane information stored in the preset database is filtered to obtain the plane information to be updated. Based on the image mask data, the semantics of each plane information to be updated are performed to obtain the updated plane information.

7. The method according to claim 1, wherein, After performing semantic segmentation processing on the environmental image data based on the respective map points to obtain the respective image mask data, the method further includes: In response to determining that the number of each map point is less than or equal to a preset map point threshold, semantic segmentation processing is performed on the environmental image data to obtain each image mask data; For each image mask data in the aforementioned image mask data, perform the following steps: Based on the preset camera intrinsic parameter matrix, preset database and the image mask data, the target historical plane information, the coordinates of each image mask and the mask depth value are generated; Based on the camera intrinsic parameter matrix and the mask depth value, the coordinates of each image mask are back-projected to obtain the three-dimensional coordinates of each mask. Based on the three-dimensional coordinates of each mask, the target historical plane information is semantically updated.

8. An augmented reality device, comprising: One or more processors; A storage device on which one or more programs are stored; One or more display screens for displaying images in front of a user; When the one or more programs are executed by the one or more processors, the one or more processors implement the method as described in any one of claims 1 to 7.

9. A computer-readable medium having a computer program stored thereon, wherein, When the program is executed by the processor, it implements the method as described in any one of claims 1 to 7.