Network device identification method in augmented reality based on visual geometry registration

By employing a visual geometric registration method and utilizing visual localization and augmented reality device data, point clouds are constructed and adaptive iterative scanning is performed. This solves the complexity and error problems of network device identification in existing technologies, enabling fast and simple network device identification.

CN115953462BActive Publication Date: 2026-03-24ZHEJIANG UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-06
Publication Date
2026-03-24

AI Technical Summary

Technical Problem

Existing technologies struggle to quickly and easily identify network devices in augmented reality scenes, especially wired and wireless network devices of similar shape, and relying on identification codes or wireless signals presents operational complexity and error issues.

Method used

A visual geometric registration-based method is adopted to obtain the 3D bounding box and point cloud of network devices through visual positioning. Combined with the electronic compass and inertial measurement unit data of augmented reality devices, a dense point cloud is constructed and a rotation matrix mapping is performed. The network devices are identified using an adaptive iterative scanning method and a depth registration network.

Benefits of technology

It enables quick and easy identification of network devices in augmented reality scenes without requiring unique shapes, identification codes, or wireless signals, thus improving identification efficiency and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure FDA0005554942450000011
    Figure FDA0005554942450000011
  • Figure FDA0005554942450000012
    Figure FDA0005554942450000012
  • Figure FDA0005554942450000013
    Figure FDA0005554942450000013
Patent Text Reader

Abstract

A network equipment identification method in augmented reality based on visual geometric structure registration, which utilizes a camera to acquire the pose of network equipment and constructs a complex geometric structure expressing the pose of individual network equipment and the geographical correlation among multiple network equipments, automatically identifies the network equipment appearing in the augmented reality picture through the depth registration of the geometric structure. The present application does not depend on the unique appearance, identification code, world coordinates and wireless signal of network equipment, and can simply, quickly and conveniently realize the identification of various wired and wireless network equipment.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the field of network device identification, and relates to a network device identification method in augmented reality. BACKGROUND

[0002] In the prior art, automatic identification of network devices appearing in the camera screen is a basic requirement for interaction between augmented reality and network devices (such as reading device status, setting device parameters, etc.). At present, the methods for identifying Internet of Things devices in augmented reality screens mainly include visual-based and wireless signal-based methods. The visual-based method, such as Y. Sun, S. N. R. Kantareddy, J. Siegel, A. Armengol-Urpi, X. Wu, H. Wang, and S. Sarma, “Towards industrial iot-ar systems using deep learning-based object pose estimation,” in Proceedings of the IEEE 38th International Performance Computing and Communications Conference (IPCCC), 2019, pp. 1-8. and D. Jo and G. J. Kim, “Iot+ar: Pervasive and augmented environments for “digi-log” shopping experience,” Human-centric Computing and Information Sciences, vol. 9, no. 1, p. 1, 2019., mainly relies on the shape and texture features of the network device shape, and it is difficult to distinguish network devices with the same shape. Therefore, it is generally necessary to paste special identification codes such as two-dimensional codes on the network devices, such as H. Subakti and J. Jiang, “A marker-based cyber-physical augmented-reality indoor guidance system for smart campuses,”

[0003] in 2016IEEE 18th International Conference on High Performance Computing and Communications; IEEE 14th International Conference on Smart City; IEEE 2nd International Conference on Data Science and Systems (HPCC / SmartCity / DSS), 2016, pp. 1373-1379. However, it is time-consuming and laborious to prepare special identification codes for each network device, and it is difficult to identify in use due to lossless. Ming et al. tried to use visual positioning combined with the GPS position of the augmented reality device to calculate a unique global coordinate for each network device (M. Xia, Y. Wang, X. Wang, Z. Cheng, and K. Chi, “Insight: An enabled user interface for vision-based markerless interaction with IoT nodes,” in Proceedings of the 2021 IEEE Wireless Communications and Networking Conference (WCNC), 2021, pp. 1-7.), and used coordinate matching to achieve network device identification. Due to the error of GPS positioning, especially the large indoor error, this method is difficult to identify network devices deployed close together. The method based on wireless signal uses an antenna array to track the source of the wireless signal and map it to the augmented display screen. Since augmented reality devices often do not have complex antenna arrays, Y. Park, S. Yun, and K. H. Kim, “When IoT met Augmented Reality: Visualizing the Source of the Wireless Signal in AR View,” in Proceedings of the the 17th Annual International Conference on Mobile Systems, Applications, and Services (MobiSys), 2019, 117-129. proposed using a two-antenna array for tracking, but required the user to rotate the augmented reality device along each axis to eliminate angular ambiguity, which was relatively complex to operate. At the same time, the method based on wireless signal cannot identify wired network devices.

[0004] Registration of visual geometry can be abstracted as a point cloud registration problem. At present, deep learning methods such as Predator (S. Huang, Z. Gojcic, M. Usvyatsov, A. Wieser, and K. Schindler, "Predator: Registration of 3d point clouds with low overlap," in Proceedings of the 2021 IEEE / CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2021, pp. 4265-4274.) are the most effective methods in point cloud registration tasks. However, in the network device identification process, adjacent network devices may be occluded or skipped during scanning, resulting in an incomplete matching problem; at the same time, the registration result will be limited by the error characteristics of the augmented display device obtaining the network device geometry. Therefore, the existing point cloud registration network cannot be directly used for network device identification applications. SUMMARY

[0005] In order to overcome the shortcomings of the prior art, the present application provides a network device identification method in augmented reality based on visual geometry registration, which does not rely on the unique shape of the network device, the identification code, the world coordinates and the wireless signal, and can simply, quickly and conveniently realize the identification of various wired and wireless network devices.

[0006] The technical scheme adopted by the present application to solve its technical problems is:

[0007] A network device identification method in augmented reality based on visual geometry registration, comprising the following steps:

[0008] (1) In the deployment phase, first, the network device i appearing in the camera screen is visually positioned to obtain its three-dimensional bounding box in the camera coordinate system Wherein, represents the position of the network device, and represents the size and orientation of the network device; then, by sampling the surface points, it is converted into a dense point cloud Again, the rotation matrix alpha N is obtained by using the electronic compass and inertial measurement unit data of the augmented reality device And by and map and to the dense point cloud in the reference coordinate system and point Because augmented reality (AR) devices have limited fields of view, it is often necessary to observe other network devices by moving and rotating the AR device. The rotation matrix α can be obtained using visual history measurement methods by monitoring the camera feeds, electronic compass, and inertial measurement unit data of the AR device. D And offset vector m, and the dense point cloud of network device j observed after the augmented reality device is moved and rotated. Mapped to a reference coordinate system consistent with u: By using multiple network devices The fusion of these elements yields a large, dense point cloud Q representing the poses of individual network devices and the geographical associations of multiple network devices; simultaneously, by utilizing the L-axis of multiple nodes... G Constructing a sparse point cloud Q * and L for each network device G It is bound to its ID and stored in the database;

[0009] (2) In the recognition stage, using the same method as in step (1), a dense point cloud P with a size smaller than Q can be obtained, representing the pose of individual network devices and the geographical associations between multiple network devices appearing in the augmented reality scene during the recognition stage. * Sparse point cloud P * Then, the transformation matrix is ​​obtained through an adaptive iterative scanning method. satisfy Where P k Let P be the k-th point; for network device i that identifies nodes appearing in the augmented reality scene, through... Get its in Q * The corresponding point j in the database is identified by querying the database to obtain the network device ID corresponding to j, thereby achieving identification.

[0010] Furthermore, in step (2), the adaptive iterative scanning method works as follows: First, the transformation matrix from P to Q is calculated using an incomplete depth registration network. And continuously estimate of like This indicates that the number of network devices scanned during the identification phase is insufficient, and more nodes need to be scanned to expand P; secondly, if the following conditions are met... Then remove the corresponding point of P from Q and re-register. If it still yields the desired result... satisfy This indicates that the number of network devices scanned during the identification phase is still insufficient, and more nodes need to be scanned to expand P; the above process is repeated until only one remains. satisfy until.

[0011] Furthermore, the incomplete deep registration network uses Predator as its backbone network and improves the registration performance of network devices in recognition applications by adding spatial attention in the edge dimension and channel attention in the frequency dimension.

[0012] Furthermore, the spatial attention process of the edge dimension is as follows: First, fully connected graphs G(Q) and G(P) are generated using Q and P, where (i, j) represents the graph in G(Q) connected to vertex Q. i and Q j The connected edges are used to initialize the feature vector using information about the edge's position, size, and angle. Where l(·) represents the vertex position, cat(·,·) represents the join operation, and h θ Similarly, for a neural network with parameter θ, the eigenvector of edge (a, b) in G(P) can be calculated. Next, the query vector is calculated. key vector and value vector Update in Among them W q W k and W v It is a learnable weight matrix; update The formula is Where att(·) represents the self-attention module, h σ Let's represent a neural network with parameter σ; then, we will... Input a linear layer and you will get its score. Finally, the eigenvalues ​​of each vertex in G(Q) are iteratively updated using the following formula: in, Q represents i The k-th round feature representation, max represents the max pooling operation, h φ Let's say we have a neural network with parameter φ. Finally, we perform this update twice with non-shared parameter φ, concatenate the results of each round, and input them into another neural network to obtain the final features.

[0013] The channel attention operation in the frequency dimension works as follows: This operation is performed on both G(P) and G(Q). For G(Q), the weight of edge (i, j) is calculated as follows: Then, using the weights of all edges in G(Q), construct the edge weight matrix A and the diagonal matrix D, where the non-zero elements in D are determined by... Calculate |Q|, representing the number of vertices in Q, and compute the normalized Laplacian matrix. Next, through L=U∧U T Decompose L to obtain U; ​​then, decompose the input features of Q. The features are divided into W groups based on the channel dimension, where C represents the feature dimension, and C is divisible by W. Each group contains [number of channels]. Following this method, U is also divided into W parts, and experiments are conducted using its frequency components sequentially to obtain the optimal frequency component ranking; then, the top m most important components in W are selected to pair U and X. Q Grouping, with each group containing [number of channels]. C is divisible by m; for the m-th... i (1≤m) i Sub-features ≤m,m≤W Its frequency dimension channel attention is It is the mth in U i The rows corresponding to important channels, yes The k-th channel in the matrix; finally, the channel attention of m sub-features is connected, input into a neural network to obtain the final channel attention in the frequency dimension, and then compared with X. Q Multiplication Update X Q .

[0014] In this invention, a camera is used to acquire the pose of network devices, and a complex geometric structure is constructed to express the pose of individual network devices and the geographical association between multiple network devices. Through depth registration of the geometric structure, network devices appearing in the augmented reality scene are automatically identified.

[0015] The main advantages of this invention are that it can easily, quickly, and conveniently identify various wired and wireless network devices without relying on the unique shape, identification code, world coordinates, and wireless signal of the network device. Detailed Implementation

[0016] The present invention will now be described in further detail.

[0017] A method for identifying network devices in augmented reality based on visual geometric structure registration includes the following steps:

[0018] (1) During the deployment phase, the network device i appearing in the camera view is first visually located to obtain its three-dimensional bounding box in the camera coordinate system. in, Indicates the location of the network device. and Indicate the size and orientation of the network device; then, by... Surface point sampling converts it into a dense point cloud Furthermore, the rotation matrix α is obtained using data from the electronic compass and inertial measurement unit of the augmented reality device. N and through and Will and Mapped to a compact point cloud in a reference coordinate system and points Because augmented reality (AR) devices have limited fields of view, it is often necessary to observe other network devices by moving and rotating the AR device. The rotation matrix α can be obtained using visual history measurement methods by monitoring the camera feeds, electronic compass, and inertial measurement unit data of the AR device. D And offset vector m, and the dense point cloud of network device j observed after the augmented reality device is moved and rotated. Mapped to a reference coordinate system consistent with i: By using the Q of multiple network devices G The fusion of these elements yields a large, dense point cloud Q representing the poses of individual network devices and the geographical associations of multiple network devices; simultaneously, by utilizing the L-axis of multiple nodes... G Constructing a sparse point cloud Q * and L for each network device G It is bound to its ID and stored in the database;

[0019] (2) In the recognition stage, using the same method as in step (1), a dense point cloud P with a size smaller than Q can be obtained, representing the pose of individual network devices and the geographical associations between multiple network devices appearing in the augmented reality scene during the recognition stage. * Sparse point cloud P * Then, the transformation matrix is ​​obtained through an adaptive iterative scanning method. satisfy Where P k Let P be the k-th point; for network device i that identifies nodes appearing in the augmented reality scene, through... Get its in Q * The corresponding point j in the database is identified by querying the database to obtain the network device ID corresponding to j, thereby achieving identification.

[0020] Furthermore, in step (2), the adaptive iterative scanning method works as follows: First, the transformation matrix from P to Q is calculated using an incomplete depth registration network. And continuously estimate of like This indicates that the number of network devices scanned during the identification phase is insufficient, and more nodes need to be scanned to expand P; secondly, if the following conditions are met... Then remove the corresponding point of P from Q and re-register. If it still yields the desired result... satisfy This indicates that the number of network devices scanned during the identification phase is still insufficient, and more nodes need to be scanned to expand P; the above process is repeated until only one remains. satisfy until.

[0021] Furthermore, the incomplete deep registration network uses Predator as its backbone network and improves the registration performance of network devices in recognition applications by adding spatial attention in the edge dimension and channel attention in the frequency dimension.

[0022] Furthermore, the spatial attention process of the edge dimension is as follows: First, fully connected graphs G(Q) and G(P) are generated using Q and P, where (i, j) represents the graph in G(Q) connected to vertex Q. i and Q j The connected edges are used to initialize the feature vector using information about the edge's position, size, and angle. Where l(·) represents the vertex position, cat(·,·) represents the join operation, and h θ Similarly, for a neural network with parameter θ, the eigenvector of edge (a, b) in G(P) can be calculated. Next, the query vector is calculated. key vector and value vector Update in Among them W q W k and W v It is a learnable weight matrix; update The formula is Where att(·) represents the self-attention module, h σ Let's represent a neural network with parameter σ; then, we will... Input a linear layer and you will get its score. Finally, the eigenvalues ​​of each vertex in G(Q) are iteratively updated using the following formula: in, Q represents i The k-th round feature representation, max represents the max pooling operation, h φ Let's say we have a neural network with parameter φ. Finally, we perform this update twice with non-shared parameter φ, concatenate the results of each round, and input them into another neural network to obtain the final features.

[0023] The channel attention operation in the frequency dimension works as follows: This operation is performed on both G(P) and G(Q). For G(Q), the weight of edge (i, j) is calculated as follows: Then, using the weights of all edges in G(Q), construct the edge weight matrix A and the diagonal matrix D, where the non-zero elements in D are determined by... Calculate |Q|, representing the number of vertices in Q, and compute the normalized Laplacian matrix. Next, through L=U∧U TDecompose L to obtain U; ​​then, decompose the input features of Q. The features are divided into W groups based on the channel dimension, where C represents the feature dimension, and C is divisible by W. Each group contains [number of channels]. Following this method, U is also divided into W parts, and experiments are conducted using its frequency components sequentially to obtain the optimal frequency component ranking; then, the top m most important components in W are selected to pair U and X. Q Grouping, with each group containing [number of channels]. C is divisible by m; for the m-th... i (1≤m) i Sub-features ≤m,m≤W Its frequency dimension channel attention is It is the mth in U i The rows corresponding to important channels, yes The k-th channel in the matrix; finally, the channel attention of m sub-features is connected, input into a neural network to obtain the final channel attention in the frequency dimension, and then compared with X. Q Multiplication Update X Q .

[0024] In this embodiment, by capturing visual signals from a small number of network devices using the camera of a mobile augmented reality device such as a smartphone, their visual geometry can be calculated. Then, a registration deep learning network is used to match these signals with the geometric structures of nodes registered during the deployment phase. This automatically identifies network devices within the augmented reality field of view without relying on their unique shape, identification code, world coordinates, or wireless signals. This method can be widely used for various wired and wireless network devices, effectively improving the efficiency of user interaction with network devices.

[0025] The embodiments described in this specification are merely examples of implementations of the inventive concept and are for illustrative purposes only. The scope of protection of this invention should not be considered limited to the specific forms described in these embodiments; rather, it extends to equivalent technical means conceived by those skilled in the art based on the inventive concept.

Claims

1. A method for identifying network devices in augmented reality based on visual geometric structure registration, characterized in that, The method includes the following steps: (1) During the deployment phase, the network device i appearing in the camera view is first visually located to obtain its three-dimensional bounding box in the camera coordinate system. in, Indicates the location of the network device. and Indicate the size and orientation of the network device; then, by... Surface point sampling converts it into a dense point cloud Furthermore, the rotation matrix α is obtained using data from the electronic compass and inertial measurement unit of the augmented reality device. N and through and Will and Mapped to a compact point cloud in a reference coordinate system and points Because augmented reality (AR) devices have limited fields of view, it is often necessary to observe other network devices by moving and rotating the AR device. The rotation matrix α can be obtained using visual history measurement methods by monitoring the camera feeds, electronic compass, and inertial measurement unit data of the AR device. D And offset vector m, and the dense point cloud of network device j observed after the augmented reality device is moved and rotated. Mapped to a reference coordinate system consistent with i: By using the Q of multiple network devices G The fusion of these elements yields a large, dense point cloud Q representing the poses of individual network devices and the geographical associations of multiple network devices; simultaneously, by utilizing the L-axis of multiple nodes... G Constructing a sparse point cloud Q * and L for each network device G It is bound to its ID and stored in the database; (2) In the recognition stage, using the same method as in step (1), a dense point cloud P with a size smaller than Q can be obtained, representing the pose of individual network devices and the geographical associations between multiple network devices appearing in the augmented reality scene during the recognition stage. * Sparse point cloud P * Then, the transformation matrix is ​​obtained through an adaptive iterative scanning method. satisfy Where P k Let P be the k-th point; for network device i that identifies nodes appearing in the augmented reality scene, through... Get its in Q * The corresponding point j in the database is identified by querying the database to obtain the network device ID corresponding to j, thereby achieving identification.

2. The method for identifying network devices in augmented reality based on visual geometric structure registration as described in claim 1, characterized in that, In step (2), the adaptive iterative scanning method works as follows: First, the transformation matrix from P to Q is calculated using an incomplete depth registration network. And continuously estimate of like This indicates that the number of network devices scanned during the identification phase is insufficient, and more nodes need to be scanned to expand P; secondly, if the following conditions are met... Then remove the corresponding point of P from Q and re-register. If it still yields the desired result... satisfy This indicates that the number of network devices scanned during the identification phase is still insufficient, and more nodes need to be scanned to expand P; the above process is repeated until only one remains. satisfy until.

3. The method for identifying network devices in augmented reality based on visual geometric structure registration as described in claim 2, characterized in that, The incomplete deep registration network uses Predator as its backbone network and improves the registration performance of network devices in recognition applications by adding spatial attention in the edge dimension and channel attention in the frequency dimension.

4. The method for identifying network devices in augmented reality based on visual geometric structure registration as described in claim 3, characterized in that, The spatial attention process of the edge dimension is as follows: First, fully connected graphs G(Q) and G(P) are generated using Q and P, where (i, j) represents the graph in G(Q) formed by vertex Q. i and Q j The connected edges are used to initialize the feature vector using information about the edge's position, size, and angle. Where l(·) represents the vertex position, cat(·,·) represents the join operation, and h θ Similarly, for a neural network with parameter θ, the eigenvector of edge (a, b) in G(P) can be calculated. Next, the query vector is calculated. key vector and value vector Update in Among them W q W k and W v It is a learnable weight matrix; update The formula is Where att(·) represents the self-attention module, h σ Let's represent a neural network with parameter σ; then, we will... Input a linear layer and you will get its score. Finally, the eigenvalues ​​of each vertex in G(Q) are iteratively updated using the following formula: in, Q represents i The k-th round feature representation, max represents the max pooling operation, h φ Let's say we have a neural network with parameter φ. Finally, we perform this update twice with non-shared parameter φ, concatenate the results of each round, and input them into another neural network to obtain the final features.

5. The method for identifying network devices in augmented reality based on visual geometric structure registration as described in claim 3, characterized in that, The channel attention operation in the frequency dimension works as follows: This channel attention operation is performed on both G(P) and G(Q). For G(Q), the weight of edge (i, j) is calculated as follows: Then, using the weights of all edges in G(Q), construct the edge weight matrix A and the diagonal matrix D, where the non-zero elements in D are determined by... Calculate |Q|, representing the number of vertices in Q, and compute the normalized Laplacian matrix. Next, through L=U∧U T Decompose L to obtain U; ​​then, decompose the input features of Q. The features are divided into W groups based on the channel dimension, where C represents the feature dimension, and C is divisible by W. Each group contains [number of channels]. Following this method, U is also divided into W parts, and experiments are conducted using its frequency components sequentially to obtain the optimal frequency component ranking; then, the top m most important components in W are selected to pair U and X. Q Grouping, with each group containing [number of channels]. C is divisible by m; for the m-th... i (1≤m) i Sub-features ≤m,m≤W Its frequency dimension channel attention is It is the mth in U i The rows corresponding to important channels, yes The k-th channel in the matrix; finally, the channel attention of m sub-features is connected, input into a neural network to obtain the final channel attention in the frequency dimension, and then compared with X. Q Multiplication Update X Q .

Citation Information

Patent Citations

  • Fishing boat mark identification method and system based on deep learning technology

    CN115424275A

  • Repositioning method and apparatus in camera pose tracking process, device, and storage medium

    US20200302615A1