Systems and methods for virtual and augmented reality

JP7901142B2Active Publication Date: 2026-08-05MAGIC LEAP INC
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
MAGIC LEAP INC
Filing Date
2024-12-30
Publication Date
2026-08-05

AI Technical Summary

Benefits of technology

【0021】 本発明はまた、深層ミドルエンドマッチャアーキテクチャのアテンショングラフニューラルネットワークの特徴点エンコーダを用いて、特徴点位置pおよびその視覚的記述子dを単一ベクトルの中にマッピングするステップと、深層ミドルエンドマッチャアーキテクチャのアテンショングラフニューラルネットワークの交互セルフおよびクロスアテンション層を用いて、ベクトルに基づいて、L回の繰り返される回数にわたって実行し、表現fを作成するステップと、深層ミドルエンドマッチャアーキテクチャのアテンショングラフニューラルネットワークの最適マッチング層を実行し、M×Nスコア行列を表現fから作成し、M×Nスコア行列に基づいて、最適部分的割当を見出すステップとを含み得る、コンピュータ実装方法を提供する。 本発明は、例えば、以下の項目を提供する。 (項目1) コンピュータシステムであって、 コンピュータ可読媒体と、 前記コンピュータ可読媒体に接続されるプロセッサと、 コンピュータ可読媒体上の命令のセットと を備え、 前記コンピュータ可読媒体上の命令のセットは、 深層ミドルエンドマッチャアーキテクチャであって、 アテンショングラフニューラルネットワークであって、前記アテンショングラフニューラルネットワークは、 特徴点位置pおよびその視覚的記述子dを単一ベクトルの中にマッピングするための特徴点エンコーダと、 前記ベクトルに基づいて、L回繰り返され、表現fを作成する交互セルフおよびクロスアテンション層と を有する、アテンショングラフニューラルネットワークと、 最適マッチング層であって、前記最適マッチング層は、M×Nスコア行列を前記表現fから作成し、前記M×Nスコア行列に基づいて、最適部分的割当を見出す、最適マッチング層と を含む、深層ミドルエンドマッチャアーキテクチャ を含む、コンピュータシステム。 (項目2) 前記特徴点エンコーダにおいて、特徴点i毎の初期表現 【数】 が、視覚的外観と場所とを組み合わせ、 【数】 のように、個別の特徴点位置が、多層パーセプトロン(MLP)とともに高次元ベクトルの中に埋め込まれる、項目1に記載のコンピュータシステム。 (項目3) 前記特徴点エンコーダは、前記アテンショングラフニューラルネットワークが、外観および位置についてともに推測することを可能にする、項目2に記載のコンピュータシステム。 (項目4) 前記特徴点エンコーダは、前記2つの画像の特徴点であるノードを伴う単一完全グラフを有する多重グラフニューラルネットワークを含む、項目1に記載のコンピュータシステム。 (項目5) 前記グラフは、多重グラフであり、前記多重グラフは、2つのタイプの非指向性エッジ、すなわち、特徴点iを同一画像内の全ての他の特徴点に接続する画像内エッジ(セルフエッジ、Eself)と、特徴点iを他の画像内の全ての特徴点に接続する画像間エッジ(クロスエッジ、Ecross)とを有し、結果として生じる多重グラフニューラルネットワークが、ノード毎に、高次元状態から開始し、各層において、全てのノードに関する全ての所与のエッジを横断してメッセージを同時に集約することによって、更新された表現を算出するように、メッセージパッシング公式を使用して、両方のタイプのエッジに沿って情報を伝搬する、項目4に記載のコンピュータシステム。 (項目6) 【数】 が、層lにおける画像A内の要素iに関する中間表現である場合、メッセージmE→iは、全ての特徴点{j:(i,j)∈E}からの集約の結果であり、E∈{Eself,Ecross}であり、A内の全てのiに関する残りのメッセージパッシング更新は、 【数】 であり、式中、[·||·]は、連結を示す、項目5に記載のコンピュータシステム。 (項目7) 異なるパラメータを伴う固定された数の層Lが、連鎖され、l=1から開始して、lが奇数である場合、E=Eselfであり、lが偶数である場合、E=Ecrossであるように、前記セルフおよびクロスエッジに沿って、交互に集約される、項目6に記載のコンピュータシステム。 (項目8) 前記交互セルフおよびクロスアテンション層は、前記メッセージmE→iを算出し、前記集約を実施するアテンション機構を用いて算出され、前記セルフエッジは、セルフアテンションに基づき、前記クロスエッジは、クロスアテンションに基づき、iの表現に関して、クエリqiが、その属性であるキーkjに基づいて、いくつかの要素の値vjを読み出し、前記メッセージは、 【数】 のように、前記値の加重平均として算出される、項目6に記載のコンピュータシステム。 (項目9) アテンションマスクαijは、 【数】 のように、前記キー·クエリ類似性にわたるソフトマックスである、項目8に記載のコンピュータシステム。 (項目10) 前記個別のキー、クエリ、および値は、前記グラフニューラルネットワークの深層特徴の線形投影として算出され、クエリ特徴点iは、画像Q内にあり、全てのソース特徴点は、画像S内にあり、方程式 【数】 において、(Q,S)∈{A,B}2である、項目8に記載のコンピュータシステム。 (項目11) 前記交互セルフおよびクロスアテンション層の最終マッチング記述子は、 【数】 の線形投影である、項目1に記載のコンピュータシステム。 (項目12) 前記最適マッチング層は、 【数】 のように、セットに関する対毎スコアをマッチング記述子の類似性として表し、式中、<·,·>は、内積であり、学習された視覚的記述子とは対照的に、前記マッチング記述子は、正規化されず、その大きさは、特徴あたりで変化し、訓練の間、予測信頼度を反映し得る、項目1に記載のコンピュータシステム。 (項目13) 前記最適マッチング層は、オクルージョンおよび可視性のために、オクルードされる特徴点を抑制し、マッチングされない特徴点がダストビンスコアに明示的に割り当てられるように、特徴点の各セットをダストビンスコアで拡張する、項目12に記載のコンピュータシステム。 (項目14) 前記スコアSは、 【数】 のように、新しい行および列に、単一学習可能パラメータで充填される点/ビンおよびビン/ビンスコアを付加することによって、S-に拡張される、項目13に記載のコンピュータシステム。 (項目15) 前記最適マッチング層は、T回の反復にわたって、シンクホーンアルゴリズムを使用して、前記M×Nスコア行列に基づいて、前記最適部分的割当を見出す、項目13に記載のコンピュータシステム。 (項目16) T回の反復後、前記最適マッチング層は、前記ダストビンスコアをドロップし、P=P-1:M,1:Nを復元し、 【数】 は、オリジナル割当であり、 【数】 は、前記ダストビンスコア拡張を伴う割当である、項目15に記載のコンピュータシステム。 (項目17) コンピュータ実装方法であって、 深層ミドルエンドマッチャアーキテクチャのアテンショングラフニューラルネットワークの特徴点エンコーダを用いて、特徴点位置pおよびその視覚的記述子dを単一ベクトルの中にマッピングすることと、 前記深層ミドルエンドマッチャアーキテクチャのアテンショングラフニューラルネットワークの交互セルフおよびクロスアテンション層を用いて、前記ベクトルに基づいて、L回の繰り返される回数にわたって実行し、表現fを作成することと、 前記深層ミドルエンドマッチャアーキテクチャのアテンショングラフニューラルネットワークの最適マッチング層を実行し、M×Nスコア行列を前記表現fから作成し、前記M×Nスコア行列に基づいて、最適部分的割当を見出すことと を含む、方法。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007901142000066
    Figure 0007901142000066
  • Figure 0007901142000067
    Figure 0007901142000067
  • Figure 0007901142000068
    Figure 0007901142000068
Patent Text Reader

Abstract

To provide a system and a method for virtual and augmented reality.SOLUTION: This invention is related to connected mobile computing systems, methods, and configurations, and more specifically to mobile computing systems, methods, and configurations featuring at least one wearable component which may be utilized for virtual and / or augmented reality operations. The description relates to feature matching. Our approach establishes pointwise correspondences between challenging image pairs. It takes off-the-shelf local features as input and uses an attentional graph neural network to solve an assignment optimization problem. The deep middle-end matcher acts as a middle-end and handles partial point visibility and occlusion elegantly, producing a partial assignment matrix.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] (Cross-reference of related applications) This application claims priority to U.S. Provisional Patent Application No. 62 / 935,597, filed on 14 November 2019, which is incorporated herein by reference in its entirety.

[0002] The present invention relates to a connected mobile computing system, method, and configuration, and more specifically to a mobile computing system, method, and configuration characterized by at least one wearable component that can be used for virtual reality and / or augmented reality operations. [Background technology]

[0003] The mixed reality or augmented reality eyepiece display is desirable to be lightweight, low-cost, have a small form factor, have a wide virtual image field of view, and be as transparent as possible. In addition, it is desirable to have a configuration that presents virtual image information in multiple focal planes (e.g., two or more) in order to be practical for a wide variety of use cases without exceeding an acceptable tolerance for mismatch between convergence / divergence motion and near / far accommodation. Referring to Figure 8, an augmented reality system is illustrated, featuring a head-mounted viewing component (2), a handheld controller component (4), and an interconnected auxiliary computing or controller component (6) which may be configured to be worn on the user as a belt pack or equivalent. Each of these components may be operably coupled (10, 12, 14, 16, 17, 18) to each other and to other connected resources (8) such as cloud computing or cloud storage resources via wired or wireless configurations, such as those defined by IEEE 802.11, Bluetooth (RTM), and other connectivity standards and configurations. For example, U.S. Patent Applications Nos. 14 / 555,585, 14 / 690,401, 14 / 331,218, 15 / 481,255, 62 / 627,155, 62 / 518,539, 16 / 229,532, 16 / 155,564, 15 / 413,284, 16 / 020,541, 62,702,322, 62 / 206,765, 15,597,694, 16 / 221,065, and 15 / 968,673 Various aspects of such components are described, such as various embodiments of two depicted optical elements (20) through which the user can see the surrounding world together with a visual component that may be produced by associated system components, for an augmented reality experience.As illustrated in Figure 8, such a system may also include, but is not limited to, various sensors configured to provide information about the user's surroundings, including various camera-type sensors (22, 24, 26) (monochrome, color / RGB, and / or thermal imaging components, etc.), a depth camera sensor (28), and / or sound sensors (30) such as a microphone. There is a need for small, continuously connected wearable computing systems and assemblies, such as those described herein, which can be used to provide users with a rich augmented reality experience. [Overview of the project] [Means for solving the problem]

[0004] This book describes an aspect of what might be called a “deep middle-end matcher,” a neural network configured to match two sets of local features by finding correspondences and eliminating points of non-matching. Such a neural network configuration can be used in conjunction with spatial computing resources, such as those illustrated in Figure 8, including cameras and processing resources, which constitute such a spatial computing system, though not limited to these. Within a deep middle-end matcher type configuration, assignments can be estimated by solving an optimal transport problem, and their cost is predicted by a graph neural network. We describe an attention-based, flexible context aggregation mechanism that enables the deep middle-end matcher configuration to infer together the underlying 3D scene and feature assignments. Compared to conventional manually designed heuristics, our technique learns initial values ​​across geometric transformations and regularities of the 3D world through end-to-end image-to-correspond training. The deep middle-end matcher is superior to other learned approaches and brings a new, state-of-the-art technique to the task of pose estimation in challenging real-world indoor and outdoor environments. These methods and configurations can be adapted in real time to modern graphical processing units ("GPUs") and easily integrated into modern Structural Analysis from Motion ("SfM") or Simultaneous Localization and Mapping ("SLAM") systems, all of which may be incorporated into systems such as those illustrated in Figure 8.

[0005] The present invention provides a computer system including a computer-readable medium, a processor connected to the computer-readable medium, and a set of instructions on the computer-readable medium. The set of instructions may include an attention graph neural network having a feature point encoder for mapping a feature point position p and its visual descriptor d into a single vector, and alternating self and cross attention layers that are repeated L times based on the vector to create a representation f, and an optimal matching layer that creates an M×N score matrix from the representation f and finds an optimal partial assignment based on the M×N score matrix, including a deep middle-end matcher architecture.

[0006] The computer system may further include, in the feature point encoder, an initial representation for each feature point i

Chemical formula

Chemical formula

[0007] The computer system may further include that the feature point encoder enables the attention graph neural network to infer both about appearance and location.

[0008] The computer system may further include that the feature point encoder includes a multi-graph neural network having a single complete graph with nodes that are feature points of two images.

[0009] The computer system may further include that the graph has two types of undirected edges, namely, in-image edges (self-edges; E that connect feature point i to all other feature points within the same image) self) and an inter-image edge (cross edge, E) that connects feature point i to all feature points in other images. cross The resulting multigraph neural network may be a multigraph that propagates information along both types of edges using a message passing formula, such that each node starts from a higher-dimensional state and, in each layer, simultaneously aggregates messages across all given edges relating to all nodes to compute an updated representation.

[0010] The computer system further, [ka] However, if it is an intermediate representation of element i in image A in layer l, then message m E→i However, this is the result of aggregation from all feature points {j:(i,j)∈E}, and E∈{E self ,E cross} and the remaining message passing updates for all i in A are as follows: [ka] In the formula, [·||·] may also indicate concatenation.

[0011] The computer system further consists of a fixed number of layers L with different parameters, starting from l=1, and if l is odd, then E=E self And if l is an even number, then E = E cross This may include alternating aggregation along the self and cross edges.

[0012] The computer system further includes alternating self and cross-attention layers, and messages m E→i The following is calculated and aggregated using an attention mechanism: self-edges are calculated based on self-attention, and cross-edges are calculated based on cross-attention, with respect to the representation of i, query qi is the key k, which is its attribute j Based on this, the values v of some elements j are read, and the message may include being calculated as a weighted average of the values as follows

Chemical formula

[0013] The computer system may further include that the attention mask α ij is a softmax over key-query similarities as follows

Chemical formula

[0014] The computer system may further include that individual keys, queries, and values are calculated as linear projections of the deep features of a graph neural network, the query feature point i is within the image Q, all source feature points are within the image S, and in the following equation, (Q, S) ∈ {A, B} 2 and may include being as follows

Chemical formula

[0015] The computer system may further include that the final matching descriptors of the alternating self and cross-attention layers are the following linear projections

Chemical formula

[0016] The computer system may further include that the optimal matching layer represents the per-pair score for the set as the similarity of the matching descriptors as follows

Chemical formula

[0017] The computer system may further include extending each set of feature points with dust bin scores so that the optimal matching layer suppresses occluded feature points for occlusion and visibility, and unmatched feature points are explicitly assigned to dust bin scores.

[0018] The computer system further adds point / bin and bin / bin scores to the score S, which are filled with a single learnable parameter in new rows and columns, as follows: - This may include being extended to include. [ka]

[0019] The computer system may further include an optimal matching layer that, over T iterations, uses the Sinkhorn algorithm to find the optimal partial assignment based on an M × N score matrix.

[0020] The computer system further determines that after T iterations, the optimal matching layer drops the dustbin score, and P=P - 1:M,1:N Restore, [ka] However, it is the original allocation. [ka] However, this may include assignments that involve extensions that are converted into dustbin scores.

[0021] The present invention also provides a computer implementation method which may include the steps of: mapping a feature point position p and its visual descriptor d into a single vector using a feature point encoder of an attention graph neural network of a deep middle-end matcher architecture; performing an alternating self and cross-attention layer of an attention graph neural network of a deep middle-end matcher architecture based on the vector over L iterations to create a representation f; and running an optimal matching layer of an attention graph neural network of a deep middle-end matcher architecture to create an M×N score matrix from the representation f, and finding the optimal partial assignment based on the M×N score matrix. The present invention provides, for example, the following items: (Item 1) A computer system, Computer-readable media and, A processor connected to the aforementioned computer-readable medium, A set of instructions on a computer-readable medium and Equipped with, The set of instructions on the computer-readable medium is It is a deep middle-end matcher architecture, An attention graph neural network, wherein the attention graph neural network is A feature point encoder for mapping feature point location p and its visual descriptor d into a single vector, Based on the aforementioned vector, alternating self and cross attention layers are repeated L times to create representation f. An attention graph neural network having, An optimal matching layer, wherein the optimal matching layer creates an M×N score matrix from the representation f and finds the optimal partial assignment based on the M×N score matrix. A deep middle-end matcher architecture, including A computer system, including a computer system. (Item 2) In the feature point encoder, the initial representation for each feature point i

number

number

number

number

number

number

number

number

number

number

number

number

[0022] The present invention will be further described by reference to the accompanying drawings.

[0023] [Figure 1] Figure 1 is a representative sketch illustrating feature matching using a deep middle-end matcher.

[0024] [Figure 2] Figure 2 shows the correspondence estimated by the deep middle-end matcher for two difficult indoor image pairs.

[0025] [Figure 3] Figure 3 is a representative sketch illustrating a formula for a deep middle-end matcher and a method for solving optimization problems.

[0026] [Figure 4] Figure 4 is an image showing the mask as a ray.

[0027] [Figure 5] Figure 5 is a graph showing indoor and outdoor posture estimation.

[0028] [Figure 6] Figure 6 shows qualitative image matching.

[0029] [Figure 7]Figure 7 shows the steps for visualizing attention within self and cross-attention masks at various layers and heads.

[0030] [Figure 8] Figure 8 shows an augmented reality system. [Modes for carrying out the invention]

[0031] Detailed explanation Finding correspondences between points within an image is an essential step for computer vision tasks to tackle 3D reconstruction or visual localization, such as simultaneous localization and mapping (SLAM) and structural analysis from motion (SfM). These are processes known as data association, which involve matching local features and then estimating 3D structure and camera orientation from such correspondences. Factors such as large viewpoint changes, occlusion, blurring, and lack of texture make 2D / 2D data association particularly challenging.

[0032] In this explanation, we present a novel approach to considering the feature matching problem. Instead of learning better task-independent local features followed by simple matching heuristics and tricks, we propose learning the matching process from existing local features using a novel neural architecture called the Deep Middle-End Matcher (DMEM). In the context of SLAM, where the problem is typically divided into a visual feature detection front-end and a bundle adjustment or pose estimation back-end, our network is directly in the middle; i.e., the Deep Middle-End Matcher is the learnable middle-end. Figure 1 illustrates feature matching using the Deep Middle-End Matcher. Our approach establishes point-by-point correspondences between difficult image pairs. This takes off-the-shelf local features as input and uses an attention graph neural network to solve the assignment optimization problem. The Deep Middle-End Matcher acts as the middle-end, accurately handling partial point visibility and occlusion and producing a partial assignment matrix.

[0033] In this study, learned feature matching is considered to be finding partial assignments between two sets of local features. We rethink the classical graph-based strategy for matching by solving a linear assignment problem that, when reduced to an optimal transport problem, can be solved discriminatorily [see references 50, 9, 31 below]. The cost function of this optimization is predicted by a graph neural network (GNN). Inspired by the success of Transformer [see reference 48 below], both the spatial relationships of feature points and their visual appearance are leveraged using self (in-image) and cross (inter-image) attention. This formula enforces the assignment structure of prediction while enabling the cost to learn complex initial values ​​and accurately handle occlusion and non-reproducible feature points. Our method is trained end-to-end from image to correspondence, i.e., it learns initial values ​​for pose estimation from a large annotated dataset, enabling a deep middle-end matcher to infer about 3D scenes and assignments. Our research can be applied to various multi-view geometric shape problems requiring high-quality feature correspondence.

[0034] The superiority of the deep middle-end matcher over both manual matchers and trained correct correspondence classifiers is demonstrated. Figure 2 shows the correspondences estimated by the deep middle-end matcher for two difficult indoor image pairs. The deep middle-end matcher successfully estimates the correct pose, while other trained or manual methods fail (correct correspondences are shown in green). The proposed method yields the most substantial improvements when combined with the deep front-end, SuperPoint [see reference 14 below], thereby advancing homography estimation and indoor and outdoor pose estimation tasks to the cutting edge and laying the groundwork for deep SLAM.

[0035] 2. Related Research

[0036] Local feature matching is generally performed by i) detecting points of interest, ii) calculating visual descriptors, iii) matching these with nearest neighbor (NN) search, iv) filtering out incorrect matches, and finally, v) estimating geometric transformations. Classical pipelines developed in the 2000s are often based on SIFT [see reference 25 below], filtering matches using heuristics such as Lowe's proportion test [see reference 25 below], cross-checking, and nearest neighbor consensus [see references 46, 8, 5, 40 below], and finding transformations using robust solvers such as RANSAC [see references 17, 35 below].

[0037] Recent research on deep learning for matching often focuses on using convolutional neural networks (CNNs) to learn better sparse detectors and local descriptors [see references 14, 15, 29, 37, 54 below] from data. To improve its discriminative ability, some studies explicitly look to broader context using region features [see reference 26 below] or log-polar patches [see reference 16 below]. Other approaches learn to filter matches by classifying them into correct and incorrect matches [see references 27, 36, 6, 56 below]. These still act on a set of matches estimated by NN search, and therefore ignore assignment structure and discard visual information. To actually learn to match, research has so far focused on dense matching [see reference 38 below] or 3D point clouds [see reference 52 below], and still exhibits such limitations. In contrast, our learnable middleend simultaneously performs context aggregation, matching, and filtering within a single end-to-end architecture.

[0038] Graph matching problems are typically formulated as quadratic assignment problems, which are NP-hard, expensive, complex, and therefore require impractical solvers [see reference 24 below]. Regarding local features, the computer vision literature of the 2000s [see references 4, 21, 45 below] uses manual costs with many heuristics, making it complex and fragile. Caetano et al. [see reference 7 below] learn the cost of optimization for simpler linear assignments but use shallow models, while our deep middle-end matcher uses neural networks to learn flexible costs. Related to graph matching is the optimal transport problem [see reference 50 below], i.e., the Sinkhorn algorithm [see references 43, 9, 31 below], which is a generalized linear assignment that is efficient but has simple approximate solutions.

[0039] Deep learning for sets such as point clouds aims to design permutation-equivalent or invariant functions by aggregating information across elements. Some studies treat all of them equally through global pooling [see references 55, 32, 11 below] or instance normalization [see references 47, 27, 26 below], while others focus on local neighborhoods within coordinate or feature space [see references 33, 53 below]. Attention [see references 48, 51, 49, 20 below] can perform both global and data-dependent local aggregation by focusing on specific elements and attributes, and is therefore more flexible. Our study uses the fact that a particular instance of a message-passing graph neural network [see references 18, 3 below] can be recognized on a complete graph. [See references 22, 57 below] By applying attention to multi-edge, i.e., multiple graphs, a deep middle-end matcher can learn complex inferences about two sets of local features.

[0040] 3. Deep Middle-End Matcher Architecture

[0041] Motivation: In image matching problems, several regularities of the world can be utilized. That is, the 3D world is primarily smooth, and sometimes planar, and all correspondences for a given pair of images are derived from a single epipolar transform when the scene is static, with some poses being more likely than others. In addition, 2D feature points are usually projections of prominent 3D points such as angles or blobs, and therefore, correspondences across images must conform to certain physical constraints. That is, i) a feature point can have at most one correspondence in another image, and ii) some feature points will not be matched due to occlusion and detector failures. An effective model for feature matching should aim to find all correspondences between reprojections of the same 3D point and identify feature points that do not have a match. Figure 3 shows how to formulate a deep middle-end matcher as a step to solve the optimization problem, and its cost is predicted by a deep neural network. The deep middle-end matcher includes two main components: an attention graph neural network (section 3a) and an optimal matching layer (section 3b). The first component uses a feature point encoder to map feature point locations p and their visual descriptors d into a single vector, and then uses alternating self and cross-attention layers (repeated L times) to create a more effective representation f. The optimal matching layer creates an M×N score matrix, expands it with a dust bin, and then uses the Sinkhorn algorithm (over T iterations) to find the optimal partial assignment. This reduces the need for specific domain expertise and heuristics, i.e., learns relevant initial values ​​directly from the data.

[0042] Formula: Consider two images A and B, each with a set of visual descriptors d associated with a feature point location p. We refer to them together as local features (p,d). A feature point has x and y image coordinates and a detection confidence c, i.e., p. i :=(x,y,c) iIt consists of visual descriptor d. i ∈R D These can be extracted by a CNN like SuperPoint or a conventional descriptor like SIFT. Images A and B have M and N local features, respectively, with the sets of feature point indices A:={1,...,M} and B:={1,...,N}.

[0043] Partial assignment: Constraints i) and ii) mean that the correspondence is derived from a partial assignment between two sets of feature points. For integration into downstream tasks and better interpretability, each possible correspondence should have a certain confidence value. As a result, the partial soft assignment matrix P ∈ [0,1] is as follows: M×N Define. [ka]

[0044] Our goal is to design a neural network that predicts assignment P from two sets of local features.

[0045] 3.1 Attention Graph Neural Networks

[0046] The first major block of the deep middle-end matcher (see Section 3a) is the attention graph neural network, whose job is as follows: given initial local features, by linking the features together, the matching descriptor is f i ∈R D The goal is to calculate the long-range feature linkage, which is essential for robust matching and requires the aggregation of information from within images and across image pairs.

[0047] Intuitively, distinctly different information about a given feature point depends not only on its visual appearance and location, but also on its spatial and visual relationships to other simultaneously visible feature points, such as neighbors or prominent ones. On the other hand, knowledge of feature points in a second image can help resolve ambiguity by comparing candidate matches or by inferring relative photometric or geometric transformations from global and unambiguous cues.

[0048] When asked to match a given set of ambiguous feature points, humans repeatedly compare both images. That is, they select tentative matching feature points, examine each of them, and look for contextual clues that help clarify the true match from other self-similarities. This suggests an iterative process in which attention can be directed to specific locations.

[0049] Feature point encoder: Initial representation for each feature point i [ka] This combines visual appearance and location. Feature point locations are embedded in a high-dimensional vector along with a multilayer perceptron (MLP), as shown below. [ka]

[0050] An encoder is an instance of a “position encoder” that is introduced into a transformer [see reference 48 below], enabling the network to make inferences about both appearance and location (particularly effective when using an attention mechanism).

[0051] Multiple Graph Neural Networks: Consider a single complete graph whose nodes are feature points of both images. The graph is a multiple graph, i.e., it has two types of non-directional edges: in-image edges, i.e., self-edges E selfThis connects feature point i to all other feature points within the same image. This creates an inter-image edge, i.e., a cross edge E. cross This connects feature point i to all feature points in other images. Information is propagated along both types of edges using the message passing formula [see references 18, 3 below]. The resulting multi-graph neural network computes an updated representation at each layer, starting from a higher-dimensional state for each node and simultaneously aggregating messages across all given edges for all nodes.

[0052] [ka] Let be an intermediate representation of element i in image A in layer l. Message m E→i This is the result of aggregation from all feature points {j:(i,j)∈E}, where E∈{E self ,E cross}. The remaining message passing updates for all i in A are as follows: [ka] In the formula, [·||·] indicates concatenation. A similar update can be performed simultaneously for all feature points in image B. A fixed number of layers L with different parameters are chained together and alternately aggregated along self and cross edges. Thus, starting from l=1, if l is odd, then E=E self And if l is an even number, then E = E cross That is the case.

[0053] Attention aggregation: The attention mechanism receives messages m E→i The following is calculated and aggregated: Self-edges are based on self-attention [see reference 48 below], and cross-edges are based on cross-attention. Similar to database reading, the representation of i is used for query q. i This is its attribute, key k j Based on this, several element values ​​vj Read the data. Calculate the message as a weighted average of its values, as shown below. [ka] During the formula, attention mask α ij This is a softmax over key-query similarity, as follows: [ka]

[0054] The key, query, and value are calculated as a linear projection of the deep features of the graph neural network. The query feature point i lies in image Q, all source feature points lies in image S, and (Q,S)∈{A,B} 2 Therefore, it can be written as follows: [ka]

[0055] Each layer l has its own projection parameters, which are shared with respect to all feature points in both images. In practice, multi-head attention is used to improve representation [see reference 48 below].

[0056] Our formula provides maximum flexibility because the network can learn to focus on a subset of feature points based on specific attributes. In Figure 4, the mask αij is shown as a ray. Attention aggregation builds a dynamic graph between feature points. Self-attention (top) can direct attention to any location within the same image, e.g., a distinctly different location, and is therefore not limited to neighboring locations. Cross-attention (bottom) directs attention to locations in other images, such as latent matches, that have similar local appearances. The deep middle-end matcher represents both appearance and feature point locations x iBecause they are encoded internally, they can be read out or attention directed to based on them. This involves the step of directing attention to neighboring feature points and reading out the relative positions of similar or prominent feature points. This enables the representation of geometric transformations and assignments. The final matching descriptor is a linear projection such as, [ka] The same applies to the feature points within B.

[0057] 3.2. Optimal Matching Layer

[0058] The second major block of the deep middle-end matcher (see Section 3b) is the optimal matching layer, which produces a partial assignment matrix. As in the standard graph matching formula, the assignment P is a score matrix S ∈ R for all possible matchings. M×N Calculate the total score under the constraints in equation 1. [ka] This can be obtained by maximizing [the value]. This is equivalent to solving the linear assignment problem.

[0059] Score prediction: Constructing separate representations for all (M+1)×(N+1) potential matches would be impractical. Instead, we represent each pair score as the similarity of the matching descriptors, as follows: [ka] In the formula, <·,·> represents the inner product. In contrast to learned visual descriptors, matching descriptors are not normalized, their size varies per feature, and may reflect predictive confidence during training.

[0060] Occlusion and Visibility: To suppress feature points that are occluded in the network, each set is extended with a dust bin so that unmatched feature points are explicitly assigned to it. This technique is common in graph matching, and dust bins are also used by SuperPoint [see reference 14 below] to account for image cells that cannot be detected. The score S is then added to the new rows and columns by adding point / bins and bin / bin scores filled with a single learnable parameter, as shown below. - Extend to... [ka]

[0061] Each feature point in A will be assigned to a single feature point or dustbin in B, but each dustbin will have a similar degree of matching to feature points in other sets, i.e., N and M matching to the dustbins in A and B. [ka] This shows the expected number of matches for each feature point and dustbin in A and B. Extended assignment P - Here, the following constraints apply. [ka]

[0062] Sinkhorn algorithm: The solution to the above optimization problem is score S -This corresponds to the optimal transport between discrete distributions a and b [see reference 31 below]. This can be approximately solved using the Sinkhorn algorithm [see references 43, 49 below], which is a discriminable version of the Hungarian algorithm [see reference 28 below] classically used for bipartite matching. This solves the regularized transport problem and inevitably leads to soft assignment. This normalization is equivalent to iteratively performing alternating softmax along rows and columns and is therefore easily parallelized on a GPU. After T iterations, the dustbin is dropped and P=P - 1:M,1:N Restore it.

[0063] 3.3.Loss

[0064] By design, both the graph neural network and the optimal matching layer are discriminable; that is, this enables backpropagation from matching to visual descriptors. The deep middle-end matcher performs ground truth matching in a supervised manner. [ka] These are trained from ground-truth relative transformations, i.e., using attitude and depth maps or homography. This also includes several feature points. [ka] If they do not have any reprojections within their vicinity, they are marked as unmatched. Based on this marking, assignment P - Minimize the negative log-likelihood of [the expression]. [ka]

[0065] This teacher also aims to maximize the accuracy and recall of the matching.

[0066] 3.4. Comparison with related studies

[0067] Deep middle-end matchers versus direct correspondence classifiers [see references 27 and 56 below]: Deep middle-end matchers benefit from strong induction bias by being permutationally equivariant overall for both image and local features. In addition, they directly embed commonly used cross-check constraints into training, i.e., probabilities P greater than 0.5. i,j Any matching involving this will inevitably be consistent with each other.

[0068] Deep Middle-End Matcher vs. Instance Normalization [See Reference 47 below]: Attention, as used by the deep middle-end matcher, treats all feature points equally and is a more flexible and effective context aggregation mechanism than instance normalization [See References 27, 56, 26 below], which has been used in previous research on feature matching.

[0069] Deep Middle-End Matcher vs. ContextDesc [see Reference 26 below]: The Deep Middle-End Matcher can infer both appearance and location, while ContextDesc processes them separately. In addition, ContextDesc is a front-end that also requires a larger region extractor and loss for feature point scoring. The Deep Middle-End Matcher only requires local features that are learned or manually processed, and therefore can be a simple drop-in replacement for existing matchers.

[0070] Deep Middle-End Matcher vs. Transformer [See Reference 48 below]: The deep middle-end matcher borrows self-attention from the transformer but embeds it within the graph neural network, and also introduces cross-attention, which is symmetric. This simplifies the architecture and results in better feature reuse across layers.

[0071] 4. Implementation Details

[0072] The deep middle-end matcher can be combined with any local feature detector and descriptor, but works particularly well with SuperPoint [see reference 14 below], which produces reproducible and sparse feature points, i.e., enables highly efficient matching. Visual descriptors are sampled bilinearly from a semi-dense feature map, which is discriminable. Both local feature extraction and subsequent "gluing" are performed directly on the GPU. During testing, confidence thresholds can be used to extract matches from soft assignments, either by reserving some or simply by using all of them and their confidence in subsequent steps such as weighted pose estimation.

[0073] Architecture Details: All intermediate representations (keys, query values, descriptors) have the same dimensions D=256 as the SuperPoint descriptor. For numerical stability, T=100 Sinkhorn iterations are performed in logarithmic space using L=9 layers of alternating multi-head self and cross-attention, each with 4 heads. The model is implemented in PyTorch [see reference 30 below] and launched in real time on the GPU. The forward pass takes an average of 150ms (7FPS).

[0074] Training Details: To enable data augmentation, SuperPoint detection and description steps are performed on the fly as batches during training. Several random feature points are further added for efficient batching and increased robustness. Further details are provided in Appendix A.

[0075] 5. Experiment

[0076] 5.1 Homography Estimation

[0077] We will conduct large-scale homography estimation experiments using both robust (RANSAC) and non-robust (DLT) estimators, employing real images and synthetic homography.

[0078] Image pairs are generated by sampling random homography and applying random photometric distortion to real images, following a recipe similar to that of the dataset [see references 12, 14, 37, and 36 below]. The base images are derived from a set of one million disruptive images in the Oxford and Paris dataset [see reference 34 below], which are split into training, validation, and test sets.

[0079] Baseline: The deep middle-end matcher is compared to several matchers applied to SuperPoint local features, namely the nearest neighbor (NN) matcher and various mismatch rejectors, namely mutual check (or cross-check), PointCN [see reference 27 below], and order-aware network (OANet) [see reference 56 below]. All learned methods, including the deep middle-end matcher, are trained on ground truth correspondences, which are found by projecting feature points from one image to another. Homography and photometric distortion are generated on the fly, i.e., image pairs are never seen twice during training.

[0080] Metrics: Matching accuracy (P) and recall (R) are calculated from ground truth correspondence. Homography estimation is performed using both RANSAC and direct linear transformation (DLT) with direct least squares solutions [see reference 19 below]. The mean reprojection error at the four corners of the image is calculated, and the area under the cumulative error curve (AUC) for up to 10 pixels is reported.

[0081] Results: The deep middle-end matcher is sufficiently expressive to master homography, achieving 98% recall and high accuracy. Table 1 shows homography estimations for the deep middle-end matcher, DLT, and RANSAC. The deep middle-end matcher reconstructs almost all possible matches while suppressing most mismatches. Because the deep middle-end matcher correspondence is of high quality, the direct linear transformation (DLT), a least-squares-based solution without a robustness mechanism, outperforms RANSAC. The estimated correspondence is so good that a robust estimator is not required; i.e., the deep middle-end matcher works even better with DLT than with RANSAC. Mismatch elimination methods such as PointCN and OANet cannot predict better matches than the NN matcher itself and rely too heavily on the initial descriptor. [Table 1]

[0082] 5.2. Indoor Posture Estimation

[0083] Indoor image matching is extremely difficult due to the lack of texture, numerous self-similarities, complex 3D geometric shapes, and large viewpoint changes. As shown below, deep middle-end matchers can effectively learn initial values ​​and overcome these challenges.

[0084] Dataset: We use ScanNet [see reference 10 below], a large indoor dataset consisting of monocular sequences with ground-truth pose and depth images, and clearly defined training, validation, and test splits corresponding to different scenes. Previous studies typically select training and evaluation pairs based on time lag [see references 29, 13 below] or SfM simultaneous visibility [see references 27, 56, 6 below], calculated using SIFT. We argue that this limits the difficulty of pairs and instead select them based on overlap scores calculated for all possible image pairs within a given sequence, using only ground-truth pose and depth. This results in a significantly broader baseline of pairs, which corresponds to the current state of affairs in real-world indoor image matching. We obtain 230 million training pairs and sample 1,500 test pairs by discarding pairs with too little or too much overlap. Further details are provided in Appendix A.

[0085] Metrics: As in previous studies [see references 27, 56, and 6 below], we report the AUC of the attitude error at thresholds (5, 10, and 20), where the attitude error is the maximum angular error in rotation and translation. Relative attitude is obtained from basic matrix estimation using RANSAC. We also report matching accuracy and matching score [see references 14 and 54 below], where matching is considered correct based on its epipolar distance.

[0086] Baseline: Deep middle-end matchers and various baseline matchers are evaluated using both square root normalized SIFT [see references 25 and 2 below] and SuperPoint [see reference 14 below] features. Deep middle-end matchers are trained with corresponding and unmatched feature points derived from ground truth attitude and depth. All baselines are based on nearest neighbor (NN) matchers and potential mismatch elimination methods. In the "Manual" category, simple cross-checking (mutual), ratio testing [see reference 25 below], descriptor distance thresholding, and the more complex GMS [see reference 5 below] are considered. Methods in the "Learned" category are PointCN [see reference 27 below] and its follow-up OANet [see reference 56 below] and NG-RANSAC [see reference 6 below]. Using the accuracy criteria defined above and their respective regression losses, PointCN and OANet are retrained on ScanNet with classification loss for both SuperPoint and SIFT. For NG-RANSAC, the original trained model is used. We do not include any graph matching methods that would slow down the number of feature points under consideration by several orders of magnitude. For reference, other local features using publicly available trained models, namely ORB with GMS [see reference 39 below], D2-Net [see reference 15 below], and ContextDesc [see reference 26 below], are also reported.

[0087] Results: The deep middle-end matcher enables significantly higher pose accuracy compared to both manual and trained matchers. Table 2 shows wide baseline indoor pose estimation on ScanNet. AUC, matching score (MS), and precision (P) of pose error are all reported in pose estimation AUC percent. The deep middle-end matcher outperforms all manual and trained matchers when applied to both SIFT and SuperPoint. These advantages are substantial when applied to both SIFT and SuperPoint. Figure 5 shows indoor and outdoor pose estimation. The deep middle-end matcher significantly improves pose accuracy over OANet, a state-of-the-art mismatch exclusion neural network. It has significantly higher accuracy than other trained matchers, demonstrating its higher expressiveness. It also produces up to 10 times more correct matches when applied to SIFT because it acts on the complete set of possible matches rather than a limited set of nearest neighbors. Both SuperPoint and the deep middle-end matcher achieve state-of-the-art results in indoor pose estimation. They complement each other well, as reproducible feature points enable the estimation of a greater number of correct matches, even in extremely challenging situations (see Figure 2). [Table 2]

[0088] Figure 6 shows qualitative image matching. A deep middle-end matcher is compared to a nearest neighbor (NN) matcher with two manually and trained mismatch rejectors in three environments. The deep middle-end matcher consistently estimates more accurate matches (green line) and fewer mismatches (red line) against recurring texture, large viewpoint, and lighting changes.

[0089] 5.3. Outdoor posture estimation

[0090] Outdoor image sequences present their own set of challenges (e.g., lighting variations and occlusion), so we train and evaluate a deep middle-end matcher for pose estimation in outdoor settings. We use the same evaluation metrics and baseline methods as those used in indoor pose estimation tasks.

[0091] The evaluation will be conducted on the PhotoTourism dataset, which is part of the CVPR'19 image matching task [see Reference 1 below]. This is a subset of the YFCC100M dataset [see Reference 44 below] and has ground truth poses and sparse 3D models obtained from off-the-shelf SfM tools [see References 29, 41, and 42 below]. For training, the MegaDepth dataset [see Reference 23 below] will be used, which also has clean depth maps calculated using multiview stereo. Scenes within the PhotoTourism test set will be removed from the training set.

[0092] Results: Table 3 shows outdoor pose estimation on the PhotoTourism dataset. Matching SuperPoint and SIFT features using a deep middle-end matcher yields significantly higher pose accuracy (AUC), precision (P), and matching score (MS) than manual or other learned methods. When applied to both SuperPoint and SIFT, the deep middle-end matcher outperforms all baselines at all relative pose thresholds. Most notably, the resulting matching accuracy is very high (84.9%), and the deep middle-end matcher enhances similarities by "stapling" local features together. [Table 3]

[0093] 5.4. Understanding Deep Middle-End Matchers

[0094] Ablation Study: To evaluate our design decisions, we repeat the indoor ScanNet experiment, but this time focusing on different deep middle-end matcher variants. Table 4 shows the ablation of deep middle-end matchers on ScanNet using SuperPoint local features. All deep middle-end matcher blocks are useful and provide substantial performance gains. The differences from the complete model are shown. The optimal matching layer alone is an improvement over the baseline nearest neighbor matcher, but the GNN accounts for the majority of the gains provided by the deep middle-end matcher. Both cross-attention and positional encoding are important for effective patching, and deeper networks further improve accuracy. [Table 4]

[0095] Visualization of Attention: Understanding the proposed technique would not be complete without attempting to visualize the attention patterns of the deep middle-end matcher throughout the matching process. The extensive diversity of self and cross-attention patterns is shown in Figure 7, reflecting the complexity of the learned behavior. Figure 7 shows the visualization of attention, i.e., the self and cross-attention masks αij at various layers and heads. The deep middle-end matcher learns the diversity of patterns and can focus on global or local contexts, self-similarity, distinctly different features, and matching candidates.

[0096] 6. Conclusion

[0097] This disclosure describes what we call a “deep middle-end matcher,” an attention graph neural network inspired by the success of transformers in NLP, for local feature matching. We believe that the data association component of 3D reconstruction pipelines has not received adequate attention from the research community, and that an effective learning-based middle-end is our solution. The deep middle-end matcher enhances the receiving field of local features, disregards features where the correspondence is missing, and effectively performs both the roles of context description and correct correspondence classification. Importantly, the internal structure of the deep middle-end matcher is learned holistically from real-world data. Our results in 2D / 2D feature matching show a significant improvement over existing state-of-the-art solutions.

[0098] Our description herein provides a sufficient demonstration of the use of a learnable middle-end in a feature matching pipeline as a modern deep learning-based alternative to manually designed heuristics. Part of our future research will focus on evaluating deep middle-end matchers within a full 3D reconstruction pipeline.

[0099] Various exemplary embodiments of the present invention are described herein. These embodiments are used for non-limiting purposes only. They are provided to illustrate broader applicable aspects of the present invention. Various modifications may be made to the described invention, and equivalents may be substituted without departing from the true spirit and scope of the invention. In addition, many modifications may be made to adapt specific situations, materials, composition of substances, processes, process actions, or steps to the object, spirit, or scope of the invention. Furthermore, as will be understood by those skilled in the art, each individual modification described and illustrated herein has discrete components and features that can be readily separated from or combined with features of any of several other embodiments without departing from the scope or spirit of the invention. All such modifications are intended to be within the scope of the claims associated with this disclosure.

[0100] The present invention includes methods that can be carried out using the subject device. The methods may include the act of providing such a suitable device. Such provision may be carried out by an end user. In other words, the act of “providing” simply requires the end user to acquire, access, approach, position, configure, activate, launch, or otherwise act to provide the device required in the subject method. The methods enumerated herein may be carried out in any logically possible order of the enumerated events, and in the order in which the events are enumerated.

[0101] Exemplary aspects of the present invention, along with details relating to the selection and manufacture of materials, are described above. Other details of the present invention are understood in relation to the patents and published documents referred to above, and are generally graspable or understandable to those skilled in the art. The same may apply to the method-based aspects of the present invention in terms of additional actions as generally or theoretically adopted.

[0102] In addition, while the present invention is described with reference to several embodiments that optionally incorporate various features, the present invention should not be limited to those described or shown as possible for each modification of the invention. Various modifications and equivalents that can be made to the invention as described (whether listed herein or not included for certain brevity) can be substituted without departing from the true spirit and scope of the invention. In addition, if a range of values ​​is provided, it should be understood that each intermediate value between the upper and lower limits of that range and any other described or intermediate values ​​within that range is included within the scope of the invention.

[0103] Furthermore, it should be assumed that any optional feature of the modified versions of the invention described herein may be described or claimed independently or in combination with any one or more of the features described herein. A reference to a single object includes the possibility that there may be multiple identical articles present. More specifically, as used herein and in the claims associated herein, the singular forms “a,” “an,” “said,” and “the” include multiple supporting objects unless otherwise specifically stated. In other words, the use of articles allows for “at least one” of the subject articles in the above description and in the claims associated with this disclosure. Furthermore, it should be noted that such claims may be drafted to exclude any optional elements. Thus, this statement is intended to serve as a precedent for the use of such exclusive terms, or “negative” restrictions, such as “alone,” “only,” and equivalents, in relation to the enumeration of claim elements.

[0104] Without the use of such exclusive terms, the term “comprising” in a claim as associated with this disclosure shall allow for the inclusion of any additional elements, regardless of whether a given number of elements are enumerated in such claim, or the addition of features may be considered as transforming the nature of the elements described in such claim. Unless specifically defined herein, all technical and scientific terms used herein should be given the broadest possible, generally understood meaning while maintaining the validity of the claims.

[0105] The scope of the present invention should not be limited to the provided examples and / or subject matter specification, but rather should be limited only to the scope of the claim language associated with this disclosure.

[0106] 7. Appendix A - Further experimental details

[0107] Homography estimation:

[0108] The test set contains 1,024 pairs of 640×480 images. Homography is generated by applying random viewpoints, scaling, rotation, and translation to the original full-size images to avoid boundary artifacts. 512 top-scoring feature points detected by SuperPoint with a 4-pixel non-maximum suppression (NMS) radius are evaluated. A correspondence is considered correct if it has a reprojection error of less than 3 pixels. When estimating homography using RANSAC, the opencv function findHomography is used, along with 3,000 iterations and a 3-pixel positive correspondence threshold.

[0109] Indoor posture estimation:

[0110] The overlap score between two images A and B is the mean ratio of pixels in A that are visible in B after accounting for missing depth values ​​and occlusion (and vice versa), by checking for consistency within depth using relative error. The overlap range of 0.4 to 0.8 is used for training and evaluation. For training, 200 pairs per scene are sampled at each baseline time point, as in

[15] . The test set is generated by subsampling sequences in groups of 15, and then sampling 15 pairs every 300 sequences. All ScanNet images and depth maps are resized to VGA 640x480. A maximum of 1,024 SuperPoint feature points (using a publicly available trained model with an NMS radius of 4) and 2,048 SIFT feature points (using an OpenCV implementation) are detected. An epipolar threshold of 5.10e-4 is used when calculating accuracy and matching scores. The pose is partitioned and computed by first estimating the fundamental matrix using OpenCv's findEssentialMat and RANSAC, with a one-pixel positive correspondence threshold divided by the average focal length, followed by recoverPose. In contrast to previous studies [28, 59, 6], a more accurate AUC is calculated using explicit integration rather than a rough histogram.

[0111] Outdoor posture estimation:

[0112] For training on Megadepth, the overlap score is the ratio of visible triangulated feature points in two images, as shown in

[15] . At each baseline time point, pairs with overlap scores within [0.1,0.7] are sampled. For evaluation on the PhotoTourism dataset, all 11 scenes and the overlap scores calculated by Ono

[30] are used, with a selection range of [0.1,0.4]. Images are resized so that their longest edges are less than 1,600 pixels. 2,048 feature points are detected for both SIFT and SuperPoint (with an NMS radius of 3). Other evaluation parameters are the same as those used for indoor evaluation.

[0113] Training for deep-level mid-end matchers:

[0114] To train on homography / indoor / outdoor data, the first 200,000 / 100,000 / 50,000 The learning rate was initially constant at 10e-4 over the number of iterations, followed by 900,000 iterations. The Adam optimizer is used, with exponential decay of 0.999998 / 0.999992 / 0.999992 continuing up to the iteration. When using SuperPoint features, batches are employed with 32 / 64 / 16 image pairs and a fixed number of feature points (512 / 400 / 1,024) per image. When using SIFT features, 1,024 feature points and 24 pairs are used. Due to the limited number of training scenes, the outdoor model is initialized with a homography model. Prior to feature point encoding, feature points are normalized by the maximum edge of the image.

[0115] Ground truth correspondences M and unmatched sets I and J are first generated by calculating an M × N reprojection matrix between all detected feature points using ground truth homography or pose and depth maps. Correspondences are cells along both rows and columns that have the minimum reprojection error, lower than given thresholds, i.e., 3, 5, and 3 pixels, for homography, indoor, and outdoor matching, respectively. For homography, unmatched feature points are simply those that do not appear in M. For indoor and outdoor matching, due to errors in depth and pose, unmatched feature points must also have a minimum reprojection error greater than 15 and 5 pixels, respectively. This still allows for ignoring markings for feature points whose correspondences are ambiguous, while still providing some guidance through sinkhorn normalization.

[0116] 8.References

[0117] The following references are incorporated into this specification as a whole by reference and are referred to in the above description. [Table 5-1] [Table 5-2] [Table 5-3] [Table 5-4] [Table 5-5] [Table 5-6]

Claims

1. A computer system, wherein the computer system is Computer-readable media and A processor connected to the aforementioned computer-readable medium, The set of instructions on the computer-readable medium and Equipped with, The set of instructions on the computer-readable medium includes a deep middle-end matcher architecture that can be executed by the processor. The aforementioned deep middle-end matcher architecture is, An attention graph neural network that creates representations, A score is created from the aforementioned expression, and an optimal matching layer finds the optimal assignment based on the score. Includes, The aforementioned optimal matching layer is [Math 1] As shown above, the pairwise score for each set is represented as the similarity of the matching descriptors, S i,j This is a ratio to each score, <・,・> represents the inner product, A is an image, B is a computer system, which is an image.

2. The aforementioned attention graph neural network is A feature point encoder for mapping a feature point location p and a visual descriptor d associated with the feature point location p into a single vector, Based on the single vector, an alternating self-attention layer and a cross-attention layer are created by repeating L times to create representation f. It has, The aforementioned optimal matching layer creates an M×N score matrix from the representation f, and finds the optimal partial assignment based on the M×N score matrix. p is the feature point location, d is a descriptor, L represents multiple occurrences. f is an expression, The computer system according to claim 1, wherein M × N is a score matrix having length M and width N.

3. In the feature point encoder, for each feature point, an initial representation [Math 2] It combines visual appearance and location. [Math 3] As shown above, the individual feature point locations are embedded in a high-dimensional vector along with a multilayer perceptron. i is a feature point, 【Number 4】 This is the initial representation for the aforementioned feature point, d i This is a descriptor for the aforementioned feature point, p i This is the feature point position relative to the aforementioned feature point, The computer system according to claim 2, wherein the MLP is a multilayer perceptron for the feature point location.

4. The computer system according to claim 3, wherein the feature point encoder enables the attention graph neural network to make inferences about both appearance and location.

5. The computer system according to claim 2, wherein the feature point encoder includes a multigraph neural network having a single complete graph with nodes which are feature points of two images.

6. The computer system according to claim 5, wherein the single complete graph is a multiple graph, the multiple graph having two types of non-directional edges, namely, in-image edges that connect a feature point i to all other feature points in the same image, and inter-image edges that connect a feature point i to all feature points in other images, and the resulting multiple graph neural network propagates information along both types of edges using a message-passing formula so that it computes an updated representation by starting from a high-dimensional state for each node and simultaneously aggregating messages across all given edges relating to all nodes in each layer. [Request Item 7] [Number 5] However, if it is an intermediate representation of feature point i in image A in layer 1, then message m E→i This is the result of aggregation from all feature points {j: (i, j) ∈ E}, where E ∈ {E self , E cross } and the remaining message passing updates for all i in A are, [Math 6] And, [・||・] indicates connection. i is a feature point, E self It is self-edged, E cross It is a cross edge, The computer system according to claim 6, wherein the MLP is a multiphase perceptron for the feature point location.

8. A fixed number of layers having a plurality of different parameters are chained, starting from l = 1, and when l is odd, E = E self and when l is even, E = E cross The computer system according to claim 7, which is alternately aggregated along the self-edge and the cross-edge so as to be.

9. The alternating self-attention layer and the cross-attention layer are calculated using an attention mechanism, and the attention mechanism is the message m E→i The following is calculated, the aggregation is performed, the self-edge is based on self-attention, and the cross-edge is based on cross-attention, with respect to the representation of i, query q i This involves the attributes and key k of several elements. j Based on this, the value v of some of the elements j The message reads, [Number 7] As shown above, it is calculated as a weighted average of the aforementioned values ​​vj, i is a feature point, I understand E This is a message, q i This is a query for the aforementioned feature point, v i This is the read value for the aforementioned feature point, j is the step variable in the summation formula, α ij This is the softmax function, The computer system according to claim 7, wherein the MLP is a multilayer perceptron for the feature point location.

10. Attention Mask α ij teeth, [Number 8] As shown above, the key k j and the query q i A computer system according to claim 9, which is a softmax with respect to similarities.

11. The final matching descriptors of the alternating self-attention layer and the cross-attention layer are: [Number 9] As shown, it is a linear projection, W represents the slope of the linear projection, b represents a constant in the linear projection, A is an image, the computer system according to claim 2.

12. The computer system according to claim 1, wherein the optimal matching layer extends each set of feature points with dust bin scores so that, for occlusion and visibility, occluded feature points are suppressed and unmatched feature points are explicitly assigned to dust bin scores.

13. The aforementioned score S for each individual is, [Number 10] As shown above, by adding points / bins and bin / bin scores filled with a single learnable parameter to the new columns and rows, [Math 11] It was expanded to, The computer system according to claim 12, wherein N+1 and M+1 represent extensions.

14. The attention graph neural network is A feature point encoder for mapping a feature point location p and a visual descriptor d associated with the feature point location p into a single vector, Based on the single vector, an alternating self-attention layer and a cross-attention layer are created by repeating L times to create representation f. It has, The computer system according to claim 12, wherein the optimal matching layer creates an M × N score matrix from the representation f and, over T iterations, uses the Sinkhorn algorithm to find the optimal partial assignment based on the M × N score matrix.

15. After T iterations, the optimal matching layer drops the dustbin score, and P = P - 1:M,1:N Restore, [Math 12] This is the original allocation, [Number 13] The computer system according to claim 14, wherein the assignment is an assignment with the dustbin score.

16. A method implemented in a computer, wherein the method is By mapping the feature point positions p into a single vector using the encoder of the attention graph neural network, a representation is created using the aforementioned attention graph neural network. Using at least one layer of the attention graph neural network, an expression f is performed based on the vector, and a score is created from the expression using the optimal matching layer. The optimal matching layer is used to find the optimal assignment based on the score, wherein the optimal matching layer is [Number 14] As shown above, the pairwise score for each set is represented as the similarity of the matching descriptors, S i,j This is a ratio to each score, <・,・> represents the inner product, A is an image, B is an image, and Methods that include...