Image and three-dimensional model matching method, device and electronic equipment
Patent Information
- Application Number
- CN202611256133.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-18
- Publication Date
- 2026-09-25
AI Technical Summary
[0005]有鉴于此,本申请提供了一种图像与三维模型的匹配方法、装置及电子设备,以解决图像与三维模型之间的匹配精度低的问题
[0012]本申请实施例提供的图像与三维模型的匹配方法,通过第一自注意力机制对各个视角图像的第一图像特征进行处理,能够建模不同视角之间的依赖关系以得到第二图像特征,进而基于各视角的第二图像特征生成全局共识嵌入,以该全局共识嵌入为查询向量,通过交叉注意力机制对各视角的第二图像特征进行加权融合得到多视角融合特征,进而基于多视角融合特征匹配得到目标物体对应的三维模型,一方面,在多视角场景下对图像和三维模型进行匹配,多视角所包含的丰富信息能够较好地区分在某一个或某几个视角下相似的三维模型,提高了图像与三维模型之间的匹配准确度;另一方面,引入交叉注意力机制对第二图像特征进行加权融合,使得与全局物体结构最一致的视角获得更高的融合权重,而包含模糊或遮挡或背景干扰的低质量视角被自适应地抑制,在保证多视角融合特征的信息丰富度的同时,减少了无效信息对多视角融合特征的干扰,提高了图像与三维模型之间的匹配效率。
Smart Images

Figure CN122821175A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing technology, specifically to methods, apparatus, and electronic devices for matching images with three-dimensional models. Background Technology
[0002] With the rapid development of 3D modeling technology and the continuous expansion of application scenarios, the number of 3D models has exploded. How to efficiently and accurately retrieve the corresponding 3D models based on existing 2D images has become an important research direction in the field of cross-modal information retrieval.
[0003] Currently, in related technologies, image encoders (such as CLIP's visual encoder, AlexNet, ResNet, etc.) are used to extract features from two-dimensional images to obtain a first feature vector. Three-dimensional model encoders (such as PointNet++, DGCNN, Point-BERT, or MVCNN, etc.) are used to extract features from three-dimensional models to obtain a second feature vector. Then, cross-modal matching of the first and second feature vectors is performed through contrastive learning, metric learning, and other methods to obtain the three-dimensional model corresponding to the two-dimensional image.
[0004] However, in the aforementioned related technologies, two-dimensional images can often only capture the features of an object from a certain perspective. When faced with a similar three-dimensional model from a certain perspective, it is impossible to accurately match the three-dimensional model corresponding to the two-dimensional image, resulting in low matching accuracy between the image and the three-dimensional model. Summary of the Invention
[0005] In view of this, this application provides a method, apparatus, and electronic device for matching images with three-dimensional models to solve the problem of low matching accuracy between images and three-dimensional models.
[0006] In a first aspect, this application provides a method for matching an image with a 3D model, the method comprising: Acquire multiple images of the target object from different perspectives; where different perspective images correspond to different acquisition perspectives; Feature extraction is performed on each of the aforementioned viewpoint images to obtain the first image features corresponding to each of the aforementioned viewpoint images; The first image features are processed by a first self-attention mechanism to model the dependency relationship between different viewpoints and obtain the second image features corresponding to each viewpoint image. Based on each of the second image features, a global consensus embedding is generated; Using the global consensus embedding as the query vector, the second image features are weighted and fused through a cross-attention mechanism to obtain multi-view fused features; The three-dimensional model features obtained by encoding the pre-constructed three-dimensional model are then matched with the multi-view fusion features to obtain a target three-dimensional model that matches the target object.
[0007] Secondly, this application provides an image-to-3D model matching device, the device comprising: The image acquisition module is used to acquire multiple images of the target object from different perspectives; different perspective images correspond to different acquisition perspectives. The feature extraction module is used to extract features from each of the viewpoint images to obtain the first image features corresponding to each of the viewpoint images. The feature processing module is used to process each of the first image features through a first self-attention mechanism, model the dependency relationship between different viewpoints, and obtain the second image features corresponding to each of the viewpoint images respectively. A global generation module is used to generate a global consensus embedding based on each of the second image features; The feature fusion module is used to use the global consensus embedding as the query vector and perform weighted fusion on each of the second image features through a cross-attention mechanism to obtain multi-view fused features. The model matching module is used to encode the 3D model features obtained from the pre-constructed 3D model, and perform similarity matching between the 3D model features and the multi-view fusion features to obtain a target 3D model that matches the target object.
[0008] Thirdly, this application provides an electronic device, including: a memory and a processor, which are communicatively connected to each other. The memory stores computer instructions, and the processor executes the computer instructions to perform the image-to-3D model matching method described in the first aspect or any corresponding embodiment.
[0009] For example, the electronic device is a matching device.
[0010] Fourthly, this application provides a computer-readable storage medium storing computer instructions for causing a computer to execute the image-to-3D model matching method described in the first aspect or any corresponding embodiment.
[0011] Fifthly, this application provides a computer program product, including computer instructions for causing a computer to execute the image-to-3D model matching method described in the first aspect or any corresponding embodiment.
[0012] The image-to-3D model matching method provided in this application processes the first image features of images from various perspectives through a first self-attention mechanism. This model can model the dependencies between different perspectives to obtain second image features. Then, a global consensus embedding is generated based on the second image features of each perspective. Using this global consensus embedding as a query vector, the second image features of each perspective are weighted and fused through a cross-attention mechanism to obtain multi-view fusion features. Finally, the 3D model corresponding to the target object is matched based on the multi-view fusion features. On the one hand, matching images and 3D models in multi-view scenarios allows the rich information contained in multiple perspectives to better distinguish similar 3D models in one or several perspectives, improving the matching accuracy between images and 3D models. On the other hand, the introduction of a cross-attention mechanism to weightedly fuse the second image features gives higher fusion weights to perspectives most consistent with the global object structure, while low-quality perspectives containing blur, occlusion, or background interference are adaptively suppressed. This ensures the richness of information in the multi-view fusion features while reducing the interference of invalid information on the multi-view fusion features, thus improving the matching efficiency between images and 3D models. Attached Figure Description
[0013] To more clearly illustrate the technical solutions in the specific embodiments or related technologies of this application, the drawings used in the description of the specific embodiments or related technologies will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0014] Figure 1 An optional application scenario of an embodiment of this application is illustrated by way of example; Figure 2 This is a schematic flowchart of an image-to-3D model matching method according to an embodiment of this application; Figure 3 This is another flowchart illustrating the method for matching images with 3D models according to embodiments of this application; Figure 4 An exemplary diagram illustrates the 3D model feature extraction method of a 3D model encoder; Figure 5 This is a structural block diagram of an image-to-3D model matching device according to an embodiment of this application; Figure 6 This is a schematic diagram of the hardware structure of an electronic device according to an embodiment of this application. Detailed Implementation
[0015] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0016] It should be noted that the information (including but not limited to user input information, such as information entered by the user into input boxes), data (including but not limited to data used for analysis, stored data, and displayed data, such as context code, all code of the current project, the service pressure corresponding to operations performed on all code of the current project, and the code development status of the current project), and signals involved in this application are all authorized by the user or fully authorized by all parties, and the collection, use, and processing of related data must comply with relevant laws, regulations, and standards. For example, the context code, operations performed on all code of the current project, the corresponding service pressure, and the code development status involved in this application were all obtained with full authorization.
[0017] The terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this application, "multiple" means two or more, unless otherwise explicitly specified.
[0018] As one optional application scenario in the embodiments of this application, such as Figure 1 As shown, the system may include a matching device 10.
[0019] The matching device 10 is used to implement the image-to-3D model matching method provided in the embodiments of this application. Exemplarily, the matching device 10 can be a standalone physical server, a server cluster or a distributed system, or a cloud server providing cloud services.
[0020] For example, such as Figure 1 As shown, the matching device 10 includes an image encoder, a 3D model encoder, and a feature fusion model.
[0021] Optionally, the matching device 10 acquires a pre-built 3D model. Further, the 3D model is sampled into a point cloud using a 3D model encoder, and normals are estimated on the surface of the 3D model to obtain the surface normal vector corresponding to each sampling point. Each point in the point cloud is encoded as a multi-dimensional vector containing spatial coordinates, color information, and a surface normal vector. The point cloud is divided into multiple local point cloud blocks, and these local point cloud blocks are processed using a second self-attention mechanism to model the interaction between spatial location, color appearance, and surface normals, thereby obtaining the 3D model features.
[0022] Optionally, the matching device 10 acquires multiple viewpoint images of the target object. Further, an image encoder extracts features from each viewpoint image to obtain first image features for each viewpoint; a feature fusion model fuses the multi-viewpoint features based on a first self-attention mechanism and a cross-attention mechanism to obtain multi-viewpoint fused features. Then, the electronic device performs similarity matching between the 3D model features and the multi-viewpoint fused features to obtain a target 3D model that matches the target object.
[0023] According to an embodiment of this application, an embodiment of a method for matching images with three-dimensional models is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.
[0024] This embodiment provides a method for matching images with 3D models, which can be used in the aforementioned matching device (hereinafter referred to as electronic device). Figure 2 This is a flowchart of an image-to-3D model matching method according to an embodiment of this application, such as... Figure 2 As shown, the process includes the following steps: Step S201: Obtain multiple perspective images of the target object.
[0025] In this embodiment of the application, when matching an image with a 3D model, the electronic device acquires multiple perspective images of the target object. Different perspective images correspond to different acquisition perspectives.
[0026] For example, a viewpoint image refers to a two-dimensional image of the same target object acquired from different angles. Optionally, multiple viewpoint images can be images acquired from different shooting angles, different frame images extracted from a video sequence, or images from multiple viewpoints extracted from an image sequence acquired synchronously by a multi-camera array.
[0027] For example, the viewing angles corresponding to multi-view images can be flexibly set and adjusted according to the actual situation, such as front view, left view, right view, top view, and rear view corresponding to multi-view images.
[0028] Step S202: Extract features from each viewpoint image to obtain the first image features corresponding to each viewpoint image.
[0029] In this embodiment of the application, after acquiring the above-mentioned multiple viewpoint images, the electronic device performs feature extraction on each viewpoint image to obtain the first image features corresponding to each viewpoint image.
[0030] For example, the electronic device extracts features from each viewpoint image by using a pre-trained visual Transformer (such as the image encoder mentioned above) to obtain the first image features corresponding to each viewpoint image.
[0031] Step S203: Process each first image feature through the first self-attention mechanism, model the dependency relationship between different viewpoints, and obtain the second image features corresponding to each viewpoint image.
[0032] In this embodiment of the application, after obtaining the first image features, the electronic device processes each first image feature through a first self-attention mechanism, models the dependency relationship between different viewpoints, and obtains the second image features corresponding to each viewpoint image.
[0033] For example, the first self-attention mechanism can be understood as an operational mechanism that allows each first image feature to "see each other." For instance, for a given first image feature, the electronic device calculates the query vector, key vector, and value vector using a learnable projection matrix, then calculates the relevance score between the query vector and all key vectors. After processing with a softmax function, the attention weight is obtained. Finally, based on this attention weight, all value vectors are weighted and summed to obtain the second image feature corresponding to the first image feature. For example, the first self-attention mechanism is a multi-head attention mechanism.
[0034] Step S204: Generate a global consensus embedding based on each second image feature.
[0035] In this embodiment, after acquiring the aforementioned second image features, the electronic device generates a global consensus embedding based on each of the second image features. The global consensus embedding is used to characterize the overall appearance features of the target object under all viewing angles. For example, the electronic device compresses and aggregates each of the second image features to obtain the global consensus embedding.
[0036] Step S205: Using the global consensus embedding as the query vector, the features of each second image are weighted and fused through a cross-attention mechanism to obtain multi-view fused features.
[0037] In this embodiment of the application, after obtaining the above-mentioned global consensus embedding, the electronic device uses the global consensus embedding as the query vector and performs weighted fusion on each second image feature through a cross-attention mechanism to obtain multi-view fusion features.
[0038] For example, the electronic device uses global consensus embedding as the query vector and each second image feature as the key vector and value vector, and adaptively determines the fusion weight of each viewpoint through a cross-attention mechanism; further, each second image feature is multiplied by its corresponding fusion weight and then summed to obtain the multi-view fusion feature.
[0039] For example, when obtaining the fusion weights, a cross-attention mechanism is used to ensure that the viewpoint most consistent with the global object structure receives a higher fusion weight, while low-quality viewpoints containing blur, occlusion, or background interference receive a lower fusion weight.
[0040] Step S206: The three-dimensional model features obtained by encoding the pre-constructed three-dimensional model are matched with the multi-view fusion features to obtain a target three-dimensional model that matches the target object.
[0041] In this embodiment of the application, after obtaining the above-mentioned multi-view fusion features, the electronic device encodes the three-dimensional model features obtained from the pre-constructed three-dimensional model, performs similarity matching between the three-dimensional model features and the multi-view fusion features, and obtains a target three-dimensional model that matches the target object.
[0042] For example, the electronic device pre-builds a database containing various 3D object models, and encodes each 3D model in the database into 3D model features through the aforementioned 3D model encoder; further, it performs similarity matching (such as cosine similarity) on the 3D model features and multi-view fusion features, and determines the target 3D model that matches the target object by identifying the most similar 3D model.
[0043] It should be noted that 3D model features refer to the features of a three-dimensional model, not that the model has three dimensions.
[0044] The image-to-3D model matching method provided in this embodiment processes the first image features of images from various perspectives through a first self-attention mechanism. This model can model the dependencies between different perspectives to obtain second image features. Then, a global consensus embedding is generated based on the second image features of each perspective. Using this global consensus embedding as a query vector, the second image features of each perspective are weighted and fused through a cross-attention mechanism to obtain multi-view fusion features. Finally, the 3D model corresponding to the target object is matched based on the multi-view fusion features. On the one hand, matching images and 3D models in multi-view scenarios allows the rich information contained in multiple perspectives to better distinguish similar 3D models in one or several perspectives, improving the matching accuracy between images and 3D models. On the other hand, the introduction of a cross-attention mechanism to weight and fuse the second image features gives higher fusion weights to perspectives most consistent with the global object structure, while low-quality perspectives containing blur, occlusion, or background interference are adaptively suppressed. This ensures the richness of information in the multi-view fusion features while reducing the interference of invalid information, thus improving the matching efficiency between images and 3D models.
[0045] The following section introduces the methods for constructing 3D model features.
[0046] This embodiment provides a method for matching images with 3D models, which can be used in the aforementioned electronic devices. Figure 3 This is a flowchart of an image-to-3D model matching method according to an embodiment of this application, such as... Figure 3 As shown, the process includes the following steps: Step S301: Sample the 3D model as a point cloud.
[0047] For example, when acquiring features of a 3D model, the electronic device samples the 3D model as a point cloud. Each point is encoded as a multi-dimensional vector containing spatial coordinates, color information, and a surface normal vector. For example, a point cloud refers to a discretized representation of a 3D model; spatial coordinates characterize the position of a point in 3D space, typically represented by three coordinate values: x, y, and z; color information refers to the appearance color of the point, typically represented by three color channel values: R, G, and B; and the surface normal vector refers to the normal direction of the local surface containing the point, typically represented by three directional components: Nx, Ny, and Nz.
[0048] For example, the multidimensional vector is composed of a three-dimensional spatial coordinate vector, a three-dimensional RGB color vector, and a three-dimensional normal vector. For example, for a textureless 3D model, the RGB color channels are filled with a uniform constant value to maintain a multidimensional vector consistent with that of a textured 3D model. Optionally, this constant value can be any value and can be flexibly set and adjusted according to the actual situation, such as 0.8.
[0049] Step S302: Divide the point cloud into multiple local point cloud blocks.
[0050] For example, after acquiring the point cloud, the electronic device divides the point cloud into multiple local point cloud blocks.
[0051] For example, the electronic device uses the Farthest Point Sampling (FPS) algorithm to obtain multiple key points from the point cloud; further, for each key point, the k-Nearest Neighbors (kNN) algorithm is used to divide the neighboring points around the key point into a local point cloud block.
[0052] For example, by employing the farthest point sampling algorithm to acquire key points and then using the k-nearest neighbor algorithm to divide the local point cloud into blocks, it is possible to ensure the uniform distribution of the local point cloud blocks in space. This allows each local point cloud block to cover different areas of the object's surface, thereby effectively capturing the local geometric structure information of the 3D model. Compared to random sampling, farthest point sampling can better represent the global distribution of the point cloud; compared to dense sampling, this farthest point sampling and k-nearest neighbor algorithm have higher computational efficiency while maintaining representational capability.
[0053] Step S303: The local point cloud blocks are processed through the second self-attention mechanism to model the interaction between spatial position, color appearance and surface normals, and obtain the three-dimensional model features.
[0054] In this embodiment of the application, after obtaining the aforementioned local point cloud blocks, the electronic device processes the local point cloud blocks through a second self-attention mechanism to model the interaction relationship between spatial position, color appearance and surface normals, thereby obtaining three-dimensional model features.
[0055] For example, the electronic device extracts features from the 3D model using the aforementioned 3D model encoder to obtain the 3D model features. For example... Figure 4 As shown, for a 3D model, the 3D model encoder obtains a point cloud through point cloud sampling and normal calculation, and each point is encoded as a 9-dimensional vector (including a 3D spatial coordinate vector, a 3D RGB color vector, and a 3D normal vector). Furthermore, the point cloud is divided into multiple local point cloud blocks through the farthest point sampling algorithm and the k-nearest neighbor algorithm. Furthermore, the interaction between spatial location, color appearance, and surface normal is jointly modeled through a second self-attention mechanism to process the local point cloud blocks and obtain the 3D model features.
[0056] For example, the second self-attention mechanism is a multi-head attention mechanism.
[0057] Step S304: Acquire multiple perspective images of the target object.
[0058] Step S305: Extract features from each viewpoint image to obtain the first image features corresponding to each viewpoint image.
[0059] Step S306: Process each first image feature through the first self-attention mechanism, model the dependency relationship between different viewpoints, and obtain the second image features corresponding to each viewpoint image.
[0060] Step S307: Generate a global consensus embedding based on each second image feature.
[0061] Step S308: Using the global consensus embedding as the query vector, the features of each second image are weighted and fused through a cross-attention mechanism to obtain multi-view fused features.
[0062] Step S309: The three-dimensional model features obtained by encoding the pre-constructed three-dimensional model are matched with the multi-view fusion features to obtain a target three-dimensional model that matches the target object.
[0063] The above steps S304-S309 and Figure 2 Steps S201-S206 in the embodiment are similar; see details below. Figure 2 Examples are not detailed here.
[0064] The image-to-3D model matching method provided in this embodiment introduces surface normal information and jointly models the interaction between spatial location, color appearance and surface normal through a second self-attention mechanism. This allows the final obtained 3D model features to contain rich geometric information related to surface orientation and curvature, thereby enabling better differentiation of objects with similar shapes but different geometric structures. In the subsequent image-to-3D model matching process, the rich geometric information significantly reduces the dependence on color and texture information during matching. As a result, even when faced with 3D models without texture or with degraded color, the 3D model corresponding to the target object can be determined through matching.
[0065] The training process of the model is described below. It includes the following two stages: The first stage involves freezing the pre-trained image encoder and training the 3D model encoder so that its output is aligned with the image encoder's output in the embedding space. The image encoder extracts features from images at each viewpoint, while the 3D model encoder encodes the pre-built 3D model. For example, the 3D model encoder processes the aforementioned local point cloud blocks using a second self-attention mechanism to encode the pre-built 3D model.
[0066] The second stage involves freezing the trained 3D model encoder and training the feature fusion model to align its output with the trained 3D model encoder's output in the embedding space. The feature fusion model is used to fuse the first image features based on the first self-attention mechanism and the cross-attention mechanism.
[0067] For example, a pre-trained image encoder refers to an image feature extraction network that is pre-trained on a large-scale dataset.
[0068] For example, the embedding space refers to the representation space of high-dimensional feature vectors. In the embodiments of this application, image features (the first image feature and the second image feature mentioned above) and three-dimensional model features are aligned in the same embedding space.
[0069] For example, freezing refers to keeping a portion of the network parameters unchanged during model training, meaning that these parameters are not involved in gradient updates. For instance, freezing a pre-trained image encoder can preserve its generalization ability during training.
[0070] For example, during alignment, an image alignment adapter (IAA) and a text alignment adapter (TAA) are introduced. Optionally, the text description of the 3D model is encoded using a pre-trained text encoder to obtain the text description features corresponding to the 3D model. Further, for the output of the 3D model encoder, the image alignment adapter aligns the output with the image features output by the image encoder, and the text alignment adapter aligns the output with the text description features. Through the image alignment adapter and the text alignment adapter, the output of the 3D model encoder is decoupled into a visual latent subspace and a semantic latent subspace, achieving joint alignment of image, text, and 3D modalities. In other words, the output of the 3D model encoder is aligned not only with the single-view image features at the visual level but also with the text description features at the semantic level, enabling the 3D model encoder to learn both visual appearance information and semantic category information simultaneously during training, thereby obtaining a more discriminative 3D model feature representation. It should be noted that the above-mentioned text alignment adapter and text description features are only used in the training phase of the 3D model encoder; users do not need to input text when actually performing image retrieval.
[0071] For example, the training process of the 3D model encoder (i.e., the first stage mentioned above) and the training process of the feature fusion model (i.e., the second stage mentioned above) are optimized using a symmetric contrastive loss function. Specifically, different modal features of the same training object are used as positive sample pairs, and features of different training objects are used as negative sample pairs. Model training is performed by maximizing the similarity between positive sample pairs and minimizing the similarity between negative sample pairs.
[0072] For example, the different modalities include image modalities and 3D modalities. The image modalities are two-dimensional view images of the training object, while the 3D modalities are the 3D models extracted during training.
[0073] For example, a positive sample refers to a feature pair consisting of different modal features of the same training object. For instance, the viewpoint image features of the training object and the 3D model features of the training object constitute a positive sample pair.
[0074] For example, a negative sample refers to a feature pair composed of features from different training objects. For instance, the viewpoint image features of the first training object and the 3D model features of the second training object constitute a negative sample pair.
[0075] For example, a symmetric contrastive loss function symmetrically constrains the mapping in two directions. For instance, the image is mapped to the model direction, and the model is mapped to the image direction.
[0076] For example, similarity refers to how close two feature vectors are in the embedding space. For instance, cosine similarity or Euclidean distance can be used to measure the closeness of two feature vectors in the embedding space. Optionally, higher similarity indicates that the two feature vectors are closer in the embedding space, and the corresponding image features and model features are more consistent.
[0077] The image-to-3D model matching method provided in this embodiment adopts a two-stage progressive training strategy. In the first stage, the pre-trained image encoder is frozen and the 3D model encoder is trained so that the output of the 3D model encoder is aligned with the output of the image encoder in the embedding space. In the second stage, the aligned 3D model encoder is frozen and the feature fusion model is trained so that the fused multi-view features are aligned with the 3D model features in the embedding space. The staged training strategy effectively avoids the problem of mutual interference between different modalities in end-to-end joint training and achieves accurate cross-modal alignment.
[0078] This embodiment also provides an image-to-3D model matching device, which is used to implement the above embodiments and preferred embodiments; details already described will not be repeated. As used below, the term "module" can be a combination of software and / or hardware that performs a predetermined function. Although the device described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.
[0079] This embodiment provides a device for matching images with 3D models, such as... Figure 5 As shown, it includes: The image acquisition module 501 is used to acquire multiple images of the target object from different perspectives; wherein, different perspective images correspond to different acquisition perspectives; Feature extraction module 502 is used to extract features from images from each viewpoint to obtain the first image features corresponding to each viewpoint. The feature processing module 503 is used to process each first image feature through a first self-attention mechanism, model the dependency relationship between different viewpoints, and obtain the second image features corresponding to each viewpoint image respectively. The global generation module 504 is used to generate a global consensus embedding based on each second image feature; The feature fusion module 505 is used to use the global consensus embedding as the query vector and perform weighted fusion of each second image feature through a cross attention mechanism to obtain multi-view fused features. The model matching module 506 is used to perform similarity matching between the 3D model features obtained by encoding the pre-built 3D model and the multi-view fusion features to obtain a target 3D model that matches the target object.
[0080] In an exemplary embodiment, the apparatus further includes: The model uses modules to sample the 3D model into a point cloud; each point is encoded as a multi-dimensional vector containing spatial coordinates, color information, and surface normal vectors. The point cloud partitioning module is used to divide a point cloud into multiple local point cloud blocks; The feature acquisition module is used to process local point cloud blocks through a second self-attention mechanism, model the interaction between spatial location, color appearance and surface normals, and obtain 3D model features.
[0081] In an exemplary embodiment, the multidimensional vector is composed of a three-dimensional spatial coordinate vector, a three-dimensional RGB color vector, and a three-dimensional normal vector.
[0082] In an exemplary embodiment, the point cloud segmentation module is used for: Multiple key points are obtained from the point cloud using the farthest point sampling algorithm; For each key point, the k-nearest neighbor algorithm is used to divide the neighboring points around the key point into a local point cloud block.
[0083] In an exemplary embodiment, the apparatus further includes: The first training module is used to freeze the pre-trained image encoder and train the 3D model encoder so that the output of the 3D model encoder is aligned with the output of the image encoder in the embedding space. The image encoder is used to extract features from images at each viewpoint, and the 3D model encoder is used to encode the pre-built 3D model. The second training module is used to freeze the trained 3D model encoder and train the feature fusion model so that the output of the feature fusion model is aligned with the output of the trained 3D model encoder in the embedding space. The feature fusion model is used to fuse various first image features based on the first self-attention mechanism and the cross-attention mechanism.
[0084] In an exemplary embodiment, the training process of the 3D model encoder and the training process of the feature fusion model are optimized using a symmetrical contrastive loss function; In this model, different modal features of the same training object are used as positive sample pairs, and features of different training objects are used as negative sample pairs. The model is trained by maximizing the similarity between positive sample pairs and minimizing the similarity between negative sample pairs.
[0085] The image-to-3D model matching apparatus provided in this application embodiment can execute the image-to-3D model matching method provided in any embodiment of this application, and has the corresponding functional modules and beneficial effects of the method execution. Further functional descriptions of the above modules and units are the same as those in the corresponding embodiments described above, and will not be repeated here.
[0086] Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application.
[0087] The following is a detailed reference. Figure 6 This diagram illustrates a suitable structural schematic for implementing the electronic device described in the embodiments of this application. The electronic device may include a processor (e.g., a central processing unit, graphics processor, etc.) 601, which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 602 or a program loaded from memory 608 into random access memory (RAM) 603. The RAM 603 also stores various programs and data required for the operation of the electronic device. The processor 601, ROM 602, and RAM 603 are interconnected via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.
[0088] Typically, the following devices can be connected to I / O interface 605: input devices 606 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 607 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; memory devices 608 including, for example, magnetic tapes, hard disks, etc.; and communication devices 609. Communication device 609 allows electronic devices to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 6Electronic devices with various devices are shown, but it should be understood that it is not required to implement or have all of the devices shown, and more or fewer devices may be implemented or have instead.
[0089] Specifically, according to embodiments of this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this application include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 609, or installed from a memory 608, or installed from a ROM 602. When the computer program is executed by the processor 601, it performs the functions defined in the image-to-3D model matching method of embodiments of this application.
[0090] Figure 6 The electronic device shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments of this application.
[0091] This application also provides a computer-readable storage medium. The methods described in this application can be implemented in hardware or firmware, or implemented as recordable on a storage medium, or implemented as computer code downloaded over a network and originally stored on a remote storage medium or a non-transitory machine-readable storage medium and subsequently stored on a local storage medium. Thus, the methods described herein can be processed by software stored on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. The storage medium can be a magnetic disk, optical disk, read-only memory, random access memory, flash memory, hard disk, or solid-state drive, etc.; further, the storage medium can also include combinations of the above types of memory. It is understood that computers, processors, microprocessor controllers, or programmable hardware include storage components capable of storing or receiving software or computer code. When the software or computer code is accessed and executed by the computer, processor, or hardware, the image-to-3D model matching method shown in the above embodiments is implemented.
[0092] A portion of this application can be applied as a computer program product, such as computer program instructions, which, when executed by a computer, can invoke or provide the methods and / or technical solutions according to this application through the operation of the computer. Those skilled in the art will understand that the forms in which computer program instructions exist in a computer-readable medium include, but are not limited to, source files, executable files, installation package files, etc. Correspondingly, the ways in which computer program instructions are executed by a computer include, but are not limited to: the computer directly executing the instructions, or the computer compiling the instructions and then executing the corresponding compiled program, or the computer reading and executing the instructions, or the computer reading and installing the instructions and then executing the corresponding installed program. Here, the computer-readable medium can be any available computer-readable storage medium or communication medium accessible to a computer.
[0093] Although embodiments of this application have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of this application, and all such modifications and variations fall within the scope defined by the appended claims.
Claims
1. A method for matching an image with a 3D model, characterized in that, The method includes: Acquire multiple images of the target object from different perspectives; where different perspective images correspond to different acquisition perspectives; Feature extraction is performed on each of the aforementioned viewpoint images to obtain the first image features corresponding to each of the aforementioned viewpoint images; The first image features are processed by a first self-attention mechanism to model the dependency relationship between different viewpoints and obtain the second image features corresponding to each viewpoint image. Based on each of the second image features, a global consensus embedding is generated; Using the global consensus embedding as the query vector, the second image features are weighted and fused through a cross-attention mechanism to obtain multi-view fused features; The three-dimensional model features obtained by encoding the pre-constructed three-dimensional model are then matched with the multi-view fusion features to obtain a target three-dimensional model that matches the target object.
2. The method according to claim 1, characterized in that, The method further includes: The three-dimensional model is sampled as a point cloud; wherein each point is encoded as a multi-dimensional vector containing spatial coordinates, color information and surface normal vector; The point cloud is divided into multiple local point cloud blocks; The local point cloud blocks are processed by a second self-attention mechanism to model the interaction between spatial location, color appearance and surface normals, thereby obtaining the features of the three-dimensional model.
3. The method according to claim 2, characterized in that, The multidimensional vector is composed of a three-dimensional spatial coordinate vector, a three-dimensional RGB color vector, and a three-dimensional normal vector.
4. The method according to claim 2, characterized in that, The step of dividing the point cloud into multiple local point cloud blocks includes: Multiple key points are obtained from the point cloud using the farthest point sampling algorithm; For each key point, the k-nearest neighbor algorithm is used to divide the neighboring points around the key point into a local point cloud block.
5. The method according to claim 1 or 2, characterized in that, The method further includes: Freeze the pre-trained image encoder and train the 3D model encoder so that the output of the 3D model encoder is aligned with the output of the image encoder in the embedding space; wherein, the image encoder is used to extract features from each of the viewpoint images, and the 3D model encoder is used to encode the pre-constructed 3D model; The trained 3D model encoder is frozen, and the feature fusion model is trained so that the output of the feature fusion model is aligned with the output of the trained 3D model encoder in the embedding space; wherein, the feature fusion model is used to fuse each of the first image features based on the first self-attention mechanism and the cross-attention mechanism.
6. The method according to claim 5, characterized in that, The training process of the 3D model encoder and the training process of the feature fusion model are optimized using a symmetrical contrastive loss function. In this model, different modal features of the same training object are used as positive sample pairs, and features of different training objects are used as negative sample pairs. The model is trained by maximizing the similarity between positive sample pairs and minimizing the similarity between negative sample pairs.
7. A device for matching an image with a three-dimensional model, characterized in that, The device includes: The image acquisition module is used to acquire multiple images of the target object from different perspectives; different perspective images correspond to different acquisition perspectives. The feature extraction module is used to extract features from each of the viewpoint images to obtain the first image features corresponding to each of the viewpoint images. The feature processing module is used to process each of the first image features through a first self-attention mechanism, model the dependency relationship between different viewpoints, and obtain the second image features corresponding to each of the viewpoint images respectively. A global generation module is used to generate a global consensus embedding based on each of the second image features; The feature fusion module is used to use the global consensus embedding as the query vector and perform weighted fusion on each of the second image features through a cross-attention mechanism to obtain multi-view fused features. The model matching module is used to encode the 3D model features obtained from the pre-constructed 3D model, and perform similarity matching between the 3D model features and the multi-view fusion features to obtain a target 3D model that matches the target object.
8. An electronic device, characterized in that, include: A memory and a processor are communicatively connected, the memory storing computer instructions, and the processor executing the computer instructions to perform the image-to-3D model matching method of any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions for causing the computer to perform the image-to-three-dimensional model matching method according to any one of claims 1 to 6.
10. A computer program product, characterized in that, Includes computer instructions for causing a computer to perform the matching method of an image and a three-dimensional model as described in any one of claims 1 to 6.