Feature point detection method and device, computer equipment and storage medium
By extracting multiple matrices from 3D mesh data and utilizing an adaptive hierarchical and spatial gradient fusion strategy for feature point detection, this method solves the problems of poor efficiency and recognition performance in traditional methods, achieving efficient and accurate feature point detection.
Patent Information
- Application Number
- CN202511073802.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-31
- Publication Date
- 2025-11-18
AI Technical Summary
Traditional feature point detection methods rely on manual extraction, lacking guidance from global information, resulting in poor efficiency and recognition performance. Furthermore, existing neural network methods either ignore connectivity relationships or rely on high-precision annotations, making them difficult to apply in real-world scenarios.
By extracting point cloud coordinates, point normals, mass, Laplacian, and gradient matrices from 3D mesh data, inputting them into a trained feature point detection model, and using adaptive hierarchical fusion and spatial gradient weight fusion strategies, the feature point detection results are output.
It improves the efficiency and recognition effect of feature point detection, reduces the need for manual intervention, enhances the model's anti-interference ability and detection accuracy, and reduces annotation costs.
Smart Images

Figure CN120976482A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of three-dimensional digitization, and in particular to a feature point detection method and device, a computer device, and a storage medium. BACKGROUND
[0002] In face beautification deformation, tooth orthodontic simulation, crown generation design, tooth measurement comparison, industrial scanning, industrial measurement, industrial design, and other scenarios, geometric feature point information plays a key role in operating three-dimensional mesh data.
[0003] Traditional feature point detection methods rely on manually extracted feature operators. Feature point detection often only considers local detail information, and not only requires manual operation, but also lacks global information guidance. The efficiency and recognition effect of feature point detection are poor. SUMMARY
[0004] The embodiments of the present application provide a feature point detection method and device, a computer device, and a storage medium, which can improve the efficiency and recognition effect of feature point detection.
[0005] In a first aspect, the embodiments of the present application provide a feature point detection method, which includes:
[0006] obtaining three-dimensional mesh data of a target object;
[0007] extracting a plurality of feature matrices of different types from the three-dimensional mesh data;
[0008] inputting the plurality of feature matrices of different types into a trained feature point detection model;
[0009] The feature point detection model performs feature point detection on the three-dimensional mesh data according to the plurality of feature matrices of different types, and outputs a feature point detection result of the target object.
[0010] In a second aspect, the embodiments of the present application also provide a feature point detection device, which includes a transceiver unit and a processing unit, wherein:
[0011] The transceiver unit is configured to obtain three-dimensional mesh data of a target object;
[0012] The processing unit is configured to extract a plurality of feature matrices of different types from the three-dimensional mesh data, input the plurality of feature matrices of different types into a trained feature point detection model, and perform feature point detection on the three-dimensional mesh data according to the plurality of feature matrices of different types, and output a feature point detection result of the target object.
[0013] In a third aspect, the embodiments of the present application further provide a computer device, which comprises a memory and a processor, the memory stores a computer program, and the processor implements the above method when executing the computer program.
[0014] In a fourth aspect, the embodiments of the present application further provide a computer readable storage medium, which stores a computer program, the computer program comprises program instructions, and the program instructions can implement the above method when executed by a processor.
[0015] The embodiments of the present application provide a feature point detection method and device, a computer device and a storage medium. The method comprises: obtaining three-dimensional mesh data of a target object; extracting a plurality of feature matrices of different types from the three-dimensional mesh data; inputting the plurality of feature matrices of different types into a trained feature point detection model; and performing feature point detection on the three-dimensional mesh data according to the plurality of feature matrices of different types by the feature point detection model, and outputting a feature point detection result of the target object. The embodiments of the present application can automatically extract a plurality of feature matrices of different types from three-dimensional mesh data, perform feature point detection processing by using a plurality of feature matrices of different types globally, and do not need manual intervention, thereby improving the efficiency and recognition effect of feature point detection. BRIEF DESCRIPTION OF DRAWINGS
[0016] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0017] Figure 1 A flowchart of a feature point detection method provided by the embodiments of the present application is shown in the figure.
[0018] Figure 2 A structure diagram of a feature point detection model provided by the embodiments of the present application is shown in the figure.
[0019] Figure 3 A sub-flowchart of a feature point detection method provided by the embodiments of the present application is shown in the figure.
[0020] Figure 4 A structure diagram of an adaptive hierarchical fusion module provided by the embodiments of the present application is shown in the figure.
[0021] Figure 5 Another structure diagram of a feature point detection model provided by the embodiments of the present application is shown in the figure.
[0022] Figure 6 Another sub-flowchart of a feature point detection method provided by the embodiments of the present application is shown in the figure.
[0023] Figure 7 A structural schematic diagram of a spatial gradient weight fusion module provided by an embodiment of the present application is shown in FIG. 1;
[0024] Figure 8 Another structural schematic diagram of a feature point detection model provided by an embodiment of the present application is shown in FIG. 2;
[0025] Figure 9 Another sub-flowchart of a feature point detection method provided by an embodiment of the present application is shown in FIG. 3;
[0026] Figure 10 A schematic block diagram of a feature point detection device provided by an embodiment of the present application is shown in FIG. 4;
[0027] Figure 11 A schematic block diagram of a computer device provided by an embodiment of the present application is shown in FIG. 5. DETAILED DESCRIPTION
[0028] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work fall within the scope of protection of the present application.
[0029] It should be understood that, when used in the specification and the appended claims, the terms "comprise" and "include" indicate the presence of the described features, integers, steps, operations, elements, and / or components, but do not exclude one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0030] It should also be understood that the terms used in the present application specification are only for the purpose of describing particular embodiments and are not intended to limit the present application. As used in the present application specification and the appended claims, unless otherwise clearly indicated by the context, the singular forms "a", "an" and "the" are intended to include the plural forms as well.
[0031] It should be further understood that the term "and / or" used in the present application specification is intended to mean one or more of any combination of the associated listed items and all possible combinations thereof.
[0032] The embodiments of the present application provide a feature point detection method, device, computer device, and storage medium.
[0033] The execution subject of the feature point detection method can be a feature point detection apparatus provided by an embodiment of the application or a computer device integrated with the feature point detection apparatus. The feature point detection apparatus can be implemented in hardware or software, and the computer device can be a terminal or a server.
[0034] In related technologies, feature point detection of a target object mainly includes the following four methods.
[0035] Firstly, feature point detection is performed by relying on a manually extracted feature operator. The feature point detection of this technology only considers local detail information, and needs manual operation and lacks global information guidance. The efficiency and recognition effect of feature point detection are poor, and the technology lacks generalization ability.
[0036] Secondly, feature point detection is performed by relying on a neural network of three-dimensional point cloud data. The neural network uses a point cloud detection method to detect and recognize the position of a feature point. This method converts three-dimensional grid data into three-dimensional point cloud data for recognition, ignores the connection relationship between points, i.e., the relationship of edges, and does not fully utilize existing data information. The recognition result often has a certain bottleneck. For some local abnormal noise points in the data, it is difficult to judge and distinguish only by the position relationship of points, and false detection is likely to occur, resulting in poor anti-interference ability of the model.
[0037] Thirdly, global pooling operation is directly used to obtain global feature information of an object in a two-dimensional manner. This technology can only obtain single-dimensional global information and lacks secondary global information. Because of the lack of downsampling and local pooling operation, it is difficult to obtain the local adjacency relationship between points. For some complex data scenes, the detection effect is poor, and problems such as missed detection and feature point coordinate offset are likely to occur, resulting in low detection precision.
[0038] Fourthly, feature point detection is performed by predicting absolute coordinates by regression. This technology has high requirements for the labeling precision of training data. A slight labeling error can cause incorrect supervision information, and the labeling cost is high, which makes it difficult to apply in actual project scenarios. Moreover, the method of directly predicting the absolute position of coordinates has a large spatial range, the training process converges slowly, the model generalization performance is poor, and the application cost of the model is indirectly increased.
[0039] To solve the at least one technical problem, the embodiment of the present application provides a feature point detection method, device, computer equipment and storage medium. First, three-dimensional mesh data of a target object is obtained. A plurality of different types of feature matrices are extracted from the three-dimensional mesh data. The plurality of different types of feature matrices are input into a trained feature point detection model. The feature point detection model performs feature point detection processing on the three-dimensional mesh data according to the plurality of different types of feature matrices, and outputs a feature point detection result of the target object. The embodiment of the present application can automatically extract a plurality of different types of feature matrices in the three-dimensional mesh data, perform feature point detection processing on the plurality of different types of feature matrices globally, and does not require manual intervention, thereby improving the efficiency and recognition effect of feature point detection.
[0040] Figure 1 is a flowchart of a feature point detection method provided by the embodiment of the present application. As shown in Figure 1 , the method comprises the following steps S110-S140.
[0041] S110, three-dimensional mesh data of a target object is obtained.
[0042] In the embodiment, the target object includes a face, teeth, an industrial part, and the like, which need to be extracted for feature points.
[0043] The three-dimensional mesh data of the target object can be data obtained in advance and stored in a cloud server or a local terminal. At this time, the three-dimensional mesh data currently required for feature point detection can be obtained directly from the cloud server or the local terminal.
[0044] In addition, the three-dimensional mesh data of the target object can also be three-dimensional mesh data constructed in real time based on scanning data. The scanning data is scanning data obtained by currently scanning the target object.
[0045] S120, a plurality of different types of feature matrices are extracted from the three-dimensional mesh data.
[0046] In the embodiment, after obtaining the three-dimensional mesh data, a plurality of different types of feature matrices are extracted from the three-dimensional mesh data. The different types of feature matrices at least include a point cloud coordinate matrix (Position Matrix, referred to as P), a point normal matrix (Normal Matrix, referred to as N), a mass matrix (Mass Matrix, referred to as M), and a Laplacian matrix (Laplacian Matrix, referred to as L); or include a point cloud coordinate matrix, a point normal matrix, and a gradient matrix (Gradient Matrix, referred to as G); or include a point cloud coordinate matrix, a point normal matrix, a mass matrix, a Laplacian matrix, and a gradient matrix.
[0047] The point cloud coordinate matrix can be directly obtained from the three-dimensional mesh data, and the absolute value of the coordinates is normalized. The point normal matrix, the quality matrix, the Laplacian matrix, and the gradient matrix are the attribute information of the three-dimensional mesh itself, and can be calculated using a third-party open source library. For example, the point normal matrix, the quality matrix, the Laplacian matrix, and the gradient matrix of the three-dimensional mesh data are calculated through the algorithm interface of the preset third-party open source library according to the needs.
[0048] For a three-dimensional mesh data with n vertices, the size of the point cloud coordinate matrix is n*3; the size of the normal matrix of the mesh point is n*3; the size of the quality matrix is n*n; the size of the Laplacian matrix is n*n; and the size of the mesh gradient matrix is n*n.
[0049] The quality matrix considers the quality of each vertex in the mesh and the connection relationship between the vertices. The Laplacian matrix is composed of the degree matrix and the adjacency matrix of the mesh, both of which reflect the global information of the three-dimensional mesh data to some extent.
[0050] The gradient matrix includes the local gradient information of the three-dimensional mesh. In a two-dimensional image, the image gradient represents the change trend of the image in different directions. The larger the gradient value is, the more intense the change of the image pixel is, and the more detailed high-frequency information is contained. The change is reflected on the image, which is often the edge or boundary area of the image. For example, the Sobel operator and the Laplacian operator detect edges and textures according to the gradient or the change of the gradient of the image. The Harris corner detection detects local extreme points according to the gradients in the x and y directions of the image to achieve the purpose of detecting corner points. In three-dimensional mesh data, the gradient of the mesh also reflects the change trend of the mesh surface, and contains some local extreme information of the mesh surface. Integrating the gradient information of the mesh into the network can increase the ability of the network to extract local detailed features, improve the position accuracy of geometric feature point detection, and reduce the position error of feature point recognition.
[0051] S130, inputting the plurality of different types of feature matrices into the trained feature point detection model.
[0052] In this embodiment, according to the different training data of the feature point detection model, the point cloud coordinate matrix, the point normal matrix, the quality matrix, and the Laplacian matrix are input into the trained feature point detection model; or the point cloud coordinate matrix, the point normal matrix, and the gradient matrix are input into the trained feature point detection model; or the cloud coordinate matrix, the point normal matrix, the quality matrix, the Laplacian matrix, and the gradient matrix are input into the trained feature point detection model.
[0053] S140, the feature point detection model detects the feature points of the three-dimensional mesh data according to the plurality of different types of feature matrices, and outputs the feature point detection result of the target object.
[0054] The present embodiment will be described in detail with respect to a plurality of different types of feature matrices, including a point cloud coordinate matrix, a point normal matrix, a mass matrix, and a Laplacian matrix; or including a point cloud coordinate matrix, a point normal matrix, and a gradient matrix; or including a point cloud coordinate matrix, a point normal matrix, a mass matrix, a Laplacian matrix, and a gradient matrix.
[0055] When the plurality of different types of feature matrices include a point cloud coordinate matrix, a point normal matrix, a mass matrix, and a Laplacian matrix, refer to Figure 2 and Figure 3 wherein, Figure 2 FIG. 1 is a structural schematic diagram of a feature point detection model provided by the present application, the feature point detection model including a feature extractor module, an adaptive hierarchical fusion module, and an output module, Figure 3 FIG. 2 is a specific flowchart of step S140, which includes:
[0056] S1401a, inputting a point cloud coordinate matrix and a point normal matrix into the feature extractor module;
[0057] S1402a, performing high-dimensional mapping processing on the point cloud coordinate matrix and the point normal matrix by the feature extractor module, and outputting first features with a first feature dimension;
[0058] In the present embodiment, the point cloud feature extractor module is composed of a group of fully connected layers and activation layers, and is used to map the coordinates and normals of the point cloud data in the point cloud coordinate matrix and the point normal matrix from low-dimensional features to high-dimensional features. Specifically, the module input is a grid point cloud coordinate matrix P and a point normal matrix N, wherein each grid point contains three-dimensional coordinates [x, y, z] and a normal [nx, ny, nz], a total of 6-dimensional features, and the output is the extracted high-dimensional features F1 (first features with a first feature dimension). The input size of the entire feature extractor is n*6, and the output size is n*c1, n is the number of vertices of the input three-dimensional grid data, wherein c1 is the first feature dimension, F1 is the first feature, and in some embodiments, c1 can be set to 128. The point cloud feature extractor module is mainly responsible for upgrading the low-dimensional shallow features of the grid, and the point cloud feature extractor module can specifically be a Deep Learning on Point Sets in Metric Spaces (pointnet) model or a Deep Hierarchical Feature Learning on Point Sets in a Metric Space (pointcnn) model.
[0059] S1403a: Input the quality matrix, Laplacian matrix, and first feature into the adaptive hierarchical fusion module;
[0060] S1404a: The quality matrix, Laplacian matrix and first feature are fused hierarchically by an adaptive hierarchical fusion module to output a second feature with a second feature dimension.
[0061] In this embodiment, the adaptive hierarchical fusion module is responsible for fusing global feature information from different levels. The global information from different levels comes from the mesh quality matrix and the mesh Laplacian matrix.
[0062] In some embodiments, such as Figure 4 As shown, Figure 4 This is a structural diagram of an adaptive hierarchical fusion module. The adaptive hierarchical fusion module includes a first residual connection submodule (residual connection) and multiple fusion layers corresponding to different feature dimensions (the fusion layer includes a first fully connected unit and a channel attention pooling unit (channel attention pooling)). Figure 4 Taking three fusion layers as an example, the number of fusion layers can be set to other numbers according to actual needs (this embodiment does not limit the number of fusion layers), and the sum of multiple different feature dimensions and the first feature dimension is the second feature dimension. Specifically, the second feature can be determined through the following steps:
[0063] Based on the preset matrix factorization formula, mass matrix, and Laplacian matrix, multiple eigenvalues and their corresponding eigenvectors are determined. The eigenvalues are then sorted in ascending order to obtain an eigenvalue sequence. For each fusion layer, a first fusion weight vector is determined based on the first feature and the corresponding feature dimension. For each fusion layer, multiple target eigenvalues are determined from the eigenvalue sequence based on the corresponding feature dimension, and the eigenvectors corresponding to these target eigenvalues are identified as target eigenvectors. For each fusion layer, feature fusion processing is performed on the first feature, target eigenvalues, and target eigenvectors based on the first fusion weight vector, outputting the fusion feature corresponding to the fusion layer. A second feature is obtained by performing residual connection processing on the fusion features and the first feature corresponding to each fusion layer through the first residual connection submodule.
[0064] The fusion layer includes a first fully connected unit and a channel attention pooling unit; the first fusion weight vector corresponding to the fusion layer is determined based on the first feature and the feature dimension corresponding to the fusion layer, including:
[0065] The first feature is fully connected and channel attention pooled according to the first fully connected unit and the channel attention pooling unit to obtain the channel attention feature corresponding to the feature dimension of the fusion layer; the first fusion weight vector corresponding to the channel attention feature is determined according to the preset activation function.
[0066] Specifically, in this embodiment, in order to obtain multi-level global information and reduce the computational load of the model (which can also accelerate the model training speed), this embodiment combines the mass matrix and the Laplacian matrix, and uses matrix factorization to solve for the corresponding eigenvectors and eigenvalues. The preset matrix factorization formula is as follows:
[0067] L·M·eVec=eVal*eVec
[0068] Where · represents matrix multiplication, eVec is the eigenvector, eVal is the eigenvalue, * represents scalar multiplication, M is the mass matrix, and L is the Laplacian matrix. The eigenvalues and eigenvectors of a matrix represent its intrinsic properties, encompassing its main information. In the eigenvalues and eigenvectors of a mesh's mass matrix and Laplacian matrix, eigenvectors with larger eigenvalues represent the mesh's high-frequency information, covering details of local adjacency between points and edges, while eigenvectors with smaller eigenvalues better represent the mesh's low-frequency information, covering global adjacency semantics between points and edges.
[0069] Then, the eigenvalues and eigenvectors are sorted in ascending order of their eigenvalues to obtain an eigenvalue sequence. Combinations of eigenvalues and eigenvectors of different dimensions represent different levels of global and subglobal information representations of the grid. In this embodiment, the fusion layer uses an adaptive weight allocation strategy to automatically select eigenvalues and eigenvectors for each level. Before fusion, a weight is calculated for each eigenvector; eigenvectors with larger weights fuse more information, while those with smaller weights fuse less information. The weights are derived from the eigenvectors themselves. Specifically, Figure 4 The feature dimensions of each fusion layer are c1, c2, and c3, where c1 can be 128, c2 can be 256, and c3 can be 512. The second feature dimension is c4, which has the value of c1+c1+c2+c3=1024. It should be noted that the values of the feature dimensions in this embodiment are for illustrative purposes only. In practical applications, other values can be set according to actual needs. This application embodiment does not limit the specific values of the feature dimensions.
[0070] like Figure 4 As shown, in Figure 4The first fusion layer (counted from top to bottom) in the first column corresponds to a feature dimension of c1. The fusion layer performs a full connection layer and channel pooling on the first feature F1 output by the feature extractor to obtain a channel attention feature of 1*c1, and then uses a sigmoid activation function (i.e., a preset activation function) to obtain a weight vector W1 of 1*c1. The weight value of each channel is limited to the range of 0-1. According to the order of feature values from small to large, the first c1 dimensions of eVal1 and eVec1 are selected from all feature values and feature vectors (i.e., the target feature value and the target feature vector are determined), and the first feature F1 is fused with the selected feature value and feature vector. Since the calculation amount of directly fusing the quality matrix and the Laplacian matrix is very large, the quality matrix M and the Laplacian matrix L are subjected to eigenvalue and eigenvector decomposition, and part of the eigenvalues and eigenvectors are fused to reduce the model calculation amount. The fusion formula is as follows:
[0071] F1^=W1*eVal1*eVec1*F1=W1*L·M·eVec1*F1;
[0072] where F1^ is the fused feature, i.e., the fusion feature corresponding to the first fusion layer. According to the above matrix decomposition formula, the fusion of the first feature F1 and the eigenvalue eVal1 and the eigenvector eVec1 is equivalent to the fusion of the quality matrix M and the Laplacian matrix L, and further the global adjacency relationship information of the three-dimensional grid point edges and surfaces is fused into the network. The weight vector W1 determines the importance of the eigenvector. The higher the weight value in W1, the larger the value after multiplication with the eigenvector eVec1, i.e., the eigenvector of that dimension contributes more information. Because these eigenvectors are non-exclusive, a sigmoid activation function is used. Similarly, the full connection layer in the second and third fusion layers of the embodiment, for example, respectively upgrades the first feature F1 to c2 and c3 to obtain new features F2 and F3, and after pooling and activation, weight vectors W2 and W3 and eigenvectors eVec2 and eVec3 and eigenvalues eVal2 and eVal3 corresponding to c2 and c3 dimensions are obtained. Repeat the above fusion steps to fuse the feature F2 and the weight vector W2, the eigenvector eVec2 and the eigenvalue eVal2 to obtain the fused feature F2^, and fuse the feature F3 and the weight vector W3, the eigenvector eVec3 and the eigenvalue eVal3 to obtain the fused feature F3^. Finally, all the fused features F1^, F2^, F3^ fused at different levels and the original first feature F1 are dimensionally spliced concat in a residual connection manner to obtain the final fused feature F4 (second feature) with a feature dimension of c4. The residual connection manner can retain more original information and improve detection accuracy. In the training stage of the model, the residual connection can simplify the gradient backpropagation path and speed up network training.
[0073] S1405a, the second feature input and output module decodes the second feature through the output module, and outputs a feature point detection result.
[0074] In this embodiment, the output module is a detection decoding head module, which is responsible for feature decoding of the finally fused features. The detection decoding head module includes a group of fully connected layers, and finally uses a sigmoid activation function to separately calculate each channel to convert the prediction result into a probability.
[0075] Specifically, the output module can output an n*c Gaussian heat map, where c is the number of categories of the geometric feature points to be predicted, and each channel is responsible for predicting the position of the geometric feature points of the corresponding category. In the Gaussian heat map, the vertices in the vicinity of the geometric feature points have a high response, and after the sigmoid function, their values are close to 1. The vertices that do not belong to the geometric feature point region have a low response or no response, and after the sigmoid function, their values are close to 0. Finally, by setting a response threshold and a maximum response number, the position of the feature points in the output Gaussian heat map is extracted to obtain the feature point detection result. For example, the response threshold is set to 0.3, and the maximum response number is set to 5.
[0076] This embodiment uses the prediction mode of the Gaussian heat map to replace the mode of regressing the absolute value of the coordinates. Compared with the fourth related technology, this embodiment can reduce the solution space of the detection network and speed up the training efficiency of the network. The Gaussian heat map label can reduce the labeling requirements, small labeling errors will not have a negative effect, enhance the robustness of the detection network, reduce the labeling cost, reduce the use cost of the detection network, and improve the ease of use of the network.
[0077] In addition, the adaptive hierarchical fusion module provided in this embodiment uses an adaptive hierarchical fusion strategy to decompose the quality matrix and the Laplacian matrix containing the global information of the three-dimensional grid data by using the feature values and the feature vectors, so that the global information representation of different levels can be obtained. Different weights are given to different feature vectors using an adaptive weight vector. Compared with the first and third related technologies, this embodiment improves the ability of the detection network to obtain global information, improves the detection accuracy of the geometric feature points from a global perspective, and reduces the large-scale global prediction error of the geometric feature points.
[0078] When the plurality of different types of feature matrices include a point cloud coordinate matrix, a point normal matrix, and a gradient matrix, please refer to Figure 5 and Figure 6 wherein, Figure 5 is a structural schematic diagram of a feature point detection model provided in this application, the feature point detection model includes a feature extractor module, a spatial gradient weight fusion module, and an output module, Figure 6For another specific flowchart of step S140, step S140 includes:
[0079] S1401b, input the point cloud coordinate matrix and the point normal matrix into the feature extractor module;
[0080] S1402b, perform high-dimensional mapping processing on the point cloud coordinate matrix and the point normal matrix by the feature extractor module, and output first features with a first feature dimension;
[0081] Wherein, the specific steps of S1401b-S1402b are similar to S1401a-S1402a, and the specific steps can refer to the description of S1401a-S1402a, which will not be repeated here.
[0082] S1403b, input the gradient matrix and the first feature into the spatial gradient weight fusion module;
[0083] S1404b, perform gradient fusion processing on the gradient matrix and the first feature by the spatial gradient weight fusion module, and output third features with a third feature dimension;
[0084] Wherein, as shown in Figure 7 The spatial gradient weight fusion module includes a first spatial gradient weight fusion layer (including a second full connection unit and a first spatial attention pooling unit (spatial attention pooling)), a second spatial gradient weight fusion layer (including a third full connection unit and a second spatial attention pooling unit), and a second residual connection sub-module (residual connection). The gradient matrix includes a first gradient matrix in the x direction of the three-dimensional grid data and a second gradient matrix in the y direction of the three-dimensional grid data. The gradient fusion processing on the gradient matrix and the first feature by the spatial gradient weight fusion module outputs the third features with the third feature dimension, including:
[0085] The first gradient matrix and the first feature are processed by the first spatial gradient weight fusion layer to obtain a first sub-feature. The second gradient matrix and the first feature are processed by the second spatial gradient weight fusion layer to obtain a second sub-feature. The first sub-feature and the first sub-feature are dimensionally spliced to obtain a spliced feature. The first feature and the spliced feature are processed by the second residual connection sub-module to obtain the third feature.
[0086] Wherein, the third feature dimension in the embodiment can be equal to the second feature dimension.
[0087] As shown in Figure 7As shown, the first spatial gradient weight fusion layer includes a second full connection unit and a first spatial attention pooling unit; the first gradient matrix and the first feature are subjected to gradient fusion processing through the first spatial gradient weight fusion layer to obtain a first sub-feature, including:
[0088] The first feature is subjected to spatial attention pooling processing through the first spatial attention pooling unit to obtain a first spatial feature;
[0089] A second fusion weight vector corresponding to the first spatial feature is determined through a preset activation function;
[0090] The first feature is subjected to dimension processing through the second full connection unit to obtain a first dimension processing feature;
[0091] The first sub-feature is obtained through gradient fusion processing according to the second fusion weight vector, the first dimension processing feature and the first gradient matrix.
[0092] As shown in Figure 7 the second spatial gradient weight fusion layer includes a third full connection unit and a second spatial attention pooling unit; the second gradient matrix and the first feature are subjected to gradient fusion processing through the second spatial gradient weight fusion layer to obtain a second sub-feature, including:
[0093] The first feature is subjected to spatial attention pooling processing through the second spatial attention pooling unit to obtain a second spatial feature;
[0094] A third fusion weight vector corresponding to the second spatial feature is determined through a preset activation function;
[0095] The first feature is subjected to dimension processing through the third full connection unit to obtain a second dimension processing feature;
[0096] The second sub-feature is obtained through gradient fusion processing according to the third fusion weight vector, the second dimension processing feature and the second gradient matrix.
[0097] Specifically, the gradient matrix G in the embodiment is a complex matrix, and the gradient matrix G is decomposed into two matrices Gx (first gradient matrix) and Gy (second gradient matrix), wherein Gx is the real part of G, Gy is the imaginary part of G, and the gradient matrices Gx and Gy represent the gradients of the x and y dimensions of the grid, respectively.
[0098] When the structure of the feature point detection model is Figure 5At this time, the input space gradient weight fusion module is the feature F1 output by the feature extractor module, at this time, the spatial attention pooling is performed respectively to obtain the spatial features Fx (first spatial feature) and Fy (second spatial feature), and the size is n*1, and then the sigmoid activation function is used to obtain the spatial weight vectors Wx and Wy in two directions. The feature value is limited to 0-1 through the sigmoid activation function, the greater the feature value, the higher the weight of the point after conversion, and the more local features obtained by the point after fusion. Before fusion, the first feature is processed in dimension through the second full connection unit to obtain the first dimension processing feature F5x, and the first feature is processed in dimension through the third full connection unit to obtain the second dimension processing feature F5y, wherein the dimension processing rules in the second full connection unit and the third full connection unit are self-adaptive dimension processing rules, for example, when the feature dimension input into the space gradient weight fusion module is less than the first preset feature dimension, the second full connection unit and the third full connection unit perform dimension increasing processing on the input feature, and the input feature is increased to the first preset feature dimension, so as to obtain more feature information, when the feature dimension input into the space gradient weight fusion module is greater than the second preset feature dimension, the second full connection unit and the third full connection unit perform dimension reduction processing on the input feature, and the input feature is reduced to the second preset feature dimension, so as to reduce the calculation amount of the model, wherein the first preset feature dimension and the second preset feature dimension can be the same feature dimension, or can be different feature dimensions, when they are different feature dimensions, the second preset feature dimension is greater than the first preset feature dimension.
[0099] For example, the first preset feature dimension can be set to 256, the second preset feature dimension can be set to 512, the dimension of the first feature F1 input into the space gradient weight fusion module is 128, the feature dimension is less than the first preset feature dimension, at this time, the second full connection unit performs dimension increasing processing on F1 to obtain the first dimension processing feature F5x with a dimension of 256, and the third full connection unit performs dimension increasing processing on F1 to obtain the second dimension processing feature F5y with a dimension of 256.
[0100] Then, the features F5x and F5y are fused with the gradient matrices Gx and Gy respectively, and the fusion formula is as follows:
[0101] F6x=Wx*Gx·F5x;
[0102] F6y=Wy*Gy·F5y;
[0103] Wherein, the dimensions of Gx and Gy are both n*n, the sizes of F5x and F5y are both n*c5, F6x is the fused feature in the x direction (such as the first sub-feature), and the size is n*c5, and F6y is the fused feature in the y direction (such as the second sub-feature). The dimensions of the fused features in the two directions are spliced to obtain the final fused feature F6, and the feature dimension is c6. Finally, the feature F6 and the original input feature F1 are connected in residual to obtain the final input feature F7 (the third feature), and the feature dimension is c7. In this embodiment, c5 is 256, c6 = c5 + c5 = 512, and c7 = 512.
[0104] In S1405b, the third feature is input into an output module, and the output module decodes the third feature to output a feature point detection result.
[0105] In this embodiment, the output module is a detection decoding head module, which is responsible for decoding the final fused feature. The detection decoding head module includes a group of fully connected layers, and finally uses a sigmoid activation function to separately calculate each channel to convert the prediction result into a probability.
[0106] Specifically, the output module outputs a Gaussian heat map of n*c, where c is the number of categories of the geometric feature points to be predicted, and each channel is responsible for predicting the position of the geometric feature points of the corresponding category. In the Gaussian heat map, the vertices in the region near the geometric feature points have a high response, and after the sigmoid function, their values are close to 1. The vertices that do not belong to the region of the geometric feature points have a low response or no response, and after the sigmoid function, their values are close to 0. Finally, by setting a response threshold and a maximum response number, the position of the feature points in the output Gaussian heat map is extracted to obtain the feature point detection result. For example, the response threshold is set to 0.3, and the maximum response number is set to 5.
[0107] This embodiment uses the prediction mode of the Gaussian heat map to replace the mode of regressing the absolute value of the coordinates. Compared with the fourth related technology, this embodiment can reduce the solution space of the detection network and speed up the training efficiency of the network. The Gaussian heat map label can reduce the labeling requirements, reduce the labeling cost, reduce the use cost of the detection network, and improve the usability of the network.
[0108] In addition, the spatial gradient weight fusion module provided in this embodiment uses a spatial gradient weight fusion strategy to fuse the local gradient information in different directions, and finally reconstructs the fused feature using a spatial weight vector. Compared with the second related technology, this embodiment improves the ability of the detection network to obtain local detailed information, improves the detection accuracy of the geometric feature points from the local perspective, and reduces the small-range local prediction error of the geometric feature points.
[0109] When the plurality of different types of feature matrices include a point cloud coordinate matrix, a point normal matrix, a quality matrix, a Laplacian matrix, and a gradient matrix, refer to Figure 8 and Figure 9 wherein, Figure 8 is another structural schematic diagram of the feature point detection model provided in the present application, a feature extractor module, an adaptive hierarchical fusion module, a spatial gradient weight fusion module, and an output module, Figure 9 is a specific flowchart of step S140, which includes:
[0110] S1401c, input the point cloud coordinate matrix and the point normal matrix into the feature extractor module;
[0111] S1402c, perform high-dimensional mapping processing on the point cloud coordinate matrix and the point normal matrix through the feature extractor module, and output first features with a first feature dimension;
[0112] S1403c, input the quality matrix, the Laplacian matrix, and the first features into the adaptive hierarchical fusion module;
[0113] S1404c, perform hierarchical fusion processing on the quality matrix, the Laplacian matrix, and the first features through the adaptive hierarchical fusion module, and output second features with a second feature dimension;
[0114] Wherein, S1401c-S1404c are similar to steps S1401aS1404a, and will not be repeated here.
[0115] S1405c, input the gradient matrix and the second features into the spatial gradient weight fusion module;
[0116] S1406c, perform gradient fusion processing on the gradient matrix and the second features through the spatial gradient weight fusion module, and output third features with a third feature dimension;
[0117] Wherein, S1405c-S1406c are also similar to steps S1403b-S1404b, except that in S1403b-S1404b, the first features and the gradient matrix are input into the spatial gradient weight fusion module, and in S1405c-S1406c, the second features and the gradient matrix are input into the spatial gradient weight fusion module.
[0118] At this time, the spatial gradient weight fusion module includes a first spatial gradient weight fusion layer, a second spatial gradient weight fusion layer, and a second residual connection sub-module, the gradient matrix includes a first gradient matrix in an x direction of the three-dimensional grid data and a second gradient matrix in a y direction of the three-dimensional grid data, and the gradient fusion processing is performed on the gradient matrix and the second feature through the spatial gradient weight fusion module, and a third feature with a third feature dimension is output, including:
[0119] The first gradient matrix and the second feature are subjected to gradient fusion processing through the first spatial gradient weight fusion layer to obtain a first sub-feature; the second gradient matrix and the second feature are subjected to gradient fusion processing through the second spatial gradient weight fusion layer to obtain a second sub-feature; and the second feature, the first sub-feature, and the second sub-feature are subjected to residual connection processing through the second residual connection sub-module to obtain the third feature.
[0120] The first spatial gradient weight fusion layer includes a second full connection unit and a first spatial attention pooling unit; the first gradient matrix and the second feature are subjected to gradient fusion processing through the first spatial gradient weight fusion layer to obtain the first sub-feature, including:
[0121] The first spatial feature is obtained by subjecting the second feature to spatial attention pooling processing through the first spatial attention pooling unit; the second fusion weight vector corresponding to the first spatial feature is determined through a preset activation function; the first dimension processing feature is obtained by subjecting the second feature to dimension processing through the second full connection unit; and the first sub-feature is obtained by performing gradient fusion processing on the second fusion weight vector, the first dimension processing feature, and the first gradient matrix.
[0122] The second spatial gradient weight fusion layer includes a third full connection unit and a second spatial attention pooling unit; the second gradient matrix and the second feature are subjected to gradient fusion processing through the second spatial gradient weight fusion layer to obtain the second sub-feature, including:
[0123] The second spatial feature is obtained by subjecting the second feature to spatial attention pooling processing through the second spatial attention pooling unit; the third fusion weight vector corresponding to the second spatial feature is determined through a preset activation function; the second dimension processing feature is obtained by subjecting the second feature to dimension processing through the third full connection unit; and the second sub-feature is obtained by performing gradient fusion processing on the third fusion weight vector, the second dimension processing feature, and the second gradient matrix.
[0124] When the structure of the feature point detection model is Figure 8At this time, the input space gradient weight fusion module is the feature F4 output by the adaptive hierarchical fusion module, at this time, spatial attention pooling is performed to obtain spatial features Fx (first spatial feature) and Fy (second spatial feature), and the size is n*1, and then a sigmoid activation function is used to obtain spatial weight vectors Wx and Wy in two directions. Before fusion, the first feature is processed in dimension by a second full connection unit to obtain a first dimension processing feature F5x, and the first feature is processed in dimension by a third full connection unit to obtain a second dimension processing feature F5y, wherein the feature dimension of the first dimension processing feature and the second dimension processing feature in the embodiment is half of the second feature dimension, and when the second feature dimension is 1024, the dimension of the first dimension processing feature and the second dimension processing feature is 512. For example, the input space gradient weight fusion module is the second feature F4 output by the adaptive hierarchical fusion module, at this time, the second feature needs to be processed in dimension by a second full connection unit to obtain a first dimension processing feature F5x, and the second feature needs to be processed in dimension by a third full connection unit to obtain a second dimension processing feature F5y, at this time, since the dimension of the second feature is high, in order to reduce the model calculation amount, the dimension processing of the second full connection unit and the third full connection unit is dimension reduction processing.
[0125] Then, the features F5x and F5y are fused with the gradient matrices Gx and Gy respectively, and the fusion formula is as follows:
[0126] F6x = Wx * Gx * F5x;
[0127] F6y = Wy * Gy * F5y;
[0128] Wherein, the dimensions of Gx and Gy are both n*n, the sizes of F5x and F5y are both n*c5, F6x is the feature fused in the x direction (such as the first sub-feature), and the size is n*c5, F6y is the fusion feature F6y in the y direction (such as the second sub-feature). The dimensions of the two direction fusion features are spliced to obtain the final fusion feature F6, and the feature dimension is c6. Finally, the feature F6 and the original input feature F4 are connected in residual to obtain the final input feature F7 (third feature), and the feature dimension is c7, c5 = 1 / 2 * c4 = 512, c6 = c5 + c5 = 1024, and c7 = 1024.
[0129] S1407c, input the third feature into the output module, decode the third feature by the output module, and output the feature point detection result.
[0130] In this embodiment, the output module is a detection decoding head module, which is responsible for feature decoding of the finally fused features. The detection decoding head module includes a set of fully connected layers, and finally uses a sigmoid activation function to perform separate calculation on each channel to convert the prediction result into a probability.
[0131] Specifically, the output of the output module can be an n*c Gaussian heat map, where c is the number of categories of geometric feature points to be predicted, and each channel is responsible for predicting the position of the geometric feature points of the corresponding category. In the Gaussian heat map, the vertices in the region near the geometric feature points have a high response, and after the sigmoid function, their values are close to 1. The vertices that do not belong to the region of the geometric feature points have a low response or no response, and after the sigmoid function, their values are close to 0. Finally, by setting a response threshold and a maximum number of responses, the position of the feature points in the output Gaussian heat map is extracted to obtain the feature point detection result. For example, the response threshold is set to 0.3, and the maximum number of responses is set to 5.
[0132] In this embodiment, the prediction method of the Gaussian heat map is used instead of the method of regressing the absolute value of the coordinates. Compared with the fourth related technology, this embodiment can reduce the solution space of the detection network and speed up the training efficiency of the network. The Gaussian heat map label can reduce the labeling requirements, reduce the labeling cost, reduce the use cost of the detection network, and improve the usability of the network.
[0133] In addition, the adaptive hierarchical fusion module provided in this embodiment uses an adaptive hierarchical fusion strategy to decompose the quality matrix and the Laplacian matrix containing the global information of the three-dimensional grid data using feature values and feature vectors, so that different levels of global information representation can be obtained. Different weights are given to different feature vectors using an adaptive weight vector. Compared with the first and third related technologies, this embodiment improves the ability of the detection network to obtain global information, improves the detection accuracy of geometric feature points from a global perspective, and reduces the large-scale global prediction error of geometric feature points.
[0134] The spatial gradient weight fusion module provided in this embodiment uses a spatial gradient weight fusion strategy to fuse the local gradient information in different directions, and finally reconstructs the fused features using a spatial weight vector. Compared with the second related technology, this embodiment improves the ability of the detection network to obtain local detailed information, improves the detection accuracy of geometric feature points from a local perspective, and reduces the small-scale local prediction error of geometric feature points.
[0135] The training process of the feature point detection model of the present application is described as follows:
[0136] Before inputting a plurality of different types of feature matrices into the trained feature point detection model, the method further comprises:
[0137] Three-dimensional mesh sample data of a sample object is acquired, and the three-dimensional mesh sample data includes sample feature point position information; a plurality of different types of feature sample matrices are extracted from the three-dimensional mesh sample data; a Gaussian heat map label corresponding to the sample feature point position information is generated; the plurality of different types of feature sample matrices are taken as input parameters of a feature point detection model to be trained, and the Gaussian heat map label is taken as an output target of the feature point detection model to be trained, the feature point detection model to be trained is trained, and a trained feature point detection model is obtained.
[0138] Specifically, three-dimensional mesh data of an object to be recognized is collected, and after a certain amount of data is collected, the data is divided, and a training set, a verification set, and a test set are divided in a ratio of about 7:2:1. For example, taking three-dimensional mesh data as three-dimensional tooth mesh data as an example, about 3000 tooth mesh data are collected, of which 2100 are training sets, 600 are verification sets, and 300 are test sets.
[0139] Then, three-dimensional mesh data labeling is performed, and the class and position of the geometric feature point to be detected for each data are labeled. According to the generation method of a two-dimensional Gaussian heat map, a corresponding Gaussian heat map label (sample feature point position information) is generated by using mesh geodesic distance instead of two-dimensional pixel distance, and sample data is obtained. A plurality of different types of feature sample matrices of each sample data are extracted, and the sample feature point position information and the feature sample matrices of each sample data are saved as a separate file for subsequent model training. Each three-dimensional mesh data corresponds to a separate file.
[0140] The size of the Gaussian heat map should be n*c, where n is the number of three-dimensional mesh vertices, and c is the number of feature point classes. The number of feature points in each class is not fixed. Specifically, the tip point, the fossa point, and the edge ridge point of the tooth are labeled and detected in this embodiment.
[0141] The trained feature point detection model is obtained, and when the verification and test pass, the trained feature point detection model is output as a subsequent feature point detection module.
[0142] Further, in some embodiments, after the feature point detection result is obtained, the method further includes:
[0143] The Gaussian heat map and / or the feature points of the target object are displayed.
[0144] It can be seen that, in the embodiment, the prediction mode of the Gaussian heat map is used to replace the mode of regressing the absolute value of the coordinates, and compared with the fourth related technology, the solving space of the detection network can be reduced and the training efficiency of the network can be improved, wherein the Gaussian heat map label can reduce the labeling requirements, reduce the labeling cost, reduce the use cost of the detection network, and improve the usability of the network.
[0145] In summary, the embodiments of the present application can automatically extract a plurality of different types of feature matrices in three-dimensional mesh data, and perform feature point detection processing through a plurality of different types of feature matrices globally, without human intervention, which can improve the efficiency and recognition effect of feature point detection.
[0146] Figure 10 is a schematic block diagram of a feature point detection device provided by the embodiments of the present application. As Figure 10 shown, corresponding to the above feature point detection method, the present application also provides a feature point detection device 1000. The feature point detection device 1000 includes units for executing the above feature point detection method. Specifically, please refer to Figure 10 , the feature point detection device 1000 includes a transceiver unit 1001 and a processing unit 1002, wherein:
[0147] The transceiver unit 1001 is configured to obtain three-dimensional mesh data of a target object.
[0148] The processing unit 1002 is configured to extract a plurality of different types of feature matrices from the three-dimensional mesh data; input the plurality of different types of feature matrices into a trained feature point detection model; and perform feature point detection on the three-dimensional mesh data according to the plurality of different types of feature matrices by using the feature point detection model, and output a feature point detection result of the target object.
[0149] In some embodiments, the plurality of different types of feature matrices include a point cloud coordinate matrix, a point normal matrix, a quality matrix, and a Laplacian matrix, the feature point detection model includes a feature extractor module, an adaptive hierarchical fusion module, and an output module; and the processing unit 1002, when performing the step of performing feature point detection on the three-dimensional mesh data according to the plurality of different types of feature matrices by using the feature point detection model, and outputting the feature point detection result of the target object, is specifically configured to:
[0150] input the point cloud coordinate matrix and the point normal matrix into the feature extractor module;
[0151] perform high-dimensional mapping processing on the point cloud coordinate matrix and the point normal matrix by using the feature extractor module, and output a first feature with a first feature dimension;
[0152] input the quality matrix, the Laplacian matrix, and the first feature into the adaptive hierarchical fusion module;
[0153] The adaptive hierarchical fusion module is used for performing hierarchical fusion processing on the quality matrix, the Laplacian matrix and the first feature, and outputting a second feature with a second feature dimension.
[0154] The second feature is input into the output module, and the output module is used for decoding the second feature, and outputting a feature point detection result.
[0155] In some embodiments, the adaptive hierarchical fusion module includes a first residual connection sub-module and a plurality of fusion layers corresponding to different feature dimensions, and the sum of the plurality of different feature dimensions and the first feature dimension is the second feature dimension; when performing the step of performing hierarchical fusion processing on the quality matrix, the Laplacian matrix and the first feature by the adaptive hierarchical fusion module, and outputting a second feature with a second feature dimension, the processing unit 1002 is specifically used for:
[0156] determining a plurality of feature values and a feature vector corresponding to each feature value according to a preset matrix decomposition formula, the quality matrix and the Laplacian matrix;
[0157] sorting the plurality of feature values according to the order from small to large to obtain a feature value sequence;
[0158] for each fusion layer, determining a first fusion weight vector corresponding to the fusion layer according to the first feature and the feature dimension corresponding to the fusion layer;
[0159] for each fusion layer, determining a plurality of target feature values from the feature value sequence according to the corresponding feature dimension, and determining a target feature vector corresponding to the target feature value as the target feature vector;
[0160] for each fusion layer, performing feature fusion processing on the first feature, the target feature value and the target feature vector according to the first fusion weight vector, and outputting a fusion feature corresponding to the fusion layer;
[0161] performing residual connection processing on the fusion feature corresponding to each fusion layer and the first feature by the first residual connection sub-module to obtain the second feature.
[0162] In some embodiments, the fusion layer includes a first full connection unit and a channel attention pooling unit; when performing the step of determining the first fusion weight vector corresponding to the fusion layer according to the first feature and the feature dimension corresponding to the fusion layer, the processing unit 1002 is specifically used for:
[0163] performing full connection and channel attention pooling processing on the first feature according to the first full connection unit and the channel attention pooling unit to obtain a channel attention feature with a feature dimension corresponding to the fusion layer;
[0164] The first fusion weight vector corresponding to the channel attention feature is determined according to a preset activation function.
[0165] In some embodiments, the plurality of different types of feature matrices include a point cloud coordinate matrix, a point normal matrix, and a gradient matrix, the feature point detection model includes a feature extractor module, a spatial gradient weight fusion module, and an output module; when the processing unit 1002 executes the step of performing feature point detection on the three-dimensional mesh data according to the plurality of different types of feature matrices by the feature point detection model, and outputs the feature point detection result of the target object, it is specifically used for:
[0166] inputting the point cloud coordinate matrix and the point normal matrix into the feature extractor module;
[0167] performing high-dimensional mapping processing on the point cloud coordinate matrix and the point normal matrix by the feature extractor module, and outputting a first feature with a first feature dimension;
[0168] inputting the gradient matrix and the first feature into the spatial gradient weight fusion module;
[0169] performing gradient fusion processing on the gradient matrix and the first feature by the spatial gradient weight fusion module, and outputting a third feature with a third feature dimension;
[0170] inputting the third feature into the output module, performing decoding processing on the third feature by the output module, and outputting the feature point detection result.
[0171] In some embodiments, the spatial gradient weight fusion module includes a first spatial gradient weight fusion layer, a second spatial gradient weight fusion layer, and a second residual connection sub-module, and the gradient matrix includes a first gradient matrix in the x direction of the three-dimensional mesh data and a second gradient matrix in the y direction of the three-dimensional mesh data; when the processing unit 1002 executes the step of performing gradient fusion processing on the gradient matrix and the first feature by the spatial gradient weight fusion module, and outputs a third feature with a third feature dimension, it is specifically used for:
[0172] performing gradient fusion processing on the first gradient matrix and the first feature by the first spatial gradient weight fusion layer to obtain a first sub-feature;
[0173] performing gradient fusion processing on the second gradient matrix and the first feature by the second spatial gradient weight fusion layer to obtain a second sub-feature;
[0174] dimensionally splicing the first sub-feature and the first sub-feature to obtain a spliced feature;
[0175] performing residual connection processing on the first feature and the spliced feature by the second residual connection sub-module to obtain the third feature.
[0176] In some embodiments, the first spatial gradient weight fusion layer includes a second full connection unit and a first spatial attention pooling unit; when performing the step of performing gradient fusion processing on the first gradient matrix and the first feature through the first spatial gradient weight fusion layer to obtain the first sub-feature, the processing unit 1002 is specifically configured to:
[0177] perform spatial attention pooling processing on the first feature through the first spatial attention pooling unit to obtain a first spatial feature;
[0178] determine a second fusion weight vector corresponding to the first spatial feature through a preset activation function;
[0179] perform dimension processing on the first feature through the second full connection unit to obtain a first dimension processing feature;
[0180] perform gradient fusion processing according to the second fusion weight vector, the first dimension processing feature, and the first gradient matrix to obtain the first sub-feature.
[0181] In some embodiments, the second spatial gradient weight fusion layer includes a third full connection unit and a second spatial attention pooling unit; when performing the step of performing gradient fusion processing on the second gradient matrix and the first feature through the second spatial gradient weight fusion layer to obtain the second sub-feature, the processing unit 1002 is specifically configured to:
[0182] perform spatial attention pooling processing on the first feature through the second spatial attention pooling unit to obtain a second spatial feature;
[0183] determine a third fusion weight vector corresponding to the second spatial feature through a preset activation function;
[0184] perform dimension processing on the first feature through the third full connection unit to obtain a second dimension processing feature;
[0185] perform gradient fusion processing according to the third fusion weight vector, the second dimension processing feature, and the second gradient matrix to obtain the second sub-feature.
[0186] In some embodiments, the plurality of different types of feature matrices include a point cloud coordinate matrix, a point normal matrix, a quality matrix, a Laplacian matrix, and a gradient matrix, and the feature point detection model includes a feature extractor module, an adaptive hierarchical fusion module, a spatial gradient weight fusion module, and an output module; when performing the step of performing feature point detection on the three-dimensional mesh data according to the plurality of different types of feature matrices to output the feature point detection result of the target object, the processing unit 1002 is specifically configured to:
[0187] input the point cloud coordinate matrix and the point normal matrix into the feature extractor module;
[0188] The point cloud coordinate matrix and the point normal matrix are high-dimensional mapping processed by the feature extractor module, and a first feature with a first feature dimension is output.
[0189] The quality matrix, the Laplacian matrix, and the first feature are input into the adaptive hierarchical fusion module.
[0190] The quality matrix, the Laplacian matrix, and the first feature are hierarchically fused by the adaptive hierarchical fusion module, and a second feature with a second feature dimension is output.
[0191] The gradient matrix and the second feature are input into the spatial gradient weight fusion module.
[0192] The gradient matrix and the second feature are gradient fused by the spatial gradient weight fusion module, and a third feature with a third feature dimension is output.
[0193] The third feature is input into the output module, and the third feature is decoded by the output module, and a feature point detection result is output.
[0194] In some embodiments, before the processing unit 1002 executes the step of inputting a plurality of different types of feature matrices into the trained feature point detection model, it is specifically used for:
[0195] Obtaining three-dimensional mesh sample data of a sample object, the three-dimensional mesh sample data including sample feature point position information;
[0196] Extracting a plurality of different types of feature sample matrices from the three-dimensional mesh sample data;
[0197] Generating a Gaussian heat map label corresponding to the sample feature point position information;
[0198] Taking the plurality of different types of feature sample matrices as input parameters of a feature point detection model to be trained, taking the Gaussian heat map label as an output target of the feature point detection model to be trained, training the feature point detection model to be trained, and obtaining a trained feature point detection model.
[0199] In some embodiments, when the processing unit 1002 executes the step of the feature point detection model performing feature point detection on the three-dimensional mesh data according to the plurality of different types of feature matrices, and outputting a feature point detection result of the target object, it is specifically used for:
[0200] The feature point detection model performs feature point detection on the three-dimensional mesh data according to the plurality of different types of feature matrices, and outputs a Gaussian heat map of the target object.
[0201] Based on the Gaussian heat map, the feature points of the target object are extracted to obtain the feature point detection result.
[0202] In some embodiments, the processing unit 1002 is further configured to, after the step of extracting feature points of the target object based on the Gaussian heat map to obtain a feature point detection result, perform the following steps:
[0203] displaying the Gaussian heat map and / or the feature points of the target object.
[0204] In summary, the embodiments of the present application can automatically extract a plurality of different types of feature matrices in three-dimensional mesh data, and perform feature point detection processing based on the plurality of different types of feature matrices globally, without human intervention, thereby improving the efficiency and recognition effect of feature point detection.
[0205] It should be noted that the specific implementation process of the feature point detection apparatus 1000 and each unit can be clearly understood by those skilled in the art, and can be referred to the corresponding description in the foregoing method embodiments. For the convenience and brevity of description, it will not be repeated here.
[0206] The feature point detection apparatus 1000 described above can be implemented in the form of a computer program, which can run on a computer device as shown in Figure 11 .
[0207] Please refer to Figure 11 , Figure 11 is a schematic block diagram of a computer device provided by an embodiment of the present application. The computer device 1100 can be a terminal or a server, wherein the terminal can be a smart phone, a tablet computer, a notebook computer, a desktop computer, a personal digital assistant, a wearable device, or other electronic devices with communication functions. The server can be a stand-alone server or a server cluster composed of multiple servers.
[0208] Please refer to Figure 11 , the computer device 1100 includes a processor 1102, a memory, and a network interface 1105 connected through a system bus 1101, wherein the memory can include a non-volatile storage medium 1103 and an internal memory 1104.
[0209] The non-volatile storage medium 1103 can store an operating system 11031 and a computer program 11032. The computer program 11032 includes program instructions, which, when executed, can cause the processor 1102 to perform a feature point detection method.
[0210] The processor 1102 is configured to provide computing and control capabilities to support the operation of the entire computer device 1100.
[0211] The memory 1104 provides an environment for running of the computer program 11032 in the nonvolatile storage medium 1103, and the computer program 11032, when executed by the processor 1102, can make the processor 1102 execute a feature point detection method.
[0212] The network interface 1105 is configured to perform network communication with other devices. Those skilled in the art can understand that, Figure 11 The structure shown in the figure is only a block diagram of part of the structure related to the scheme of the present application, and does not constitute a limitation on the computer device 1100 to which the scheme of the present application is applied. The specific computer device 1100 can include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.
[0213] The processor 1102 is configured to run the computer program 11032 stored in the memory to implement the following steps:
[0214] Obtain three-dimensional mesh data of a target object;
[0215] Extract a plurality of feature matrices of different types from the three-dimensional mesh data;
[0216] Input the plurality of feature matrices of different types into a trained feature point detection model;
[0217] The feature point detection model performs feature point detection on the three-dimensional mesh data according to the plurality of feature matrices of different types, and outputs a feature point detection result of the target object.
[0218] It should be understood that, in the embodiments of the present application, the processor 1102 can be a central processing unit (CPU), and the processor 1102 can also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor.
[0219] Those skilled in the art can understand that all or part of the processes in the method of implementing the above embodiments can be completed by instructing the relevant hardware by a computer program. The computer program includes program instructions, and the computer program can be stored in a storage medium, which is a computer readable storage medium. The program instructions are executed by at least one processor in the computer system to implement the process steps of the above method embodiments.
[0220] Therefore, the application also provides a storage medium. The storage medium can be a computer readable storage medium. The storage medium stores a computer program, wherein the computer program includes program instructions. The program instructions are executed by a processor to make the processor perform the following steps:
[0221] Obtaining three-dimensional mesh data of a target object;
[0222] Extracting a plurality of feature matrices of different types from the three-dimensional mesh data;
[0223] Inputting the plurality of feature matrices of different types into a trained feature point detection model;
[0224] The feature point detection model performs feature point detection on the three-dimensional mesh data according to the plurality of feature matrices of different types, and outputs a feature point detection result of the target object.
[0225] The storage medium can be a U disk, a mobile hard disk, a read-only memory (ROM), a magnetic disk or an optical disk, and various computer readable storage media that can store program codes.
[0226] Those skilled in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be realized by electronic hardware, computer software or a combination of both. In order to clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been described in general terms in the above description. Whether the functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. A person skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the application.
[0227] In several embodiments provided in the application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of each unit is only a logical function division, and actual implementation can have another division manner. For example, a plurality of units or components can be combined or integrated into another system, or some features can be omitted or not executed.
[0228] The steps in the method embodiments of the present application can be adjusted in sequence, combined and reduced according to actual needs. The units in the device embodiments of the present application can be combined, divided and reduced according to actual needs. In addition, each functional unit in each embodiment of the present application can be integrated in one processing unit, or each unit can exist physically, or two or more units can be integrated in one unit.
[0229] The integrated unit, if realized in the form of a software functional unit and sold or used as an independent product, can be stored in a storage medium. Based on such understanding, the technical solutions of the present application essentially or say the part that contributes to the prior art, or the whole or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a plurality of instructions for causing a computer device (which can be a personal computer, a terminal or a network device, etc.) to execute all or part of the steps of the method embodiments of the present application.
[0230] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited to this. Any person skilled in the art can easily think of various equivalent modifications or replacements within the technical scope disclosed in the present application, and these modifications or replacements should be covered in the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A feature point detection method, characterized in that, include: Obtain the 3D mesh data of the target object; Extract multiple different types of feature matrices from the three-dimensional mesh data; Input multiple different types of feature matrices into the trained feature point detection model; The feature point detection model performs feature point detection on the 3D mesh data based on multiple different types of feature matrices, and outputs the feature point detection results of the target object.
2. The method according to claim 1, characterized in that, The feature matrices of various types include point cloud coordinate matrix, point normal matrix, mass matrix and Laplacian matrix, and the feature point detection model includes feature extractor module, adaptive hierarchical fusion module and output module; The feature point detection model performs feature point detection on the 3D mesh data based on multiple different types of feature matrices, and outputs the feature point detection results of the target object, including: The point cloud coordinate matrix and the point normal matrix are input into the feature extractor module; The feature extractor module performs high-dimensional mapping processing on the point cloud coordinate matrix and the point normal matrix to output a first feature with a first feature dimension. The quality matrix, the Laplacian matrix, and the first feature are input into the adaptive hierarchical fusion module; The adaptive hierarchical fusion module performs hierarchical fusion processing on the quality matrix, the Laplacian matrix, and the first feature to output a second feature with a second feature dimension. The second feature is input into the output module, and the output module decodes the second feature to output the feature point detection result.
3. The method according to claim 2, characterized in that, The adaptive hierarchical fusion module includes a first residual connection submodule and multiple fusion layers corresponding to different feature dimensions, wherein the sum of the multiple different feature dimensions and the first feature dimension is the second feature dimension; the hierarchical fusion processing of the quality matrix, the Laplacian matrix, and the first feature through the adaptive hierarchical fusion module to output a second feature with the second feature dimension includes: Based on the preset matrix decomposition formula, the mass matrix, and the Laplace matrix, multiple eigenvalues and eigenvectors corresponding to each eigenvalue are determined. Sort the feature values in ascending order to obtain a feature value sequence; For each of the fusion layers, a first fusion weight vector corresponding to the fusion layer is determined based on the first feature and the feature dimension corresponding to the fusion layer; For each of the fusion layers, multiple target feature values are determined from the feature value sequence according to the corresponding feature dimension, and the feature vector corresponding to the target feature value is determined as the target feature vector; For each of the fusion layers, feature fusion processing is performed on the first feature, the target feature value, and the target feature vector according to the first fusion weight vector, and the fusion feature corresponding to the fusion layer is output. The second feature is obtained by performing residual connection processing on the fusion features and the first feature corresponding to each fusion layer through the first residual connection submodule.
4. The method according to claim 3, characterized in that, The fusion layer includes a first fully connected unit and a channel attention pooling unit; determining the first fusion weight vector corresponding to the fusion layer based on the first feature and the feature dimension corresponding to the fusion layer includes: The first feature is fully connected and channel attention pooled according to the first fully connected unit and the channel attention pooling unit to obtain the channel attention feature corresponding to the feature dimension of the fusion layer; The first fusion weight vector corresponding to the channel attention feature is determined according to a preset activation function.
5. The method according to claim 1, characterized in that, The feature matrices of various types include point cloud coordinate matrix, point normal matrix and gradient matrix, and the feature point detection model includes feature extractor module, spatial gradient weight fusion module and output module; The feature point detection model performs feature point detection on the 3D mesh data based on multiple different types of feature matrices, and outputs the feature point detection results of the target object, including: The point cloud coordinate matrix and the point normal matrix are input into the feature extractor module; The feature extractor module performs high-dimensional mapping processing on the point cloud coordinate matrix and the point normal matrix to output a first feature with a first feature dimension. The gradient matrix and the first feature are input into the spatial gradient weight fusion module; The gradient matrix and the first feature are subjected to gradient fusion processing by the spatial gradient weight fusion module to output a third feature with a third feature dimension. The third feature is input into the output module, and the output module decodes the third feature to output the feature point detection result.
6. The method according to claim 5, characterized in that, The spatial gradient weight fusion module includes a first spatial gradient weight fusion layer, a second spatial gradient weight fusion layer, and a second residual connection submodule. The gradient matrix includes a first gradient matrix in the x-direction of the 3D mesh data and a second gradient matrix in the y-direction of the 3D mesh data. The step of performing gradient fusion processing on the gradient matrix and the first feature through the spatial gradient weight fusion module to output a third feature with a third feature dimension includes: The first sub-feature is obtained by performing gradient fusion processing on the first gradient matrix and the first feature through the first spatial gradient weight fusion layer. The second gradient matrix and the first feature are subjected to gradient fusion processing by the second spatial gradient weight fusion layer to obtain the second sub-feature; The first sub-feature and the first sub-feature are concatenated dimensionally to obtain the concatenated feature; The third feature is obtained by performing residual connection processing on the first feature and the spliced feature through the second residual connection submodule.
7. The method according to claim 6, characterized in that, The first spatial gradient weight fusion layer includes a second fully connected unit and a first spatial attention pooling unit; the step of performing gradient fusion processing on the first gradient matrix and the first feature through the first spatial gradient weight fusion layer to obtain the first sub-feature includes: The first spatial feature is obtained by performing spatial attention pooling on the first feature through the first spatial attention pooling unit; A second fusion weight vector corresponding to the first spatial feature is determined by a preset activation function; The first feature is processed by the second fully connected unit to obtain the first dimension-processed feature; The first sub-feature is obtained by performing gradient fusion processing based on the second fusion weight vector, the first dimension processing feature, and the first gradient matrix.
8. The method according to claim 7, characterized in that, The second spatial gradient weight fusion layer includes a third fully connected unit and a second spatial attention pooling unit; the step of performing gradient fusion processing on the second gradient matrix and the first feature through the second spatial gradient weight fusion layer to obtain the second sub-feature includes: The first feature is processed by spatial attention pooling through the second spatial attention pooling unit to obtain the second spatial feature; A third fusion weight vector corresponding to the second spatial feature is determined by a preset activation function; The first feature is processed by the third fully connected unit to obtain the second dimension-processed feature; The second sub-feature is obtained by performing gradient fusion processing based on the third fusion weight vector, the second dimension processing feature, and the second gradient matrix.
9. The method according to claim 1, characterized in that, The feature matrices of various types include point cloud coordinate matrix, point normal matrix, mass matrix, Laplacian matrix and gradient matrix, and the feature point detection model includes feature extractor module, adaptive hierarchical fusion module, spatial gradient weight fusion module and output module; The feature point detection model performs feature point detection on the 3D mesh data based on multiple different types of feature matrices, and outputs the feature point detection results of the target object, including: The point cloud coordinate matrix and the point normal matrix are input into the feature extractor module; The feature extractor module performs high-dimensional mapping processing on the point cloud coordinate matrix and the point normal matrix to output a first feature with a first feature dimension. The quality matrix, the Laplacian matrix, and the first feature are input into the adaptive hierarchical fusion module; The adaptive hierarchical fusion module performs hierarchical fusion processing on the quality matrix, the Laplacian matrix, and the first feature to output a second feature with a second feature dimension. The gradient matrix and the second feature are input into the spatial gradient weight fusion module; The gradient matrix and the second feature are subjected to gradient fusion processing by the spatial gradient weight fusion module to output a third feature with a third feature dimension. The third feature is input into the output module, and the output module decodes the third feature to output the feature point detection result.
10. The method according to claim 1, characterized in that, Before inputting multiple feature matrices of different types into the trained feature point detection model, the method further includes: Acquire three-dimensional mesh sample data of the sample object, wherein the three-dimensional mesh sample data includes sample feature point location information; Extract multiple feature sample matrices of different types from the three-dimensional mesh sample data; Generate Gaussian heatmap labels corresponding to the location information of the sample feature points; Multiple feature sample matrices of different types are used as input parameters of the feature point detection model to be trained, and the Gaussian heatmap labels are used as output targets of the feature point detection model to be trained. The feature point detection model to be trained is then trained to obtain a trained feature point detection model.
11. The method according to claim 1, characterized in that, The feature point detection model performs feature point detection on the 3D mesh data based on multiple different types of feature matrices, and outputs the feature point detection results of the target object, including: The feature point detection model performs feature point detection on the three-dimensional mesh data based on multiple different types of feature matrices, and outputs a Gaussian heatmap of the target object. Feature points of the target object are extracted based on the Gaussian heatmap to obtain the feature point detection results.
12. The method according to claim 11, characterized in that, After extracting feature points of the target object based on the Gaussian heatmap to obtain the feature point detection results, the method further includes: Display the Gaussian heatmap and / or the feature points of the target object.
13. A feature point detection device, characterized in that, The feature point detection device includes a transceiver unit and a processing unit, wherein: The transceiver unit is used to acquire the three-dimensional mesh data of the target object; The processing unit is configured to extract multiple different types of feature matrices from the three-dimensional mesh data; input the multiple different types of feature matrices into a trained feature point detection model; the feature point detection model performs feature point detection on the three-dimensional mesh data based on the multiple different types of feature matrices, and outputs the feature point detection result of the target object.
14. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the feature point detection method as described in any one of claims 1-12.
15. A storage medium, characterized in that, The storage medium stores a computer program, which includes program instructions that, when executed by a processor, cause the processor to perform the feature point detection method as described in any one of claims 1-12.