Three-dimensional object local detection method based on interactive interest

Through the improved point cloud annotation and feature extraction method, the three-dimensional object local detection method based on interactive interests is adopted to solve the problems of complexity and limitations in the prior art, and more accurate and efficient identification of local interest areas of three-dimensional object is achieved, improving detection performance and reliability.

CN120236118APending Publication Date: 2025-07-01DALIAN MARITIME UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510263050.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-06
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

The existing three-dimensional object region of interest recognition technology has complexity and limitations, and it is difficult to accurately capture the structural shape of objects and accurately extract implicit structural information from explicit shapes.

Method used

Through the improved point cloud annotation and feature extraction method, a three-dimensional object local detection method based on interactive interests is adopted, including an encoder mapping point clouds to high-dimensional space, constructing multi-type interest data sets, using the cuboid primitive feature extraction module and the 3D slot attention module to extract local structural features, and identifying local interest areas through the interest area identification module.

Benefits of technology

It realizes more accurate and efficient local area of ​​interest annotation and feature extraction of three-dimensional point cloud objects, improves the performance and reliability of three-dimensional object detection, and is suitable for single-region and multi-region detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120236118A_ABST
    Figure CN120236118A_ABST
Patent Text Reader

Abstract

The invention provides a three-dimensional object local detection method based on interactive interest, and relates to the technical field of computer vision, and the method comprises the steps: constructing an interest data set through a point cloud labeling interface based on interactive interest; constructing a three-dimensional object local detection model; the three-dimensional object local detection model comprises a cuboid feature extraction module used for extracting cuboid primitive features containing object shape structure local information, and a 3D groove attention module used for extracting tighter structure edge shape features and based on a 3D groove attention mechanism; a region of interest recognition module for recognizing a region of interest of the object; training network parameters based on the interest data set; and using the trained network model to obtain a single-region interest region identification result and a multi-region interest region identification result. According to the invention, the structure-based local region of interest of the three-dimensional object can be identified and extracted more accurately and efficiently, so that the performance and reliability of three-dimensional object detection are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and particularly to a method for local detection of three-dimensional objects based on interaction interest. Background Art

[0002] The recognition of regions of interest (ROIs) in three-dimensional objects is an important branch in the fields of computer vision and image processing. Its core objective is to accurately identify and extract the regions of interest from three-dimensional point cloud data for further tasks such as object detection, classification, and tracking. This technology has broad application prospects in fields such as autonomous driving, robot navigation, industrial automation, and medical imaging.

[0003] Early methods for ROI division generally followed the following steps: First, the center point of the region was selected, and then different distance calculation formulas (such as Euclidean distance and geodesic distance) were used to determine the neighborhood range. These methods can achieve ROI division to a certain extent, but there are obvious limitations. For example, PointNet is a pioneering work that successfully applied deep neural networks to point cloud datasets, using a multi-layer perceptron (MLP) with parameter sharing and a max-pooling layer constructed as a symmetric function to ensure the permutation invariance of the point cloud. However, PointNet only learns point-wise or global features and is thus limited in capturing local features of points. PointNet++ is built on top of PointNet and can learn hierarchical point cloud features in a hierarchical manner, capable of extracting local features in local neighborhoods. Although the PointNet series of methods have achieved certain results in three-dimensional object analysis, it is still difficult to learn local features based on object structure. In recent years, with the in-depth research, many studies have extended various methods for calculating local features on this basis, such as convolution-based methods, graph-based methods, and attention-based local feature extraction methods, and these works have made significant progress in extracting local features.

[0004] Although the existing technologies have made some progress in the recognition of ROIs in three-dimensional objects, there are still some significant defects. Traditional local-region-based division methods are not only complex and cumbersome but also difficult to accurately capture the structural shape of objects. It is difficult to operate when obtaining some points that are in the same structural part but far apart, so these methods face great challenges in detecting the shape of independent objects. In addition, accurately extracting implicit structural information from the explicit shape of an object remains a difficult task. A possible solution is to obtain prior information about the structure through functional annotation, that is, manually annotating each part of the object to implicitly express its structural information, but this requires accurate object structure segmentation results. And traditional three-dimensional object segmentation tasks mainly rely on spatial information and cannot obtain good shape-structure-based segmentation results. Summary of the Invention

[0005] In view of this, the present invention provides a method for local detection of three-dimensional objects based on interactive interest. The present invention aims to solve the problems existing in the existing three-dimensional object region of interest recognition technology, including the complexity and limitations of traditional region division methods, the deficiency of local feature extraction, and the challenge of accurately extracting implicit structure information from explicit shapes. Through improved point cloud annotation and feature extraction methods, the present invention can more accurately and efficiently perform the annotation task of local regions of interest of three-dimensional point cloud objects, extract local structural features of three-dimensional objects, and identify local regions of interest of three-dimensional objects, thereby improving the performance and reliability of three-dimensional object detection.

[0006] To this end, the present invention provides the following technical solutions:

[0007] The present invention provides a method for local detection of three-dimensional objects based on interactive interest, including:

[0008] Given an initial input point cloud, use an encoder to map the point cloud into a high-dimensional space to obtain point-by-point features;

[0009] Construct a multi-type interest dataset through an interest-based point cloud annotation interface; the point cloud annotation interface constrains the user-annotated region of interest by applying the results of object cuboid primitive abstraction;

[0010] Construct a three-dimensional object local detection model; the three-dimensional object local detection model includes: a cuboid feature extraction module for extracting cuboid primitive features containing local information of the object shape structure, a 3D slot attention module based on a 3D slot attention mechanism for extracting more compact structural edge shape features; an interest region recognition module for identifying the region of interest of the object;

[0011] Train network parameters based on the interest dataset;

[0012] Based on the obtained point-by-point features, use the trained network model to obtain the local region of interest recognition result. Further, it includes: using the trained network model to obtain the local region of interest recognition result, including:

[0013] Input the point cloud point-by-point features into the cuboid feature extraction module to obtain the abstract representation of the object based on the cuboid primitive and the features of the cuboid primitive;

[0014] Take the obtained abstract representation of the object based on the cuboid primitive and the features of the cuboid primitive as the initial query matrix, and take the point cloud features as the key and value, and input them into the 3D slot attention module together. Through the slot attention mechanism, use the geometric information of the point cloud to learn and correct the cuboid features to obtain more accurate local features;

[0015] The region of interest recognition module fuses the point-by-point information of the point cloud and the local features of its structure through a transformation matrix, and obtains the point-by-point interest value score of the point cloud object through a trainable classifier. The higher the score, the greater the probability of belonging to the region of interest, and the local region of interest recognition result is obtained.

[0016] Furthermore, a multi-type interest dataset is constructed through an interest-based point cloud annotation interface, including:

[0017] The interest-based point cloud annotation interface receives the region of interest that the user wants to select by clicking the mouse, records the position of the mouse click, and highlights the entire region of interest with reference to the cuboid primitive abstraction result of the point cloud object, generating point-by-point interest labels.

[0018] Furthermore, given an initial input point cloud, an encoder is used to map the point cloud into a high-dimensional space to obtain point-by-point features, including:

[0019] The input point cloud extracts the point-by-point feature fp of N points through an encoder epp containing two EdgeConv, where N is the total number of points in the point cloud.

[0020] Furthermore, the point cloud point-by-point features are input into the cuboid feature extraction module. The cuboid feature extraction module inputs the point-by-point feature fp into a fully connected layer with max pooling to extract a 1024-dimensional global feature fg, and then maps fg into the latent space to obtain a latent vector z that follows a Gaussian distribution; an initial one-hot vector feature is designed for each cuboid, where the vector vm of the m-th cuboid, when m = j, vmj = 1, when m ≠ j, vmj = 0, j ∈ (0, m); the vector is input into the encoder ecb to extract a 64-dimensional initial cuboid feature, which is respectively linked with z and input into the feature encoder ecf with shared weights to generate the cuboid feature fc.

[0021] Furthermore, through the slot attention mechanism, the geometric information of the point cloud is used to learn and correct the cuboid features to obtain more accurate local features, including:

[0022] The cuboid feature fc extracted in the cuboid feature extraction module is used as the initial value of the slot slots and input into the 3D slot attention network. The cuboid feature fc and the point-by-point feature fp are mapped into the same space through three learnable linear transformation matrices k, q, and v, and the mapped dimension is D;

[0023] In the 3D slot attention module, the attention score between the slot and the input is calculated, and the formula is:

[0024]

[0025] The update value is updated according to the attention score, and the formula is:

[0026] Among them

[0027] Update the slots through a gated recurrent unit (GRU):

[0028] slots = GRU(state = slots, inputs = updates);

[0029] Add a multi-layer perceptron with residual to improve the update performance:

[0030] slots += MLP(LayerNorm(slots)).

[0031] Furthermore, the region of interest recognition module links the updated cuboid features obtained from the 3D slot attention module with the point-wise features of the point cloud; applies a multi-layer perceptron to obtain the final interest score score ∈ (0, 1) for each point, representing the possibility that the point belongs to the region of interest. The higher the score, the more likely the point belongs to the region of interest, and vice versa; the region of interest recognition is constrained by an improved L1 loss function, and the formula is as follows:

[0032]

[0033] where y i is the labeled ground truth, is the region of interest score predicted by the network.

[0034] Furthermore, local region of interest recognition includes: the region of interest recognition task for a single structure of an object and the region of interest recognition task for multiple structures of an object; the region of interest recognition task for a single structure refers to that when the user selects the region of interest annotation, it only includes a single shape structure of the object; the region of interest recognition task for multiple structures refers to that the region of interest annotation selected by the user includes multiple shape structures of the object.

[0035] Advantages and positive effects of the present invention: The present invention proposes a three-dimensional object local detection method based on interactive interest. Its point cloud annotation interface can achieve accurate annotation of local regions of interest, with simple operation and rapid annotation, greatly improving the efficiency of constructing a three-dimensional object interest dataset. Based on the region of interest annotated by the user, this method can accurately detect the same regions of other objects of the same category from the dataset, with good visual effects, effectively completing the region of interest detection task and correctly screening out the similar regions of other objects in the dataset. In addition, this solution is not only applicable to single-region detection, but also performs well in multi-region detection, expanding the application scope and enhancing the practicability and universality in the field of three-dimensional object detection. Description of the Drawings

[0036] To more clearly illustrate the technical solutions in the embodiments of the present invention or in the prior art, the following will briefly introduce the drawings required for describing the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0037] Figure 1 It is a flowchart of the method for local detection of three-dimensional objects based on interaction interest in the embodiments of the present invention;

[0038] Figure 2 It is a network structure diagram of the three-dimensional object local detection model in the embodiments of the present invention;

[0039] Figure 3 It is a schematic diagram of point cloud annotation in the embodiments of the present invention;

[0040] Figure 4 It is a diagram of the recognition result of the point cloud region of interest for the single-region recognition task in the embodiments of the present invention; (a) is the ground truth of the annotated region of interest, and (b) is the result diagram generated by the network;

[0041] Figure 5 It is a diagram of the recognition result of the point cloud region of interest for the multi-region recognition task in the embodiments of the present invention; (a) is the ground truth of the annotated region of interest, and (b) is the result diagram generated by the network. Detailed implementation manners

[0042] In order to enable those skilled in the art to better understand the solutions of the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0043] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention described here can be implemented in an order different from those illustrated or described here. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.

[0044] In the interest-based 3D object local detection task, the core lies in accurately identifying the region of interest subjectively specified by the user. The main process is that the user selects the region of interest or the region to be recognized in the object, and by using a neural network to replace manual recognition of the same region of other similar objects in the dataset, the recognition efficiency can be greatly improved. The task of region recognition of 3D objects plays a crucial role in many fields. In the field of intelligent robots, the accurate recognition of the local structure of 3D objects in the surrounding environment can help the intelligent agent better understand the function of the object, and then help the intelligent agent complete tasks.

[0045] The main difficulty in the recognition of the region of interest of 3D objects is the difficulty in constructing the dataset. Man-made objects often show strong structural characteristics, that is, man-made objects are an organic whole composed of different partial structures. Therefore, the region of interest that people are interested in is often a certain part of the object. Traditional methods for selecting the region of interest cannot accurately select the local structure of the object. The present invention designs a local interest region annotation interface for point clouds based on the abstraction of cuboid primitives. A point cloud is a representation form of a 3D object, which is a set of points in a spatial rectangular coordinate system used to represent a 3D object. Compared with other representation forms of 3D objects, point cloud files can be obtained quickly and in large quantities. Therefore, many researchers have constructed a large number of 3D object databases based on point clouds and widely applied them in the field of 3D object analysis. However, point clouds also have problems such as disordered organization forms and inability to represent structural shapes. The present invention introduces the technology of cuboid primitive abstraction of 3D objects to constrain the local annotation of objects, so as to solve the problem of inaccurate region of interest in annotation.

[0046] Another difficulty in the recognition of the region of interest of 3D objects is that it is impossible to learn good structure-based local features during feature extraction. Since the organization form of point clouds is disordered and there is no correlation information between points, traditional methods for extracting local features of point clouds are mainly based on the Euclidean distance of point clouds, which makes it difficult to capture points that are far apart but belong to the same structural part. The present invention introduces cuboid primitive features to constrain the feature extraction of objects, so as to learn features with better structural information for downstream tasks.

[0047] As Figure 1 shown, a method for local detection of 3D objects based on interactive interest in an embodiment of the present invention includes the following steps:

[0048] S1. Construct an interest dataset through point cloud annotation;

[0049] In the embodiment of the present invention, a point cloud annotation interface is designed. By applying the results of cuboid primitive abstraction of point clouds to constrain the annotated region of interest, users can easily implement the annotation task of the region of interest, realizing fast and accurate annotation of the point cloud interest dataset.

[0050] S2. Build a 3D object local detection model based on interaction interest;

[0051] In the embodiments of the present invention, a new 3D point cloud structure feature extraction network is designed, which can learn the structure-based local features of the input point cloud and finally realize the task of identifying the local interest regions of the 3D point cloud. Given the initial input point cloud, first use the encoder to map the point cloud into a high-dimensional space to obtain point-wise features. Then input the point cloud features into a predefined cuboid feature extraction module to obtain the abstract representation of the object based on the cuboid primitive and the features of the cuboid primitive. Subsequently, take the obtained cuboid features containing the local structure information of the object as the initial query matrix, and take the point cloud features obtained in the first step as the keys and values, and input them into the slot attention module together. Through the slot attention mechanism, further learn and correct the cuboid features through the geometric information of the point cloud to obtain more accurate local features. Finally, fuse the point-wise information of the point cloud and the local features of its structure through a transformation matrix, and obtain the point-wise interest value score of the point cloud object through a single-layer MLP. The higher the score, the greater the probability of belonging to the interest region.

[0052] S3. Train the network parameters based on the interest dataset;

[0053] S4. Use the trained network model to obtain the local interest region recognition result.

[0054] The specific implementation process is as follows:

[0055] Annotation of the interest dataset: The present invention designs a more efficient point cloud annotation interface and can achieve accurate annotation of the interest regions based on the local structure. As Figure 3 shown, given a point cloud object, the interface displays the entire point cloud to the user in a visual way. The user clicks on the interest region to be selected. The interface will select the points clicked by the mouse and, taking the abstract result of the cuboid primitive of the point cloud object as a reference, highlight all the points belonging to a structural region with this point and label these points as the interested regions.

[0056] As Figure 2 shown, the specific structure of the 3D object local detection model based on interaction interest includes: a cuboid feature extraction module, a 3D slot attention module, and an interest region recognition module.

[0057] Given a point cloud object as the network input. First, extract the point-wise features fp of N points through an encoder epp containing two EdgeConv, where N is the total number of points in the point cloud.

[0058] Cuboid Feature Extraction Module: The cuboid feature module is used to extract the cuboid primitive features containing the local information of the object based on the structure. First, the obtained fp is input into the cuboid feature extraction network. Specifically, fp is input into the fully connected layer with max pooling to extract the 1024-dimensional global feature fg. Then, fg is mapped to the latent space to obtain the latent vector z that follows the Gaussian distribution. A one-hot vector is designed for each cuboid, where the vector vm of the m-th cuboid (when m = j, vmj = 1; when m ≠ j, vmj = 0, j ∈ (0, m)). The vector is input into the encoder ecb to extract the 64-dimensional initial cuboid feature, which is respectively linked with z and input into the feature encoder ecf with shared weights to generate the cuboid feature fc.

[0059] 3D Slot Attention Module: The 3D slot attention mechanism updates the cuboid features with the point-wise features of the point cloud through the iterative attention mechanism to extract more compact structured features. The cuboid feature fc extracted in the cuboid feature extraction module is used as the initial value of the slots and input into the 3D slot attention network. The cuboid feature fc and the point-wise feature fp are mapped to the same space through three learnable linear transformation matrices k, q, and v, and the mapped dimension is D. In the 3D slot attention module, first, the attention scores between the slots and the input are calculated, and the formula is:

[0060]

[0061] Then, the update value update is updated according to the attention scores, and the formula is:

[0062] where

[0063] Finally, the slots are updated through the gated recurrent unit (GRU):

[0064] slots = GRU(state = slots, inputs = updates);

[0065] A multi-layer perceptron with residuals is also added to improve the update performance:

[0066] slots += MLP(LayerNorm(slots));

[0067] Region of Interest Recognition Module: The region of interest recognition module is used to recognize the region of interest of an object. First, the updated cuboid features obtained from the 3D slot attention module are linked with the point-by-point features of the point cloud, and an attention matrix is used to determine which cuboid feature each point is linked to. Then, a multi-layer perceptron (MLP) is applied to obtain the final interest score score∈(0,1) for each point, representing the possibility that the point belongs to the region of interest. The higher the score, the more likely the point belongs to the region of interest, and vice versa. The region of interest recognition is constrained by an improved L1 loss function, and the formula is as follows:

[0068]

[0069] where y i is the labeled ground truth, is the region of interest score predicted by the network.

[0070] The local region of interest recognition results are as shown in Figure 4 and Figure 5 . Figure 4 This is the point cloud region of interest recognition result graph for the single-region recognition task in the embodiment of the present invention, that is, the region of interest specified by the user to be recognized is a single shape of the object; (a) is the labeled ground truth of the region of interest, and (b) is the region of interest result graph generated by the network; Figure 5 This is the point cloud region of interest recognition result graph for the multi-region recognition task in the embodiment of the present invention, that is, the region of interest specified by the user to be recognized is a set of multiple local shapes on the object; (a) is the labeled ground truth of the region of interest, and (b) is the result graph generated by the network.

[0071] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for local detection of three-dimensional objects based on interactive interest, characterized in that: include: Given the initial input point cloud, use the encoder to map the point cloud into a high-dimensional space to obtain point-by-point features; Constructing a multi-type interest data set through an interest-based point cloud annotation interface; the point cloud annotation interface constrains the user-annotated interest area by applying the result of the object cuboid primitive abstraction; Constructing a local detection model for a three-dimensional object; the local detection model for a three-dimensional object includes: a cuboid feature extraction module for extracting cuboid primitive features containing local information of the object shape structure, a 3D slot attention module based on a 3D slot attention mechanism for extracting more compact structural edge shape features; and an interest region recognition module for identifying the interest region of the object; Train network parameters based on the dataset of interest; Based on the obtained point-by-point features, the local region of interest recognition results are obtained using the trained network model.

2. The method for local detection of three-dimensional objects based on interactive interest according to claim 1, characterized in that: include: The trained network model is used to obtain the local interest region recognition results, including: Input the point-by-point features of the point cloud into the cuboid feature extraction module to obtain the abstract representation of the object based on the cuboid primitive and the features of the cuboid primitive; The obtained object is represented by a cuboid primitive and its features are used as the initial query matrix. The point cloud features are used as keys and values ​​and are input into the 3D slot attention module. Through the slot attention mechanism, the geometric information of the point cloud is used to learn and correct the cuboid features to obtain more accurate local features. The interest region recognition module fuses the point-by-point information of the point cloud and the local features of its structure through a transformation matrix, and obtains the point-by-point interest value score of the point cloud object through a trainable classifier. The higher the score, the greater the probability of belonging to the interest region, and the local interest region recognition result is obtained.

3. The method for local detection of three-dimensional objects based on interactive interest according to claim 1, characterized in that: Construct multi-type interest datasets through interest-based point cloud annotation interfaces, including: The interest-based point cloud annotation interface receives the interest area that the user wants to select by clicking the mouse, records the position of the mouse click, highlights the entire interest area with reference to the cuboid primitive abstraction result of the point cloud object, and generates point-by-point interest tags.

4. The method for local detection of three-dimensional objects based on interactive interest according to claim 1, characterized in that: Given the initial input point cloud, use the encoder to map the point cloud into a high-dimensional space to obtain point-by-point features, including: The input point cloud passes through an encoder epp containing two EdgeConv to extract point-by-point features fp of N points, where N is the total number of points in the point cloud.

5. The method for local detection of three-dimensional objects based on interactive interest according to claim 2, characterized in that: The point-by-point features of the point cloud are input into the cuboid feature extraction module. The cuboid feature extraction module inputs the point-by-point features fp into the fully connected layer with maximum pooling to extract the 1024-dimensional global features fg, and then maps fg to the latent space to obtain the latent vector z that obeys the Gaussian distribution; an initial one-hot vector feature is designed for each cuboid, where the vector vm of the mth cuboid is vmj=1 when m=j, and vmj=0 when m≠j, j∈(0,m); the vector is input into the encoder ecb to extract the 64-dimensional initial cuboid features, and is linked to z respectively, and input into the feature encoder ecf with shared weights to generate the cuboid features fc.

6. The method for local detection of three-dimensional objects based on interactive interest according to claim 5, characterized in that: Through the slot attention mechanism, the geometric information of the point cloud is used to learn and correct the cuboid features to obtain more accurate local features, including: The cuboid feature fc extracted from the cuboid feature extraction module is input into the 3D slot attention network as the initial value of slots. The cuboid feature fc and the point-by-point feature fp are mapped to the same space through three learnable linear transformation matrices k, q, and v. The dimension after mapping is D. In the 3D slot attention module, the attention scores of the slot and the input are calculated as follows: Update the value according to the attention score. The formula is: in Update the slot through the gated recurrent unit GRU: slots=GRU(state=slots, inputs=updates); Add a multi-layer perceptron with residuals to improve update performance: slots+=MLP(LayerNorm(slots)).

7. The method for local detection of three-dimensional objects based on interactive interest according to claim 6, characterized in that: The region of interest recognition module links the updated cuboid features obtained in the 3D slot attention module with the point-by-point features of the point cloud; A multi-layer perceptron is applied to obtain the final interest score of each point score∈(0,1), which represents the possibility that the point belongs to the region of interest. The higher the score, the more likely the point is to belong to the region of interest, and vice versa. The improved L1 loss function is used to constrain the region of interest recognition, and the formula is as follows: where y i is the true value of the annotation, Score the region of interest predicted by the network.

8. The method for local detection of three-dimensional objects based on interactive interest according to claim 1, characterized in that: Local interest region recognition includes: an interest region recognition task containing a single structure of an object and an interest region recognition task containing multiple structures of an object; the single structure interest region recognition task refers to an interest region annotation selected by a user that only contains a single shape structure of the object; the multiple structure interest region recognition task refers to an interest region annotation selected by a user that contains multiple shape structures of the object.