Three-dimensional point cloud zero-sample classification method and device based on mutex cooperation network

By constructing a mutex collaborative neural network, the problems of data structure differences and weak representation ability in zero-shot learning of 3D point clouds are solved, and the classification performance is improved, especially the classification accuracy of unseen categories.

CN118097274BActive Publication Date: 2025-10-17SOUTH CHINA UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410252712.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-03-06
Publication Date
2025-10-17
Estimated Expiration
2044-03-06

AI Technical Summary

Technical Problem

In the zero-shot learning of 3D point clouds, existing technologies suffer from poor classification performance due to differences in data structure and weak 3D representation capabilities, which lead to model bias and weak feature discrimination.

Method used

A method based on mutex collaborative network is adopted. By constructing a mutex collaborative neural network, including a point cloud encoder, a mutex learning module and a dual-branch learning module, end-to-end training is performed using mutual exclusion loss, classification loss and regularization loss functions. Soft labels for unseen categories and τ-norm regularization loss are designed to prevent the model from being biased towards visible categories.

Benefits of technology

The accuracy and robustness of zero-shot classification of three-dimensional point clouds are improved, especially the classification accuracy of unseen categories. The problems of data structure differences and weak representation capabilities are solved, and better feature discrimination and classification effects are achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118097274B_ABST
    Figure CN118097274B_ABST
Patent Text Reader

Abstract

The application discloses a three-dimensional point cloud zero sample classification method and device based on a mutex cooperation network, which comprises the following steps: data set preprocessing, dividing the data set into visible category and invisible category samples, and obtaining semantic feature vectors of each category; constructing a mutex cooperation neural network, including a point cloud encoder, a mutex learning module, a double-branch learning module and a classification module; in the training stage, a mutex loss, a classification loss and a regularization loss function are used to update the gradient propagation, and the network is trained in an end-to-end manner; in the test stage, the double-branch features are fused, and are matched with the category semantic feature vectors to obtain a classification result. The application learns 3D features divergently through mutex learning, so that more discriminative parts of an object can be found clearly, and in order to enhance the learning features of these different parts of the object, further design is made through the utilization of a graph to propagate cooperation information between activation regions, so that the performance and accuracy of three-dimensional point cloud zero sample classification are improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of three-dimensional point cloud classification, and particularly relates to a three-dimensional point cloud zero-shot classification method and device based on a mutex cooperation network. BACKGROUND

[0002] In computer vision, image recognition or classification is an important basic task, and deep learning networks have made significant progress in learning discriminative features for this task. Despite this, these methods show a strong dependence on the quality of supervision, and the process of collecting labels requires a lot of time and labor. In addition, as new incremental classes continue to appear, new classification networks need to be retrained. Therefore, traditional image recognition networks are becoming increasingly impractical. Zero-shot learning (ZSL) is inspired by human visual perception (which can use shared attributes of visible and invisible classes to recognize new classes), and the goal is to identify targets from unseen classes that have no data available in the training phase. More realistically, generalized zero-shot learning (GZSL) attempts to train a classifier to distinguish between visible and invisible samples at the same time, and is a hot topic that has received much attention in recent years.

[0003] To implement (G)ZSL, a commonly used method involves transferring semantic knowledge between visible and invisible objects by using natural language vectors such as attributes, word vectors, and sentences. In the field of 2D images, previous work focuses on establishing a relationship between visual features of visible classes and semantic attribute embedding features. Later, some feature generation methods using generative models aim to directly optimize the problem between real data and generated data. However, its practicality in the field of three-dimensional point clouds is largely unknown.

[0004] One option is to directly apply existing image ZSL methods to 3D point clouds, however, this can be subject to some limitations, resulting in poor performance: (1) data structure difference: unlike image data in standard mesh form, three-dimensional point clouds are irregular, with less data information, only coordinate information of N points. For example, some 2D works use attention maps at the last layer of the network, which is not suitable for 3D point clouds. This forced application will change the spatial characteristics of 3D point cloud data, resulting in misplacement of features. (2) weak 3D representation ability: 2D feature extractor structures are diverse, have extensive large dataset pre-training, and have very good pre-training feature expression ability, while in the 3D field, point cloud feature encoder structure research is not deep, and there is no large general-purpose dataset pre-training, which makes it difficult to derive high discriminative features. Therefore, 3D features often exhibit less obvious clustering patterns, resulting in closer feature distances between different classes, that is, the model learned from the seen samples is biased. SUMMARY

[0005] The main purpose of the present application is to overcome the shortcomings and deficiencies of the prior art, provide a three-dimensional point cloud zero sample classification method and device based on mutual exclusion body cooperative network, aiming to solve the problems of model deviation caused by 3D data structure difference, weak 3D representation ability, and poor zero sample classification performance caused by weak feature discrimination, hoping to divergently learn 3D features, to clearly find more discriminant parts of the object, and to help zero sample learning to summarize the general features and discriminant features of the sample through divergent information.

[0006] In order to achieve the above purpose, the present application adopts the following technical scheme:

[0007] In a first aspect, the present application provides a three-dimensional point cloud zero sample classification method based on mutual exclusion body cooperative network, comprising the following steps:

[0008] Data set preprocessing, dividing the data set into visible category and invisible category samples, and obtaining the semantic feature vector of the visible category and invisible category samples;

[0009] Constructing a mutual exclusion body cooperative neural network, the mutual exclusion body cooperative neural network comprising a point cloud encoder, a mutual exclusion body learning module, a double-branch learning module and a classification module; the point cloud encoder extracts features from the three-dimensional point cloud; the mutual exclusion body learning module activates the extracted features in the mutual exclusion part to obtain K activated features; the double-branch learning module embeds the mutual exclusion body features through mutual exclusion and cooperation; the classification module is used for classifying the three-dimensional point cloud;

[0010] Updating the gradient propagation by adopting mutual exclusion loss, classification loss and regularization loss function, and training the mutual exclusion body cooperative neural network in an end-to-end manner;

[0011] Based on the trained mutual exclusion body cooperative neural network, the double-branch feature fusion is performed, and the classification result is obtained by matching with the category semantic feature vector.

[0012] As a preferred technical scheme, the visible category and invisible category samples have no intersection, and the visible category and invisible category samples respectively refer to the sample categories that the model can receive and cannot receive in training.

[0013] As a preferred technical scheme, the point cloud encoder extracts features from the three-dimensional point cloud, specifically:

[0014] The three-dimensional point cloud x is input into the point cloud encoder for semantic feature extraction, and the point cloud feature is obtained c is the dimension of the semantic feature; the point cloud encoder is a pre-trained PointNet, PointAug, PointConv or DGCNN network.

[0015] As a preferred technical solution, the mutual exclusion body learning module performs mutual exclusion partial activation on the extracted features to obtain K activated features, specifically:

[0016] For the extracted three-dimensional point cloud feature f(x), the feature is expanded to K layers through a convolution layer to construct an activated vector set The similarity loss L between the constructed activated vectors is constructed mut is:

[0017]

[0018] Where i and j represent the indices of the vector set P, and the size is in the range of 1 to K; the above formula is optimized through network training, and the K activated vectors are mutually exclusive, so the entire point cloud feature is divergently learned, and then the entire object is activated by fusing the complementary semantics from multiple mutually exclusive point cloud features; finally, the original point cloud feature is updated by element multiplication as Here Each p i in the set is subjected to a tensor multiplication operation with f(x) to obtain f′ i (x), and the K mutually exclusive features are collected to obtain F′.

[0019] As a preferred technical solution, the double-branch learning module specifically includes:

[0020] The mutual exclusion branch does not perform collaborative processing on the mutually exclusive features, but directly cascades the mutually exclusive features to obtain Then a linear layer is used to obtain features that can be matched with semantic attribute vectors The purpose is to learn the influence of the direct combination of mutually exclusive parts on classification, so as to learn the relationship between the class attribute and the combination of object parts;

[0021] The collaborative branch, since the mutually exclusive parts come from the same object, the learned features of the K activated parts of the object are highly collaborative, so a graph structure is used to explicitly model the spatial and appearance relationship information between different parts of the point cloud, specifically:

[0022] K mutual exclusion body features are used as nodes of the graph to construct the graph, and the adjacency matrix of the graph represents the pairwise relationship of the mutually exclusive parts, and the construction method is to use dot product to calculate the collaborative information in the embedding space, that is:

[0023]

[0024] Where, and θ are two learnable linear projections that combine the ReLU activation function to map the feature correlation value to a new space, and then use L-layer graph convolution for information propagation:

[0025]

[0026] wherein represents the node feature representation in the l-th layer graph convolution, i.e. D (l) the weights of the l-th layer graph convolution, and σ represents the ReLU activation function; the features after the collaborative information transmission through the L-layer graph convolution are also passed through a linear layer to obtain

[0027] As a preferred technical solution, the mutual exclusion loss, the classification loss and the regularization loss function are used to update the gradient propagation, specifically:

[0028] The mutual exclusion loss refers to the similarity loss L between the activation vectors mut , and by optimizing the loss, the similarity between the activation vectors is reduced to obtain K mutually exclusive features of different regions of the activation object;

[0029] The classification loss is different from the standard full-supervised classification, because under the zero-shot condition, there is a difference in the sample categories used during training and testing. During training, the category s corresponding to the training sample can be obtained, and the relationship W between the i-th point cloud sample feature and the corresponding category attribute feature vector is established, and a fixed weight linear layer is used to replace the dot product similarity to represent the relationship between the two, and cross-entropy loss is used for training:

[0030]

[0031] To avoid the problem that the model prediction is biased towards the visible class due to only visible class samples during training, resulting in low classification accuracy of the invisible class, a soft label V is constructed for the invisible class sample:

[0032]

[0033] wherein the j-th row of V represents the similarity between the visible class i and the invisible class j, and are the semantic attribute feature matrices of the visible and invisible classes (i.e., the set of category attribute feature vectors), and M and N represent the number of visible and invisible classes, respectively, and γ is a weighting coefficient; by simplifying the above formula to a Sylvester equation, the Bartels-Stewart algorithm can be used to obtain then the pseudo-soft label is used in a supervised manner to transfer knowledge from the visible class to the invisible class:

[0034]

[0035] Here is the obtained and The closest unseen class attribute vector makes the network learn the relevant information of the unseen class in advance in an approximate manner, preventing precision deviation; in addition, the learnable classifier weight is also biased towards the source visible class, inspired by the long-tail problem strategy, and proposes to apply the tau norm to the classifier weight:

[0036]

[0037] Finally, the loss function of the overall end-to-end training is:

[0038] L total = L mut + μ1L seen + μ2L unseen + μ3L balnorm ;

[0039] Where μ1, μ2 and μ3 are loss function hyperparameters.

[0040] As a preferred technical solution, the based on the trained interbody cooperative neural network, the double-branch feature fusion is carried out, and the classification result is obtained by matching with the class semantic feature vector, specifically:

[0041] The final classification score is obtained using the trained network weight, and the index of the maximum value is the final result of classification:

[0042]

[0043] Where γ1 and γ2 are fusion hyperparameters, Y S and Y U represent the class set of visible classes and unseen classes.

[0044] Secondly, the present application provides a three-dimensional point cloud zero sample classification system based on an interbody cooperative network, which is applied to the three-dimensional point cloud zero sample classification method based on the interbody cooperative network, and comprises a preprocessing module, a network construction module, a network training module and a classification module.

[0045] The preprocessing module is used for data set preprocessing, and the data set is divided into visible class and unseen class samples, and the semantic feature vectors of the visible class and the unseen class samples are obtained.

[0046] The network construction module is configured to construct a mutex cooperative neural network, the mutex cooperative neural network comprising a point cloud encoder, a mutex learning module, a double-branch learning module and a classification module; the point cloud encoder is configured to extract features from a three-dimensional point cloud; the mutex learning module is configured to activate exclusive parts of the extracted features to obtain K activated features; the double-branch learning module is configured to embed the mutex features through two branches of mutual exclusion and cooperation; and the classification module is configured to classify the three-dimensional point cloud.

[0047] The network training module is configured to update gradient propagation by using a mutex loss, a classification loss and a regularization loss function, and train the mutex cooperative neural network in an end-to-end manner.

[0048] The classification module is configured to fuse the double-branch features based on the trained mutex cooperative neural network, match the fused double-branch features with a category semantic feature vector, and obtain a classification result.

[0049] In a third aspect, the present application provides an electronic device, which comprises:

[0050] at least one processor; and

[0051] a memory in communication with the at least one processor; wherein

[0052] the memory stores computer program instructions executable by the at least one processor, and the computer program instructions are executed by the at least one processor to enable the at least one processor to execute the three-dimensional point cloud zero-sample classification method based on the mutex cooperative network.

[0053] In a fourth aspect, the present application provides a computer readable storage medium storing a program, and the program is executed by a processor to implement the three-dimensional point cloud zero-sample classification method based on the mutex cooperative network.

[0054] Compared with the prior art, the present application has the following advantages and beneficial effects:

[0055] The three-dimensional point cloud zero-sample classification method based on the mutex cooperative network has the following advantages and beneficial effects: BRIEF DESCRIPTION OF DRAWINGS

[0056] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0057] Figure 1 This is a flowchart of a three-dimensional point cloud zero-shot classification method based on a mutex collaborative network according to an embodiment of the present invention;

[0058] Figure 2 A schematic diagram of the structure of a mutex cooperation network in an embodiment of the present invention;

[0059] Figure 3 Schematic diagram of the principle of the mutex learning module in an embodiment of the present invention;

[0060] Figure 4 Schematic diagram of the structure of a three-dimensional point cloud zero-shot classification system based on a mutex collaborative network according to an embodiment of the present invention;

[0061] Figure 5 Schematic diagram of the structure of an electronic device according to an embodiment of the present invention. DETAILED DESCRIPTION

[0062] In order to enable those skilled in the art to better understand the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present invention.

[0063] References to "embodiments" in this application mean that a particular feature, structure, or characteristic described in connection with the embodiment may be included in at least one embodiment of the application. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor does it constitute an independent or alternative embodiment that is mutually exclusive of other embodiments. It is understood, both explicitly and implicitly, by those skilled in the art that the embodiments described in this application may be combined with other embodiments.

[0064] like Figure 1 As shown, a three-dimensional point cloud zero-shot classification method based on a mutex collaborative network in this embodiment includes the following steps:

[0065] S1. Dataset preprocessing: split the dataset into visible and invisible category samples, and obtain the semantic feature vector of each category.

[0066] Furthermore, the overall point cloud data is represented as and respectively represent the label set of M visible categories and N invisible categories, and there is no intersection between them, and visible and invisible respectively refer to the categories of samples that the model can receive and cannot receive in training. In addition, and represent the visible category and the invisible semantic attribute vector, and Q is the dimension of the feature vector, which is generated by an embedding function method such as Word2Vec.

[0067] In the embodiments of the present application, the overall point cloud data adopts the public data sets ModelNet40, McGill and ScanObjectNN, which include synthetic three-dimensional scenes and real-world scenes. When experiments are conducted on the ModelNet40 and McGill benchmarks, the present application considers that the additional 30 categories introduced in ModelNet40 are visible categories compared with ModelNet10 (an earlier smaller version of ModelNet). The corresponding invisible categories are composed of samples in the ModelNet10 and McGill test sets. In experiments involving ScanObjectNN, the method of the present application utilizes samples belonging to 26 categories in ModelNet40, which do not overlap with instances in ScanObjectNN, forming a visible set. Subsequently, the present application uses 11 overlapping categories as invisible categories for testing to perform comparative analysis. The semantic attribute vectors of the categories of each data set are generated by the Word2Vec method, and the Q dimension is 300.

[0068] S2, a mutex cooperative neural network is constructed, as shown in Figure 2 The mutex cooperative neural network includes a point cloud encoder, a mutex learning module, a double-branch learning module and a classification module. The point cloud encoder extracts features from three-dimensional point clouds. The mutex learning module activates the extracted features in the mutex part to obtain K activated features. The double-branch learning module embeds the mutex feature fusion through the mutex and cooperation two branches respectively.

[0069] Further, the point cloud encoder specifically inputs the three-dimensional point cloud x into the point cloud encoder for semantic feature extraction to obtain the point cloud feature c is the dimension of the semantic feature. The point cloud encoder can be a pre-trained PointNet, PointAug, PointConv, DGCNN, etc. In the embodiments of the present application, it is deployed on four popular point cloud stems, including PointNet, PointConv, PointAug and DGCNN, and the corresponding feature dimensions c are 1024, 1024, 1024 and 2048 respectively.

[0070] Further, the mutual exclusion body learning module, please see Figure 3 , Figure 3 is a schematic diagram of the principle of the mutual exclusion body learning module in the embodiment of the application, specifically: for the extracted three-dimensional point cloud feature f(x), expand it to K layers through a convolution layer to construct an activation vector set Construct the similarity loss L between the activation vectors mut is:

[0071]

[0072] Where i and j represent the indices of the vector set P, which are in the range of 1~K. Through network training optimization of the above formula, the K activation vectors are mutually exclusive, so the entire point cloud feature can be divergently learned, and then the entire object can be activated by fusing the complementary semantics from multiple mutually exclusive point cloud features. Subsequently, the original point cloud feature is updated by element multiplication as Here represents that each p i in the set is subjected to a tensor multiplication operation with f(x) to obtain f ′i (x), and the K mutually exclusive features are collected to obtain F ′ In the embodiment of the application, K is set to 5.

[0073] Further, the dual-branch learning module, specifically: the mutually exclusive branch, the mutually exclusive branch does not perform collaborative processing on the mutually exclusive features, and directly performs a concatenation operation on the mutually exclusive features to obtain After that, a linear layer is used to obtain features that can be matched with semantic attribute vectors This branch does not perform collaborative processing on the mutually exclusive features, and the purpose is to learn the influence of the direct combination of mutually exclusive parts on classification, that is, to learn the relationship between the combination of class attributes and object parts.

[0074] The collaborative branch, since the mutually exclusive parts come from the same object, the learning features of the K activated parts of the object can be highly collaborative, so a graph structure is used to explicitly model the spatial and appearance relationship information between different parts of the point cloud. Specifically, the K mutually exclusive features are used as nodes of the graph to construct the graph, and the adjacency matrix of the graph is Indicates the pairwise relationship of the mutually exclusive parts, and the construction method is to use the dot product to calculate the collaborative information in the embedding space, that is:

[0075]

[0076] Where, and θ are two learnable linear projections that combine with a ReLU activation function to map the feature correlation values into a new space. Then the information propagation is performed by L-layer graph convolution as follows:

[0077]

[0078] wherein denotes the node feature representation in the l-th layer, i.e. denotes the node feature representation in the l-th layer, i.e.

[0079] S3, the classification module in the training stage adopts mutual exclusion loss, classification loss and regularization loss function for updating gradient propagation, so as to train the network in an end-to-end manner.

[0080] Further, the mutual exclusion loss refers to the similarity loss L between the activation vectors of the mutual exclusion body learning module mut By optimizing the loss, the similarity between the activation vectors can be reduced to obtain K mutual exclusion body features of different regions of the activation object.

[0081] Further, the classification loss is different from the standard full-supervised classification. Since there is a difference between the sample categories used in training and testing under the zero sample condition, first, the relationship W between the i-th point cloud sample feature Z p , Z g and the corresponding category attribute feature vector is established, which is represented by a fixed weight linear layer instead of dot product similarity, and second, cross-entropy loss is used for training:

[0082]

[0083] To avoid the problem that the model prediction is biased towards the visible class due to only visible class samples during training, resulting in low classification accuracy of the invisible class, a soft label V is constructed for the invisible class sample:

[0084]

[0085] wherein the j-th row of V represents the similarity between the visible class i and the invisible class j, and γ is a weighting coefficient, which is set to 5 in the embodiment of the present application. By simplifying the above formula into a Sylvester equation and using the Bartels-Stewart algorithm to solve, the following can be obtained: Then, the pseudo-soft label can be used in a supervised manner to transfer knowledge from the visible class to the invisible class:

[0086]

[0087] Herein is referred to as the obtained and The most similar invisible category attribute vector makes the network learn the relevant information of the invisible category in advance in an approximate manner, prevents precision deviation. In addition, the learnable classifier weight is also biased towards the source visible class, inspired by the long tail problem strategy, proposes to apply the tau norm to the classifier weight:

[0088]

[0089] Finally, the loss function of the overall end-to-end training is:

[0090] L total = L mut + μ1L seen + μ2L unseen + μ3L balnorm

[0091] In the embodiments of the present application, each hyperparameter is set to μ1=1, μ2=1, μ3=0.2.

[0092] S4, the test stage classification module fuses the double-branch feature and matches the category semantic feature vector to obtain a classification result.

[0093] Specifically, the final classification score is obtained using the trained network weight, and the index of the maximum value is taken as the final result of classification:

[0094]

[0095] In the embodiments of the present application, each hyperparameter is set to γ1=0.5, γ2=0.5.

[0096] Please refer to Table 1, which is an experimental effect table of the three-dimensional point cloud zero sample classification method based on the mutex cooperation network in the embodiment of the application, introduces the comparative analysis of the embodiment of the application and other most advanced 3D methods when applied to different backbone networks. It emphasizes the consistent excellent performance of the embodiment of the application in most cases. In the context of ZSL, the embodiment of the application obtains the most advanced results on various backbone encoders and datasets, and the performance is improved by about 5% compared with previous methods, thereby confirming its effectiveness. In addition, in the GZSL environment, our method performs well in improving the classification accuracy of invisible classes, while effectively solving the challenge of domain transfer. Compared with CGRL (a method specially tailored for GZSL), our method shows significant performance improvement in most scenarios. It is worth noting that when using PointConv and DGCNN as the backbone encoder, our method is about 50% higher than the existing method, highlighting its potential to use special 3D feature extractors and its superior ability to extract semantically rich information from data.

[0097] Table 1

[0098]

[0099] The application achieves the beneficial effects of designing a mutex learning module, which can more discriminatively distinguish object features by aggregating the exclusive activation regions of objects, and embedding them together with their attributes into a more independent space; designing a double-branch learning module, which propagates collaborative information between activation regions through graph modeling to obtain spatial and appearance texture relationships from different mutex local bodies; designing invisible class soft labels and τ norm regularization loss to prevent the model from excessively biasing towards visible classes during classification. It provides a good solution for zero-shot three-dimensional point cloud classification.

[0100] It should be noted that for each of the above method embodiments, in order to simplify the description, they are all expressed as a series of action combinations, but those skilled in the art should know that the application is not limited by the order of the described actions, because according to the application, certain steps can be performed in other orders or simultaneously.

[0101] Based on the same concept as the mutex collaborative network-based three-dimensional point cloud zero-shot classification method described in the aforementioned embodiment, the present invention also provides a mutex collaborative network-based three-dimensional point cloud zero-shot classification system, which can be used to implement the aforementioned mutex collaborative network-based three-dimensional point cloud zero-shot classification method. For ease of illustration, the schematic diagram of the embodiment of the mutex collaborative network-based three-dimensional point cloud zero-shot classification system only shows the parts relevant to the embodiment of the present invention. Those skilled in the art will understand that the illustrated structure does not constitute a limitation of the device, and the device may include more or fewer components than shown, or combine certain components, or have different component arrangements.

[0102] See also Figure 4 In another embodiment of the present application, a three-dimensional point cloud zero-shot classification system 100 based on a mutex collaborative network is provided, the system comprising a pre-processing module 101, a network construction module 102, a network training module 103 and a classification module 104;

[0103] The preprocessing module 101 is used for data set preprocessing, dividing the data set into visible category samples and invisible category samples, and obtaining semantic feature vectors of the visible category samples and the invisible category samples;

[0104] The network construction module 102 is used to construct a mutex cooperative neural network, which includes a point cloud encoder, a mutex learning module, a dual-branch learning module, and a classification module; the point cloud encoder extracts features from a three-dimensional point cloud; the mutex learning module performs mutually exclusive partial activation on the extracted features to obtain K activated features; the dual-branch learning module fuses and embeds the mutex features through pooling and collaboration branches respectively; and the classification module is used to classify the three-dimensional point cloud;

[0105] The network training module 103 is used to adopt the mutual exclusion loss, classification loss and regularization loss functions to update the gradient propagation and train the mutual exclusion cooperative neural network in an end-to-end manner;

[0106] The classification module 104 is used to fuse the dual-branch features based on the trained mutex cooperative neural network and match them with the category semantic feature vector to obtain a classification result.

[0107] It should be noted that the three-dimensional point cloud zero-sample classification system based on the mutex collaborative network of the present invention corresponds one-to-one to the three-dimensional point cloud zero-sample classification method based on the mutex collaborative network of the present invention. The technical features and beneficial effects described in the above-mentioned embodiment of the three-dimensional point cloud zero-sample classification method based on the mutex collaborative network are all applicable to the embodiment of the three-dimensional point cloud zero-sample classification based on the mutex collaborative network. For specific contents, please refer to the description in the embodiment of the method of the present invention. No further details will be given here. This is hereby declared.

[0108] Further, in the implementation of the three-dimensional point cloud zero sample classification system based on the mutex cooperative network of the above-mentioned embodiments, the logical division of each program module is only illustrative. In actual applications, the above-mentioned function distribution can be completed by different program modules according to needs, for example, for the configuration requirements of corresponding hardware or the convenience of software implementation. That is, the internal structure of the three-dimensional point cloud zero sample classification system based on the mutex cooperative network is divided into different program modules to complete all or part of the functions described above.

[0109] Please refer to Figure 5 In one embodiment, an electronic device implementing a three-dimensional point cloud zero sample classification method based on a mutex cooperative network is provided. The electronic device 200 can include a first processor 201, a first memory 202, and a bus. It can also include a computer program stored in the first memory 202 and executable on the first processor 201, such as a three-dimensional point cloud zero sample classification program based on a mutex cooperative network 203.

[0110] The first memory 202 includes at least one type of readable storage medium, including a flash memory, a mobile hard disk, a multimedia card, a card-type memory (such as an SD or DX memory, etc.), a magnetic memory, a magnetic disk, an optical disk, etc. In some embodiments, the first memory 202 can be an internal storage unit of the electronic device 200, such as a mobile hard disk of the electronic device 200. In other embodiments, the first memory 202 can also be an external storage device of the electronic device 200, such as a plug-in mobile hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. Further, the first memory 202 can include both an internal storage unit and an external storage device of the electronic device 200. The first memory 202 can be used not only to store application software and various data installed on the electronic device 200, such as the code of the three-dimensional point cloud zero sample classification program based on a mutex cooperative network 203, but also to temporarily store data that has been or will be output.

[0111] The first processor 201 may, in some embodiments, be composed of integrated circuits, for example, may be composed of a single packaged integrated circuit, or may be composed of multiple packaged integrated circuits of the same function or different functions, including one or more central processing units (CPUs), microprocessors, digital processing chips, graphics processors, and combinations of various control chips, etc. The first processor 201 is the control unit of the electronic device, which connects various components of the entire electronic device through various interfaces and lines, and executes programs or modules stored in the first memory 202 and calls data stored in the first memory 202 to perform various functions and process data of the electronic device 200.

[0112] Figure 5 Only the electronic device with components is shown, and those skilled in the art can understand that, Figure 5 The structure shown does not constitute a limitation on the electronic device 200, and can include fewer or more components than shown, or combine certain components, or different component arrangements.

[0113] The first memory 202 in the electronic device 200 stores a three-dimensional point cloud zero sample classification program 203 based on a mutex cooperative network, which is a combination of multiple instructions and can realize:

[0114] Data set preprocessing, dividing the data set into visible category and invisible category samples, and obtaining semantic feature vectors of the visible category and invisible category samples;

[0115] Constructing a mutex cooperative neural network, the mutex cooperative neural network includes a point cloud encoder, a mutex learning module, a double-branch learning module, and a classification module; the point cloud encoder extracts features from the three-dimensional point cloud; the mutex learning module activates the extracted features in the mutex part to obtain K activated features; the double-branch learning module embeds the mutex feature fusion through pooling and cooperation of the two branches; the classification module is used for classifying the three-dimensional point cloud;

[0116] Update the gradient propagation by taking mutex loss, classification loss, and regularization loss functions to train the mutex cooperative neural network in an end-to-end manner;

[0117] Based on the trained mutex cooperative neural network, the double-branch feature fusion is matched with the category semantic feature vector to obtain the classification result.

[0118] Further, the modules / units integrated in the electronic device 200, if realized in the form of software function units and sold or used as independent products, can be stored in a nonvolatile computer-readable storage medium. The computer-readable medium can include any entity or device capable of carrying the computer program code, recording medium, U disk, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM).

[0119] Those of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by a computer program instructing related hardware, and the program can be stored in a non-volatile computer-readable storage medium. When the program is executed, it can include the processes of the above-mentioned embodiments of each method. In the embodiments provided in the present application, any reference to memory, storage, database or other medium can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration but not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0120] The technical features of the above embodiments can be combined in any way. To make the description concise, not all possible combinations of the technical features in the above embodiments are described, but as long as the combinations of the technical features do not contradict, they should be considered within the scope of the present application.

[0121] The above embodiments are the preferred embodiments of the present application, but the embodiments of the present application are not limited to the above embodiments, and any changes, modifications, substitutions, combinations, simplifications made without departing from the spirit and principles of the present application should be considered as equivalent replacement modes and should be included in the protection scope of the present application.

Claims

1. A three-dimensional point cloud zero-shot classification method based on a mutex collaborative network, characterized by: The steps include: Dataset preprocessing: split the dataset into visible and invisible category samples, and obtain the semantic feature vectors of visible and invisible category samples; Constructing a mutex cooperative neural network, the mutex cooperative neural network includes a point cloud encoder, a mutex learning module, a dual-branch learning module, and a classification module; the point cloud encoder extracts features from a three-dimensional point cloud; the mutex learning module performs mutually exclusive partial activation on the extracted features to obtain K activated features; The dual-branch learning module embeds the mutually exclusive features through mutually exclusive and collaborative branches respectively; the classification module is used to classify the three-dimensional point cloud; Mutex loss, classification loss, and regularization loss functions are used to update gradient propagation and train the mutex cooperative neural network in an end-to-end manner. Based on the trained mutex cooperative neural network, the dual-branch features are fused and matched with the category semantic feature vector to obtain the classification result; The mutex learning module performs mutually exclusive partial activation on the extracted features to obtain K activated features, specifically: For the extracted 3D point cloud features f ( x ), expand it to K layers through a convolutional layer to build an activation vector set , construct the similarity loss between activation vectors L mut for: ; in, c is the dimension of the semantic features of the point cloud features, i and j Refers to a vector set P The index of the object is in the range of 1~K. Through network training to optimize the above formula, the K activation vectors are mutually exclusive, so the entire point cloud features are learned divergently, and then the entire object is activated by fusing the complementary semantics from multiple mutually exclusive point cloud features. Finally, the original point cloud features are updated by element-wise multiplication. , here Represents each p i Both f ( x ) performs tensor multiplication, and obtains , aggregate K mutually exclusive features to obtain ; The dual-branch learning module is specifically: Mutually exclusive branches do not perform collaborative processing on mutually exclusive features, but directly perform cascade operations on mutually exclusive features to obtain Then a linear layer is used to obtain features that can match the semantic attribute vector , Q is the dimension of the feature vector; Collaborative branch: Since the mutually exclusive parts come from the same object, the learned features of these K activated parts of the object are highly collaborative. Therefore, a graph structure is used to explicitly model the spatial and appearance relationship information between different parts of the point cloud. Specifically: The K mutex features are used as nodes to construct a graph, and its adjacency matrix Represents the pairwise relationship of mutually exclusive parts, which is constructed by using dot product to calculate the collaborative information in the embedding space, namely: ; in, and It is two learnable linear projections combined with the ReLU activation function to map the feature correlation value to a new space, and then use L layers of graph convolution to propagate information: ; in Indicates the l Node feature representation in layer graph convolution, i.e. , is the weight of the graph convolution of this layer, σ Represents the ReLU activation function; then the features after L layers of graph convolution collaborative information transmission are also obtained after a linear layer .

2. The three-dimensional point cloud zero-shot classification method based on a mutex collaborative network according to claim 1, characterized in that: There is no intersection between the visible category samples and the invisible category samples. The visible category samples and the invisible category samples refer to the sample categories that the model can accept and the sample categories that the model cannot accept during training, respectively.

3. The three-dimensional point cloud zero-shot classification method based on a mutex collaborative network according to claim 1, characterized in that: The point cloud encoder extracts features from the three-dimensional point cloud, specifically: 3D point cloud x Input the point cloud encoder to extract semantic features and obtain point cloud features , c is the dimension of semantic features; The point cloud encoder is a pre-trained PointNet, PointAug, PointConv or DGCNN network.

4. A three-dimensional point cloud zero-shot classification system based on a mutex collaborative network, characterized by: A three-dimensional point cloud zero-shot classification method based on a mutex collaborative network applied to any one of claims 1-3, comprising a preprocessing module, a network construction module, a network training module, and a classification module; The preprocessing module is used for data set preprocessing, dividing the data set into visible category and invisible category samples, and obtaining semantic feature vectors of visible category and invisible category samples; The network construction module is used to construct a mutex cooperative neural network, which includes a point cloud encoder, a mutex learning module, a dual-branch learning module and a classification module; the point cloud encoder extracts features from a three-dimensional point cloud; the mutex learning module performs mutually exclusive partial activation on the extracted features to obtain K activated features; The dual-branch learning module embeds the mutually exclusive features through mutually exclusive and collaborative branches respectively; the classification module is used to classify the three-dimensional point cloud; The network training module is used to adopt mutual exclusion loss, classification loss and regularization loss functions to update gradient propagation and train the mutex cooperative neural network in an end-to-end manner; The classification module is used to fuse the dual-branch features based on the trained mutex cooperative neural network and match them with the category semantic feature vector to obtain a classification result.

5. An electronic device, characterized in that: The electronic device comprises: at least one processor; and, a memory communicatively connected to the at least one processor; wherein, The memory stores computer program instructions that can be executed by the at least one processor, and the computer program instructions are executed by the at least one processor so that the at least one processor can execute the three-dimensional point cloud zero-shot classification method based on the mutex collaborative network as described in any one of claims 1-3.

6. A computer-readable storage medium storing a program, characterized in that: When the program is executed by a processor, the three-dimensional point cloud zero-shot classification method based on a mutex collaborative network according to any one of claims 1 to 3 is implemented.

Citation Information

Patent Citations

  • 3D point cloud data zero sample classification method, apparatus and device, and storage medium

    CN117456236A

  • Zero-shot image classification method, system, device and medium for supple-menting lacking features

    GB202317251D0