Zero-shot 3D model classification method guided by discriminative features
By introducing discriminant visual feature extraction module and pseudo-vision generation module in the three-dimensional model classification, the problem of ignoring local features in the prior art is solved, and a more accurate zero-sample three-dimensional model classification is achieved.
Patent Information
- Application Number
- CN202210716713.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-23
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2042-06-23
AI Technical Summary
The existing three-dimensional model classification method has the problem of paying attention to global features and ignoring local features in the zero-sample learning scenario, resulting in low performance.
A zero-sample three-dimensional model classification method based on discriminant feature guidance is proposed. Through the discriminant visual feature extraction module, cross-view attention map is learned, information interaction with local features is enhanced, and good alignment of semantic-visual features is achieved through the pseudo-visual generation module and the joint loss module.
The local discriminant feature acquisition of the three-dimensional model is improved, and more accurate zero-sample three-dimensional model classification is achieved, which improves performance.
Smart Images

Figure CN115131781B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer graphics, computer vision and intelligent recognition, and in particular to a zero-sample three-dimensional model classification method guided by discriminative features. Background Art
[0002] Compared with two-dimensional images, three-dimensional models have richer geometric information and spatial structural features, and are closer to the display scenes of human life. They are widely used in medical modeling, film entertainment, intelligent navigation and other fields. Thanks to the rapid development of artificial intelligence technology, three-dimensional model classification methods based on deep learning have achieved remarkable results. The three-dimensional model classification algorithm with views and point clouds as input has achieved a classification accuracy of more than 90% on the dataset ModelNet10 / ModelNet40. However, these methods are based on supervised learning and require that the training set is a large-scale, detailed annotated dataset containing all the classes to be identified. In fact, with the continuous growth of the types of three-dimensional models, the three-dimensional models used for training cannot contain all categories; and training annotation requires great manpower and material costs. Therefore, how to use existing knowledge to identify unknown categories when sample label data is insufficient or even completely missing has become an urgent problem to be solved in current research. To this end, scholars have proposed zero-shot learning to imitate humans to accurately identify unseen objects based only on conceptual descriptions. Three-dimensional model classification based on zero-shot learning is an emerging topic in the field of 3D vision, which aims to correctly classify untrained three-dimensional models. For the input 3D model and its class label, the existing methods mainly extract the global feature descriptor of the 3D model through the visual extraction network, extract the semantic feature vector of the class label through the semantic feature learning network, and then map the two to the same feature space based on the consistency constraint to capture the semantic-visual cross-domain connection, and then complete the recognition of unknown classes. This type of method has achieved certain results, but there are problems such as focusing on the global and ignoring the local, forcing constraints and ignoring the semantic-visual cross-domain differences, resulting in low overall performance. Summary of the invention
[0003] The purpose of the present invention is to overcome the shortcomings and deficiencies of the prior art and propose a zero-sample 3D model classification method based on discriminative feature guidance. For the zero-sample 3D model classification task, the important role of local discriminative features is analyzed and demonstrated, thereby achieving better performance and completing the accurate classification of zero-sample 3D models.
[0004] To achieve the above object, the technical solution provided by the present invention is: a zero-sample three-dimensional model classification method guided by discriminative features, comprising the following steps:
[0005] 1) Data input and initial feature extraction. The input is divided into two parts. One part takes the multi-view representation of the 3D model dataset as input, and then passes through the initial visual feature extraction network to obtain the multi-view feature map; the other part takes the class label of the 3D model as input, and passes through the initial semantic feature extraction network to obtain its word vector;
[0006] 2) Input the multi-view feature map into the discriminative visual feature extraction module to obtain the final discriminative visual features of the 3D model, i.e., the real visual features;
[0007] 3) Input the word vector into the pseudo-visual generation module to obtain the pseudo-visual features of the 3D model;
[0008] 4) The discriminative visual features and pseudo-visual features of the obtained 3D model are jointly constrained through a joint loss module to achieve a good alignment of semantic-visual features, thereby narrowing the differences between semantic-visual domains.
[0009] Further, in step 1), the three-dimensional model dataset Where: tr is the training set, Γ te is the test set, N=N tr +N te is the total number of 3D models, N tr is the number of 3D models in the training set, N te is the number of 3D models in the test set; x i represents the i-th three-dimensional model, y i ∈{1,2,…,C} is the three-dimensional model x i The corresponding class label; C = C tr +C te is the total number of categories, C tr is the number of categories in the training set, C te is the number of test set categories; the 3D model is represented as a multi-view form, I v,i Represents the three-dimensional model x i The v-th view, N v Refers to the number of multiple views of a 3D model;
[0010] Input the 3D model and class label in the training set, indicating is the i-th 3D model in the training set, For 3D models Corresponding class labels; first, the three-dimensional model Input the initial visual feature extraction network to extract each view I v,i The initial visual feature map is a matrix representation of a feature map, h, w and d represent the height, width and number of channels of the feature map respectively; wherein the initial visual feature extraction network adopts Resnet50;
[0011] The class label The input is passed through the initial semantic feature extraction network to obtain its word vector representation n is the dimension of the word vector; wherein the initial semantic feature extraction network adopts Word2Vec.
[0012] Further, in step 2), the specific situation of the discriminative visual feature extraction module is as follows:
[0013] a. Multi-view feature fusion: 3D model N v The feature maps of the views are spliced in the channel dimension to obtain the fused features The process is as follows:
[0014]
[0015] In the formula, is the feature of the i-th 3D model after multi-view feature fusion, concat is the splicing operation, is the initial visual feature map of the i-th 3D model multi-view, v is the value of the number of views, and d is the channel dimension of the feature map;
[0016] b. Cross-view attention generation: input fused features After M 1×1 convolutions, the information interaction between channels is completed, and M cross-view discriminative attention maps are obtained. The process is as follows:
[0017]
[0018] In the formula, represents the kth discriminative attention map of the i-th 3D model, is a 1×1 convolution operation, and k is the number of attention maps.
[0019] c. Single-view discriminative feature generation: In order to synchronize the M discriminative features obtained to each view, a bilinear attention pooling operation is introduced to enhance the information interaction of local features. and the discriminative attention map of the 3D model Perform a dot multiplication operation to obtain M discriminative features in N v Response area on the view The process is as follows:
[0020]
[0021] In the formula, ⊙ is the dot product operation, is the response area of the k discriminative features of the i-th 3D model on v views;
[0022] d. Synthesis of cross-view discriminative features: For each discriminative feature, further integrate the information of each view to obtain the cross-view discriminative features. First, global average pooling is used to merge spatial information, then maximum pooling is used to merge channel information, and finally the kth cross-view discriminative visual feature of the 3D model is obtained by splicing. The process is as follows:
[0023]
[0024] In the formula, is the k-th cross-view discriminative visual feature of the i-th 3D model, For splicing operation, To perform the maximum pooling operation in the channel dimension, To perform global average pooling operation in the spatial dimension, h is the height of the feature map spatial dimension, and w is the width of the feature map spatial dimension;
[0025] e. Discriminative feature generation: M independent discriminative visual features are concatenated to obtain the final discriminative visual features of the 3D model. The process is as follows:
[0026]
[0027] In the formula, F i is the final discriminative visual feature of the i-th 3D model, i.e., the real visual feature, It is a concatenation operation on k dimensions.
[0028] Further, in step 3), the specific situation of the pseudo-vision generation module is as follows:
[0029] a. Associated semantics extraction: In order to support the smooth mapping of semantics and visual features and better capture the associated semantic features between objects, the associated semantic features F corresponding to the visual discriminative features are first obtained through the semantic description screening submodule composed of full connections. r i , the process is as follows formula (6):
[0030] F r i =f 1 (W i )=δ(ω 0 W i +b 0 ) (6)
[0031] In the formula, F r i is the associated semantic feature corresponding to the i-th 3D model, W i is the word vector representation of the i-th three-dimensional model, f 1 is a semantic description filtering submodule consisting of a single fully connected layer, δ is the ReLU activation function, ω 0 is the network weight, b 0 is bias;
[0032] b. Pseudo visual feature generation: The associated semantic feature F r i Input to the generator to generate pseudo visual feature distribution The generator is composed of a three-layer fully connected network, and its process is as follows:
[0033]
[0034] In the formula, is the pseudo visual feature of the i-th 3D model, f 2 is a pseudo-visual generator consisting of a three-layer fully connected network, ω 1 ,ω 2 ,ω 3 are the network weights of each layer, b 1 、b 2 、b 3 are the biases for each layer respectively.
[0035] Further, in step 4), the joint loss module includes semantic discrimination loss and content perception loss, and the specific situation is as follows:
[0036] a. Semantic discrimination loss: Semantic discrimination loss aims to promote the consistency of the pseudo-visual features and real visual features of the 3D model in global cognition. and the real visual feature F i The input discriminator performs 0 / 1 discrimination, so that Continuously approaching the distribution of real visual features, thus encouraging pseudo visual features to be close to real visual features at the semantic level, the process is as follows formula (8):
[0037]
[0038] Where, L sd is the semantic discrimination loss, y i is the true label, is the predicted label; when the true label y i With predicted labels 1 if they are equal, 0 if they are not equal;
[0039] b. Content-aware loss: Content-aware loss aims to achieve fine-grained alignment of pseudo-visual features and real visual features on local features. This loss constrains the local detail information of the features by calculating the difference between feature vectors bit by bit, requiring the local features at corresponding positions to have high similarity. The process is as follows:
[0040]
[0041] In the formula, l refers to the feature dimension of pseudo-visual features and real features, L cp is the content-aware loss, F i The value in the jth dimension, express The value in the jth dimension.
[0042] Compared with the prior art, the present invention has the following advantages and beneficial effects:
[0043] 1. Zero-shot learning is a process of generalizing from known classes to unknown classes. It requires that the known classes and unknown classes have a certain correlation, and this correlation is more reflected in the local fine-grained level. Existing methods often use various feature extraction networks to capture the global descriptors of three-dimensional models, which makes it difficult to characterize their local discriminative attribute features, and there is a problem of insufficient visual feature extraction. To address this problem, the present invention proposes a discriminative visual feature extraction module, which first learns and generates an attention map across views, then synchronizes to each view using bilinear pooling, and finally fuses the discriminative features of multiple views to enhance the acquisition of local discriminative visual features of the three-dimensional model and generate the real visual features of the three-dimensional model.
[0044] 2. In terms of visual-semantic feature mapping, existing methods simply use consistency loss to achieve mandatory alignment of semantic features and visual features, ignoring the huge inter-domain differences between semantic features and visual features (information redundancy and feature alignment), resulting in poor mapping effect and poor recognition performance. To address this problem, the present invention designs a pseudo-visual generation module, analogizing the principles of human cognition, establishing a semantic description screening submodule, and automatically capturing the associated semantic features between objects; establishing a pseudo-visual generator of semantic features-visual images, generating pseudo-visual features describing objects based on associated semantic features, and supporting smooth mapping of semantic-visual features.
[0045] 3. The present invention constructs a joint loss module of semantic-content dual-layer perception, which includes semantic discrimination loss and content-aware loss. Among them, the semantic discrimination loss ensures the consistency of pseudo-visual features and real visual features in global cognition; the content-aware loss further realizes the fine-grained alignment of local features of pseudo-visual features and real visual features. The two work together to achieve good alignment of semantic-visual features, thereby narrowing the differences between semantic-visual domains. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] Figure 1 Schematic diagram of the method of the present invention (called DFG-ZS3D).
[0047] Figure 2 Schematic diagram of the discriminative visual feature extraction module. DETAILED DESCRIPTION
[0048] The present invention is further described in detail below in conjunction with embodiments and drawings, but the embodiments of the present invention are not limited thereto.
[0049] like Figure 1 and Figure 2 As shown, this embodiment provides a zero-sample 3D model classification method based on discriminative feature guidance, and its specific situation is as follows:
[0050] 1) Data input and initial feature extraction. The input is divided into two parts. One part takes the multi-view representation of the 3D model dataset as input, and then passes through the initial visual feature extraction network to obtain the multi-view feature map; the other part takes the class label of the 3D model as input, and passes through the initial semantic feature extraction network to obtain its word vector. The details are as follows:
[0051] 3D model dataset Where: tr is the training set, Γ te is the test set, N=N tr +N te is the total number of 3D models, N tr is the number of 3D models in the training set, N te is the number of 3D models in the test set; x i represents the i-th three-dimensional model, y i ∈{1,2,…,C} is the three-dimensional model x i The corresponding class label; C = C tr +C te is the total number of categories, C tr is the number of categories in the training set, C te is the number of test set categories; the 3D model is represented as a multi-view form, I v ,i Represents the three-dimensional model xi The v-th view, N v Refers to the number of multiple views of a 3D model. Generally, 12 views are selected to represent a 3D model.
[0052] Input the 3D model and class label in the training set, indicating is the i-th 3D model in the training set, For 3D models Corresponding class labels; first, the three-dimensional model Input the initial visual feature extraction network to extract each view I v,i The initial visual feature map is a matrix representation of a feature map, h, w and d represent the height, width and number of channels of the feature map respectively; wherein the initial visual feature extraction network adopts Resnet50;
[0053] The class label The input is passed through the initial semantic feature extraction network to obtain its word vector representation n is the dimension of the word vector; wherein the initial semantic feature extraction network adopts Word2Vec.
[0054] 2) Inputting the multi-view feature map into the discriminative visual feature extraction module to obtain the final discriminative visual features of the three-dimensional model, that is, the real visual features; wherein the specific situation of the discriminative visual feature extraction module is as follows:
[0055] a. Multi-view feature fusion: 3D model N v The feature maps of the views are spliced in the channel dimension to obtain the fused features The process is as follows:
[0056]
[0057] In the formula, is the feature of the i-th 3D model after multi-view feature fusion, concat is the splicing operation, is the initial visual feature map of the i-th 3D model multi-view, v is the value of the number of views, and d is the channel dimension of the feature map;
[0058] b. Cross-view attention generation: input fused features After M 1×1 convolutions, the information interaction between channels is completed, and M cross-view discriminative attention maps are obtained. The process is as follows:
[0059]
[0060] In the formula, represents the kth discriminative attention map of the i-th 3D model, is a 1×1 convolution operation, and k is the number of attention maps.
[0061] c. Single-view discriminative feature generation: In order to synchronize the M discriminative features obtained to each view, a bilinear attention pooling operation is introduced to enhance the information interaction of local features. and the discriminative attention map of the 3D model Perform a dot multiplication operation to obtain M discriminative features in N v Response area on the view The process is as follows:
[0062]
[0063] In the formula, ⊙ is the dot product operation, is the response area of the k discriminative features of the i-th 3D model on v views;
[0064] d. Synthesis of cross-view discriminative features: For each discriminative feature, further integrate the information of each view to obtain the cross-view discriminative features. First, global average pooling is used to merge spatial information, then maximum pooling is used to merge channel information, and finally the kth cross-view discriminative visual feature of the 3D model is obtained by splicing. The process is as follows:
[0065]
[0066] In the formula, is the k-th cross-view discriminative visual feature of the i-th 3D model, For splicing operation, To perform the maximum pooling operation in the channel dimension, To perform global average pooling operation in the spatial dimension, h is the height of the feature map spatial dimension, and w is the width of the feature map spatial dimension;
[0067] e. Discriminative feature generation: M independent discriminative visual features are concatenated to obtain the final discriminative visual features of the 3D model. The process is as follows:
[0068]
[0069] In the formula, F i is the final discriminative visual feature of the i-th 3D model, i.e., the real visual feature, It is a concatenation operation on k dimensions.
[0070] 3) Inputting the word vector into the pseudo-visual generation module to obtain the pseudo-visual features of the three-dimensional model; wherein the specific situation of the pseudo-visual generation module is as follows:
[0071] a. Associated semantic extraction: word vector W constructed by the initial semantic feature extraction network i It contains some non-discriminative features and has information redundancy. Directly using it as input will introduce too much noise to model learning. In order to support the smooth mapping of semantic-visual features and better capture the associated semantic features between objects, the associated semantic features F corresponding to the visual discriminative features are first obtained through the semantic description screening submodule composed of full connections. r i , the process is as follows formula (6):
[0072] F r i =f 1 (W i )=δ(ω 0 W i +b 0 ) (6)
[0073] In the formula, F r i is the associated semantic feature corresponding to the i-th 3D model, W i is the word vector representation of the i-th three-dimensional model, f 1 is a semantic description filtering submodule consisting of a single fully connected layer, δ is the ReLU activation function, ω 0 is the network weight, b 0 is bias;
[0074] b. Pseudo visual feature generation: The associated semantic feature F r i Input to the generator to generate pseudo visual feature distribution The generator is composed of a three-layer fully connected network, and its process is as follows:
[0075]
[0076] In the formula, is the pseudo visual feature of the i-th 3D model, f 2 is a pseudo-visual generator consisting of a three-layer fully connected network, ω 1 ,ω 2 ,ω 3 are the network weights of each layer, b 1 、b 2 、b 3 are the biases for each layer respectively.
[0077] 4) The discriminative visual features and pseudo-visual features of the obtained 3D model are jointly constrained by a joint loss module to achieve good alignment of semantic and visual features, thereby reducing the difference between semantic and visual domains; wherein the joint loss module includes semantic discriminative loss and content-aware loss, and the specific situation is as follows:
[0078] a. Semantic discrimination loss: Semantic discrimination loss aims to promote the consistency of the pseudo-visual features and real visual features of the 3D model in global cognition. and the real visual feature F i Input discriminator performs 0 / 1 discrimination, so that Continuously approaching the distribution of real visual features, thus encouraging pseudo visual features to be close to real visual features at the semantic level, the process is as follows formula (8):
[0079]
[0080] Where, L sd is the semantic discrimination loss, y i is the true label, is the predicted label; when the true label y i With predicted labels 1 if they are equal, 0 if they are not equal;
[0081] b. Content-aware loss: Content-aware loss aims to achieve fine-grained alignment of pseudo-visual features and real visual features on local features. This loss constrains the local detail information of the features by calculating the difference between feature vectors bit by bit, requiring the local features at corresponding positions to have high similarity. The process is as follows:
[0082]
[0083] In the formula, l refers to the feature dimension of pseudo-visual features and real features, L cp is the content-aware loss, F i The value in the jth dimension, express The value in the jth dimension.
[0084] Experimental configuration: The hardware environment of this experiment is Intel Core i7 2600k+Tesla V100 32GB+16GBRAM, and the software environment is Windows10 x64+CUDA 10.0+CuDNN 7.1+Pytorch 1.4.0+python 3.6+Matlab.
[0085] Dataset:
[0086] 3D datasets,The currently public zero-sample 3D model datasets are ZS3D and Ali. In order to fully test the effectiveness and universality of the algorithm, the above datasets are selected in the experiment.
[0087] ZS3D dataset, ZS3D is a zero-shot 3D model dataset built based on Shrec2014 and Shrec2015. It contains 1677 rigid 3D models from 41 classes, of which 1493 models belonging to 33 classes are used for training and 184 models belonging to another 8 classes are used for testing.
[0088] Ali dataset,Ali contains three sub-datasets, all of which use 5976 3D models of 30 classes in ModelNet40 as training sets, 908 3D models of 10 classes in ModelNet10, 301 3D models of 14 classes in McGill, and 720 3D models of 30 classes in Shrec2015 as test sets.
[0089] Semantic dataset, Google News corpus covers about 3 million words and phrases, providing sufficient semantic data source for zero-shot learning. In the experiment, Google News corpus is first used as a benchmark to train the Word2Vec model, and then the representation of all classes in the corresponding three-dimensional model dataset is input into the Word2vec model to obtain the word vector representation of the class, capturing the correlation between word vectors and establishing the semantic association between known and unknown classes.
[0090] The effectiveness and universality of this method are fully demonstrated through comparative experiments on ZS3D and Ali datasets. The experimental results are shown in Tables 1 and 2.
[0091] Table 1 Comparative experiments on the ZS3D dataset
[0092]
[0093]
[0094] Table 2 Comparative experiments on the Ali dataset (using ModelNet40 as the training set)
[0095]
[0096] The above embodiments are preferred implementation modes of the present invention, but the implementation modes of the present invention are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications that do not deviate from the spirit and principles of the present invention should be equivalent replacement methods and are included in the protection scope of the present invention.
Claims
1. A zero-sample 3D model classification method based on discriminative feature guidance, characterized in that: The following steps are involved: 1) Data input and initial feature extraction. The input is divided into two parts. One part takes the multi-view representation of the 3D model dataset as input, and then passes through the initial visual feature extraction network to obtain the multi-view feature map; the other part takes the class label of the 3D model as input, and passes through the initial semantic feature extraction network to obtain its word vector; 2) Input the multi-view feature map into the discriminative visual feature extraction module to obtain the final discriminative visual features of the 3D model, i.e., the real visual features; The specific situation of the discriminative visual feature extraction module is as follows: a. Multi-view feature fusion: 3D model N v The feature maps of the views are spliced in the channel dimension to obtain the fused features b. Cross-view attention generation: input fused features After M 1×1 convolutions, the information interaction between channels is completed, and M cross-view discriminative attention maps are obtained; c. Single-view discriminative feature generation: In order to synchronize the M discriminative features obtained to each view, a bilinear attention pooling operation is introduced to enhance the information interaction of local features. and the discriminative attention map of the 3D model Perform a dot multiplication operation to obtain M discriminative features in N v Response area on the view d. Cross-view discriminative feature synthesis: First, the response area Global average pooling is used to merge spatial information, and then maximum pooling is used to merge channel information. Finally, the kth cross-view discriminative visual feature of the 3D model is obtained by splicing. e. Discriminative feature generation: M independent discriminative visual features are concatenated to obtain the final discriminative visual features of the 3D model; 3) Input the word vector into the pseudo-visual generation module to obtain the pseudo-visual features of the 3D model; The specific situation of the pseudo-vision generation module is as follows: a. Associated semantics extraction: In order to support the smooth mapping of semantics and visual features and better capture the associated semantic features between objects, the associated semantic features F corresponding to the visual discriminative features are first obtained through the semantic description screening submodule composed of full connections. r i ; b. Pseudo visual feature generation: The associated semantic feature F r i Input to the generator to generate pseudo visual feature distribution The generator is composed of a three-layer fully connected network; 4) The discriminative visual features and pseudo-visual features of the obtained 3D model are jointly constrained through a joint loss module to achieve a good alignment of semantic-visual features, thereby narrowing the differences between semantic-visual domains.
2. The zero-sample 3D model classification method based on discriminative feature guidance according to claim 1, characterized in that: In step 1), the 3D model dataset Where: tr is the training set, Γ te is the test set, N=N tr +N te is the total number of 3D models, N tr is the number of 3D models in the training set, N te is the number of 3D models in the test set; x i represents the i-th 3D model, y i ∈{1,2,…,C} is the three-dimensional model x i The corresponding class label; C = C tr +C te is the total number of categories, C tr is the number of categories in the training set, C te is the number of test set categories; the 3D model is represented as a multi-view form, I v,i Represents the three-dimensional model x i The v-th view, N v Refers to the number of multiple views of a 3D model; Input the 3D model and class label in the training set, indicating is the i-th 3D model in the training set, For 3D models Corresponding class labels; first, the three-dimensional model Input the initial visual feature extraction network to extract each view I v,i The initial visual feature map is a matrix representation of a feature map, h, w and d represent the height, width and number of channels of the feature map respectively; wherein the initial visual feature extraction network adopts Resnet50; The class label The input passes through the initial semantic feature extraction network to obtain its word vector representation n is the dimension of the word vector; wherein the initial semantic feature extraction network adopts Word2Vec.
3. The zero-sample 3D model classification method based on discriminative feature guidance according to claim 1, characterized in that: In step 4), the joint loss module includes semantic discrimination loss and content perception loss, the details of which are as follows: a. Semantic discrimination loss: Semantic discrimination loss aims to promote the consistency of the pseudo-visual features and real visual features of the 3D model in global cognition. and the real visual feature F i The input discriminator performs 0 / 1 discrimination, so that Continuously approaching the distribution of real visual features, thus encouraging pseudo visual features to be close to real visual features at the semantic level, the process is as follows formula (8): Where, L sd is the semantic discrimination loss, y i is the true label, is the predicted label; when the true label y i With predicted labels 1 if they are equal, 0 if they are not equal; b. Content-aware loss: Content-aware loss aims to achieve fine-grained alignment of pseudo-visual features and real visual features on local features. This loss constrains the local detail information of the features by calculating the difference between feature vectors bit by bit, requiring the local features at corresponding positions to have high similarity. The process is as follows: In the formula, l refers to the feature dimension of pseudo-visual features and real features, L cp is the content-aware loss, F i The value in the jth dimension, express The value in the jth dimension.
Citation Information
Patent Citations
Zero sample sketch retrieval method based on semantic adversarial network
CN110175251A
Zero-sample image recognition method and system based on generative adversarial network
CN111476294A