Multi-view fly recognition method based on bilinear enhancement and supervised contrastive learning
Patent Information
- Application Number
- CN202610868766.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-16
- Publication Date
- 2026-09-25
AI Technical Summary
[0008]本申请的目的是提供基于双线性增强与监督对比学习的多视角蝇类识别方法,以解决现有技术中一阶统计量表征能力不足和视角特征分布不一致的问题,提升多视角蝇类识别准确率
本申请提供了基于双线性增强与监督对比学习的多视角蝇类识别方法,针对现有MVCNN方法仅利用一阶统计特征、忽略特征通道间二阶交互信息的问题,本申请通过双线性特征增强模块对多视角融合特征进行逐元素平方操作,得到二阶特征,并将所述多视角融合特征与所述二阶特征在特征维度上进行拼接,形成双线性增强特征,由此引入二阶交互统计量,增强了对蝇类近缘种细粒度形态差异的表达能力,提升了近缘种区分能力。针对现有方法缺少视角一致性约束的问题,本申请通过监督对比损失函数约束同一样本的不同视角图像对应的视觉特征在特征空间中相互接近,并约束不同类别样本对应的视觉特征在特征空间中相互远离,增强了模型对视角变化的鲁棒性。双线性特征增强与监督对比学习分别从特征表征和特征空间分布两个维度形成协同互补,配合交叉熵损失函数与监督对比损失函数联合优化的训练策略,显著提升了多视角蝇类识别准确率;同时,双线性特征增强模块仅包含逐元素平方操作与特征维度上的拼接操作,未引入额外可学习参数,未显著增加模型复杂度。
Smart Images

Figure CN122821586A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the fields of computer vision and pattern recognition technology, and in particular to a multi-view fly recognition method based on bilinear enhancement and supervised contrastive learning. Background Technology
[0002] Flies are important public health pests, and accurate identification of their species is crucial for disease control and ecological research. Traditional fly identification relies on expert experience, which is inefficient and limited by the availability of specialized knowledge. In recent years, deep learning-based automatic image recognition technology has provided a new technical approach for fly identification, but it still faces the following technical challenges: Limitations of single-view recognition: Most existing methods are based on a single image for recognition. However, fly specimens often have problems such as posture changes, partial occlusion, and differences in shooting angle. Single-view images often cannot fully express the key morphological features for species identification of flies, resulting in limited recognition accuracy.
[0003] Multi-view feature fusion is difficult: Although multi-view learning methods can improve recognition performance by using images from multiple perspectives of the same specimen, in the feature fusion stage, existing methods mostly use simple max pooling or average pooling, which fails to fully explore the complementarity and higher-order statistical correlation between features from different perspectives, resulting in limited representational ability of the fused features.
[0004] Viewpoint inconsistency problem: Due to the different acquisition conditions of images from different viewpoints, the distribution of visual features of the same sample from different viewpoints may vary greatly. Existing methods extract visual features from different viewpoints independently, lacking an explicit constraint mechanism to ensure the consistency of visual features of the same sample from different viewpoints in the feature space, which affects the effective fusion of multi-viewpoint features.
[0005] Difficulty in distinguishing similar species: Some closely related fly species are highly similar in morphology, and it is difficult to capture their subtle differences by relying solely on first-order statistical features, resulting in insufficient ability to distinguish closely related species.
[0006] The closest prior art to this invention is the Multi-view Convolutional Neural Network (MVCNN), whose basic idea is to use a convolutional neural network with shared parameters as a view feature extractor to extract features from each view image; after global average pooling of all view feature maps, view max pooling is used to aggregate the multi-view features into a single feature vector; and the aggregated feature vector is input into a fully connected classifier for class prediction.
[0007] The above scheme has the following inherent defects: the feature fusion level is shallow, pooling and aggregation are only performed at the final feature vector level, and the interaction information between perspectives is not introduced in the feature extraction process; the first-order statistical representation ability is insufficient, only the feature mean obtained by global average pooling is used, ignoring the second-order interaction information between feature channels; there is a lack of perspective consistency constraints, deep features from different perspectives are extracted independently, and there is a lack of explicit constraints to ensure the consistency of features from different perspectives for the same sample; the distinguishing power of closely related species is weak, and it is difficult to effectively distinguish closely related fly species with highly similar morphology by relying solely on first-order features. Summary of the Invention
[0008] The purpose of this application is to provide a multi-view fly identification method based on bilinear enhancement and supervised contrastive learning, so as to solve the problems of insufficient first-order statistical representation ability and inconsistent distribution of view features in the existing technology, and improve the accuracy of multi-view fly identification.
[0009] To achieve the above objectives, this application provides the following solution: Firstly, this application provides a multi-view fly identification method based on bilinear enhancement and supervised contrastive learning, including: A multi-view learning model is constructed, which includes a backbone network, a view pooling layer, a bilinear feature enhancement module, and a classifier. Acquire images of the fly specimen to be identified from multiple perspectives; The multiple viewpoint images are input into the same backbone network, and the visual features of each viewpoint image are extracted. The view pooling layer aggregates the visual features of the multiple view images along the view dimension to obtain multi-view fused features. The multi-view fusion feature is input into the bilinear feature enhancement module, and the multi-view fusion feature is squared element-wise to obtain a second-order feature. The multi-view fusion feature and the second-order feature are then concatenated along the feature dimension to obtain the bilinear enhanced feature. The bilinear enhancement features are input into the classifier to obtain the identification results of the fly specimen; The training process of the multi-view learning model includes: joint optimization using a cross-entropy loss function and a supervised contrastive loss function; the cross-entropy loss function is calculated based on the recognition result and the true label; the supervised contrastive loss function is calculated based on the visual features of each view image, used to constrain the visual features corresponding to different view images of the same sample to be close to each other in the feature space, and to constrain the visual features corresponding to different categories of samples to be far apart in the feature space.
[0010] Optionally, the view pooling layer is a view max pooling layer, and the pooling aggregation method of the view max pooling layer is as follows: the visual features of the multiple view images are aligned by dimension, and the maximum value of the feature values of the visual features of different view images in each dimension is taken, and the maximum value obtained in each dimension constitutes the multi-view fusion feature.
[0011] Optionally, the bilinear enhancement feature is represented as: ; in, It is a bilinear enhancement feature. For multi-view fusion features, This indicates element-wise multiplication.
[0012] Optionally, the supervised contrastive loss function is calculated within a training batch consisting of multiple samples, and the calculation method is as follows: The visual features of each viewpoint image in the training batch are normalized to obtain normalized visual features. The supervised contrastive loss function for the i-th sample in the training batch at the t-th viewpoint is: ; in, Let be the supervised contrast loss value for the i-th sample at the t-th viewpoint. The total number of samples in the training batch. Let be the number of viewpoints for the a-th sample in the training batch. Let represent the normalized visual features of the i-th sample at the t-th viewpoint. Let be the set of all viewpoints in the i-th sample except for the t-th viewpoint. Represents a set Normalized visual features of the p-th viewpoint This represents the normalized visual feature of the b-th viewpoint of the a-th sample in the training batch. For temperature parameters, For indicator functions, when When the value is 1, The value is 0.
[0013] Optionally, the training process of the multi-view learning model includes two stages: In the first stage, a single-view learning model is constructed, which includes the backbone network and a single-view classifier. The single-view learning model is trained with single-view images and corresponding real labels, and optimized using the cross-entropy loss function to obtain the pre-trained backbone network parameters. In the second stage, the backbone network of the multi-view learning model is initialized with the pre-trained backbone network parameters, the multi-view learning model is trained with multi-view images and corresponding real labels, and the cross-entropy loss function and the supervised contrastive loss function are used for joint optimization.
[0014] Optionally, in the second stage of training, the learning rate of the backbone network is set to 0.1 times the learning rate of the classifier.
[0015] Optionally, the backbone network is a residual neural network.
[0016] Optionally, the fly specimens may include at least one of the following: green bottle fly, copper bottle fly, housefly, and bighead fly.
[0017] Optionally, the multiple perspective images are images of the same fly specimen taken from multiple different angles.
[0018] According to the specific embodiments provided in this application, this application has the following technical effects: This application provides a multi-view fly identification method based on bilinear enhancement and supervised contrastive learning. Addressing the issue that existing MVCNN methods only utilize first-order statistical features and ignore second-order interaction information between feature channels, this application uses a bilinear feature enhancement module to perform element-wise squaring operations on the multi-view fused features to obtain second-order features. These second-order features are then concatenated with the multi-view fused features along the feature dimension to form bilinear enhanced features. This introduces second-order interaction statistics, enhancing the ability to express fine-grained morphological differences in closely related fly species and improving the ability to distinguish closely related species. To address the lack of viewpoint consistency constraints in existing methods, this application uses a supervised contrastive loss function to constrain visual features corresponding to different viewpoint images of the same sample to be close to each other in the feature space, and to constrain visual features corresponding to different categories of samples to be far apart in the feature space, thus enhancing the model's robustness to viewpoint changes. Bilinear feature enhancement and supervised contrastive learning complement each other from the two dimensions of feature representation and feature space distribution, respectively. Combined with the training strategy of joint optimization of cross-entropy loss function and supervised contrastive loss function, the accuracy of multi-view fly identification is significantly improved. At the same time, the bilinear feature enhancement module only includes element-wise squaring operation and concatenation operation on feature dimension, without introducing additional learnable parameters, and does not significantly increase model complexity. Attached Figure Description
[0019] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0020] Figure 1 This is a flowchart illustrating the multi-view fly identification method based on bilinear enhancement and supervised contrastive learning proposed in this application. Figure 2 This is a schematic diagram of the bilinear feature enhancement module structure; Figure 3 This is a comparison chart of SupCon feature distributions. Detailed Implementation
[0021] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0022] To make the objectives, features and advantages of this application more apparent and understandable, the application will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0023] In an exemplary embodiment, a multi-view fly identification method based on bilinear enhancement and supervised contrastive learning is provided. This method can be applied to automatic fly identification in scenarios such as customs quarantine, disease transmission monitoring, and ecological monitoring. In this embodiment, the method includes steps 101 to 106. Wherein: Step 101: Construct a multi-view learning model, which includes a backbone network, a view pooling layer, a bilinear feature enhancement module, and a classifier.
[0024] In a preferred embodiment, the backbone network is a residual neural network. This application uses ResNet-18 pre-trained on the large-scale image classification dataset ImageNet as a feature extractor to extract feature maps of dimension [B, 512, 7, 7] from the input image, where B represents the batch size.
[0025] Step 102: Acquire multiple perspective images of the fly specimen to be identified. The multiple perspective images are images of the same fly specimen taken from multiple different angles. In this application, the same fly specimen includes images taken from N angles, where N ranges from 2 to 6. The fly specimen species include at least one of *Vibrio vulgaris*, *Vibrio pulcherrima*, *Housefly*, and *Golden Fly*.
[0026] Step 103: Input the multiple viewpoint images into the same backbone network and extract the visual features of each viewpoint image.
[0027] Step 104: The visual features of the multiple view images are pooled and aggregated along the view dimension through the view pooling layer to obtain multi-view fused features.
[0028] Specifically, the view pooling layer is a view max pooling layer, and the pooling aggregation method of the view max pooling layer is as follows: the visual features of the multiple view images are aligned by dimension, and the maximum value of the feature values of the visual features of different view images in each dimension is taken. The maximum values obtained in each dimension constitute the multi-view fusion feature.
[0029] Step 105: Input the multi-view fusion feature into the bilinear feature enhancement module, perform element-wise squaring on the multi-view fusion feature to obtain second-order features, and concatenate the multi-view fusion feature and the second-order feature in the feature dimension to obtain bilinear enhanced features.
[0030] Step 106: Input the bilinear enhancement feature into the classifier to obtain the identification result of the fly specimen.
[0031] The training process of the multi-view learning model includes: joint optimization using a cross-entropy loss function and a supervised contrastive loss function; the cross-entropy loss function is calculated based on the recognition result and the true label; the supervised contrastive loss function is calculated based on the visual features of each view image, used to constrain the visual features corresponding to different view images of the same sample to be close to each other in the feature space, and to constrain the visual features corresponding to different categories of samples to be far apart in the feature space.
[0032] By implementing steps 101 to 106 above, on the one hand, the multi-view fusion features are subjected to element-wise squaring operation through the bilinear feature enhancement module to obtain second-order features, and the multi-view fusion features and the second-order features are concatenated in the feature dimension to form bilinear enhanced features. This introduces second-order interaction statistics between feature channels, which makes up for the shortcomings of the existing MVCNN method that only uses first-order statistical features and ignores second-order interaction information. By amplifying fine-grained differences through second-order interaction statistics between feature channels, the ability to express fine-grained morphological differences of closely related fly species is enhanced, and the ability to distinguish closely related species is improved.
[0033] On the other hand, by introducing a supervised contrastive loss function during training, the visual features corresponding to different viewpoint images of the same sample are constrained to be close to each other in the feature space, and the visual features corresponding to different categories of samples are constrained to be far apart in the feature space. This makes up for the lack of viewpoint consistency constraints in existing methods and enhances the robustness of the model to viewpoint changes.
[0034] The aforementioned bilinear feature enhancement module and supervised contrastive loss function complement each other from the two dimensions of feature representation and feature space distribution, respectively. Combined with the training strategy of joint optimization of cross-entropy loss function and supervised contrastive loss function, the accuracy of multi-view fly identification is significantly improved. At the same time, the bilinear feature enhancement module only includes element-wise squaring operation and concatenation operation on the feature dimension, without introducing additional learnable parameters, and does not significantly increase the model complexity.
[0035] Furthermore, in step 104, for the first... A sample, which contains Images from multiple viewpoints were used. Features were extracted from all images through a backbone network with shared weights, and then global average pooling was applied to obtain a 512-dimensional feature vector. Subsequently, a view-max pooling layer was used for fusion. ; in, Indicates the first The first sample The perspective in the first Characteristic responses on each channel Indicates the first This method utilizes multi-view fusion features from individual samples to retain the most discriminative response information from different perspectives.
[0036] In step 105, the multi-view fusion feature is defined as follows: ; The bilinear enhancement feature is represented as follows: ; Right now: ; in, It is a bilinear enhancement feature. For multi-view fusion features, This represents element-wise multiplication (Hadamard product). Step 105 introduces second-order autocorrelation while preserving the original first-order features to enhance the ability to express local fine-grained differences.
[0037] For fine-grained fly identification tasks, some species are highly similar in overall outline, but there are subtle differences in local areas such as wing vein texture and abdominal markings. Bilinear enhancement can amplify the response differences between different feature channels, thereby improving the ability of the multi-view learning model to distinguish highly similar categories. Step 105 only includes squaring and concatenation operations and does not introduce additional learnable parameters, so it does not significantly increase the complexity of the multi-view learning model.
[0038] To address the issue of inconsistent feature distributions from different perspectives, this paper introduces Supervised Contrastive Loss during the training phase. This loss constrains features from different perspectives of the same sample to be close to each other in the embedding space, while simultaneously widening the distance between different categories.
[0039] Furthermore, the supervised contrastive loss function is calculated within a training batch consisting of multiple samples, and the calculation method is as follows: The visual features of each viewpoint image in the training batch are normalized to obtain normalized visual features. Assume there are a total of The nth sample, the nth Each sample contains Visual features of each viewpoint image are obtained by extracting them from the backbone network and then performing global average pooling. .
[0040] To eliminate the influence of feature scale, L2 normalization is performed on the visual features of each viewpoint image: ; in, Let be the normalized visual features of the i-th sample at the t-th viewpoint. Let be the visual features of the i-th sample at the t-th viewpoint.
[0041] Define the set of positive samples: ; That is, the first All other viewpoints in the sample, except the current viewpoint, are considered positive samples. The supervised contrastive loss function for the t-th viewpoint of the i-th sample in the training batch is: ; in, Let be the supervised contrast loss value for the i-th sample at the t-th viewpoint. The total number of samples in the training batch. Let be the number of viewpoints for the a-th sample in the training batch. Let represent the normalized visual features of the i-th sample at the t-th viewpoint. Let be the set of all viewpoints in the i-th sample except for the t-th viewpoint. Represents a set Normalized visual features of the p-th viewpoint The cosine similarity represents the similarity between different viewpoints of the same sample. This represents the normalized visual feature of the b-th viewpoint of the a-th sample in the training batch. Temperature is a parameter used to control the smoothness of the similarity distribution; This is an indicator function used to exclude the current feature itself. When the value is 1, The value is 0; the numerator represents the similarity between positive sample pairs (different perspectives of the same sample), and the denominator represents the similarity between the t-th perspective and all other perspectives in the batch.
[0042] Finally, the supervised contrastive loss function for the batch is: ; in, This represents the average number of viewing angles in the current batch.
[0043] This supervised contrastive loss function narrows the feature distance between different perspectives of the same sample while increasing the feature margin between samples of different classes, enabling the multi-view learning model to learn a more compact and discriminative multi-view feature space. Compared with traditional methods that rely solely on classification loss, supervised contrastive learning can further enhance the model's robustness to perspective changes and improve the multi-view feature fusion effect.
[0044] Furthermore, the training process of the multi-view learning model includes two stages: In the first stage, a single-view learning model is constructed, which includes the backbone network and a single-view classifier. The single-view learning model is trained with single-view images and corresponding real labels, and optimized using the cross-entropy loss function to obtain the pre-trained backbone network parameters.
[0045] Specifically, the first stage trains a single-view convolutional neural network (SVCNN) to enable the backbone network to learn basic fine-grained visual features, providing a stable initialization for subsequent multi-view fusion. If multi-view joint training is performed directly, the network needs to learn both visual features and view fusion strategies simultaneously, which can easily lead to unstable convergence. Therefore, pre-training with a single view can effectively improve the final performance of the model.
[0046] In a preferred embodiment, the backbone network uses ResNet-18 from the Residual Network (ResNet) as a feature extractor to extract features of size [size missing]. Features are extracted from the input image and output as 512-dimensional visual features after global average pooling. The single-view classifier uses a fully connected classification layer to classify four fly species based on the 512-dimensional visual features. After training, the parameters of the backbone network are saved as initial weights for the second stage.
[0047] The first stage uses the standard cross-entropy loss function for optimization: ; in, The value of the cross-entropy loss function. For batch size, The number of fly species. Let i be the true label of the i-th sample. This represents the probability of the i-th sample predicted by the single-view learning model. After training, the parameters of the single-view learning model are saved as initial weights for the second stage.
[0048] In the second stage, the backbone network of the multi-view learning model is initialized with the pre-trained backbone network parameters, the multi-view learning model is trained with multi-view images and corresponding real labels, and the cross-entropy loss function and the supervised contrastive loss function are used for joint optimization.
[0049] The joint optimization is expressed as: ; in, This is the balancing coefficient. Cross-entropy loss primarily optimizes the classification boundary, while supervised contrastive loss optimizes the feature space distribution structure. The two complement each other and jointly improve model performance.
[0050] The training strategy is set as follows: in the second stage, the pre-trained weights from the first stage are loaded and fine-tuned using a smaller learning rate. The learning rate of the backbone network is set to 0.1 times that of the classifier to preserve the pre-trained features.
[0051] Furthermore, in the second phase of training, the learning rate is adjusted using a cosine annealing strategy.
[0052] The cosine annealing strategy: ; in, Let be the learning rate at the e-th training step. The preset maximum learning rate, Here, e represents the preset minimum learning rate, and e represents the current training step number. This represents the total number of training steps.
[0053] In summary, the complete forward propagation process is as follows: Input images from multiple perspectives: ; Panoramic view: ; Backbone network feature extraction: ; Reshaping into a perspective dimension: ; View max pooling: ; Global average pooling: ; Bilinear enhancement: ; Classified output: ; in, For classifier weights, These are the classifier bias parameters.
[0054] Meanwhile, features from each viewpoint are extracted before max pooling of the view for supervised contrastive learning: ; The final multi-view learning model outputs classification probabilities. Losses compared with supervision This enables fine-grained multi-view recognition of flies.
[0055] This application achieves significant performance improvements on multi-view fly identification tasks by introducing a bilinear feature enhancement module (Bilinear) and supervised contrastive learning (SupCon) loss into the standard multi-view convolutional neural network MVCNN, i.e., this application is MVCNN_BI_SUP, as shown in Table 1 below: Table 1
[0056] It can be seen that the bilinear feature enhancement module and the supervised contrastive loss have limited effects when used alone (the bilinear module alone even leads to a slight decrease in MVA), but when used together, they produce a significant synergistic gain, improving MVA by 1.12%. This indicates that the two modules complement each other from different dimensions: the bilinear module enhances the higher-order statistical representation ability of features, and the supervised contrastive loss enhances the consistency of perspectives. The synergistic effect of the two significantly improves fine-grained recognition performance.
[0057] The beneficial effects achieved by this application at the technical level include: First, by introducing a second-order interaction statistic, namely feature squares, between feature channels, the multi-view learning model can capture fine-grained morphological differences that cannot be expressed by first-order statistics, such as stripe shapes and wing vein branching patterns. This is of great significance for distinguishing closely related fly species with highly similar morphologies. Second, supervised contrastive loss explicitly constrains the feature representations of the same sample from different perspectives to be close to each other in the embedding space, effectively alleviating the feature inconsistency problem caused by perspective differences and significantly improving the robustness of the multi-view learning model. Finally, ablation experiments show that after introducing bilinear enhancement, the multi-view learning model's ability to distinguish morphologically similar closely related species is further enhanced; for example, the accuracy of distinguishing between *Bufo gargarizans* and *Bufo salina* is improved. The synergistic effect of these three improvements comprehensively enhances the performance of the multi-view learning model in fine-grained fly identification.
[0058] like Figure 1 The diagram illustrates the complete technical flow of the multi-view fly identification method based on bilinear enhancement and supervised contrastive learning proposed in this application, which consists of two training phases. Phase 1: A ResNet-18 backbone network is trained using single-view images and optimized using a cross-entropy loss function to obtain pre-trained weights. Phase 2: These pre-trained weights are loaded as initialization parameters for the backbone network, and multiple view images of the same fly specimen are input. All view images are processed by a backbone network with shared weights to extract view-level visual features. These features are then pooled and aggregated through a view max-pooling layer to obtain multi-view fused features, which are input to a bilinear feature enhancement module. Through element-wise squaring and concatenation operations, the 512-dimensional first-order feature vector is expanded into a 1024-dimensional bilinear enhanced feature vector. Finally, the fly species identification result is output through a fully connected classification layer. On the other hand, the view-level visual features extracted from the backbone network are used to calculate the supervised contrastive loss. This loss is weighted and summed with the cross-entropy loss to form the total loss function, jointly optimizing the model parameters.
[0059] like Figure 2 As shown, the bilinear feature enhancement module feeds the input first-order feature vector (512 dimensions) into two branches: the first branch keeps the original first-order features unchanged, and the second branch performs element-wise squaring on the feature vector to obtain a second-order feature vector; then, the two feature vectors are concatenated along their feature dimensions to obtain a 1024-dimensional bilinear enhanced feature vector. The bilinear feature enhancement module only includes element-wise squaring and concatenation operations, without introducing any additional learnable parameters. Its essence is a simplified version of bilinear features (i.e., only taking the diagonal elements of the covariance matrix).
[0060] like Figure 3As shown in the figure, the difference in feature distribution before and after the introduction of supervised contrastive loss is illustrated. It can be seen from the figure that after introducing supervised contrastive loss, the clustering of viewpoint-level visual features of the same class of samples in the embedding space is significantly improved, and the feature intervals between samples of different classes are also clearer. This indicates that supervised contrastive loss effectively enhances the intra-class compactness and inter-class separability of the feature space.
[0061] In summary, this application designs the multi-view feature fusion and training strategy as follows: I. Bilinear Feature Enhancement After obtaining the multi-view fused feature map through the view max pooling layer, this fused feature is directly used as a 512-dimensional first-order feature vector. To enhance the model's ability to capture fine-grained differences, this method introduces a bilinear feature enhancement module. Specifically, the first-order feature vector is concatenated with its element-wise product along the feature dimension, i.e., constructing... After this operation, the feature dimension is expanded to 1024 dimensions, resulting in a bilinear enhanced feature vector. This module only involves square and concatenation operations, without introducing any learnable parameters, yet it can effectively utilize the second-order interaction statistics between feature channels, significantly improving the model's ability to express fine-grained morphological differences such as abdominal markings and wing vein bifurcation, which is particularly beneficial for distinguishing closely related fly species with highly similar morphologies.
[0062] II. Multi-perspective supervision and comparative constraints To address the issue of inconsistent feature distributions caused by differences in lighting, pose, and background across images from different viewpoints, this method introduces supervised contrastive loss during training. For each training batch, the viewpoint-level visual features (512-dimensional vectors extracted by the backbone network and globally averaged) corresponding to images of the same fly specimen from different viewpoints are L2 normalized, and then the supervised contrastive loss is calculated. This loss explicitly narrows the distance between the features of different viewpoints within the same sample in the embedding space, while simultaneously widening the feature distributions of samples from different categories. In this embodiment, the temperature parameter τ of the supervised contrastive loss is set to 0.07, and its balance coefficient λ in the total loss is set to 0.5. Through this constraint, the robustness of the model to viewpoint changes is significantly enhanced, and the multi-view fusion effect is further improved.
[0063] III. Two-stage progressive training To avoid interference from multi-view fusion tasks in the backbone network's learning of basic visual features, this method employs a two-stage training strategy. In the first stage, a backbone network (ResNet-18) and a single-view classifier are trained using single-view images, employing only the cross-entropy loss function to enable the backbone network to master the basic fine-grained features of fly images. The pre-trained weights of the backbone network are saved after training. In the second stage, these pre-trained weights are loaded to initialize the backbone network of the multi-view recognition model, and then joint training is performed using multi-view images, while simultaneously optimizing the cross-entropy loss and supervised contrastive loss. In the second stage, the learning rate of the backbone network is set to 0.1 times that of the classifier layer to protect the feature extraction capabilities obtained through pre-training. Furthermore, the learning rate is dynamically adjusted using a cosine annealing strategy. This two-stage strategy effectively avoids mutual interference between feature learning and fusion strategy learning, ensuring that the model, with good visual representations, focuses on learning the optimal fusion method between multiple views, thereby achieving better recognition performance.
[0064] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0065] This document uses specific examples to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the methods and core ideas of this application. Furthermore, those skilled in the art will recognize that, based on the ideas of this application, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of this application.
Claims
1. A multi-view fly identification method based on bilinear enhancement and supervised contrastive learning, characterized in that, The method includes: A multi-view learning model is constructed, which includes a backbone network, a view pooling layer, a bilinear feature enhancement module, and a classifier. Acquire images of the fly specimen to be identified from multiple perspectives; The multiple viewpoint images are input into the same backbone network, and the visual features of each viewpoint image are extracted. The view pooling layer aggregates the visual features of the multiple view images along the view dimension to obtain multi-view fused features. The multi-view fusion feature is input into the bilinear feature enhancement module, and the multi-view fusion feature is squared element-wise to obtain a second-order feature. The multi-view fusion feature and the second-order feature are then concatenated along the feature dimension to obtain the bilinear enhanced feature. The bilinear enhancement features are input into the classifier to obtain the identification results of the fly specimen; The training process of the multi-view learning model includes: joint optimization using a cross-entropy loss function and a supervised contrastive loss function; the cross-entropy loss function is calculated based on the recognition result and the true label; the supervised contrastive loss function is calculated based on the visual features of each view image, used to constrain the visual features corresponding to different view images of the same sample to be close to each other in the feature space, and to constrain the visual features corresponding to different categories of samples to be far apart in the feature space.
2. The multi-view fly identification method based on bilinear enhancement and supervised contrastive learning according to claim 1, characterized in that, The view pooling layer is a view max pooling layer. The pooling aggregation method of the view max pooling layer is as follows: the visual features of the multiple view images are aligned by dimension, and the maximum value of the feature values of the visual features of different view images in each dimension is taken. The maximum values obtained in each dimension constitute the multi-view fusion feature.
3. The multi-view fly identification method based on bilinear enhancement and supervised contrastive learning according to claim 1, characterized in that, The bilinear enhancement feature is represented as follows: ; in, It is a bilinear enhancement feature. For multi-view fusion features, This indicates element-wise multiplication.
4. The multi-view fly identification method based on bilinear enhancement and supervised contrastive learning according to claim 1, characterized in that, The supervised contrastive loss function is calculated within a training batch consisting of multiple samples, and the calculation method is as follows: The visual features of each viewpoint image in the training batch are normalized to obtain normalized visual features. The supervised contrastive loss function for the i-th sample in the training batch at the t-th viewpoint is: ; in, Let be the supervised contrast loss value for the i-th sample at the t-th viewpoint. The total number of samples in the training batch. Let be the number of viewpoints for the a-th sample in the training batch. Let represent the normalized visual features of the i-th sample at the t-th viewpoint. Let be the set of all viewpoints in the i-th sample except for the t-th viewpoint. Represents a set Normalized visual features of the p-th viewpoint This represents the normalized visual feature of the b-th viewpoint of the a-th sample in the training batch. For temperature parameters, For indicator functions, when When the value is 1, The value is 0 at that time.
5. The multi-view fly identification method based on bilinear enhancement and supervised contrastive learning according to claim 1, characterized in that, The training process of the multi-view learning model includes two stages: In the first stage, a single-view learning model is constructed, which includes the backbone network and a single-view classifier. The single-view learning model is trained with single-view images and corresponding real labels, and optimized using the cross-entropy loss function to obtain the pre-trained backbone network parameters. In the second stage, the backbone network of the multi-view learning model is initialized with the pre-trained backbone network parameters, the multi-view learning model is trained with multi-view images and corresponding real labels, and the cross-entropy loss function and the supervised contrastive loss function are used for joint optimization.
6. The multi-view fly identification method based on bilinear enhancement and supervised contrastive learning according to claim 5, characterized in that, In the second stage of training, the learning rate of the backbone network is set to 0.1 times the learning rate of the classifier.
7. The multi-view fly identification method based on bilinear enhancement and supervised contrastive learning according to claim 5, characterized in that, In the second phase of training, the learning rate is adjusted using a cosine annealing strategy.
8. The multi-view fly identification method based on bilinear enhancement and supervised contrastive learning according to claim 1, characterized in that, The backbone network is a residual neural network.
9. The multi-view fly identification method based on bilinear enhancement and supervised contrastive learning according to claim 1, characterized in that, The fly specimens include at least one of the following species: green bottle fly, copper bottle fly, housefly, and bighead fly.
10. The multi-view fly identification method based on bilinear enhancement and supervised contrastive learning according to claim 1, characterized in that, The multiple perspective images are images of the same fly specimen taken from multiple different angles.