Paleontological fossil image retrieval method combining global self-attention and local attention
By combining the feature extraction network of global self-attention and local attention, the features of paleontological fossil images are extracted and fused, and the problem of insufficient distinction ability of local detail features and texture similar features in small sample paleontological fossil images in the prior art is solved, achieving higher retrieval accuracy and accuracy.
Patent Information
- Application Number
- CN202411898361.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-23
- Publication Date
- 2025-05-27
AI Technical Summary
In paleontological fossil images, the existing image search method focuses on extracting multi-scale image features, and ignores the distinction between local detail features and texture similar features in small sample paleontological fossil images, resulting in lower generalization and accuracy and poor retrieval accuracy.
A feature extraction network combining global self-attention and local attention is adopted to extract features of paleontological fossil images through local feature branches, global feature branches and feature fusion modules, fuse features from multiple stages to enhance description capabilities, and improve retrieval accuracy through weighted similarity calculations.
It improves the accuracy of paleontological fossil image retrieval, reduces the impact of rock background noise, reduces the interference of negative samples on sorting, and improves the accuracy of image retrieval.
Smart Images

Figure CN120047713A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to image retrieval, and more particularly to a method for retrieving paleontological fossil images that combines global self-attention and local attention. Background Art
[0002] Paleontological fossils are the remains and traces of organisms formed and existing in strata during prehistoric geological periods, including fossils of plants, invertebrates, vertebrates, etc. and their trace fossils, providing unique opportunities for scientists to study the evolution of life. Scientists can understand the origin, evolution process of ancient organisms, the main life forms of each era, trace the evolutionary history of different biological groups, and reveal the origin, evolution and extinction events of species through paleontological fossils. With the development of fields such as paleontology, geology and archaeology, more and more paleontological fossil images are recorded, digitized and stored in databases, and these images contain rich information on the evolution of life on Earth.
[0003] When paleontologists conducting fieldwork discover unknown fossils in the wild, they need to identify the content of the unknown fossils and thus need to compare them with similar fossil features to initially infer information such as the morphological characteristics of the unknown fossils, the biodiversity of organisms in the geological history period, and the interrelationships between organisms. However, in the face of a huge number of paleontological fossil images, relying on manual retrieval is not only slow, but also has a large degree of human subjectivity in the retrieval results, and cannot meet the requirements for the real-time and accuracy of retrieving the same or similar pictures.
[0004] With the rise of machine learning technology in recent years, the application of artificial intelligence in various fields has gradually matured and plays an important role in many fields. By using image recognition methods in the field of computer images to identify and retrieve the content of fossil images, it can not only effectively reduce the error rate and subjectivity in the process of fossil image retrieval, but also improve the retrieval speed. Traditional content-based image retrieval (CBIR) refers to retrieving relevant images containing a certain object in the query image from an image library. The similar images obtained by CBIR query are semantically similar to the query image. Through CBIR technology, paleontologists can quickly and accurately retrieve fossil images related to the unknown fossil image and more detailed information descriptions, so as to infer the morphological characteristics and biological information of the unknown fossils faster.
[0005] The classic content-based image retrieval method provided by Video-Google mainly includes feature extraction and nearest neighbor search. After that, many image retrieval systems have been developed based on this idea, and feature extraction and nearest neighbor search have been optimized in different business environments. The key step of these systems is feature extraction. By extracting key features from a given image, then converting these features into a vector of a fixed size, and then using these vectors for fast visual content search. However, the semantic gap between the high-level semantics of images and their low-level visual features is a huge challenge for existing image retrieval.
[0006] In order to make the similarity measure of image representations in space between two images generated by a deep learning model semantically reflect their correlation, in 2018, Liu et al. used an improved SIFT algorithm to extract features from fossil images. The improved SIFT algorithm suppresses the generation of local multiple extreme points during the extreme value detection process and uses the Harris corner detection operator to screen feature points; in 2020, Ross Marchant et al. designed a network based on recurrent CNNs to train on a large foraminifera fossil image set. The features extracted by the CNN network contain more semantic information than SIFT. However, the fossil image data used in these two methods are pure-color background images that only contain the main body of the fossil and cannot be applied to fossil images with complex backgrounds taken on-site; therefore, in 2021, Hou et al. proposed a fossil image retrieval method based on the fusion of salient features and global features to address the lack of salient features in complex fossil images. Although it can effectively improve the retrieval performance, the way of fusing salient features and global features is a channel splicing method, and there is still a lot of background noise in the global features extracted by a general classification network.
[0007] In summary, although existing deep learning methods have played an important role under data-driven conditions, in the case of limited data or low data quality, such as in the field of paleontology, since existing image retrieval methods focus on extracting multi-scale image features, cascading shallow and deep image features, or obtaining spatial and channel semantic information, ignoring the distinction between local detail features and texture similarity features in small-sample paleontological fossil images, the generalization and accuracy are relatively low, and the precision of the retrieval results calculated using similarity is not high, and the retrieval accuracy is poor. Summary of the Invention
[0008] Aiming at the deficiencies of the existing technology, the purpose of the present invention is to provide a paleontological fossil image retrieval method that combines global self-attention and local attention, which improves the retrieval accuracy.
[0009] To achieve the above object, the present invention adopts the following technical solutions to implement:
[0010] A method for retrieving paleontological fossil images combining global self-attention and local attention, comprising the following steps:
[0011] Step 1, construct a paleontological fossil image dataset, preprocess the paleontological fossil image dataset, and divide it into a training set, a test set, and a validation set;
[0012] Step 2, construct a feature extraction network combining global self-attention and local attention, which includes four stages, and each stage includes a local feature branch LFEB, a global feature branch GFEB, and a feature fusion module FFB, where: the processing process of the local feature branch LFEB for the input feature map l i-1 is expressed as:
[0013] l i = f 1×1 (f 3×3 (LN(f 1×1 (l i-1 )))) + l i-1
[0014] In the formula, f 3×3 ( ) represents a 3×3 convolution operation, f 1×1 ( ) represents a 1×1 convolution operation, LN( ) represents layer normalization processing, and l i is the output feature map of the local feature branch LFEB;
[0015] The processing process of the global feature branch GFEB for the input feature map g i-1 is expressed as:
[0016] g i = f 1×1 (f 1×1 (MHSA(LN(G i-1 )))) + g i-1
[0017] In the formula, MHSA( ) represents the multi-head self-attention mechanism, and g i represents the output feature map of the global feature branch GFEB;
[0018] The processing process of the feature fusion module FFB for the input feature map is expressed as:
[0019] F i = IRMLP(Concat[G i , L i , Fi i ″]) + F i ′
[0020] F i = f3×3 (Concat[g i ,l i ,F i ′])
[0021] G i =CA(g i )
[0022] L i =SA(l i )
[0023] F i ′=Avgpool(f 1×1 (F i-1 ))
[0024] In the formula, G i represents the output feature map of the channel attention mechanism CA respectively, L i represents the output feature map of the spatial attention mechanism SA,
[0025] F i-1 represents the output feature map of the feature fusion module FFB in the previous stage, Avgpool( ) represents the global average pooling operation, and after performing the average pooling operation on F i-1 , F i ′ is obtained. The feature map g i output by the global feature extraction branch GFEB, the feature map l i output by the local feature extraction branch LFEB, and F i ′ are aggregated to obtain F i . IRMLP( ) represents the aggregation of G i , L i , and F i by the inverse residual module IRMLP, and is expressed by the formula as:
[0026] IRMLP(x) = f 1×1 (f 1×1 (f 3×3 (LN(x)) + LN(x)))
[0027] In the formula, x represents the input feature map of the inverse residual module IRMLP;
[0028] The feature map output by the feature fusion module FFB in the last stage is successively passed through the global average pooling layer and layer normalization, and then input into the linear classifier for classification;
[0029] Step 3: First, initialize the parameters of the feature extraction network that combines global self-attention and local attention, and then input the training set into the feature extraction network that combines global self-attention and local attention for training to obtain a trained network model;
[0030] Step 4: Use the trained model to extract the feature vectors of the paleontological fossil images and the feature vector of the fossil image p to be queried;
[0031] Step 5: Generate a list of paleontological fossil images closest to the fossil image p to be queried by calculating the weighted similarity between the feature vector of the fossil image p to be queried and the feature vectors of the paleontological fossil images in the image database, which is the retrieval result.
[0032] Further, the process of constructing the paleontological fossil image dataset in Step 1 is as follows: Collect paleontological fossil images of different species, add labels to the paleontological fossil images in combination with prior knowledge to obtain the paleontological fossil image dataset.
[0033] Further, the preprocessing of the paleontological fossil image dataset in Step 1 refers to cropping, randomly rotating by 90°, horizontally flipping, vertically flipping, enhancing the brightness and contrast, adjusting the random gray coefficient, and filtering the paleontological fossil images.
[0034] Further, the processing processes of the channel attention CA and the spatial attention SA on the input feature map in Step 2 are respectively expressed as:
[0035] CA(x) = σ(MLP(AvgPool(x)) + MLP(MaxPool(x)))
[0036] SA(x) = σ(f 7×7 (Concat[AvgPool(x), MaxPool(x)]))
[0037] In the formula, σ is the Sigmoid function, Avgpool( ) represents the global average pooling operation, MaxPool( ) represents the maximum pooling operation, MLP( ) represents the multi-layer perception mechanism, Concat( ) represents feature splicing, and f 7×7 ( ) represents the 7×7 convolution operation.
[0038] Further, the process of initializing the parameters of the feature extraction network combining global attention and local self-attention in Step 3 is as follows: Set the weight decay to 1e- 2 , set the initial learning rate to 1e- 4 , set the number of iterations to 100, set the batch size to 16, set the size of the input image to 224×224, and use the Adam optimizer and the gradient accumulation strategy.
[0039] Further, the specific process of Step 5 is as follows:
[0040] Step 5.1: Calculate the Euclidean distance between the fossil image p to be queried and the paleontological fossil image q in the image databasei The original distance d(p, d i ), sort them in ascending order of distance, and take the top k sorting results as the initial query results. Among them, the original distance d(p, q i ) is expressed as:
[0041]
[0042] In the formula, x p and x q respectively represent the feature vectors of the fossil image p to be queried and the paleontological fossil image q in the image database i ;
[0043] Step 5.2: First, define k paleontological fossil images adjacent to the fossil image p to be queried in the image database, which is expressed as:
[0044]
[0045] Then, the set of k mutually nearest neighbor images is expressed as:
[0046] R(p, k) = {q i |(q i ∈ N(p, k)) ∩ (p ∈ N(q i , k))}
[0047] In the formula, N(q i , k) represents the k nearest neighbor images of the paleontological fossil image q i in the image database;
[0048] Step 5.3: Calculate the Jaccard distance d J (p, q i ), which is expressed as:
[0049]
[0050] Step 5.4: Encode the k mutually nearest neighbor images into a vector V p , which is expressed as:
[0051]
[0052] In the formula, is defined as a binary function, which is expressed as:
[0053]
[0054] Step 5.5: Through the Gaussian kernel of the pairwise distance, is redefined as:
[0055]
[0056] Step 5.6: Weight the original distance and the Jaccard distance to obtain the final distance d * , which is the weighted similarity and is expressed as:
[0057] d * (p, q i ) = (1 - λ)d J (p, q i ) + λd(p, q i )
[0058] In the formula, λ ∈ [0, 1], representing the penalty factor;
[0059] Step 5.7: Generate a list of paleontological fossil images closest to the fossil image p to be queried according to the final distance d * , which is the retrieval result.
[0060] Compared with the prior art, the present invention has the following technical effects:
[0061] (1) The present invention uses the feature fusion module FFB to fuse the global features and local features extracted by the feature extraction network encoder at multiple stages, and uses them as the description features of the final image. This description feature is used as the final feature for fossil image retrieval, strengthening the feature description of the main part and main details in the fossil image, reducing the influence of noise such as rock backgrounds, and not only solving the problem of low retrieval accuracy caused by feature extraction and similarity matching in the prior feature extraction model for paleontological fossil image retrieval tasks.
[0062] (2) The present invention calculates the Euclidean distance, obtains the initial sorted list according to the image similarity, and on this basis, defines the set of k mutually nearest neighbor images for each image. If there are more repeated samples in the set, it means that the two images are more similar. By encoding the set of k mutually nearest neighbor images and reallocating weights and calculating the Jaccard distance according to the original distance between the predicted image and the k-nearest neighbor images, and then weighting the original distance and the Jaccard distance to correct the initial sorted list, it can reduce the interference of negative samples on sorting and improve the accuracy of image retrieval. BRIEF DESCRIPTION OF THE DRAWINGS
[0063] Figure 1 : Schematic diagram of the overall structure of the feature extraction network combining global self-attention and local attention of the present invention;
[0064] Figure 2 : Schematic diagram of the structure of the feature fusion module FFB of the present invention;
[0065] Figure 3 : Flowchart for solving the weighted similarity of the present invention;
[0066] Figure 4 : Some examples of the paleontological fossil image dataset of the present invention;
[0067] Figure 5 : The PR curves of the feature extraction network combining global self-attention and local attention of the present invention for images of various species;
[0068] Figure 6 : The ROC curves of the feature extraction network combining global self-attention and local attention of the present invention for images of various species. Detailed implementation manners
[0069] The following further elaborates on the specific content of the present invention in conjunction with embodiments.
[0070] A paleontological fossil image retrieval method combining global self-attention and local attention includes the following steps:
[0071] Step 1: Construct a paleontological fossil image dataset, preprocess the paleontological fossil image dataset, and divide it into a training set, a test set, and a validation set. The process is as follows:
[0072] Step 1.1: The process of constructing the paleontological fossil image dataset is as follows: Collect paleontological fossil images of different species, add labels to the paleontological fossil images in combination with the prior knowledge of researchers to obtain the paleontological fossil image dataset. Some images in the paleontological fossil image dataset are as Figure 4 shown;
[0073] Step 1.2: Preprocessing the paleontological fossil image dataset means cropping, randomly rotating by 90°, horizontally flipping, vertically flipping, enhancing the brightness and contrast, adjusting the random gray coefficient, and filtering the paleontological fossil images, so as to obtain standardized input image data adapted to the feature extraction network model, which helps to improve the training efficiency and effect of the network model;
[0074] Step 1.3: Divide the preprocessed paleontological fossil image dataset into a training set, a validation set, and a test set, which are used to batch import into the feature extraction network model for training, testing, and validating the feature extraction network model;
[0075] Step 2: Construct as Figure 1The feature extraction network that combines global self-attention and local attention as shown includes four stages, named S1, S2, S3, and S4 in sequence. Each stage includes a local feature extraction branch (LFEB), a global feature extraction branch (GFEB), and a feature fusion module (FFB). The parameter dimensions of the input feature maps in each stage are different to obtain semantic information at different scales, so as to extract features at different scales. The parameter dimensions of the feature maps input to the four stages S1, S2, S3, and S4 are shown in Table 1;
[0076] The local feature extraction branch (LFEB) is based on the CNN network and is used to efficiently extract the local features of paleontological fossil images. It includes a 1×1 convolutional layer, a 3×3 convolutional layer, batch normalization processing, and a ReLU activation function. The processing of the input feature map l by the local feature extraction branch (LFEB) is expressed as: i-1 as follows:
[0077] l i = f 1×1 (f 3×3 (LN(f 1×1 (l i-1 )))) + l i-1
[0078] In the formula, f 3×3 ( ) represents the 3×3 convolution operation, f 1×1 ( ) represents the 1×1 convolution operation, LN( ) represents layer normalization processing, and l i is the output feature map of the local feature extraction branch (LFEB);
[0079] The global feature extraction branch (GFEB) is used to extract the global features of paleontological fossil images. Based on the transformer, a multi-head self-attention mechanism (MHSA) is introduced as the encoder, which enhances the perception ability of the local details of the fossil images. Thus, by segmenting the sub-semantic space, the model can focus on information in different dimensions, thereby improving the expression ability and attention distribution of the feature extraction network model and enabling it to obtain more fine-grained image features. The processing of the input feature map g by the global feature extraction branch (GFEB) is expressed as: i-1 as follows:
[0080] g i = f 1×1 (f 1×1 (MHSA(LN(G i-1 )))) + g ii-1
[0081] In the formula, MHSA( ) represents the multi-head self-attention mechanism, and g i represents the output feature map of the global feature extraction branch (GFEB);
[0082] Table 1 Parameter dimensions of the feature maps input in each stage
[0083]
[0084] The structure of the feature fusion module FFB is as follows Figure 2 shown. It adaptively fuses local features, global features at different levels, and the feature map output by the feature fusion module FFB of the previous stage according to the input feature map. During this period, by inputting the output feature map l i of the local feature branch LFEB into the spatial attention mechanism SA to enhance local details and suppress irrelevant regions, and inputting the output feature map g i of the global feature branch GFEB into the channel attention mechanism CA; then, connecting the feature maps L i and G i output by the spatial attention mechanism SA and the channel attention mechanism CA, and the feature map F i output by the feature fusion module FFB of the previous stage to the inverse residual module IRMLP, and aggregating L i , G i and F i through IRMLP, effectively preventing the problems of gradient disappearance, explosion and network degradation, and thus effectively capturing global and local feature information at each level;
[0085] The processing process of the feature fusion module FFB for the input feature map is expressed as:
[0086] F i = IRMLP(Concat[G i , L i , F i ″](+ F i ′
[0087] F i = f 3×3 (Concat[g i , l i , F i ′])
[0088] G i = CA(g i )
[0089] L i = S(l i )
[0090] F i ′ = Avgpool(f 1×1 (F i-1 ))
[0091] where Gi respectively represent the output feature maps of the channel attention CA, L i represents the output feature map of the spatial attention SA,
[0092] F i-1 represents the output feature map of the feature fusion module FFB in the previous stage. Avgpook() represents the global average pooling operation. After performing the average pooling operation on F i-1 we get F i ′. The feature map g i output by the global feature branch GFEB, the feature map l i output by the local feature branch LFEB, and i F i ′ are aggregated to obtain F i 、L i and F i ′. IRMLP() represents the aggregation of G
[0093] LRMLP(x) = f 1×1 (f 1×1 (f 3×3 (LN(x)) + LN(x)))
[0094] In the formula, x represents the input feature map of the inverse residual module IRMLP;
[0095] The processing processes of the channel attention CA and the spatial attention SA on the input feature map are respectively expressed as:
[0096] CA(x) = σ(MLP(AvgPool(x)) + MLP(MaxPool(x)))
[0097] SA(x) = σ(f 7×7 (Concat[AvgPool(x), MaxPool(x)]))
[0098] In the formula, σ is the Sigmoid function, Avgpool() represents the global average pooling operation, MaxPool() represents the maximum pooling operation, MLP() represents the multi-layer perception mechanism, Concat() represents the feature concatenation, and f 7×7 () represents the 7×7 convolution operation;
[0099] The feature map output by the feature fusion module FFB in the last stage S4 passes through the global average pooling layer Avgpool and the layer normalization Layer Norm in sequence, and then is input into the linear classifier Linear for classification. The output feature vector is used as the final description feature of the image;
[0100] The paleontological fossil image and the fossil image p to be queried respectively pass through the global feature extraction branch GFEB and the local branch LFEB. A set of output feature vectors is obtained at each stage. There will be losses in the features extracted layer by layer at each stage. The feature fusion block FEB is used to fuse at each layer to reduce the losses. The feature fusion block FEB includes the channel attention CA, the spatial attention SA and the inverted residual module IRMLP, which adaptively fuse the semantic information between the features of different scales of each branch;
[0101] Step 3: First, initialize the parameters of the feature extraction network that combines global attention and local self-attention, and then input the training set into the feature extraction network that combines global attention and local self-attention for training to obtain a trained network model;
[0102] The process of initializing the parameters of the feature extraction network that combines global attention and local self-attention is as follows: Set the weight decay to 1e- 2 , set the initial learning rate to 1e- 4 , set the number of iterations to 100, set the batch size to 16, set the size of the input image to 224×224, use the Adam optimizer, and adopt the gradient accumulation strategy;
[0103] Step 4: Use the trained model to extract the feature vector of the paleontological fossil image and the feature vector of the fossil image p to be queried, where: the fossil image p to be queried is selected from the test set;
[0104] Step 5: As Figure 3 shown, by calculating the weighted similarity between the feature vector of the fossil image p to be queried and the feature vector of the paleontological fossil image in the image database, a list of the paleontological fossil images closest to the fossil image p to be queried is generated, which is the retrieval result. The specific process is as follows:
[0105] Step 5.1: Calculate the original distance d(p, q i ) between the fossil image p to be queried and the paleontological fossil image q in the image database by the Euclidean distance, sort them in ascending order of distance, and take the top k sorting results as the initial query results, that is, the initial sorted list; where, the original distance d(p, q i ) is expressed as: i )
[0106]
[0107] In the formula, x p and x q respectively represent the feature vectors of the fossil image p to be queried and the paleontological fossil image q in the image database; i
[0108] Step 5.2. First, define k paleontological fossil images adjacent to the fossil image p to be queried in the image database, denoted as:
[0109]
[0110] Then, the set of k mutually nearest neighbor images is denoted as:
[0111] R(p, k) = {q i | (q i ∈ N(p, k)) ∩ (p ∈ N(q i , k))}
[0112] In the formula, N(q i , k) represents the k nearest neighbor images of the paleontological fossil image q i in the image database;
[0113] If two images in the image database are similar, then there is an intersection in their corresponding sets of k mutually nearest neighbors, that is, there are some pictures that appear in both sets. If the proportion of the number of intersections is large enough, then the two pictures are very similar. Therefore, the Jaccard metric is used to calculate the similarity between the sets of k mutually nearest neighbors;
[0114] Step 5.3. Calculate the Jaccard distance d J (p, p i ), denoted as:
[0115]
[0116] Step 5.4. Encode the k mutually nearest neighbor images into a vector V p , denoted as:
[0117]
[0118] In the formula, is defined as a binary function, denoted as:
[0119]
[0120] Since the Jaccard distance calculation assigns equal weights to all neighbors, the resulting neighbor set is not discriminative. Generally speaking, the closest neighbor image should be more similar to the image to be queried. Therefore, the weights need to be recalculated based on the original distance. Closer neighbors are assigned larger weights, while farther neighbors are assigned smaller weights. Therefore, it is necessary to redefine
[0121] Step 5.5. Through the Gaussian kernel of pairwise distances, The redefinition formula is:
[0122]
[0123] Step 5.6: To fully reflect the importance of the original distance in re - sorting, the original distance and the Jaccard distance are weighted to modify the initial sorted list to obtain the final distance d * , which is the weighted similarity and is expressed as:
[0124] d * (p, q i )=(1 - λ)d J (p, q i )+λd(p, q i )
[0125] In the formula, λ ∈ [0, 1] represents the penalty factor. When λ = 0, only the Jaccard distance is calculated. When λ = 1, only the original distance is calculated;
[0126] Step 5.7: According to the final distance d * , generate a list of paleontological fossil images closest to the fossil image p to be queried, which is the final retrieval result.
[0127] To verify the effectiveness of the feature extraction network proposed in this embodiment, the platform is a computer with GPU as NVIDIA GeForce RTX 3080. The experimental parameters are set to be optimized using the Adam optimizer, with the weight decay set to 1e - 2 , the initial learning rate set to 1e - 4 , the number of iterations set to 100, the batch size set to 16, the size of the input image set to 224×224, and the gradient accumulation strategy is used to achieve the effect of large batches. The performance of the retrieval method in this embodiment on the test set is evaluated using Precision and Recall. The results are as Figure 5 and Figure 6 shown. It can be seen that the curves of different categories in the ROC curve tend to the upper left, and the curves of different categories in the P - R curve tend to the upper right, indicating that the classification performance of the feature extraction network model in this embodiment is good, and the effectiveness and accuracy of feature extraction are high.
[0128] To verify the effectiveness and advantages of the proposed method for paleontological fossil image retrieval that combines global self-attention and local attention in paleontological fossil image retrieval, retrieval experiments were conducted on a self-built paleontological fossil image dataset to compare it with existing methods, including R-MAC, NetVLAD, GCCL, DPSH, DHN, DSHSD, CSQ, and CSCE. The evaluation criteria used were mAP, F1, and Topk. As shown in Table 2, it can be seen that the proposed image retrieval method in this embodiment has higher retrieval efficiency and retrieval accuracy.
[0129] Table 2 Retrieval performance of different methods on the self-built paleontological fossil image dataset
[0130]
[0131]
Claims
1. A paleontological fossil image retrieval method combining global self-attention and local attention, characterized in that: The steps include: Step 1: construct a paleontological fossil image dataset, preprocess the paleontological fossil image dataset, and divide it into a training set, a test set, and a validation set; Step 2: Construct a feature extraction network combining global self-attention and local attention, which includes four stages, each of which includes a local feature branch LFEB, a global feature branch GFEB and a feature fusion module FFB, wherein: the local feature branch LFEB performs a self-attention on the input feature map l i-1 The processing process is expressed as: l i =f 1×1 (f 3×3 (LN(f 1×1 (l i-1 ))))+l i-1 In the formula, f 3×3 ( ) represents a 3×3 convolution operation, f 1×1 ( ) represents 1×1 convolution operation, LN( ) represents layer normalization processing, l i It is the output feature map of the local feature branch LFEB; The global feature branch GFEB inputs the feature map g i-1 The processing process is expressed as: g i =f 1×1 (f 1×1 (MHSA(LN(G i-1 ))))+g i-1 In the formula, MHSA( ) represents the multi-head self-attention mechanism, g i Represents the output feature map of the global feature branch GFEB; The processing process of the feature fusion module FFB on the input feature map is expressed as: F i =IRMLP(Concat[G i ,L i ,F i ″])+F i ′ F i ″=f 3×3 (Concat[g i ,l i ,F i ′]) G i =CA(g i ) IT i =SA(l i ) F i ′=Avgpool(f 1×1 (F i-1 )) In the formula, G i They represent the output feature maps of the channel attention mechanism CA, L i represents the output feature map of the spatial attention mechanism SA, F i-1 represents the output feature map of the feature fusion module FFB in the previous stage, Avgpool( ) represents the global average pooling operation, and F i-1 After the average pooling operation, we get F i ′, the feature map g output by the global feature branch GFEB i , the feature map l output by the local feature branch LFEB i and F i 'After polymerization, F i ″, IRMLP( ) represents the inverse residual module IRMLP on G i , L i and F i The aggregation is expressed as: IRMLP(x)=f 1×1 (f 1×1 (f 3×3 (LN(x))+LN(x))) Where x represents the input feature map of the reverse residual module IRMLP; In the last stage, the feature map output by the feature fusion module FFB is sequentially passed through the global average pooling layer and layer normalization, and then input into the linear classifier for classification; Step 3: Initialize the parameters of the feature extraction network combining global self-attention and local attention, and then input the training set into the feature extraction network combining global self-attention and local attention for training to obtain a trained network model; Step 4: Use the trained model to extract the feature vector of the paleontological fossil image and the feature vector of the fossil image p to be queried; Step 5: By calculating the weighted similarity between the feature vector of the fossil image p to be queried and the feature vector of the paleontological fossil image in the image database, a list of paleontological fossil images closest to the fossil image p to be queried is generated, which is the retrieval result.
2. The paleontological fossil image retrieval method combining global self-attention and local attention according to claim 1, characterized in that: The process of constructing the paleontological fossil image dataset in step 1 is: collecting paleontological fossil images of different species, adding labels to the paleontological fossil images in combination with prior knowledge, and obtaining the paleontological fossil image dataset.
3. The paleontological fossil image retrieval method combining global self-attention and local attention according to claim 1, characterized in that: The preprocessing of the paleontological fossil image data set in step 1 refers to cropping, 90° random rotation, horizontal flipping, vertical flipping, enhancing light and dark contrast, adjusting random grayscale coefficient and filtering the paleontological fossil images.
4. The paleontological fossil image retrieval method combining global self-attention and local attention according to claim 1, characterized in that: The processing process of the channel attention CA and the spatial attention SA on the input feature map in step 2 is respectively expressed as: CA(x)=σ(MLP(AvgPool(x))+MLP(MaxPool(x))) SA(x)=σ(f 7×7 (Concat[AvgPool(x),MaxPool(x)])) Where σ is the Sigmoid function, Avgpool( ) represents the global average pooling operation, MaxPool( ) represents the maximum pooling operation, MLP( ) represents the multi-layer perception mechanism, Concat( ) represents feature concatenation, and f 7×7 ( ) represents a 7×7 convolution operation.
5. The paleontological fossil image retrieval method combining global self-attention and local attention according to claim 1, characterized in that: The process of initializing the parameters of the feature extraction network combining global attention and local self-attention in step 3 is as follows: the weight decay is set to 1e -2 , the initial learning rate is set to 1e -4 , the number of iterations is set to 100, the batch size is set to 16, the size of the input image is set to 224×224, and the Adam optimizer and gradient accumulation strategy are used.
6. The paleontological fossil image retrieval method combining global self-attention and local attention according to claim 1, characterized in that: The specific process of step 5 is as follows: Step 5.1: Calculate the Euclidean distance between the fossil image p to be queried and the paleontological fossil image q in the image database i The original distance d(p,q i ), sort them from small to large according to the distance, and take the first k sorted results as the initial query results, where the original distance d(p,q i ) is expressed as: In the formula, x p and x q Represent the fossil image to be queried p and the paleontological fossil image q in the image database respectively i The eigenvector of Step 5.2: First define k paleontological fossil images in the image database that are adjacent to the fossil image p to be queried, expressed as: Then, the set of k nearest neighbor images is expressed as: R(p,k)={q i |(q i ∈N(p,j))∩(p∈N(q i ,k))} In the formula, N(q i ,k) represents the paleontological fossil image q in the image database i k nearest neighbor images of; Step 5.3: Calculate the Jaccard distance d between k nearest neighbor image sets J (p,q i ), expressed as: Step 5.4: Encode k nearest neighbor images into a vector V p , expressed as: In the formula, It is defined as a binary function, expressed as: Step 5.5: Use the Gaussian kernel of pairwise distance to transform The redefinition formula is: Step 5.6: Weight the original distance and Jaccard distance to get the final distance d * , which is the weighted similarity, expressed as: d * (p,q i )=(1-λ)d J (p,q i )+λd(p,q i ) In the formula, λ∈[0,1] represents the penalty factor; Step 5.7: According to the final distance d * , generate a list of paleontological fossil images that are closest to the query fossil image p, which is the retrieval result.