Multi-scale Fusion Landmark Image Retrieval Method and System with Feature Consistency Suggestions

By constructing a multi-scale fusion landmark image retrieval network with feature consistency suggestions, the problem of low retrieval accuracy caused by changes in landmark image scale under different shooting conditions is solved, the dependence on fine-grained label information is reduced, the accuracy of landmark image matching is improved, and the practical application in the field of smart tourism is promoted.

CN114579794BActive Publication Date: 2025-06-10XI'AN UNIVERSITY OF ARCHITECTURE AND TECHNOLOGY
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202210334948.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-31
Publication Date
2025-06-10
Estimated Expiration
2042-03-31

AI Technical Summary

Technical Problem

When the existing landmark image retrieval technology processes the changes in landmark information scale under different camera equipment shooting conditions, the search accuracy is low and relies on a large amount of fine-grained label information, which limits practical applications.

Method used

A multi-scale fusion landmark image retrieval method with feature consistency suggestions is proposed. By constructing a multi-scale fusion landmark image retrieval network with feature consistency suggestions, including ResNet50 network, feature self-attention fusion network and regional feature consistency suggestions, it reduces dependence on fine-grained label information and improves retrieval accuracy.

Benefits of technology

It effectively solves the problem of low retrieval accuracy caused by changes in landmark image scales under different shooting conditions, reduces the dependence on manual annotation, improves the accuracy of landmark image matching, and promotes the practical application in the field of smart tourism.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114579794B_ABST
    Figure CN114579794B_ABST
Patent Text Reader

Abstract

The present invention discloses a multi-scale fusion landmark image retrieval method and system for feature consistency recommendation, which collects landmark image data and constructs a landmark retrieval training data set T r and a test data set T e ; constructs a multi-scale fusion landmark image retrieval network for feature consistency recommendation; calculates a loss value by constructing a total loss function, and uses the landmark retrieval training data set T r to train the multi-scale landmark image retrieval network to obtain a multi-scale fusion landmark image retrieval model for feature consistency recommendation; inputs the test data set T e into the multi-scale fusion landmark image retrieval model for feature consistency recommendation and outputs the retrieval result of the landmark image. The present invention solves the problem of low retrieval accuracy caused by scale differences under different shooting conditions, reduces the dependence on a large amount of fine-grained label information, improves the matching accuracy of landmark images, and is beneficial to the actual application deployment in the field of intelligent tourism.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image retrieval, and more specifically, to a multi-scale fusion landmark image retrieval method and system with feature consistency recommendation. Background Art

[0002] With the rapid development of social networking sites, communication and multimedia technologies, digital image devices, etc., the use of digital images involves various aspects such as national defense and military, medical and health, public entertainment, and family life. Data such as images and videos are growing at an alarming rate every day. For these massive pictures containing rich visual information, how to conveniently, quickly, and accurately query and retrieve the images required or interested by users in these vast image libraries has become a research hotspot in the field of multimedia information retrieval. Landmark image retrieval refers to finding images containing the same landmark building instances from database images, which can achieve intuitive geographical exploration and navigation of local landmarks, further provide route optimization and recommendations for similar tourist attractions, and has important application value in the field of smart tourism.

[0003] Currently, driven by the excellent performance of neural networks, landmark image retrieval technology has achieved excellent results in dealing with problems such as illumination changes and shooting angle changes. However, when actually applied to Internet recommendation systems, due to the different shooting distances between different camera devices, the landmark information captured by the camera will have serious scale changes in the image.

[0004] To address the above problems, the commonly used solution is to extract discriminative and representative fixed local visual features of all landmark buildings to solve the scale difference problem of landmark images caused by different shooting conditions, thereby improving the accuracy of landmark image retrieval. However, the above method relies heavily on additional landmark annotation information, such as central building, top, window information, etc., and a large amount of manpower is required to make additional label information for the landmark data set, which greatly limits the practical application of the landmark retrieval method. Summary of the Invention

[0005] In order to solve the problems existing in the prior art, the present invention provides a multi-scale fusion landmark image retrieval method with feature consistency recommendation, which solves the problem of low retrieval accuracy caused by scale differences under different shooting conditions, reduces the dependence on a large amount of fine-grained label information, improves the matching accuracy of landmark images, and is conducive to the actual application deployment in the field of smart tourism.

[0006] To achieve the above object, the present invention provides the following technical solution: A multi-scale fusion landmark image retrieval method with feature consistency recommendation, and the specific steps are as follows:

[0007] S1 Collect landmark image data and construct a landmark retrieval training data set T rWith the test dataset T e ;

[0008] S2 constructs a multi-scale fusion landmark image retrieval network with feature consistency suggestions, including a ResNet50 network connected with a multi-scale information extraction module, a feature self-attention fusion network, and a regional feature consistency suggestion item;

[0009] S3 constructs a total loss function through a feature consistency suggestion function, a triplet loss function, and a classification function, calculates the loss value, and uses the landmark retrieval training dataset T r Train the multi-scale landmark image retrieval network to obtain a multi-scale fusion landmark image retrieval model;

[0010] S4 uses the landmark retrieval test dataset T e Input the multi-scale fusion landmark image retrieval model with feature consistency suggestions, and output the retrieval results of the landmark images.

[0011] Furthermore, in step S1, the method of manual annotation is used to label the corresponding categories of the same kind of landmarks in the landmark images as prefixes, and an independent number is assigned after the category. Among them, the category prefixes between different landmarks are different, and the numbers of the same landmark are different.

[0012] Furthermore, in step S2, a multi-scale information extraction module is connected after the maximum pooling layer of the ResNet50 network. The ResNet50 network is used to obtain the initial local feature map of the landmark image; the multi-scale information extraction module extracts multiple local feature blocks of the initial local feature map in the order from the upper left to the lower right through a tensor recombination function, and obtains N local feature blocks f i N and M local feature blocks f i M .

[0013] Furthermore, in step S2, a feature self-attention fusion network is constructed after ResNet50. The feature self-attention fusion network includes two feature self-attention fusion branches, and each of the two feature self-attention fusion branches is composed of a single layer of Transformer encoding layer.

[0014] Furthermore, in step S2, the specific steps of the processing process of the feature self-attention fusion network are as follows:

[0015] 1) Two Transformer encoding layers are respectively initialized to generate initial global feature maps C′ 0 、C″ 0 , and the local feature blocks f i N 、f i M and the initial global feature map C′0 , C″ 0 are grouped in pairs to obtain f i N and C′ 0 groups, f i M and C″ 0 groups; f i N and C′ 0 groups, f i M and C″ 0 are respectively input into the Transformer encoding layer, and in the Transformer encoding layer, the standard learnable position vector E pos is embedded into f i N and C′ 0 groups, f i M and C″ 0 groups to obtain two preliminarily fused landmark global feature maps F i 1 , F i 2 ;

[0016] 2) Use the result vector sequences z′ 0 and z" 0 to respectively represent the specific information of the two landmark global feature maps F i 1 , F i 2 . The result vector sequences z′ 0 and z" 0 are respectively input into two Transformer encoding layers to perform self-attention learning on the important information in the two landmark global feature maps F i 1 , F i 2 to obtain the weights of each part in the result vector sequences z′ 0 and z" 0 , obtain the distribution probability of the weights, and update the weights of the specific information of the landmark global feature maps F i 1 , F i 2 ;

[0017] 3) Concatenate the two landmark global feature maps F i 1 , F i 2 to obtain the joint global feature map F i 3 .

[0018] Further, in step S2, in the regional feature consistency proposal item, the global feature map F generated by the feature self-attention fusion branch is obtained by constructing a feature consistency proposal function i 1 , F i 2 respectively focuses on the same regions of landmark buildings with different category prefixes. Specifically, the feature consistency proposal function is as follows:

[0019]

[0020] In the formula, represents the Euclidean norm, represents the landmark global feature map F generated by the feature self-attention fusion branch i 1 , F i 2 , K = 2, c k is the cluster center vector.

[0021] Further, in step S2, the cluster center vector c k is randomly initialized according to the content learned from the global feature map and is updated by moving its average value:

[0022]

[0023] where α controls the update rate of c k , represents the landmark global feature map F generated by the feature self-attention fusion branch i 1 , F i 2 .

[0024] Further, in step S3, the total loss function is:

[0025]

[0026] where L is the feature consistency proposal function, is the classification loss function, is the triplet loss function. Specifically:

[0027] The classification loss function designs a batch normalization layer BN(), a linear layer W, and a Softmax layer after any global feature map , specifically:

[0028]

[0029] The triplet loss function is used to enhance any global feature map The discriminability is specifically as follows:

[0030]

[0031] In the formula, F represents the global feature map, k = 1, 2, 3 represents different global feature maps, A is the Anchor representing the sample itself, N is the negative representing the sample of a different class from A, and P is the Positive representing the sample of the same class as A. respectively represent the feature vectors of the source sample, negative sample, and positive sample that make up the triplet. and respectively represent the Euclidean distances of the positive sample pair and the negative sample pair, and m represents the margin threshold of the triplet loss. + represents taking the positive value.

[0032] Furthermore, in step S4, the landmark retrieval test dataset T e is input into the landmark image retrieval model to obtain the joint global feature map F of the test landmark images. i 3 , and the similarity of the global feature maps of the landmark images in the landmark retrieval test dataset T e is calculated through the cosine distance function, and the image retrieval results are sorted and output according to the similarity. The cosine distance function is specifically as follows:

[0033]

[0034] In the formula, and where F represents the global feature map, and j 1 and j 2 represent any one of the test samples and non-test samples in the landmark retrieval test dataset T e , and |||| represents the modulus.

[0035] The present invention also provides a multi-scale fusion landmark image retrieval system with feature consistency suggestions, including:

[0036] A data acquisition module for collecting landmark image data to construct a landmark retrieval training dataset T r and a test dataset T e ;

[0037] A network construction module for a multi-scale fusion landmark image retrieval network with feature consistency suggestions, including a ResNet50 network connected with a multi-scale information extraction module, a feature self-attention fusion network, and a regional feature consistency suggestion item;

[0038] A network training module, which is used to construct a total loss function through a feature consistency recommendation function, a triplet loss function, and a classification function, calculate the loss value, and use landmark retrieval to train the dataset T r Train the multi-scale landmark image retrieval network to obtain a multi-scale fusion landmark image retrieval model with feature consistency recommendation;

[0039] A retrieval module, which is used to retrieve the landmark retrieval test dataset T e Input it into the multi-scale fusion landmark image retrieval model with feature consistency recommendation, and output the retrieval result of the landmark image.

[0040] Compared with the prior art, the present invention has at least the following beneficial effects:

[0041] The present invention proposes a multi-scale fusion landmark image retrieval method with feature consistency recommendation. The network backbone selects the ResNet50 structure. By designing a multi-scale information extraction module, it obtains multiple medium-sized local feature blocks in the order from the upper left to the lower right to complete the extraction of multi-scale information; a feature self-attention fusion network is proposed. Through the feature self-attention fusion branch Transformer encoding layer, self-attention learning is performed on the important information in multiple local feature blocks, and the local feature blocks are fused with the initial global feature map randomly initialized by the Transformer encoding layer to generate two landmark global feature maps, and then the two landmark global feature maps are concatenated to form a joint global feature map, thereby establishing a high-dimensional feature information chain; a regional feature consistency recommendation item is designed to respectively restrict the regional features that the global feature map focuses on. The multi-scale landmark image retrieval method of the present invention solves the problems of unbalanced and insufficient extraction of different-scale features, reduces the loss caused by manual annotation, improves the retrieval ability of the multi-scale landmark image retrieval network for multi-scale landmark images, realizes a more accurate retrieval matching rate, and promotes the deployment and application of multi-scale landmark image retrieval in real scenarios. Brief Description of the Drawings

[0042] Figure 1 It is a flowchart of the implementation of the present invention;

[0043] Figure 2 It is the overall structure diagram of the network of the present invention;

[0044] Figure 3 It is a schematic diagram of the retrieval result of the retrieval method of the present invention in the landmark building dataset Paris6k. Detailed Embodiments

[0045] The present invention will be further described below with reference to the accompanying drawings and specific embodiments.

[0046] Such as Figure 1Shown as follows: The present invention proposes a multi-scale fusion landmark image retrieval method with feature consistency suggestions, and the specific steps are as follows:

[0047] 1. Obtain landmark image data captured by different camera devices, and construct a landmark retrieval training dataset T r and a test dataset T e , and perform image preprocessing operations on T r . The specific steps include:

[0048] Obtain a large number of landmark images from multiple camera devices, and use the manual annotation method to label the same type of landmark in the landmark image with the corresponding category as the prefix, and assign an independent number after the category, (such as Big Wild Goose Pagoda - 0001, Big Wild Goose Pagoda - 0002) Repeat the above steps to construct a landmark image retrieval dataset. After construction, divide the data into a landmark retrieval training dataset T r and a test dataset T e , and perform preprocessing on the images in the detection training set T r . The landmark retrieval training dataset T r and the test dataset T e are respectively used to train the network and test the network.

[0049] Preferably, when performing landmark annotation, the category prefixes between different landmarks are different, and the numbers of the same landmark are different;

[0050] Preferably, the preprocessing includes: during training, perform random flipping up, down, left, and right and random erasing image preprocessing operations on all landmark images in the training set T r . The image size is uniformly scaled to a fixed size of 256×256, and a normalization operation is performed.

[0051] 2. Construct a multi-scale fusion landmark image retrieval network with feature consistency suggestions, which specifically includes a backbone network ResNet50 (Backbone), a multi-scale information extraction module (Multi-scale Information Extraction Module), a feature self-attention fusion network (Feature Self-attention Fusion Network), and a regional feature consistency suggestion item. After the maximum pooling layer of the backbone network ResNet50, there is a multi-scale information extraction module connected. The backbone network ResNet50 obtains an initial local feature map of the input image. The multi-scale information extraction module extracts multiple local feature blocks of the initial local feature map. The feature self-attention fusion network is used for self-attention learning and fusing multiple local feature blocks and the initial global feature mapping generated by the Transformer encoding layer to obtain two landmark global feature mappings F i 1 , Fi 2 , map the two landmark global feature maps F i 1 , F i 2 Perform a splicing operation to obtain the combined global feature map F i 3 ; The regional feature consistency proposal item is used to restrict the regional features that the landmark global feature maps F i 1 , F i 2 focus on.

[0052] As Figure 2 shown, the specific steps include:

[0053] This network mainly consists of three parts:

[0054] ① Use the max pooling layer of the backbone network ResNet50 on the input image to establish the initial local feature map, and use the multi-scale information extraction module to extract multiple local feature blocks of the initial local features through the recombined tensor function;

[0055] ② Construct a feature self-attention fusion network after the backbone network ResNet50. The feature self-attention fusion network includes two feature self-attention fusion branches. Each of the two feature self-attention fusion branches consists of a Transformer encoding layer. The Transformer encoding layer first randomly initializes to generate the global feature map, and then effectively fuses the initial global feature map with the local feature blocks to obtain two landmark global feature maps F i 1 , F i 2 , and splice F i 1 , F i 2 to get the third global feature map F i 3 , F i 1 , F i 2 , F i 3 to form a high-dimensional feature information chain.

[0056] ③ Design a regional feature consistency proposal item to respectively restrict the regional features that the landmark global feature maps F i 1 , F i 2 focus on, so that the landmark global feature maps F i 1 , F i2 Landmark buildings with different category prefixes can spontaneously focus on the same area, such as F i 1 , F i 2 and focus on windows, doors, etc. respectively.

[0057] 3. Multi-scale information extraction module, as Figure 2 shown, the specific steps include:

[0058] The backbone network ResNet50 neural network is composed of several batch normalization layers, several convolutional layers, and several non-linear activation layers. The images I r in the landmark retrieval training dataset T i , i ∈ 1, 2, 3…, are input into the ResNet50 neural network in batches of size n to generate an initial local feature map. The size of this initial local feature map is 16×16×2048. The multi-scale information extraction module extracts local feature blocks f i n and f i M of different scales from the initial local feature map through a tensor reorganization function, prompting the network to focus on local information at different scales. Among them, the size of f i n is set to 2×2×2048 and is divided into N blocks, where N is 64. The size of f i M is 4×4×2048 and is divided into M blocks, where M is 16. Thus, N local feature blocks f i N and M local feature blocks f i M are obtained.

[0059] 4. Feature self-attention fusion network, as Figure 2 shown, the specific steps include:

[0060] 1) Two Transformer encoding layers are respectively initialized to generate initial global feature maps C′ 0 , C″ 0 . The initial global feature maps C′ 0 , C″ 0 are used to extract global classification features The local feature blocks f i N , f i M and the initial global feature maps C′ 0 , C″ 0 are paired in pairs, that is, f i N and C′0 Group, f i M With C″ 0 Group, and input them into the Transformer encoding layers of two feature self-attention fusion branches respectively. In the Transformer encoding layer, the standard learnable position vector E pos Is embedded into f i N With C′ 0 Group and f i M With C″ 0 In the group, and obtain two initially fused landmark global feature maps F i 1 , F i 2 ;

[0061] The embedding of the position vector E pos Can retain the position information of the initial global feature map and the local feature blocks, where the position information of the initial global feature maps in the two branches is defined as 0, and the local feature blocks f i N , f i M Are 1 to N, and 1 to M respectively.

[0062] Through the result vector sequences z′ 0 And z" 0 Respectively represent the specific information of the two landmark global feature maps F i 1 , F i 2 The result vector sequences z′ 0 And z" 0 Are respectively defined as formulas (1) and (2):

[0063]

[0064] In the formula, E represents processing the local feature block into a computable vector, E pos Represents the embedding of the standard position information, (p, p) is the resolution of the feature block, C is the number of channels, and D is the dimension

[0065] 2) Input the result vector sequences z′ 0 And z" 0 Into the two Transformer encoding layers respectively. The multi-head self-attention module in the Transformer encoding layer processes the two landmark global feature maps F i 1 , F i 2Perform self-attention learning on the important information in it. Specifically, the result vector sequence z 0 is multiplied by the randomly initialized matrix W Q , W K , W V to generate the Q, K, and V matrices. Then, the similarity between Q and all K is calculated, and the result vector sequence z' 0 and z" 0 The weights of each part are obtained. The weights are transformed into a probability distribution using the softmax regression function. The calculation process is shown in Equation (3):

[0066]

[0067] In the formula, is used to transform the attention matrix into a standard normal distribution.

[0068] The obtained probability distribution is fed back into the result vector sequence z' 0 and z" 0 to update the weights of the landmark global feature map F i 1 , F i 2 for the specific information, and a new landmark global feature map F i 1 , F i 2 is obtained, such that the information with high weights in the landmark global feature map F i 1 , F i 2 is more important, and the information with low weights is less important.

[0069] 3) Concatenate the two landmark global feature maps F i 1 , F i 2 to obtain the combined global feature map F i 3 . The process of the concatenation operation is defined by Equation (4):

[0070]

[0071] In the formula, represents the concatenation operation

[0072] F i 1 , F i 2 , F i 3 constitute a high-dimensional feature information chain.

[0073] 5. Execution of the regional feature consistency suggestion items, and the specific steps include:

[0074] First, design a feature consistency suggestion function so that the global feature map F generated by the feature self-attention fusion branch i 1 , F i 2 respectively focuses on the same regions of landmark buildings with different category prefixes. For the global feature map F i 1 , F i 2 Initialize a cluster center vector c k , and then use the feature consistency suggestion function to gradually constrain the features learned by this global feature map around the cluster center vector c k so that this global feature map can perceive the same regional features, as shown in formula (5):

[0075]

[0076] In the formula, represents the Euclidean norm. In this model represents the global feature map generated by the feature self-attention fusion branch, K = 2, where the cluster center vector c k is randomly initialized according to the content learned by the global feature map and updated by moving its average value:

[0077]

[0078] where α controls the update rate of c k . By optimizing this feature consistency suggestion function, gradually approaching the cluster center vector c k can make the regional features captured by this global feature map similar.

[0079] 6. Loss calculation, and the specific steps include:

[0080] Use the triplet loss function to enhance the discriminability of any global feature map :

[0081]

[0082] In the formula, F represents the global feature map, k = 1, 2, 3 represent different global feature maps, A is the Anchor representing the sample itself, N is the negative representing the sample of a different class from A, and P is the Positive representing the sample of the same class as A. respectively represent the feature vectors of the source sample, negative sample, and positive sample that make up the triple, and respectively represent the Euclidean distances of the positive sample pair and the negative sample pair, m represents the margin threshold of the triple loss, + represents taking the positive value.

[0083] Meanwhile, after any global feature mapping design a batch normalization layer BN(), a linear layer W, and a Softmax layer to obtain a classification loss function for calculating the classification loss:

[0084]

[0085] In the formula, is the probability distribution of the correct prediction of sample i, p i is the probability of the correct prediction for each sample, N represents the total number of possible situations, represents the classification loss.

[0086] The final total loss function is jointly composed of the feature consistency recommendation function, the triple loss function, and the classification loss function, and can be expressed as formula (9):

[0087]

[0088] 7. Landmark image retrieval network, the specific steps include:

[0089] First, input the data of the landmark retrieval training dataset T r obtained in step 1 into the network for training in batches of a certain size n. After determining the total loss according to step 5, use the adaptive gradient descent algorithm to train the landmark image retrieval network to obtain a landmark building image retrieval model.

[0090] Then test the trained model. For the landmark retrieval dataset T e obtained in step 1, obtain the global feature mapping F i 1 , F i 2 , splice them to obtain the joint global feature mapping F i 3 = [F i 1 , F i 2 , for the landmark image I e in the landmark retrieval dataset T j, j = 1, 2, 3…, calculate the similarity of the global feature maps of pairwise landmark images through the cosine distance function, and finally output the sorting result according to the similarity size to complete the landmark image retrieval. The calculation of the cosine function is shown in formula (10), and the calculation of the cosine distance is shown in formula (11).

[0091]

[0092]

[0093] In the formula, and where F represents the global feature map, j 1 and j 2 represent any one image of the test sample and the non-test sample in the landmark retrieval test dataset, |||| represents the modulus, and represents the dot product.

[0094] The working principle of the present invention:

[0095] Step 1, collect landmark image data from different shooting devices and construct a landmark retrieval training dataset T r for training the network designed by the present invention.

[0096] Step 2, construct a multi-scale fusion landmark image retrieval network for feature consistency recommendation.

[0097] 2.1, use the maximum pooling layer of the ResNet50 network to obtain the initial local feature map of the landmark image as 16×16×2048;

[0098] 2.2, execute the multi-scale information extraction module, divide the initial local feature map into 64 local feature blocks of size 2×2×2048 and 16 local feature blocks of size 4×4×2048 in the order from the upper left to the lower right through the recombined tensor function;

[0099] 2.3, execute the feature self-attention fusion network. First, initialize and generate the initial global feature maps C′ 0 、C″ 0 , the initial global feature maps C′ 0 、C″ 0 are used to extract global classification features Then, pair the local feature blocks f i N 、f i M obtained in step 2.2 and the global feature maps C′ 0 、C″ 0 pairwise to get f i N and C′ 0Group f i M With C″ 0 Group;

[0100] Bring f i n With C′ 0 Group and f i M With C″ 0 The groups are respectively input into the multi - head self - attention module in the Transformer encoding layer for learning to obtain two landmark global feature maps F i 1 ,F i 2 , and a splicing operation is performed to obtain F i 3 ;

[0101] Design area - feature consistency suggestion items to respectively restrict the area features that F i 1 ,F i 2 focuses on.

[0102] Step 3, loss calculation. Calculate the loss values of the three global feature maps F i 1 ,F i 2 ,F i 3 through the feature - consistency suggestion function, the triplet loss function, and the classification function, and select the gradient - descent algorithm to train the landmark image retrieval network to obtain the optimal model of the network;

[0103] Step 4, input the landmark retrieval dataset T e to the landmark image retrieval model to obtain the three global feature maps of the test landmark images, calculate the similarity between pairs of images in the landmark retrieval test dataset T e through the cosine - distance function and output according to the similarity to complete the retrieval of landmark images.

[0104] The present invention also provides a multi - scale fusion landmark image retrieval system with feature - consistency suggestion, including:

[0105] A data acquisition module for collecting landmark image data to construct the landmark retrieval training dataset T r and the test dataset T e ;

[0106] A network construction module, a multi-scale fusion landmark image retrieval network for feature consistency recommendation, includes a ResNet50 network connected with a multi-scale information extraction module, a feature self-attention fusion network, and a regional feature consistency recommendation item;

[0107] A network training module, which is used to construct a total loss function through a feature consistency recommendation function, a triplet loss function, and a classification function, calculate the loss value, and use the landmark retrieval training dataset T r to train the multi-scale landmark image retrieval network to obtain a multi-scale fusion landmark image retrieval model for feature consistency recommendation;

[0108] A retrieval module, which is used to input the landmark retrieval test dataset T e into the multi-scale fusion landmark image retrieval model for feature consistency recommendation and output the retrieval result of the landmark image.

[0109] The present invention also provides a computer device, which includes a computer, a server, or other terminal devices with computing functions. The device includes a processor and a memory connected through a bus. The memory stores a program, and this program is configured to be executed by the processor. The program includes a method for executing the above-mentioned multi-scale fusion landmark image retrieval method for feature consistency recommendation.

[0110] The present invention also provides a computer storage medium, in which a computer program is stored. When the program is executed by a processor, the processor implements the above-mentioned multi-scale fusion landmark image retrieval method for feature consistency recommendation when executing the computer program.

[0111] Figure 3 This is the retrieval result of the method of the present invention on the landmark dataset Paris6k. Among them Figure 3 the first column in represents the query image, and the 2nd - 6th images in each row represent the query results. It can be found from the query results that the matching accuracy of the method of the present invention is relatively high. Through Figure 3 the sixth image in row (a), the third image in row (b), and the third image in row (e) of the retrieval result graph, it can be seen that the method of the present invention can be accurately retrieved when the landmark appears at a relatively small scale. Moreover, through Figure 3 the sixth image in row (b) and the third image in row (c) of the retrieval result graph, it can be seen that the method of the present invention has good retrieval effects in the case of perspective change and illumination change.

[0112] Compare the CMC (Cumulative Match Characteristic) result performance of the method of the present invention and other existing excellent retrieval methods on the dataset Paris6k. The results are shown in Table 1:

[0113] CMC Performance Comparison of the Landmark Building Dataset Paris6k in Table 1

[0114] Method mAP R-MAC 82.8% DELF+FT+ATT 84.9% siaMAC+QE* 85.7% R-MAC+R+QE 86.3% The method of the present invention 87.0%

[0115] As can be seen from Table 1, compared with other advanced algorithms, the mAP of the method of the present invention is 87.0%, and the mAP is improved by 0.7% compared with the R-MAC+R+QE method. The retrieval effect is in the leading position, further proving the effectiveness of the method of the present invention.

Claims

1. A multi-scale fusion landmark image retrieval method with feature consistency recommendation, characterized in that, the specific steps are as follows: S1 Collect landmark image data and construct a landmark retrieval training data set and the test data set ; S2 Construct a multi-scale fusion landmark image retrieval network with feature consistency recommendation, including a ResNet50 network connected with a multi-scale information extraction module, a feature self-attention fusion network, and a regional feature consistency recommendation item; S3 constructs a total loss function through a feature consistency recommendation function, a triplet loss function, and a classification function, calculates the loss value, and uses landmark retrieval to train the dataset Train the multi-scale landmark image retrieval network to obtain a multi-scale fusion landmark image retrieval model; S4 inputs the landmark retrieval test data set into the multi-scale fusion landmark image retrieval model with input feature consistency suggestions, and outputs the retrieval results of landmark images; In step S2, a multi-scale information extraction module is connected after the maximum pooling layer of the ResNet50 network, and the ResNet50 network is used to obtain an initial local feature map of the landmark image; the multi-scale information extraction module extracts a plurality of local feature blocks of the initial local feature map in the order from the upper left to the lower right through a tensor recombination function, and obtains N local feature blocks of different scales and M local feature blocks ; Among them, the local feature block is a small-scale feature block, and is a large-scale feature block; In step S2, a feature self-attention fusion network is constructed after ResNet50. The feature self-attention fusion network includes two feature self-attention fusion branches, and each of the two feature self-attention fusion branches is composed of one Transformer encoding layer; In step S2, the specific steps of the processing process of the feature self-attention fusion network are: 1) Two Transformer encoding layers are respectively initialized to generate initial global feature maps , and the local feature blocks , and the initial global feature maps are paired up two by two to obtain groups, groups; groups, are respectively input into the Transformer encoding layer, and the standard learnable position vectors are embedded into groups, groups to obtain two preliminarily fused landmark global feature maps ; 2) Using the result vector sequences and to represent the specific information of two landmark global feature maps respectively, input the result vector sequences and into two Transformer encoding layers respectively to perform self-attention learning on the important information in the two landmark global feature maps , and obtain the weights of each part in the result vector sequences and to get the distribution probability of the weights, and update the weights of the specific information of the landmark global feature map ; 3) Concatenate the global feature maps of the two landmarks to obtain the combined global feature map .

2. A multi-scale fusion landmark image retrieval method with feature consistency recommendation according to claim 1, characterized in that, in step S1, the method of manual annotation is used to label the same type of landmark in the landmark image with the corresponding category as the prefix, and an independent number is assigned after the category. Among them, the category prefixes between different landmarks are different, and the numbers of the same landmark are different.

3. A multi-scale fusion landmark image retrieval method with feature consistency recommendation according to claim 1, characterized in that, In step S2, in the regional feature consistency suggestion item, the global feature map generated by the feature self-attention fusion branch is made to focus on the same regions of landmark buildings with different category prefixes by constructing a feature consistency suggestion function respectively, and the specific feature consistency suggestion function is as follows: (4) In the formula, represents the Euclidean norm, represents the landmark global feature map generated by the feature self-attention fusion branch , K = 2, is the cluster center vector.

4. A multi-scale fusion landmark image retrieval method with feature consistency recommendation according to claim 3, characterized in that, In step S2, the cluster center vector is randomly initialized according to the content learned by the global feature mapping and updated by moving its average value: Among them Control of the update rate represents the landmark global feature map generated by the feature self-attention fusion branch .

5. A multi-scale fusion landmark image retrieval method with feature consistency recommendation according to claim 4, characterized in that, in step S3, the total loss function is: (8) Among them, is the feature consistency recommendation function, is the classification loss function, is the triplet loss function. Specifically: The classification loss function is to design a batch normalization layer after any global feature map a linear layer and a layer, specifically: ​ (6) (7) The triplet loss function is used to enhance the discriminability of any global feature map , specifically as follows: (5) In the formula, F represents the global feature mapping, k = 1, 2, 3 represent different global feature mappings, A is the Anchor representing the sample itself, N is the negative representing the sample of a different class from A, and P is the Positive representing the sample of the same class as A. respectively represent the feature vectors of the source sample, negative sample, and positive sample that make up the triplet. and respectively represent the Euclidean distances of the positive sample pair and the negative sample pair. represents the margin threshold of the triplet loss. represents taking the positive value.

6. A multi-scale fusion landmark image retrieval method with feature consistency recommendation according to claim 1, characterized in that, In step S4, the landmark retrieval test data set is input into the landmark image retrieval model to obtain the joint global feature map of the test landmark images , and the similarity between the global feature maps of the landmark images in the landmark retrieval test data set is calculated through the cosine distance function, and the image retrieval results are sorted and output according to the similarity. The cosine distance function is specifically: (9) (10) In the formula, and where F represents the global feature map, j 1 and j 2 denote any one image of the test samples and non-test samples in the landmark retrieval test dataset, and || || represents the modulus. in the test dataset 7. A multi-scale fusion landmark image retrieval system with feature consistency recommendation, characterized in that, runs a multi-scale fusion landmark image retrieval method with feature consistency recommendation according to claim 1. The system includes: A data acquisition module, which is used to acquire landmark image data and construct a landmark retrieval training data set and the test data set ; A network construction module for a multi-scale fusion landmark image retrieval network with feature consistency recommendation, including a ResNet50 network connected with a multi-scale information extraction module, a feature self-attention fusion network, and a regional feature consistency recommendation item; A network training module, configured to construct an overall loss function through a feature consistency recommendation function, a triplet loss function, and a classification function, calculate a loss value, and utilize a landmark retrieval training data set Train a multi-scale landmark image retrieval network to obtain a multi-scale fusion landmark image retrieval model with feature consistency recommendation; A retrieval module for retrieving the landmark retrieval test data set Input the multi-scale fusion landmark image retrieval model with suggestions for input feature consistency, and output the retrieval results of landmark images.

Citation Information

Cited By

  • Data fusion method for testing quality consistency of industrial data

    CN121636620A