Visual scene recognition positioning method and positioning device, storage medium and computer equipment
By introducing Transformer feature aggregator and proxy attention mechanism into the visual positioning method, the problems of long inference time and poor global feature robustness in the visual positioning method are solved, and more efficient visual scene recognition and position recognition are achieved.
Patent Information
- Application Number
- CN202510313107.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-17
- Publication Date
- 2025-07-25
AI Technical Summary
The existing deep learning-based visual positioning methods are difficult to effectively extract effective information in visual scenes and reduce redundant information when facing environmental changes and dynamic object occlusion.
Using a feature aggregator architecture based on Transformer, combining the proxy attention mechanism and the gated feedforward network, we learn image feature context association through the self-attention mechanism, highlight the characteristics of the region of interest, reduce redundant information, and extract more robust global features.
It improves the accuracy of visual scene recognition and the real-timeness of position recognition, reduces the amount of calculation, and enhances the recognition ability and robustness of the model.
Smart Images

Figure CN120374932A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the technical field of visual positioning. Specifically, it relates to a visual scene recognition and positioning method, a positioning device, a storage medium, and a computer device. Background Art
[0002] For mobile energy storage devices in service parks, the accuracy of map construction and positioning is related to the efficiency and safety of automatic mobile devices performing tasks. Visual position recognition technology plays an important role in this, mainly used for loop detection in the SLAM system to eliminate the cumulative error of the odometer. However, visual position recognition faces many challenges: environmental changes caused by climate and light, occlusion by dynamic objects, different angles of the content obtained by the camera, the real-time performance of the system, etc.
[0003] Currently, there are many visual positioning methods based on deep learning. However, the robustness of these models depends on two-stage Re-rank, which uses local matching to achieve precise positioning after global feature rough matching, resulting in long inference time of the model and more computing resources occupied. And relying only on global features and using convolutional operations is not conducive to representing the internal correlation of image features, resulting in poor feature robustness. To solve the above problems, a feature aggregator architecture based on Transformer is introduced into visual position recognition. Transformer is a model architecture without recurrence, which completely relies on the attention mechanism to model the global dependence between the input and the output, and is composed of an encoder and a decoder. The encoder compresses the input of the model into a deep representation, and then the decoder restores it to the representation of the input. Only the encoder part is used in the feature aggregation module to encode the input features, and the compressed representation is used as the global feature for image matching.
[0004] In current Transformer-based methods, one approach is to use a shallow CNN to encode an image, which is then fed as input into a backbone network based on Transformer to extract image features at multiple scales. These features incorporate multi-level attention, focusing on different semantically meaningful regions, thereby generating a global feature representation. After obtaining candidate images through matching, the output tokens of the Transformer layer filtered by the fused attention mask are used as block-level descriptors to perform geometric verification and re-ranking. Although local features perform well in re-ranking, they rely too much on two-stage Re-rank, resulting in an overly long inference time, and the global descriptors generated by this network are less effective than previous NetVLAD or CosPlace algorithms. Another approach is a global feature aggregator based on a multi-layer perceptron (MLP). This model uses the feature maps extracted from a pre-trained backbone and inputs them into an isotropic global feature aggregator composed of multi-layer perceptrons (MLPs), and iteratively incorporates global relationships into each individual feature map to replace traditional local or pyramidal aggregation. However, treating global information equally in this way results in a large amount of redundancy being retained in the features, and it is unable to pay more attention to the information of regions of interest that contribute to recognition. Summary of the Invention
[0005] The technical problem solved by this application is: how to fully extract the effective information in the visual scene, reduce redundant information, and reduce the amount of computation.
[0006] This application provides a visual scene recognition and localization method, and the visual scene recognition and localization method includes:
[0007] Feature extraction stage: Use the backbone network of the recognition and localization model to extract the preliminary image features of the original image, and sequentially process the preliminary image features through the proxy attention mechanism module and the gated feed-forward network of the recognition and localization model to obtain the global image features of the original image;
[0008] Model training stage: Use the global image features as training samples to train the recognition and localization model;
[0009] Recognition and localization stage: Use the trained recognition and localization model to extract the feature vectors of the image to be retrieved and the database images, calculate the similarity between the feature vector of the image to be retrieved and the feature vectors of each database image, and determine the matching image of the image to be retrieved according to each similarity.
[0010] Optionally, the method of sequentially processing the preliminary image features through the proxy attention mechanism module and the gated feed-forward network of the recognition and localization model includes:
[0011] Obtain a query matrix, a key matrix, and a value matrix respectively according to the preliminary image features;
[0012] Perform a pooling operation on the query matrix to obtain a proxy matrix, and perform self-attention calculation on the proxy matrix, key matrix, and value matrix to obtain a self-attention calculation result;
[0013] Perform self-attention calculation on the query matrix, proxy matrix, and self-attention calculation result to obtain a proxy self-attention result;
[0014] Obtain the output of the proxy attention mechanism module according to the proxy self-attention result and the preliminary image features.
[0015] Optionally, the method of sequentially passing the preliminary image features through the proxy attention mechanism module and the gated feed-forward network of the recognition and localization model further includes:
[0016] Perform convolution processing on the output of the proxy attention mechanism module on two branches of the gated feed-forward network respectively;
[0017] Perform activation processing on the convolution result of one of the branches using an activation function, and perform dot multiplication on the result of the activation processing and the convolution result of the other branch;
[0018] Add the dot multiplication result to the output of the proxy attention mechanism module.
[0019] Optionally, the method of training the recognition and localization model using the global image features as training samples includes:
[0020] Use the hybrid feature data augmentation model of the recognition and localization model to perform data augmentation on the global image features to obtain new global image features;
[0021] Use the global image features and the new global image features together as training samples to train the recognition and localization model.
[0022] Optionally, the method of using the hybrid feature data augmentation model of the recognition and localization model to perform data augmentation on the global image features to obtain new global image features includes:
[0023] Calculate the average value of the global image features at the same location;
[0024] Calculate new global image features according to the average value and the trainable offset boundary margin.
[0025] Optionally, the similarity is cosine similarity.
[0026] Optionally, the method of determining the matching image of the image to be retrieved according to each similarity includes:
[0027] When the cosine similarity is greater than the threshold, the corresponding database image is used as the matching image of the image to be retrieved; when the cosine similarities of multiple database images all exceed the threshold, the database image with the largest cosine similarity is used as the matching image.
[0028] The present application also provides a visual scene recognition and positioning device, which includes:
[0029] A feature extraction module, which is used to extract preliminary image features of the original image by using the backbone network of the recognition and positioning model, and process the preliminary image features through the proxy attention mechanism module and the gated feed-forward network of the recognition and positioning model in sequence to obtain the global image features of the original image;
[0030] A model training module, which is used to train the recognition and positioning model by using the global image features as training samples;
[0031] A recognition and positioning module, which uses the trained recognition and positioning model to extract the feature vectors of the image to be retrieved and the database images, calculates the similarities between the feature vector of the image to be retrieved and the feature vectors of each database image, and determines the matching image of the image to be retrieved according to each similarity.
[0032] The present application also provides a computer-readable storage medium, which stores a visual scene recognition and positioning program, and when the visual scene recognition and positioning program is executed by a processor, the above-mentioned visual scene recognition and positioning method is realized.
[0033] The present application also provides a computer device, which includes a computer-readable storage medium, a processor, and a visual scene recognition and positioning program stored in the computer-readable storage medium, and when the visual scene recognition and positioning program is executed by the processor, the above-mentioned visual scene recognition and positioning method is realized.
[0034] A visual scene recognition and positioning method, a positioning device, a storage medium, and a device provided by the present application have the following technical effects:
[0035] Through the self-attention mechanism, this solution enables the model to learn the association of image feature context, highlights the features of more interesting regions, reduces redundant information, and can extract more representative and robust features in the scene picture, not only effectively improving the accuracy in scene recognition, but also meeting the real-time requirement of location recognition. Description of the Drawings
[0036] Figure 1 It is the main step flowchart of the visual scene recognition and positioning method according to one or more embodiments;
[0037] Figure 2 Schematic diagram of the detailed process of a visual scene recognition and localization method according to one or more embodiments;
[0038] Figure 3 Schematic diagram of the principle of data augmentation according to one or more embodiments;
[0039] Figure 4 Schematic diagram of the recognition and localization process according to one or more embodiments;
[0040] Figure 5 Principle block diagram of a visual scene recognition and localization device according to one or more embodiments;
[0041] Figure 6 Schematic diagram of a computer device according to one or more embodiments. Detailed implementation manners
[0042] In order to make the objectives, technical solutions and advantages of the present application more clear and understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0043] Before describing the various embodiments of the present application in detail, the technical concept of the present application will be briefly described first: In the current visual recognition and localization methods, there are problems such as a large amount of redundant information during feature extraction, difficulty in obtaining effective information, and excessive computational complexity. Therefore, the visual scene recognition and localization method provided by the present application uses the proxy attention mechanism module and gated feed-forward network of the recognition and localization model to process the preliminary image features extracted by the backbone network, and obtains global image features, which can relieve the computational pressure of the self-attention mechanism. At the same time, the gated mechanism is used to screen the regions of interest, reduce redundant information, and obtain more accurate, more robust, and stronger representation ability features, which is beneficial to improving the recognition ability of the model. The specific principle of the visual scene recognition and localization method of the present application will be described below with more embodiments.
[0044] Specifically, as Figure 1 shown, the visual scene recognition and localization method of the first embodiment includes the following steps:
[0045] Step S10, feature extraction stage: Use the backbone network of the recognition and localization model to extract the preliminary image features of the original image, and process the preliminary image features through the proxy attention mechanism module and the gated feed-forward network of the recognition and localization model in sequence to obtain the global image features of the original image;
[0046] Step S20, model training stage: Use the global image features as training samples to train the recognition and localization model;
[0047] Step S30, Recognition and Location Stage: Use the trained recognition and location model to extract the feature vectors of the image to be retrieved and the database images, calculate the similarity between the feature vector of the image to be retrieved and the feature vectors of each database image, and determine the matching image of the image to be retrieved based on each similarity.
[0048] In one or more embodiments, as Figure 2 shown, the backbone network in Step S10 uses ResNet18, and is cropped at the penultimate ResNet residual block. The backbone network extracts preliminary image features with a spatial dimension of 20×20. Exemplarily, a 1×1 convolutional layer is applied to the preliminary image features for channel reduction. In the actual processing, a Dimension Flatten process is also required, that is, the multi-dimensional array is reduced to a one-dimensional array through a flattening operation.
[0049] Furthermore, the data processing process of the proxy attention mechanism module can be expressed by the following formula:
[0050]
[0051] Specifically, the data processing process of the proxy attention mechanism module is as follows: According to the preliminary image feature X, the query matrix key matrix and value matrix are obtained respectively. A pooling operation is performed on the query matrix to obtain the proxy matrix A self-attention calculation is performed on the proxy matrix key matrix and value matrix to obtain the self-attention calculation result Value; A self-attention calculation is performed on the query matrix proxy matrix and the self-attention calculation result Value to obtain the proxy self-attention result; The output of the proxy attention mechanism module is obtained according to the proxy self-attention result and the preliminary image feature X
[0052] Furthermore, the data processing process of the gated feed-forward network can be expressed by the following formula:
[0053]
[0054] Specifically, in order to standardize the local information transmitted to the next layer to facilitate adapting to the change of perspective, we designed a gated feed-forward network (Gate-FN) to replace the feed-forward network (FFN) layer. The data processing process of the gated feed-forward network is as follows: The output of the proxy attention mechanism module Convolution processing is performed separately on two branches of the gated feed-forward network; the convolution result of one branch is activated using an activation function, and the result of the activation process is multiplied element-wise with the convolution result of the other branch; the element-wise multiplication result is added to the output of the proxy attention mechanism module. Exemplarily, the two branches respectively perform convolution operations using 1×1 and 3×3 convolution blocks, as shown in shown, where P 0 represents performing a convolution operation on the element-wise multiplication result. A 1×1 sized convolutional kernel is used to perform convolution on the channels to enrich the features using context information, and a DWC (depthwise separable convolution) with a convolutional kernel size of 3×3 is used to encode spatially adjacent pixel positions to facilitate the learning of local image structure information. At the same time, a gating mechanism is introduced to control the information flow, further enhancing the information of the region of interest flowing to the next layer.
[0055] Furthermore, during the actual calculation process, the output result of the gated feed-forward network is subjected to projection processing, flattening processing, and normalization processing to meet the format requirements of the training process.
[0056] In one or more embodiments, during the training process, in order to improve the generalization performance of the model, we designed a data augmentation method to enhance the diversity of features. This method only performs linear interpolation on the extracted global image features and labels, thereby generating new training samples without overly increasing the burden of the training process. Specifically, the data augmentation method is as follows: calculate the average value of each global image feature at the same location, and calculate the new global image feature based on the average value and the trainable offset boundary margin.
[0057] Exemplarily, as Figure 3 shown, in each training batch, after being processed by the feature extraction and aggregation network, we obtain n global image features F: (b, n, c), where b represents the number of locations, n represents the number of pictures describing the same location, and c represents the dimension. First, calculate the average value of the n global image features describing the same location in a batch. The average value has the same label as them. Subsequently, we introduce a random and trainable offset boundary margin margin, which represents the offset range of the new global image feature. The new global image feature is generated within a specific range of the average value and then input into the subsequent loss calculation process. To facilitate training and enhance the fitting of the classification boundary, we introduce a random scale α, and the range is defined as [-0.5, 0.5]. Finally, this process can be represented by the following formula:
[0058]
[0059] Furthermore, during training, the GSV-Cities dataset is selected as the training set, which contains 67,000 different geographical locations depicted in 560,000 images. This dataset covers a large number of scenarios, including changes in lighting and angles at the same location, and different states of similar scenarios. It enables the model to encode significant features in the same scene during training, resulting in more robust global features and stronger generalization ability. The size of the input original image is adjusted to 320×320, and a batch containing 40 positions is used, with each position depicted by 4 images, thus obtaining a small batch of 160 images. Stochastic Gradient Descent (SGD) is used for optimization, with a momentum of 0.9 and a weight decay of 0.001. The initial learning rate is 0.05 and is divided by 3 every 5 epochs.
[0060] In one or more embodiments, the similarity used in the recognition and positioning stage is the cosine similarity. For the mobile energy storage device equipped with this model, the recognition and positioning process in the mobile mapping process is as Figure 4 shown. The current scene image is obtained through the camera sensor as the current image to be retrieved. The shape and size of the image to be retrieved are adjusted to N×320×320×3 and then input into the model to obtain a feature vector with a shape of N×1024, denoted as Q. i To compare with the stored database images to determine whether they are from the same location, the database images are also input into the model to obtain a feature vector R with a shape of M×1024. i The feature vector obtained from each retrieved image is filled into the database. The calculation process of the cosine similarity is as follows:
[0061]
[0062] where n is 1024. When the value is close to 1, the two images are more likely to represent the same location. If it is close to -1, the two images indicate different locations. We represent the matching results with a cosine similarity matrix, and the maximum value in each row is the optimal match. In the experiment, when the cosine approximation is greater than the set threshold, the database image corresponding to the image to be retrieved is retrieved. If there are multiple database images that meet the conditions, the database image with the largest cosine similarity is selected as the recognition result.
[0063] To verify the performance of the global feature aggregation model, the recall metrics of other visual scene recognition methods were compared. When the feature extractor is the lightweight backbone network ResNet18, the aggregation model with the gated proxy attention layer has improved the global retrieval recall compared to other methods, whether using CNN or Transformer as the aggregator architecture. For example, in the MSLS_VAL dataset, the recall@1 metric of the model reached 83.81%, exceeding the previous state-of-the-art algorithm MixVPR by 2.47% under the same training conditions (epoch: 80, Tesla K80, dim: 1024), and improving by 1.11% compared to 82.7% reported in the literature. In addition, the image retrieval results of the model designed in this project were visually compared with those of other algorithms on the MSLS_VAL dataset. The experiments show that even in challenging environments such as occlusion, illumination changes, seasonal changes, and repetitive structures, the model based on the gated proxy attention layer in this embodiment can maintain excellent retrieval performance. This is because after the global information is interacted using the self-attention mechanism, Gating Forward further suppresses the features with less information, enabling each layer to pay more attention to the overall connection and enhancing the understanding of the global scene.
[0064] As Figure 5 shown, in one or more embodiments, the visual scene recognition and localization device includes a feature extraction module 100, a model training module 200, and an identification and localization module 300. The feature extraction module 100 is configured to extract preliminary image features of the original image using the backbone network of the identification and localization model, and process the preliminary image features through the proxy attention mechanism module and the gated feed-forward network of the identification and localization model in sequence to obtain the global image features of the original image. The model training module 200 is configured to use the global image features as training samples to train the identification and localization model. The identification and localization module 300 extracts the feature vectors of the image to be retrieved and the database images using the trained identification and localization model, calculates the similarity between the feature vector of the image to be retrieved and the feature vectors of each database image, and determines the matching image of the image to be retrieved based on each similarity. For a more detailed working process of the feature extraction module 100, the model training module 200, and the identification and localization module 300, reference can be made to the relevant description of the embodiments of the visual scene recognition and localization method, which will not be elaborated here.
[0065] In one or more embodiments, a computer-readable storage medium stores a visual scene recognition and localization program, and when the visual scene recognition and localization program is executed by a processor, it implements the visual scene recognition and localization method of the above embodiments.
[0066] In one or more embodiments, a computer device includes a computer-readable storage medium, a processor, and a visual scene recognition and positioning program stored in the computer-readable storage medium. When the visual scene recognition and positioning program is executed by the processor, the above-described visual scene recognition and positioning method is implemented. At the hardware level, as Figure 6 shown, the computer device includes a processor 12, an internal bus 13, a network interface 14, and a computer-readable storage medium 11. The processor 12 reads the corresponding computer program from the computer-readable storage medium and then runs it, forming a request processing device at the logical level. Of course, in addition to the software implementation, one or more embodiments of this specification do not exclude other implementation manners, such as logic devices or a combination of software and hardware, etc. That is to say, the execution subject of the following processing flow is not limited to each logical unit, and can also be hardware or a logic device.
[0067] Exemplarily, a computer-readable storage medium includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer-readable storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory, or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD), or other optical storage, magnetic cassette tapes, disk storage, quantum memory, graphene-based storage media, or other magnetic storage devices, or any other non-transmission media that can be used to store information that can be accessed by a computing device.
[0068] The specific implementation manners of the present application have been described in detail above. Although some embodiments have been shown and described, those skilled in the art should understand that these embodiments can be modified and perfected without departing from the principles and spirit of the present application defined by the claims and their equivalents, and these modifications and perfections should also be within the protection scope of the present application.
Claims
1. A visual scene recognition and positioning method, characterized in that, The visual scene recognition and positioning method includes: Feature extraction stage: Using the backbone network of the recognition and positioning model to extract the preliminary image features of the original image, and processing the preliminary image features through the proxy attention mechanism module and the gated feed-forward network of the recognition and positioning model in sequence to obtain the global image features of the original image; Model training stage: Using the global image features as training samples to train the recognition and positioning model; Recognition and positioning stage: Using the trained recognition and positioning model to extract the feature vectors of the image to be retrieved and the database images, calculating the similarity between the feature vector of the image to be retrieved and the feature vectors of each database image, and determining the matching image of the image to be retrieved according to each similarity.
2. The visual scene recognition and positioning method according to claim 1, wherein The method of processing the preliminary image features through the proxy attention mechanism module and the gated feed-forward network of the recognition and positioning model in sequence includes: Obtaining a query matrix, a key matrix, and a value matrix respectively according to the preliminary image features; Performing a pooling operation on the query matrix to obtain a proxy matrix, and performing self-attention calculation on the proxy matrix, the key matrix, and the value matrix to obtain a self-attention calculation result; Performing self-attention calculation on the query matrix, the proxy matrix, and the self-attention calculation result to obtain a proxy self-attention result; Obtaining the output of the proxy attention mechanism module according to the proxy self-attention result and the preliminary image features.
3. The visual scene recognition and positioning method according to claim 2, wherein The method of processing the preliminary image features through the proxy attention mechanism module and the gated feed-forward network of the recognition and positioning model in sequence further includes: Performing convolution processing on the output of the proxy attention mechanism module on two branches of the gated feed-forward network respectively; Activating the convolution result of one branch using an activation function, and performing dot multiplication on the activation result and the convolution result of the other branch; Adding the dot multiplication result to the output of the proxy attention mechanism module.
4. The visual scene recognition and positioning method according to claim 1, characterized in that The method of using the global image features as training samples to train the recognition and positioning model includes: Using the hybrid feature data augmentation model of the recognition and positioning model to perform data augmentation on the global image features to obtain new global image features; Using the global image features and the new global image features together as training samples to train the recognition and positioning model.
5. The visual scene recognition and positioning method according to claim 4, wherein The method of using the hybrid feature data augmentation model of the recognition and positioning model to perform data augmentation on the global image features to obtain new global image features includes: Calculating the average value of the global image features at the same location; Calculating new global image features according to the average value and the trainable offset boundary margin.
6. The visual scene recognition and positioning method according to claim 1, characterized in that The similarity is cosine similarity.
7. The visual scene recognition and positioning method according to claim 6, wherein The method of determining the matching image of the image to be retrieved according to each similarity includes: When the cosine similarity is greater than the threshold, the corresponding database image is used as the matching image of the image to be retrieved; when the cosine similarities of multiple database images all exceed the threshold, the database image with the largest cosine similarity is used as the matching image.
8. A visual scene recognition and positioning device, characterized in that, The visual scene recognition and positioning device includes: A feature extraction module, which is used to extract preliminary image features of the original image by using the backbone network of the recognition and localization model, and process the preliminary image features through the proxy attention mechanism module and the gated feed-forward network of the recognition and localization model in sequence to obtain the global image features of the original image; A model training module, which is used to train the recognition and localization model by using the global image features as training samples; A recognition and localization module, which uses the trained recognition and localization model to extract the feature vectors of the image to be retrieved and the database images, calculates the similarity between the feature vector of the image to be retrieved and the feature vectors of each database image, and determines the matching image of the image to be retrieved according to each similarity.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a visual scene recognition and localization program, and when the visual scene recognition and localization program is executed by a processor, it implements the visual scene recognition and localization method according to any one of claims 1 to 7.
10. A computer device, characterized in that, The computer device includes a computer-readable storage medium, a processor, and a visual scene recognition and localization program stored in the computer-readable storage medium. When the visual scene recognition and localization program is executed by the processor, it implements the visual scene recognition and localization method according to any one of claims 1 to 7.