Visual position recognition method and device, computer device, and storage medium
Patent Information
- Application Number
- CN202610883686.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-17
- Publication Date
- 2026-09-04
AI Technical Summary
[0004]本申请实施例的主要目的在于提出一种视觉位置识别方法和装置、计算机设备、存储介质,能够解决动态变化环境下视觉位置识别准确性下降的技术问题
[0009] The visual position recognition method, apparatus, computer device, and storage medium proposed in this application first acquire a target image and perform multi-scale feature extraction on the target image to obtain multiple initial feature maps; wherein any two initial feature maps have different scales; then, channel dimension unification is performed on each of the multiple initial feature maps to obtain a target feature map; next, each of the multiple target feature maps is used as the current feature map in turn, and attention fusion is performed based on the current feature map and all target feature maps to obtain a current attention feature map; the preset historical position memory matrix is updated according to the current attention feature map to obtain a current position memory matrix; then, feature fusion is performed according to the current position memory matrix and the current attention feature map to obtain a position enhancement feature map of the current feature map; then, feature fusion is performed on multiple position enhancement feature maps to obtain a position embedding vector; finally, position detection is performed based on the position embedding vector to obtain the target position; wherein, the target position is the shooting position corresponding to the target image. In this application, on the one hand, a current attention feature map is generated for each of the multiple target feature maps, enabling the automatic adjustment of the weights of features at different scales based on the environment of the target image, thereby effectively addressing local occlusion or viewpoint changes. On the other hand, a historical position memory matrix is introduced, and the current attention features are combined with the historical position matrix to generate a position-enhanced feature map, realizing the combination of multi-scale analysis of spatial features and temporal contextual information, thus achieving a comprehensive understanding of the dynamic environment. It is evident that this application achieves multi-scale feature analysis and temporal contextual modeling capabilities not found in traditional VPR systems, significantly improving the robustness of recognition in dynamically changing environments.
Smart Images

Figure CN122695211A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the fields of artificial intelligence, computer vision and other technologies, and is applicable to the financial and medical fields. In particular, it relates to a visual position recognition method and device, computer equipment and storage medium. Background Technology
[0002] Visual Place Recognition (VPR) primarily determines the shooting location by analyzing image content. For example, in compliance audits of bank branches in the financial sector, VPR technology is used to identify the location of branch photos to verify the placement of security facilities and advertisements. Another example is in medical equipment asset management in the healthcare sector, where VPR technology is used to identify the location of photos taken by cameras on equipment in corridors / wards to verify that the equipment is in its designated location and prevent misuse.
[0003] Related technologies primarily rely on manually designed features (such as Scale-Invariant Feature Transform (SIFT) and Speeded-Up Robust Features (SURF)) for location recognition. However, these methods are susceptible to changes in lighting, viewing angles, and dynamic occlusion in complex indoor environments. Related technologies have also proposed hierarchical feature extraction techniques using multi-scale feature pyramids or U-Net structures for visual location recognition. However, current multi-scale methods (such as Asymmetric Pyramid Aggregation Network (APANet)) employ fixed weights to fuse features from different scales, making it difficult to adapt to changes in feature importance caused by local occlusion in dynamic environments. Summary of the Invention
[0004] The main objective of this application is to provide a visual position recognition method, apparatus, computer device, and storage medium that can solve the technical problem of decreased accuracy of visual position recognition under dynamic changing environments.
[0005] To achieve the above objectives, a first aspect of this application proposes a visual position recognition method, the method comprising: A target image is acquired, and multi-scale feature extraction is performed on the target image to obtain multiple initial feature maps; wherein any two of the initial feature maps have different scales. The channel dimension of each of the multiple initial feature maps is unified to obtain the target feature map; Each of the multiple target feature maps is used as the current feature map in turn, and attention fusion is performed based on the current feature map and all the target feature maps to obtain the current attention feature map; The preset historical position memory matrix is updated based on the current attention feature map to obtain the current position memory matrix; Feature fusion is performed based on the current position memory matrix and the current attention feature map to obtain the position-enhanced feature map of the current feature map; The multiple location-enhanced feature maps are fused to obtain a location embedding vector; Position detection is performed based on the position embedding vector to obtain the target position; wherein, the target position is the shooting position corresponding to the target image.
[0006] To achieve the above objectives, a second aspect of this application provides a visual position recognition device, the device comprising: An image encoding module is used to acquire a target image and perform multi-scale feature extraction on the target image to obtain multiple initial feature maps; wherein any two of the initial feature maps have different scales; The feature projection module is used to unify the channel dimension of each of the multiple initial feature maps to obtain the target feature map; The attention fusion module is used to take each of the multiple target feature maps in turn as the current feature map, and perform attention fusion based on the current feature map and all the target feature maps to obtain the current attention feature map; The memory update module is used to update the preset historical position memory matrix according to the current attention feature map to obtain the current position memory matrix; The position enhancement module is used to perform feature fusion based on the current position memory matrix and the current attention feature map to obtain the position enhancement feature map of the current feature map; The feature fusion module is used to fuse multiple location-enhanced feature maps to obtain a location embedding vector; A position detection module is used to perform position detection based on the position embedding vector to obtain the target position; wherein, the target position is the shooting position corresponding to the target image.
[0007] To achieve the above objectives, a third aspect of this application provides a computer device, the computer device including a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the visual position recognition method described in the first aspect.
[0008] To achieve the above objectives, a fourth aspect of the present application provides a storage medium, which is a computer-readable storage medium storing a computer program that, when executed by a processor, implements the visual position recognition method described in the first aspect.
[0009] The visual position recognition method, apparatus, computer device, and storage medium proposed in this application first acquire a target image and perform multi-scale feature extraction on the target image to obtain multiple initial feature maps; wherein any two initial feature maps have different scales; then, channel dimension unification is performed on each of the multiple initial feature maps to obtain a target feature map; next, each of the multiple target feature maps is used as the current feature map in turn, and attention fusion is performed based on the current feature map and all target feature maps to obtain a current attention feature map; the preset historical position memory matrix is updated according to the current attention feature map to obtain a current position memory matrix; then, feature fusion is performed according to the current position memory matrix and the current attention feature map to obtain a position enhancement feature map of the current feature map; then, feature fusion is performed on multiple position enhancement feature maps to obtain a position embedding vector; finally, position detection is performed based on the position embedding vector to obtain the target position; wherein, the target position is the shooting position corresponding to the target image. In this application, on the one hand, a current attention feature map is generated for each of the multiple target feature maps, enabling the automatic adjustment of the weights of features at different scales based on the environment of the target image, thereby effectively addressing local occlusion or viewpoint changes. On the other hand, a historical position memory matrix is introduced, and the current attention features are combined with the historical position matrix to generate a position-enhanced feature map, realizing the combination of multi-scale analysis of spatial features and temporal contextual information, thus achieving a comprehensive understanding of the dynamic environment. It is evident that this application achieves multi-scale feature analysis and temporal contextual modeling capabilities not found in traditional VPR systems, significantly improving the robustness of recognition in dynamically changing environments.
[0010] Additional aspects and advantages of this application will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of this application. Attached Figure Description
[0011] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments of this application will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0012] Figure 1 This is a schematic diagram of an application environment for a visual position recognition method according to an embodiment of this application; Figure 2 This is a flowchart of a visual position recognition method provided in an embodiment of this application; Figure 3 yes Figure 2 The flowchart for step 203 in the document; Figure 4 yes Figure 2 The flowchart for step 204 in the document; Figure 5 yes Figure 2 The flowchart for step 205 in the document; Figure 6 This is a schematic diagram of the structure of a visual position recognition device in one embodiment of this application; Figure 7 This is a schematic diagram of the structure of a computer device according to one embodiment of this application; Figure 8 This is another structural schematic diagram of a computer device in one embodiment of this application. Detailed Implementation
[0013] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0014] It should be noted that although functional modules are divided in the device schematic diagram and a logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than the module division in the device or the order in the flowchart. The terms "first," "second," etc., in the specification, claims, and the aforementioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence.
[0015] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.
[0016] First, let's analyze some of the terms used in this application: Artificial intelligence (AI) is a new branch of computer science that studies, develops, and applies theories, methods, technologies, and systems to simulate, extend, and expand human intelligence. It aims to understand the essence of intelligence and produce intelligent machines that can react in a way similar to human intelligence. Research in this field includes robotics, speech recognition, image recognition, natural language processing, and expert systems. AI can simulate the information processes of human consciousness and thought. Furthermore, AI utilizes digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceiving the environment, acquiring knowledge, and using that knowledge to achieve optimal results.
[0017] Natural Language Processing (NLP): NLP uses computers to process, understand, and utilize human language (such as Chinese and English). It is a branch of artificial intelligence and an interdisciplinary field of computer science and linguistics, often referred to as computational linguistics. NLP includes syntactic analysis, semantic analysis, and discourse understanding. It is commonly used in machine translation, handwritten and printed character recognition, speech recognition and text-to-speech conversion, information and image processing, information extraction and filtering, text classification and clustering, sentiment analysis, and opinion mining. It involves data mining, machine learning, knowledge acquisition, knowledge engineering, artificial intelligence research, and linguistic research related to language computation.
[0018] Indoor visual scene recognition is one of the core problems in computer vision, aiming to determine the shooting location by analyzing image content. In related technologies, attention mechanisms have been introduced to enhance the model's ability to focus on key regions, such as self-attention in Transformers and spatial attention modules in Convolutional Neural Networks (CNNs). Meanwhile, to handle temporal data, Long Short-Term Memory (LSTM) networks and gated recurrent units are used to model the temporal continuity of scenes, while memory networks further extend the ability to preserve long-term context. Furthermore, cross-scale feature fusion methods and nonlocal neural networks improve the robustness of scene representation by aggregating multi-resolution information.
[0019] Despite the progress made in previous research, existing methods still have the following limitations: 1. Static feature fusion: Most multi-scale methods use fixed weights to fuse features at different scales, which makes it difficult to adapt to changes in feature importance caused by local occlusion in dynamic environments; 2. Short-term memory dependence: Temporal modeling based on long short-term memory networks only retains limited historical information, while semantically guided attention methods do not explicitly establish long-term memory associations across frames; 3. Context separation: Although existing hierarchical architectures can extract multi-scale features, they do not dynamically combine spatial hierarchy with temporal context, resulting in insufficient robustness to transient occlusion (such as pedestrian movement).
[0020] This application proposes an indoor visual location recognition method based on cross-scale recurrent context memory (RCM). The method enhances its robustness in occluded and dynamically changing environments by designing a cross-scale RCM network. The RCM network employs a hierarchical processing mechanism, combining multi-scale visual features with dynamically accumulated temporal context, and optimizes the location embedding vector through a recurrent attention mechanism. Specifically, the RCM network acts between the feature extraction backbone network and the similarity measurement network, transforming the raw convolutional features output by the backbone network into a discriminative context-aware representation. The RCM network includes a cross-scale attention mechanism and a gated memory update strategy. The former models the dependencies between multi-scale features, while the latter preserves historical location information through a learnable historical location memory matrix. This method effectively identifies transient changes in indoor environments caused by moving objects or occlusion, and integrates multi-scale context through max pooling to generate compact location embedding vectors. The method is compatible with mainstream CNN or visual Transformer backbone networks, is computationally efficient and easy to integrate, and significantly improves the accuracy and stability of indoor location recognition.
[0021] This method addresses location recognition in dynamically changing environments (such as occluded environments) by fusing multi-scale feature modeling with a temporal context memory mechanism to achieve robust location embedding learning. This application can be widely used in fields requiring real-time environmental understanding, such as intelligent monitoring systems and video risk control scene recognition. It maintains high-precision positioning capabilities even in densely populated areas or when some areas are temporarily obscured.
[0022] This application provides a visual position recognition method, a visual position recognition device, a computer device, and a computer-readable storage medium, which are specifically described through the following embodiments. First, the visual position recognition method in the embodiments of this application is described.
[0023] The visual position recognition method provided in this application can be applied to, for example, Figure 1In this application environment, the client communicates with the server via a network. The server first acquires the target image from the client and performs multi-scale feature extraction on the target image, obtaining multiple initial feature maps; the scales of any two initial feature maps are different. The server unifies the channel dimensions of each of the multiple initial feature maps to obtain the target feature map. The server uses each of the multiple target feature maps in turn as the current feature map and performs attention fusion based on the current feature map and all target feature maps to obtain the current attention feature map. The server updates the preset historical position memory matrix based on the current attention feature map to obtain the current position memory matrix. The server performs feature fusion based on the current position memory matrix and the current attention feature map to obtain the position enhancement feature map of the current feature map. The server performs feature fusion on multiple position enhancement feature maps to obtain the position embedding vector. The server performs position detection based on the position embedding vector to obtain the target position; the target position is the shooting position corresponding to the target image. In this application, on the one hand, a current attention feature map is generated for each of the multiple target feature maps, enabling automatic adjustment of the weights of features at different scales based on the current environment of the target image, thereby effectively addressing local occlusion or viewpoint changes. On the other hand, a historical position memory matrix is introduced, and the current attention features are combined with the historical position matrix to generate a position-enhanced feature map, realizing the combination of multi-scale analysis of spatial features and temporal contextual information, thus achieving a comprehensive understanding of the dynamic environment. It is evident that this application achieves multi-scale feature analysis and temporal contextual modeling capabilities not found in traditional VPR systems, significantly improving the robustness of recognition in dynamically changing environments. The client can be, but is not limited to, various personal computers, laptops, smartphones, tablets, and portable wearable devices. The server can be implemented using a standalone server or a server cluster consisting of multiple servers. The invention will be described in detail below through specific embodiments.
[0024] It should be noted that in each specific embodiment of this application, when it is necessary to process data related to the user's identity or characteristics, such as text data or voice data, the user's permission or consent will be obtained first. Moreover, the collection, use and processing of this data will comply with relevant laws, regulations and standards.
[0025] Reference Figure 2 , Figure 2 This is an optional flowchart of a visual position recognition method provided in an embodiment of this application, which may include, but is not limited to, steps 201 to 207.
[0026] Step 201: Acquire the target image and perform multi-scale feature extraction on the target image to obtain multiple initial feature maps. The scales of any two initial feature maps are different. The scale includes spatial dimension and channel dimension, and both the spatial and channel dimensions of any two initial feature maps are different. The spatial dimension includes width and height, and at least one of the width and height of any two initial feature maps is different.
[0027] A target image refers to an image for which location identification is performed. For example, in compliance audits of bank branches in the financial sector, the target image is a photograph of the branch. Visual location recognition technology can be used to identify the location of security facilities and advertisements in these photographs. Similarly, in medical equipment asset management in the healthcare sector, the target image is a photograph of a corridor / ward taken by a camera on the equipment. Visual location recognition technology can be used to identify the location of these photographs to verify that the equipment is in its designated location and prevent misuse.
[0028] In one embodiment, a pre-trained convolutional neural network (such as ResNet or EfficientNet) or a visual Transformer (such as Swin Transformer) may be used as the infrastructure to extract multi-scale features from the target image.
[0029] In one example, given a target image I, the pre-trained backbone network extracts L initial feature maps at different scales from the target image I. ,in ∈ Indicates the first An initial feature map, having spatial dimension × and channel dimensions .
[0030] Step 202: Unify the channel dimensions of each of the multiple initial feature maps to obtain the target feature map. Notably, any two target feature maps have the same channel dimensions.
[0031] For example, in order to unify feature representations at different scales, each initial feature map Projecting onto a shared latent space via 1×1 convolution: , in, Indicates learnable parameters, It is the first The initial feature map has a spatial dimension, D, which is a uniform channel dimension. After projection, the target feature map is obtained. .
[0032] Step 203: Take each of the multiple target feature maps in turn as the current feature map, and perform attention fusion based on the current feature map and all target feature maps to obtain the current attention feature map.
[0033] Specifically, L target feature maps Each target feature map in the sequence is used as the current feature map in turn, such as the first... The current feature map can also be represented as Utilizing a cross-scale attention mechanism on the current feature map and L target feature maps Perform attention fusion to obtain the first Current attention feature map .
[0034] In one embodiment, reference is made to Figure 3 Step 203 may include: Step 301: Project the current feature map using the first projection matrix to obtain the query projected features; Step 302: Concatenate all target feature maps along the spatial dimension to obtain the target concatenated feature map; Step 303: Project the target stitched feature map using the second projection matrix to obtain the key projection feature; Step 304: Project the target stitched feature map using the third projection matrix to obtain the value projection feature; Step 305: Calculate the weights based on the query projection features and key projection features to obtain the cross-attention weights; Step 306: Modulate the value projection features according to the cross-attention weights to obtain the current attention feature map.
[0035] In step 301, the query projection feature can be represented as: ,in, For the first The query projection features corresponding to each initial feature map. For the first Each target feature map ∈ D×D This is the first learnable projection matrix.
[0036] In step 302, the target splicing features can be represented as: Concat(·) means concatenating features across all scales along the spatial dimension. This represents the feature map of the k-th target.
[0037] In step 303, the key projection feature can be represented as: ,in, For the first The key projection features corresponding to the initial feature maps ∈ D×D This is the learnable second projection matrix.
[0038] In step 304, the value projection feature can be expressed as: ,in, For the first The values corresponding to the initial feature maps are projected features. ∈ D×D It is a learnable third projection matrix.
[0039] In step 305, the cross-attention weights can be expressed as: , For the first The cross-attention weights corresponding to the current feature map, where D is the uniform channel dimension and Softmax is the normalization function. It is the first The spatial dimension of the current feature map It is the first The spatial dimension of the target feature map.
[0040] In step 306, the current attention feature map can be represented as: = * ,in, For the first The current attention feature map corresponding to each current feature map.
[0041] The advantage of the above embodiments is that they enable features at each scale to adaptively focus on relevant information at other scales.
[0042] Step 204: Update the preset historical position memory matrix according to the current attention feature map to obtain the current position memory matrix.
[0043] A historical location memory matrix is used to store historical environmental information and can specifically include multiple reference location vectors. For example, a historical location memory matrix... ∈ D×N D is the dimension of the reference position vector, and N is the number of reference position vectors.
[0044] In one example, before capturing the target image, the robot observes the entire corridor in the indoor office environment (including posters on the left wall and potted plants on the right) and obtains multiple reference images. Reference position vectors are extracted and stored from these reference images to construct a historical position memory matrix. Then, the robot returns to the indoor office environment to capture the target image. This time, considering potential occlusions (such as pedestrians walking by) or changes in lighting, the robot updates the historical position memory matrix based on the current attention feature map of the target image, thus obtaining the current position memory matrix.
[0045] It should be noted that the specific processing procedure for extracting the reference position vector from the reference image can be referred to the processing procedure for extracting the position embedding vector from the target image.
[0046] In one embodiment, reference is made to Figure 4 Step 204 may include: Step 401: Concatenate the current attention feature map and the historical position memory matrix to obtain the initial hybrid memory matrix; Step 402: Modulate the initial hybrid memory matrix based on the current attention feature map to obtain the target hybrid memory matrix; Step 403: The current location memory matrix is obtained by fusing the historical location memory matrix and the target hybrid memory matrix.
[0047] In step 401, t represents the time step. The first time step at time t A current attention feature map, For historical location memory matrix, [·;·] is the feature concatenation operation.
[0048] In step 402, based on right Modulation can be performed to obtain the target hybrid memory matrix. This modulation process can be accomplished based on a cross-attention mechanism.
[0049] In one embodiment, step 402 may include: generating weights on the current attention feature map using a gating network to obtain a gating weight vector; performing feature perception on the initial hybrid memory matrix using a multilayer perceptron to obtain a memory perception matrix; and multiplying the memory perception matrix element-wise using the gating weight vector to obtain a target hybrid memory matrix.
[0050] Specifically, the target hybrid memory matrix can be represented as: ,in, Let be the gated weight vector at time step t. = * +b, The first time step at time t The gating weight matrix of a gating network, where b is the bias parameter of the gating network. Here, ⊙ represents the activation function, ⊙ denotes element-wise multiplication, and MLP stands for Multilayer Perceptron.
[0051] The advantage of the above embodiments is that gating memory updates can be achieved by combining gating networks and multilayer perceptrons, and the historical location memory matrix can be updated based on the environmental information corresponding to the target image, which helps to improve the robustness of location recognition.
[0052] In step 403, the current position memory matrix can be represented as: , The current position is the memory matrix, and LayerNorm is the layer normalization operation.
[0053] The advantage of the embodiments of steps 401 to 403 described above is that they can preserve the temporal contextual information of the environment, cope with dynamic changes in the environment, and help improve the robustness of location recognition.
[0054] Step 205: Perform feature fusion based on the current position memory matrix and the current attention feature map to obtain the position-enhanced feature map of the current feature map. This feature fusion process can be completed based on a cross-attention mechanism.
[0055] In one embodiment, reference is made to Figure 5 Step 205 may include: Step 501: Use the current attention feature map as the query input feature, and the current position memory matrix as the key input feature and value input feature; Step 502: Perform attention calculation on the query input features, key input features, and value input features to obtain memory attention features; Step 503: Fuse the current attention feature map and the memory attention feature map to obtain the location enhancement feature map.
[0056] In one embodiment, step 502 may include: obtaining a multi-head attention network, the multi-head attention network including at least two attention heads; performing attention calculations on query input features, key input features, and value input features using the attention heads to obtain individual memory attention features; and concatenating the individual memory attention features from all attention heads to obtain the memory attention features. Specifically, the memory attention features can be represented as: , This represents a multi-head attention network. To query input features, Features are input as both keys and values. For example, a multi-head attention network is set to have 8 heads. Attention is calculated through different attention heads, and then the individual memory attention features output by the 8 attention heads are concatenated as vectors to obtain the memory attention features.
[0057] In step 503, the location-enhanced feature map can be represented as: ,in, The first time step at time t Enhanced feature maps at each location.
[0058] The advantage of the above embodiments is that by combining multi-scale analysis of spatial features with contextual information in the temporal dimension, a comprehensive understanding of the dynamically changing environment is achieved.
[0059] Step 206: Perform feature fusion on multiple location-enhanced feature maps to obtain a location embedding vector.
[0060] In one embodiment, L location-enhanced feature maps are concatenated to obtain a location embedding vector.
[0061] In another embodiment, step 206 may include: flattening each of the multiple location enhancement feature maps to obtain a target location enhancement vector; and performing max pooling on the target location enhancement vector to obtain a location embedding vector.
[0062] Specifically, the position embedding vector can be represented as: Where Flatten is the matrix flattening operation. It contains multi-scale spatial information and temporal context, which can effectively represent the essential features of a location. The location embedding vector R is used for similarity calculation and location recognition.
[0063] Step 207: Perform position detection based on the position embedding vector to obtain the target position. The target position is the shooting position corresponding to the target image. Alternatively, the target position can be described as the location of the device that captured the target image. The target position can be obtained by mapping the position embedding vector through a fully connected layer.
[0064] In one embodiment, step 207 may include: obtaining candidate position embedding vectors of candidate positions from a reference database; calculating the similarity between the position embedding vectors and the candidate position embedding vectors to obtain position similarity; and filtering candidate positions based on position similarity to obtain the target position.
[0065] Steps 201 to 207 of this application, on the one hand, generate a current attention feature map for each of the multiple target feature maps, enabling automatic adjustment of the weights of features at different scales based on the current environment of the target image, thereby effectively addressing local occlusion or viewpoint changes; on the other hand, they introduce a historical position memory matrix and generate a position-enhanced feature map by combining the current attention features with the historical position matrix, realizing the combination of multi-scale analysis of spatial features with temporal contextual information, thus achieving a comprehensive understanding of the dynamic environment. It is evident that this achieves multi-scale feature analysis and temporal contextual modeling capabilities not possessed by traditional VPR systems, significantly improving the robustness of recognition in dynamically changing environments.
[0066] In one example, consider applying the visual position recognition method of this application to a position recognition scenario in an indoor office environment: 1. Initial situation: The robot observes the entire corridor, including the poster on the left wall and the potted plants on the right, obtains multiple reference images, and extracts and stores the historical location memory matrix from the reference images.
[0067] 2. Occlusion Scenario: A pedestrian walks by, obstructing 70% of the field of vision. Traditional methods may fail to identify the location, but this application is able to: match visible ceiling lights (coarse-grained features) with the environmental layout in memory; supplement the feature information of the occluded area using a historical location memory matrix; and correctly identify the location by combining currently visible features and historical memory.
[0068] 3. Change of perspective: The robot turns, and part of the poster moves out of its field of vision. This application is able to: activate relevant memories through the new perspective of the potted plant; use cross-scale attention to connect the partially visible poster with the complete poster in memory; and maintain the continuity of position recognition.
[0069] This example demonstrates how this application addresses the challenges posed by occlusion and changing viewpoints by combining multi-scale analysis and temporal memory.
[0070] In summary, the present application achieves at least the following beneficial effects: 1. Dynamic adaptation to environmental changes: Through a cross-scale attention mechanism, the weights of features at different scales can be automatically adjusted according to the current environment, thereby effectively responding to local occlusion or changes in viewpoint. 2. Long-term memory retention: The historical location memory matrix can store and update historical environmental information, overcoming the limitation of limited memory capacity in traditional recurrent neural networks. 3. Spatiotemporal joint modeling: By combining multi-scale analysis of spatial features with temporal contextual information, a comprehensive understanding of dynamically changing environments is achieved.
[0071] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0072] Please see Figure 6 One embodiment of this application also provides a visual position recognition device, which can implement the above-described visual position recognition method. Figure 6 The block diagram of the visual position recognition device provided in this application embodiment is shown. The device includes: an image encoding module 601, a feature projection module 602, an attention fusion module 603, a memory update module 604, a position enhancement module 605, a feature fusion module 606, and a position detection module 607. Detailed descriptions of each functional module are as follows: Image encoding module 601 is used to acquire a target image and perform multi-scale feature extraction on the target image to obtain multiple initial feature maps; wherein any two of the initial feature maps have different scales; The feature projection module 602 is used to unify the channel dimension of each of the multiple initial feature maps to obtain the target feature map; The attention fusion module 603 is used to take each of the multiple target feature maps in turn as the current feature map, and perform attention fusion based on the current feature map and all the target feature maps to obtain the current attention feature map; The memory update module 604 is used to update the preset historical position memory matrix according to the current attention feature map to obtain the current position memory matrix; The position enhancement module 605 is used to perform feature fusion based on the current position memory matrix and the current attention feature map to obtain the position enhancement feature map of the current feature map; Feature fusion module 606 is used to fuse multiple location-enhanced feature maps to obtain a location embedding vector; The position detection module 607 is used to perform position detection based on the position embedding vector to obtain the target position; wherein the target position is the shooting position corresponding to the target image.
[0073] In one embodiment, the attention fusion module 603 is specifically used for: The current feature map is projected using the first projection matrix to obtain the query projected features; All the target feature maps are stitched together along the spatial dimension to obtain the target stitched feature map; The target stitched feature map is projected using a second projection matrix to obtain key projection features; The target stitched feature map is projected using a third projection matrix to obtain the value projection feature; The cross-attention weights are obtained by calculating the weights based on the query projection features and the key projection features. The value projection features are modulated according to the cross-attention weights to obtain the current attention feature map.
[0074] In one embodiment, the memory update module 604 is specifically used for: The current attention feature map and the historical position memory matrix are concatenated to obtain an initial hybrid memory matrix; The initial hybrid memory matrix is modulated based on the current attention feature map to obtain the target hybrid memory matrix; The current location memory matrix is obtained by fusing the historical location memory matrix and the target hybrid memory matrix.
[0075] In one embodiment, the memory update module 604 is specifically used for: A gating network is used to generate weights on the current attention feature map to obtain a gating weight vector; The initial hybrid memory matrix is subjected to feature perception using a multilayer perceptron to obtain the memory perception matrix; The target hybrid memory matrix is obtained by multiplying the memory perception matrix element by element using the gating weight vector.
[0076] In one embodiment, the position enhancement module 605 is configured to: The current attention feature map is used as the query input feature, and the current position memory matrix is used as the key input feature and value input feature; Attention is calculated on the query input features, the key input features, and the value input features to obtain memory attention features; The location enhancement feature map is obtained by fusing the current attention feature map and the memory attention feature map.
[0077] In one embodiment, the position enhancement module 605 is configured to: Obtain a multi-head attention network, wherein the multi-head attention network includes at least two attention heads; The attention head is used to perform attention calculations on the query input features, the key input features, and the value input features to obtain individual memory attention features; The memory attention features are obtained by concatenating the individual memory attention features of all the attention heads.
[0078] In one embodiment, the feature fusion module 606 is used for: Flatten each of the multiple location enhancement feature maps to obtain the target location enhancement vector; The target position enhancement vector is subjected to max pooling to obtain the position embedding vector.
[0079] It should be noted that the specific implementation of this visual position recognition device is basically the same as the specific implementation of the visual position recognition method described above, and will not be repeated here.
[0080] This invention provides a visual position recognition device. On one hand, it generates a current attention feature map for each of multiple target feature maps, enabling automatic adjustment of the weights of features at different scales based on the current environment of the target image, thus effectively addressing local occlusion or viewpoint changes. On the other hand, it introduces a historical position memory matrix and generates a position-enhanced feature map by combining the current attention features with the historical position matrix. This achieves a combination of multi-scale spatial feature analysis and temporal contextual information, thereby enabling a comprehensive understanding of dynamic environments. It is evident that this device achieves multi-scale feature analysis and temporal contextual modeling capabilities not found in traditional VPR systems, significantly improving the robustness of recognition in dynamically changing environments.
[0081] Specific limitations regarding the visual position recognition device can be found in the limitations of the visual position recognition method described above, and will not be repeated here. Each module in the aforementioned visual position recognition device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device in hardware form, or stored in the memory of a computer device in software form, so that the processor can call and execute the operations corresponding to each module.
[0082] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 7 As shown, the computer device includes a processor, memory, network interface, and database connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile and / or volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The network interface is used to communicate with external clients via a network connection. When executed by the processor, the computer program implements the functions or steps of a visual position recognition method on the server side.
[0083] In one embodiment, a computer device is provided, which may be a client, and its internal structure diagram may be as follows: Figure 8As shown, the computer device includes a processor, memory, network interface, display screen, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The network interface is used to communicate with an external server via a network connection. When executed by the processor, the computer program implements the functions or steps of a visual position recognition method on the client side.
[0084] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to perform the following steps: A target image is acquired, and multi-scale feature extraction is performed on the target image to obtain multiple initial feature maps; wherein any two of the initial feature maps have different scales. The channel dimension of each of the multiple initial feature maps is unified to obtain the target feature map; Each of the multiple target feature maps is used as the current feature map in turn, and attention fusion is performed based on the current feature map and all the target feature maps to obtain the current attention feature map; The preset historical position memory matrix is updated based on the current attention feature map to obtain the current position memory matrix; Feature fusion is performed based on the current position memory matrix and the current attention feature map to obtain the position-enhanced feature map of the current feature map; The multiple location-enhanced feature maps are fused to obtain a location embedding vector; Position detection is performed based on the position embedding vector to obtain the target position; wherein, the target position is the shooting position corresponding to the target image.
[0085] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, the computer program performing the following steps when executed by a processor: A target image is acquired, and multi-scale feature extraction is performed on the target image to obtain multiple initial feature maps; wherein any two of the initial feature maps have different scales. The channel dimension of each of the multiple initial feature maps is unified to obtain the target feature map; Each of the multiple target feature maps is used as the current feature map in turn, and attention fusion is performed based on the current feature map and all the target feature maps to obtain the current attention feature map; The preset historical position memory matrix is updated based on the current attention feature map to obtain the current position memory matrix; Feature fusion is performed based on the current position memory matrix and the current attention feature map to obtain the position-enhanced feature map of the current feature map; The multiple location-enhanced feature maps are fused to obtain a location embedding vector; Position detection is performed based on the position embedding vector to obtain the target position; wherein, the target position is the shooting position corresponding to the target image.
[0086] It should be noted that the functions or steps that can be implemented by the computer-readable storage medium or computer device described above can be referred to the relevant descriptions on the server side and client side in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.
[0087] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0088] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0089] It should be noted that any AI models, software tools, or components not belonging to this company appearing in the embodiments of this application are merely illustrative examples and do not represent actual use. All user personal information involved in the embodiments of this application has been authorized (with the knowledge and consent) by the relevant parties or has been fully authorized by all parties, and the executing entity may obtain it through various legal and compliant means. The collection, storage, use, processing, transmission, provision, and disclosure of the information, data, and signals involved all comply with relevant laws and regulations and do not violate public order and good morals.
[0090] The embodiments described in this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided by the embodiments of this application. As those skilled in the art will know, with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of this application are also applicable to similar technical problems.
[0091] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of this application, and may include more or fewer steps than shown, or combine certain steps, or different steps.
[0092] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0093] Those skilled in the art will understand that all or some of the steps in the methods disclosed above, as well as the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, or suitable combinations thereof.
[0094] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0095] It should be understood that in this application, "at least one (item)" means one or more, and "more than" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.
[0096] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0097] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0098] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0099] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions to cause an electronic device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing programs, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0100] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.
Claims
1. A visual position recognition method, characterized in that, The method includes: A target image is acquired, and multi-scale feature extraction is performed on the target image to obtain multiple initial feature maps; wherein any two of the initial feature maps have different scales. The channel dimension of each of the multiple initial feature maps is unified to obtain the target feature map; Each of the multiple target feature maps is used as the current feature map in turn, and attention fusion is performed based on the current feature map and all the target feature maps to obtain the current attention feature map; The preset historical position memory matrix is updated based on the current attention feature map to obtain the current position memory matrix; Feature fusion is performed based on the current position memory matrix and the current attention feature map to obtain the position-enhanced feature map of the current feature map; The multiple location-enhanced feature maps are fused to obtain a location embedding vector; Position detection is performed based on the position embedding vector to obtain the target position; wherein, the target position is the shooting position corresponding to the target image.
2. The method according to claim 1, characterized in that, The step of performing attention fusion based on the current feature map and all the target feature maps to obtain the current attention feature map includes: The current feature map is projected using the first projection matrix to obtain the query projected features; All the target feature maps are stitched together along the spatial dimension to obtain the target stitched feature map; The target stitched feature map is projected using a second projection matrix to obtain key projection features; The target stitched feature map is projected using a third projection matrix to obtain the value projection feature; The cross-attention weights are obtained by calculating the weights based on the query projection features and the key projection features. The value projection features are modulated according to the cross-attention weights to obtain the current attention feature map.
3. The method according to claim 1, characterized in that, The step of updating the preset historical position memory matrix based on the current attention feature map to obtain the current position memory matrix includes: The current attention feature map and the historical position memory matrix are concatenated to obtain an initial hybrid memory matrix; The initial hybrid memory matrix is modulated based on the current attention feature map to obtain the target hybrid memory matrix; The current location memory matrix is obtained by fusing the historical location memory matrix and the target hybrid memory matrix.
4. The method according to claim 3, characterized in that, The process of modulating the initial hybrid memory matrix based on the current attention feature map to obtain the target hybrid memory matrix includes: A gating network is used to generate weights on the current attention feature map to obtain a gating weight vector; The initial hybrid memory matrix is subjected to feature perception using a multilayer perceptron to obtain the memory perception matrix; The target hybrid memory matrix is obtained by multiplying the memory perception matrix element by element using the gating weight vector.
5. The method according to any one of claims 1 to 4, characterized in that, The step of fusing features based on the current position memory matrix and the current attention feature map to obtain the position-enhanced feature map of the current feature map includes: The current attention feature map is used as the query input feature, and the current position memory matrix is used as the key input feature and value input feature; Attention is calculated on the query input features, the key input features, and the value input features to obtain memory attention features; The location enhancement feature map is obtained by fusing the current attention feature map and the memory attention feature map.
6. The method according to claim 5, characterized in that, The process of performing attention calculations on the query input features, the key input features, and the value input features to obtain memory attention features includes: Obtain a multi-head attention network, wherein the multi-head attention network includes at least two attention heads; The attention head is used to perform attention calculations on the query input features, the key input features, and the value input features to obtain individual memory attention features; The memory attention features are obtained by concatenating the individual memory attention features of all the attention heads.
7. The method according to any one of claims 1 to 4, characterized in that, The step of fusing features from multiple location-enhanced feature maps to obtain a location embedding vector includes: Flatten each of the multiple location enhancement feature maps to obtain the target location enhancement vector; The target position enhancement vector is subjected to max pooling to obtain the position embedding vector.
8. A visual position recognition device, characterized in that, The device includes: An image encoding module is used to acquire a target image and perform multi-scale feature extraction on the target image to obtain multiple initial feature maps; wherein any two of the initial feature maps have different scales; The feature projection module is used to unify the channel dimension of each of the multiple initial feature maps to obtain the target feature map; The attention fusion module is used to take each of the multiple target feature maps in turn as the current feature map, and perform attention fusion based on the current feature map and all the target feature maps to obtain the current attention feature map; The memory update module is used to update the preset historical position memory matrix according to the current attention feature map to obtain the current position memory matrix; The position enhancement module is used to perform feature fusion based on the current position memory matrix and the current attention feature map to obtain the position enhancement feature map of the current feature map; The feature fusion module is used to fuse multiple location-enhanced feature maps to obtain a location embedding vector; A position detection module is used to perform position detection based on the position embedding vector to obtain the target position; wherein, the target position is the shooting position corresponding to the target image.
9. A computer device, characterized in that, The computer device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the visual position recognition method according to any one of claims 1 to 7.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the visual position recognition method according to any one of claims 1 to 7.