Robot position recognition method, device, equipment, medium and product
Patent Information
- Application Number
- CN202611017442.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-09
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2046-07-09
AI Technical Summary
[0004]本申请提供一种机器人位置识别方法、装置、设备、介质及产品,用以解决相关技术中如何提高复杂场景下机器人位置识别的准确性的技术问题
[0015]本申请实施例提供的机器人位置识别方法、装置、设备、介质及产品,通过全局特征提取网络进行多尺度特征提取和跨图像各区域间进行特征交互,得到最终的全局描述符,可以使不同视角的对应场景结构能够自适应建立关联,有效克服了复杂场景下由局部遮挡,视角差异以及动态环境干扰带来的表观差异问题,打破了相关技术仅依赖同序号区域交互的局限,提高了在复杂环境中对机器人进行位置识别的鲁棒性和准确性。
Smart Images

Figure CN122530609B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of robotics technology, and more specifically, to a robot position recognition method, apparatus, device, medium, and product. Background Technology
[0002] With the development of robotics, visual sensing, and computing power, mobile robots, service robots, and related intelligent devices are increasingly being applied to indoor and outdoor navigation, inspection, delivery, and human-machine services. During autonomous movement, robots often need to perform position recognition. Position recognition refers to searching a database of images with similar appearances in a pre-built environmental map based on the images currently collected by the robot, in order to determine the robot's approximate location on the map. In practical applications, factors such as partial occlusion, changes in lighting, differences in viewing angle, and changes in the appearance of the environment can cause significant differences in appearance between the current image and the database image, thus affecting the accuracy of the position recognition results.
[0003] Therefore, improving the accuracy of robot position recognition in complex scenarios has become a pressing technical problem for the industry. Summary of the Invention
[0004] This application provides a robot position recognition method, apparatus, device, medium, and product to solve the technical problem of how to improve the accuracy of robot position recognition in complex scenarios in related technologies.
[0005] In a first aspect, this application provides a robot position recognition method, including: Obtain a set of images of the robot's current environment; The image set is input into a global feature extraction network to obtain a global descriptor corresponding to each image in the image set output by the global feature extraction network; The robot's position is determined based on the distance between the global descriptor of the latest image in the image set and the global descriptor of historical images in the database; The global feature extraction network includes a multi-scale feature extraction module, a similarity bias generation module, and a cross-image full-region interaction module. The multi-scale feature extraction module is used to extract the whole image features and region features of each image in the image set, and to concatenate the whole image features and region features of each image to obtain the first aggregated feature of the image set; The similarity bias generation module is used to generate an attention bias matrix based on the first aggregated feature. The attention bias matrix is used to characterize the similarity between different regions of different images in the image set. The cross-image full-region interaction module is used to perform feature interaction between regions of different images in the image set based on the first aggregated feature and the attention bias matrix, and generate a global descriptor for each image in the image set.
[0006] In some embodiments, the multi-scale feature extraction module includes a self-attention unit, a multi-scale generalized mean pooling unit, and a channel concatenation operation unit; extracting the whole-image features and region features of each image in the image set, and concatenating the whole-image features and region features of each image to obtain the first aggregated feature of the image set, including: The image set is input into the self-attention unit, and feature extraction is performed on each image in the image set based on the self-attention unit to obtain the whole image feature and spatial feature map; The spatial feature map is input into the multi-scale generalized mean pooling unit. Based on the multi-scale generalized mean pooling unit, the spatial feature map is divided into regions of different scales to obtain multiple region blocks. The multiple region blocks are then pooled to obtain the region features. The whole image features and the region features are input to the channel stitching operation unit. Based on the channel stitching operation unit, the whole image features and multiple region features of each image are stitched together to obtain the first aggregated feature.
[0007] In some embodiments, the similarity bias generation module includes a first dimension rearrangement unit, a first Euclidean norm unit, and a first matrix unit; generating an attention bias matrix based on the first aggregated feature includes: The first aggregated feature is linearly mapped and normalized to obtain the second aggregated feature; The second aggregated feature is input into the first dimension rearrangement unit, and the second aggregated feature is rearranged in dimension based on the first dimension rearrangement unit to obtain the third aggregated feature; The third aggregated feature is input into the first Euclidean norm unit, and the third aggregated feature is normalized by the second norm based on the first Euclidean norm unit to obtain the fourth aggregated feature. The fourth aggregated feature is input into the first matrix unit, and the attention bias matrix is output by performing matrix multiplication on the fourth aggregated feature and its transpose based on the first matrix unit.
[0008] In some embodiments, the cross-image full-region interaction module includes a first cross-image attention unit and a second cross-image attention unit; the step of performing feature interaction on regions between different images in the image set based on the first aggregated feature and the attention bias matrix to generate a global descriptor for each image in the image set includes: The fifth aggregate feature is obtained by rearranging the dimensions of the first aggregate feature; The fifth aggregated feature and the attention bias matrix are input into the first cross-image attention unit. Multi-head attention calculation is performed on the fifth aggregated feature based on the first cross-image attention unit, and the attention bias matrix is introduced in the multi-head attention calculation process to obtain the first attention output feature. The first attention output feature and the fifth aggregated feature are fused to obtain the sixth aggregated feature. The sixth aggregated feature and the attention bias matrix are input into the second cross-image attention unit. The multi-head attention calculation is performed on the sixth aggregated feature based on the second cross-image attention unit, and the attention bias matrix is introduced in the multi-head attention calculation process to obtain the second attention output feature. The second attention output feature and the sixth aggregated feature are fused to obtain the seventh aggregated feature. A global descriptor for each image in the image set is generated based on the seventh aggregation feature.
[0009] In some embodiments, the cross-image full-region interaction module further includes a third-dimensional rearrangement unit and a second Euclidean norm unit; generating a global descriptor for each image based on the seventh aggregated feature includes: The seventh aggregated feature is input into the third dimension rearrangement unit, and the seventh aggregated feature is rearranged in dimension based on the third dimension rearrangement unit to obtain the initial descriptor of each image; The initial descriptor is input into the second Euclidean norm unit, and the initial descriptor is normalized to L2 based on the second Euclidean norm unit to obtain the global descriptor of each image.
[0010] In some embodiments, the global feature extraction network further includes an image pair weak supervision module, which includes a fourth-dimensional rearrangement unit, a temperature scaling unit, and a label comparison unit; the training phase of the global feature extraction network includes the following steps: The sample attention bias matrix is input into the fourth dimension rearrangement unit. Based on the fourth dimension rearrangement unit, the sample attention bias matrix is rearranged in dimensions, and the part associated with the whole image features of the sample image set is extracted to obtain the global local correlation tensor. The global local correlation tensor is input to the temperature scaling unit. The temperature scaling unit takes the maximum value of the region dimension of the global local correlation tensor and performs scaling processing to obtain the image pair response matrix. The location identifier of each sample image in the sample image set is input into the label comparison unit, and a weak label matrix of image pairs is generated based on the comparison of location identifiers between different sample images by the label comparison unit. The parameters of the global feature extraction network are updated based on the loss value between the image pair response matrix and the image pair weak label matrix.
[0011] Secondly, this application provides a robot position recognition device, comprising: The acquisition module is used to acquire a set of images of the robot's current environment; The network module is used to input the image set into the global feature extraction network to obtain the global descriptor corresponding to each image in the image set output by the global feature extraction network; The recognition module is used to determine the position of the robot based on the distance between the global descriptor of the latest image in the image set and the global descriptor of the historical images in the database; The global feature extraction network includes a multi-scale feature extraction module, a similarity bias generation module, and a cross-image full-region interaction module. The multi-scale feature extraction module is used to extract the whole image features and region features of each image in the image set, and to concatenate the whole image features and region features of each image to obtain the first aggregated feature of the image set; The similarity bias generation module is used to generate an attention bias matrix based on the first aggregated feature. The attention bias matrix is used to characterize the similarity between different regions of different images in the image set. The cross-image full-region interaction module is used to perform feature interaction between regions of different images in the image set based on the first aggregated feature and the attention bias matrix, and generate a global descriptor for each image in the image set.
[0012] Thirdly, embodiments of this application provide an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor is configured to implement the above-described method when executing the program through the computer program.
[0013] Fourthly, embodiments of this application provide a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the above-described method.
[0014] Fifthly, embodiments of this application provide a computer program product, including a computer program that, when executed by a processor, implements the above-described method.
[0015] The robot position recognition method, apparatus, device, medium, and product provided in this application embodiment perform multi-scale feature extraction and feature interaction across different regions of an image through a global feature extraction network to obtain the final global descriptor. This enables the corresponding scene structures from different perspectives to adaptively establish associations, effectively overcoming the appearance differences caused by local occlusion, perspective differences, and dynamic environmental interference in complex scenes. It breaks through the limitation of related technologies that only rely on interaction of regions with the same sequence number, and improves the robustness and accuracy of robot position recognition in complex environments. Attached Figure Description
[0016] To more clearly illustrate the technical solutions in this application or related technologies, the drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0017] Figure 1 This is one of the flowcharts illustrating the robot position recognition method provided in the embodiments of this application.
[0018] Figure 2 This is a second schematic flowchart of the robot position recognition method provided in the embodiments of this application.
[0019] Figure 3 This is the third flowchart illustrating the robot position recognition method provided in this application embodiment.
[0020] Figure 4 This is the fourth flowchart illustrating the robot position recognition method provided in the embodiments of this application.
[0021] Figure 5 This is a schematic diagram of the robot position recognition device provided in an embodiment of this application.
[0022] Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0023] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below with reference to the accompanying drawings of the embodiments. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0024] It should be noted that the terms "first," "second," etc., used in this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or device that includes a series of steps or modules is not necessarily limited to those steps or modules explicitly listed, but may include other steps or modules not explicitly listed or inherent to such processes, methods, products, or devices.
[0025] Figure 1 This is one of the flowcharts illustrating the robot position recognition method provided in the embodiments of this application, such as... Figure 1 As shown, the method includes steps 110 to 130. These method steps are merely one possible implementation of this application.
[0026] Step 110: Obtain a set of images of the robot's current environment.
[0027] Specifically, in this embodiment, a robot refers to an intelligent device that has autonomous mobility, perceives the environment through visual sensors, and performs navigation or service tasks; the image set here refers to a set of photos taken within a continuous time period, and each image in the image set carries a corresponding timestamp.
[0028] Specifically, the robot can acquire a set of original color images of its surroundings through its onboard visual sensors, thereby obtaining an image set of the robot's current environment. In practical applications, the Kinect v2 can be used as the visual sensor.
[0029] Step 120: Input the image set into the global feature extraction network to obtain the global descriptor corresponding to each image in the image set output by the global feature extraction network; The global feature extraction network includes a multi-scale feature extraction module, a similarity bias generation module, and a cross-image full-region interaction module. The multi-scale feature extraction module is used to extract the whole image features and region features of each image in the image set, and concatenates the whole image features and region features of each image to obtain the first aggregated feature of the image set.
[0030] The similarity bias generation module is used to generate an attention bias matrix based on the first aggregated feature. The attention bias matrix is used to characterize the similarity between different regions of different images in the image set.
[0031] The cross-image full-region interaction module is used to perform feature interaction between regions of different images in the image set based on the first aggregated feature and the attention bias matrix, and generate a global descriptor for each image in the image set.
[0032] Specifically, in this embodiment, the image set can be processed by a global feature extraction network, which includes a multi-scale feature extraction module, a similarity bias generation module, and a cross-image full-region interaction module.
[0033] In one example, training samples can be input into a global feature extraction network to obtain global descriptors and image pair response predictions. Weak labels for image pairs are generated based on location identifiers, and the difference between the image pair response predictions and the weak labels is calculated according to the loss function. At the same time, constraints are imposed on the distance relationship of global descriptors. The network parameters are iteratively optimized by continuously updating the network parameters through backpropagation using an optimizer, and finally a trained global feature extraction network is obtained.
[0034] The training samples are a collection of sample images sampled from a public dataset based on real locations. Each location contains multiple sample images with the same location identifier. The loss function is a combination of a multi-similarity loss used to optimize the global descriptor distance distribution and an image pair weak supervision loss used to constrain the difference between the image pair response prediction and the image pair weak label.
[0035] Whole-image features refer to feature vectors obtained after extracting global information from an entire input image, used to summarize the global semantic content of the image; region features refer to feature vectors extracted after dividing the input image into local parts, reflecting the appearance or structural information of specific local patches of the image.
[0036] The image set can be input into the multi-scale feature extraction module of the global feature extraction network. The multi-scale feature extraction module extracts the whole image features of each image in the image set and divides each image into multiple regions to obtain multiple region features of each image.
[0037] In one example, an image can be divided into multiple parts at different scales to obtain regional features at different scales. The whole image features of the same image and these regional features are then combined to obtain the first aggregated feature of each image. The set of the first aggregated features of all images in the image set is the first aggregated feature of the image set.
[0038] The first aggregated feature of the image set is input into the similarity bias generation module. The similarity bias generation module performs feature mapping and inner product operation on the first aggregated feature of the image set to generate an attention bias matrix. The attention bias matrix can adaptively reflect the correlation between different regions of different images in the current image set.
[0039] The first aggregated feature and the attention bias matrix are input into the cross-image full-region interaction module. This module uses the attention bias matrix as prior guidance to trigger a cross-image feature interaction mechanism, enabling any region of one image to interact with every region of another image. This ensures that even with spatial displacement, semantically similar region features can interact with each other, outputting a global descriptor for each image in the image set. The global descriptor is a final vector form that highly summarizes environmental information and is easily compared directly.
[0040] For example, acquiring a set of raw color images of the surrounding environment through vision sensors installed on the robot itself. IS Image collection IS The pre-trained global feature extraction network is input and processed sequentially through its multi-scale feature extraction module, similarity bias generation module, and cross-image full-region interaction module. Finally, principal component analysis (PCA) is used for dimensionality reduction to obtain the image set. IS Images in I Corresponding global descriptor f q Here, PCA dimensionality reduction refers to a dimensionality reduction technique that uses orthogonal transformation to convert high-dimensional features into low-dimensional principal component features.
[0041] In related technologies, the regional features of a certain index in an image mainly interact with the regional features of the same index in other images. When the viewing angle, distance, or imaging range changes, stable regions with corresponding relationships in the same location image may appear in different spatial partitions. In addition, under conditions of occlusion, interference from dynamic objects, or changes in environmental appearance, some regional features may be missing, subject to noise interference, or significantly changed. Relying solely on the interaction between regional features of the same index is insufficient to fully capture the true and effective regional correspondences between images. Therefore, this embodiment associates each region between different images in the image set, enabling the regional features of a certain index in an image to interact with the regional features of each index in other images, thereby obtaining a more accurate global descriptor and improving the accuracy of location recognition.
[0042] Step 130: Determine the robot's position based on the distance between the global descriptor of the latest image in the image set and the global descriptor of historical images in the database.
[0043] Specifically, a database refers to a pre-built collection of environmental maps and related features of known locations. Historical images refer to images that have been pre-collected and labeled with location information, used as reference benchmarks for localization.
[0044] In practical applications, an environmental map database can be obtained, which stores a collection of historical images. This collection can be a set of images previously taken by the robot or a set of images collected through other means. The historical image collection includes multiple historical images.
[0045] Specifically, the global descriptor of the latest image currently acquired by the robot can be compared with the global descriptors of all historical images in the database.
[0046] All images in the image set of the robot's current environment belong to one batch. When calculating the global descriptor of each image in the historical image set, the historical image set can be divided into multiple batches according to the number of images in the image set of the robot's current environment. All historical images in each batch are then input into the trained global feature extraction network to obtain the global descriptor of each historical image. Here, the method of obtaining the global descriptor of the historical image is the same as the method of obtaining the global descriptor of the image in the image set, and will not be described in detail here.
[0047] For example, the historical image collection in the database includes DB 1. DB 2、...、 DB z These historical images, a collection of historical images { DB 1, DB 2,..., DB z Images from} are input into a pre-trained global feature extraction network in batches to obtain a global descriptor set { f d1 , f d2 ,..., f dz},in z The number of historical images in the historical image set. f d1 , f d2 ... f dz They are respectively DB 1. DB 2、...、 DB z The corresponding global descriptor. Some historical images may have had their global descriptors calculated before. These global descriptors are stored in the database. If the database stores the global descriptor of the historical image, it can be retrieved directly from the database.
[0048] Each image in the image set of the robot's current environment carries a timestamp. Therefore, the global descriptor of the latest image can be identified by the timestamp. The Euclidean distance between the global descriptor of the latest image and the global descriptors of each historical image in the database is calculated. The historical image corresponding to the global descriptor with the smallest distance is selected as the result of robot position recognition.
[0049] The robot position recognition method provided in this application uses a global feature extraction network to perform multi-scale feature extraction and feature interaction across different regions of an image to obtain the final global descriptor. This enables the corresponding scene structures from different perspectives to adaptively establish associations, effectively overcoming the appearance differences caused by local occlusion, perspective differences, and dynamic environmental interference in complex scenes. It breaks through the limitation of related technologies that rely solely on interaction between regions with the same sequence number, and improves the robustness and accuracy of robot position recognition in complex environments.
[0050] It should be noted that each implementation method of this application can be freely combined, rearranged, or executed individually, and does not need to rely on or depend on a fixed execution order.
[0051] In some embodiments, the multi-scale feature extraction module includes a self-attention unit, a multi-scale generalized mean pooling unit, and a channel concatenation operation unit; it extracts the whole-image features and region features of each image in the image set, and concatenates the whole-image features and region features of each image to obtain the first aggregated feature of the image set, including: The image set is input into the self-attention unit, and features are extracted from each image in the image set based on the self-attention unit to obtain the whole image features and spatial feature map; The spatial feature map is input into a multi-scale generalized mean pooling unit. Based on the multi-scale generalized mean pooling unit, the spatial feature map is divided into regions of different scales to obtain multiple region blocks. The multiple region blocks are then pooled to obtain region features. The whole image features and region features are input into the channel stitching operation unit. Based on the channel stitching operation unit, the whole image features and multiple region features of each image are stitched together to obtain the first aggregated feature.
[0052] Specifically, the self-attention unit is the TransformerBackbone unit of the adaptive visual self-attention mechanism. Images from the image set are input into the adaptive visual Transformer Backbone unit, which extracts overall image features and spatial feature maps. The spatial feature map refers to the local features extracted at each spatial location in the image. The spatial feature map is input into the Multi-Scale Generalized Mean Pooling (MSGeMPooling) unit. The MSGeMPooling unit divides the spatial feature map into different spatial scales, uniformly dividing it into 4 regions using a 2×2 grid and 9 regions using a 3×3 grid. Pooling is performed on the feature maps of these 13 regions, compressing the spatial features within each region into a single feature vector, thus outputting the regional features of each image in parallel. These features and the overall image features are concatenated by the channel concatenation unit to obtain the first aggregated feature.
[0053] Figure 2 This is a second schematic flowchart of the robot position recognition method provided in the embodiments of this application, as shown below. Figure 2 As shown, the dimension of the image set is ,in H and W These are the height and width of the image, respectively. B This represents the number of images in the image set. The original color image set is input into the multi-scale feature extraction module, and after being processed sequentially by the Adaptive Vision Transformer Backbone unit and the MSGeMPooling unit, a result of size [size missing] is obtained. B Full-image features of ×1×768 C 1. Regional characteristics R 1. Regional characteristics R 2. Regional characteristics R 3. Regional characteristics R 4. Regional characteristics R 5. Regional characteristics R 6. Regional characteristicsR 7. Regional characteristics R 8. Regional characteristics R 9. Regional characteristics R 10 Regional characteristics R 11 Regional characteristics R 12 and regional characteristics R 13 , C 1. R 1. R 2. R 3. R 4. R 5. R 6. R 7. R 8. R 9. R 10 , R 11 , R 12 and R 13 After channel splicing operation by the channel splicing operation unit, the resulting size is... B The first aggregation feature of ×14×768 A 1.
[0054] The robot position recognition method provided in this application embodiment achieves feature fusion of global image generalization and local details through multi-scale feature extraction and stitching, thereby improving the multi-scale expressive capability of the first aggregated feature.
[0055] In some embodiments, the similarity bias generation module includes a first dimension rearrangement unit, a first Euclidean norm unit, and a first matrix unit; generating an attention bias matrix based on the first aggregated feature includes: The first aggregated feature is linearly mapped and normalized to obtain the second aggregated feature; The second aggregated feature is input into the first dimension rearrangement unit, and the second aggregated feature is rearranged in dimension based on the first dimension rearrangement unit to obtain the third aggregated feature. The third aggregated feature is input into the first Euclidean norm unit, and the third aggregated feature is normalized by the second norm based on the first Euclidean norm unit to obtain the fourth aggregated feature. The fourth aggregated feature is input into the first matrix unit, and the fourth aggregated feature is multiplied by the transpose of the fourth aggregated feature based on the first matrix unit to output the attention bias matrix.
[0056] Specifically, linear mapping refers to the mathematical operation of transforming the feature space of the input features using a fully connected layer; normalization is the scaling of the vector along the feature dimension; dimensional rearrangement refers to the operation of changing the shape and size of a multidimensional tensor; L2 normalization is the calculation of the Euclidean length of the input vector along the feature dimension and the division of each element in the vector by this length, thereby scaling the overall scale of the feature vector to 1; attention bias matrix refers to the prior weight matrix that represents the correlation across different regions of the image and is used to break the interaction restrictions of regions with the same index.
[0057] The steps in this embodiment can be executed by a similarity bias generation module, which includes a first linear mapping unit Linear1, a first normalization unit LayerNorm1, a first dimension reshape unit Reshape1, a first Euclidean norm unit L2Norm1, and a first matrix unit MatMul1. Linear1 includes a fully connected layer, and the number of input and output channels of the fully connected layer can be set according to the dimension of the processed features. The normalization dimension of LayerNorm1 can be set according to the dimension of the processed features. Reshape1 is obtained using the reshape operation in the torch library. L2Norm1 is used to perform L2 norm normalization on the input vector along the feature dimension. The torch library refers to an open-source library for tensor computation.
[0058] Specifically, Figure 3 This is the third flowchart illustrating the robot position recognition method provided in the embodiments of this application, as shown below. Figure 3 As shown, the first aggregated feature can be sent to the similarity bias generation module, which then performs linear mapping through Linear1, normalization through LayerNorm1, merges or flattens the set and region dimensions through Reshape1, performs L2Norm1 for L2 normalization, and finally performs matrix multiplication through MatMul1 to output the attention bias matrix.
[0059] For example, Linear1 has 768 input channels and 768 output channels, and LayerNorm1 has a normalization dimension of 768. The first aggregated feature... A 1. The sample is fed into the similarity bias generation module, and then passes through Linear1 and LayerNorm1 in sequence to obtain a sample of size 1. B ×14×768 Second aggregation feature A 2. Second aggregation feature A 2. After dimensional rearrangement using Reshape1, the resulting size is 14. B ×768 Third Aggregation Feature A 3. Third aggregation feature A3. After L2Norm1 normalization, the size is 14. B ×768 Fourth Aggregation Feature A 4. The fourth aggregation feature A 4 and its transpose are input into MatMul1 to obtain a matrix of size 14. B ×14 B Attention bias matrix M 1.
[0060] The robot position recognition method provided in this application uses an attention bias matrix adaptively generated by matrix multiplication to represent the correlation between image regions, thus solving the limitation of interaction between regions with the same index in traditional methods.
[0061] In some embodiments, the cross-image full-region interaction module includes a first cross-image attention unit and a second cross-image attention unit; based on the first aggregated features and the attention bias matrix, it performs feature interaction on regions between different images in the image set to generate a global descriptor for each image in the image set, including: The fifth aggregate feature is obtained by rearranging the dimensions of the first aggregate feature; The fifth aggregated feature and the attention bias matrix are input into the first cross-image attention unit. Multi-head attention calculation is performed on the fifth aggregated feature based on the first cross-image attention unit. The attention bias matrix is introduced in the multi-head attention calculation process to obtain the first attention output feature. The first attention output feature and the fifth aggregated feature are fused to obtain the sixth aggregated feature. The sixth aggregated feature and the attention bias matrix are input into the second cross-image attention unit. Multi-head attention calculation is performed on the sixth aggregated feature based on the second cross-image attention unit. The attention bias matrix is introduced in the multi-head attention calculation process to obtain the second attention output feature. The second attention output feature and the sixth aggregated feature are fused to obtain the seventh aggregated feature. A global descriptor for each image in the image set is generated based on the seventh aggregated feature; The cross-image full-region interaction module also includes a third-dimensional rearrangement unit and a second Euclidean norm unit; it generates a global descriptor for each image based on the seventh aggregated feature, including; The seventh aggregated feature is input into the third dimension rearrangement unit, and the seventh aggregated feature is rearranged in dimension based on the third dimension rearrangement unit to obtain the initial descriptor of each image. The initial descriptor is input into the second Euclidean norm unit, and the initial descriptor is normalized by the second Euclidean norm unit to obtain the global descriptor for each image.
[0062] Specifically, multi-head attention computation is a feature fusion mechanism that generates query vectors, key vectors, and value vectors based on input features, maps the query vectors, key vectors, and value vectors to multiple independent feature subspaces, performs scaling dot product attention computation in each feature subspace, and then concatenates and linearly maps the attention outputs of each feature subspace. Through this mechanism, the model can simultaneously capture the complex correlations between input features from different representation subspaces.
[0063] The steps in this embodiment can be executed through a cross-image full-region interaction module, which includes a second-dimensional rearrangement unit Reshape2, a first cross-image attention (Transformer) unit Encoder1, a second cross-image Transformer unit Encoder2, a third-dimensional rearrangement unit Reshape3, and a second Euclidean norm unit L2Norm2.
[0064] Reshape2 and Reshape3 are obtained using the reshape operation in the torch library; L2Norm2 is used to normalize the input vector along the feature dimension using the L2 norm.
[0065] Encoder1 and Encoder2 have the same structure. l Indicates the first l Cross-image Transformer unit ( l =1,2), Encoder l Including the l Multi-head attention unit (MHA) across images l , No. (2 l -1) Residual Normalization Unit AddNorm 2l-1 , No. l Feedforward Network Unit (FFN) l and the 2nd l Residual Normalization Unit AddNorm 2l .
[0066] Specifically, MHA l Including the l Query the linear mapping layer Q_Linear l , No. l Key linear mapping layer K_Linear l , No. l Value linear mapping layer V_Linear l , No. l Query dimension reshape operation layer Reshape_Q l , No. l Key dimension rearrangement operation layer Reshape_K l, No. l Value dimension rearrangement operation layer Reshape_V l , No. l Scaling matrix multiplication layer MatMul_QK l , No. l Attention Bias Fusion Operation Layer AddBias l , No. l Normalization operation layer Softmax l , No. l Attention random deactivation (dropout) operation layer Dropout_Attn l , No. l Output Linear Mapping Layer (Out_Linear) l and the l Output dropout operation layer Dropout_Out l .
[0067] Among them, Q_Linear l K_Linear l V_Linear l and Out_Linear l Each includes a fully connected layer, with both the input and output feature dimensions of the fully connected layer being 768; Reshape_Q l Reshape_K l and Reshape_V l Obtained using the reshape operation in the torch library; MatMul_QK l After the matrix multiplication operation, divide by the scaling factor, which is the square root of 48; AddBias l Includes a softplus operation layer, a multiplication operation layer, and an addition operation layer; Dropout_Attn l and Dropout_Out l The dropout parameter is 0.1 for all values; AddNorm 2l-1 Includes a residual connection operation layer Add 2l-1 A layer normalization operation layer LayerNorm 2l ;FFN l Including feedforward linear layer FFN_Linear 2l-1 The Gaussian Error Linear Unit (GELU) activation layer GELU l , Feedforward dropout operation layer Dropout_FFN l and feedforward linear layer FFN_Linear2l AddNorm 2l Includes a residual connection operation layer Add 2l A layer normalization operation layer LayerNorm 2l+1 .
[0068] In one example, LayerNorm 2l and LayerNorm 2l+1 The normalized dimension is 768; FFN_Linear 2l-1 The number of input channels is 768 and the number of output channels is 2048. FFN_Linear 2l The number of input channels is 2048, and the number of output channels is 768. Dropout_FFN l The dropout parameter is 0.1. l Substituting 1 and 2 into the Encoder respectively l This will give you the structures of Encoder1 and Encoder2.
[0069] The first aggregation feature A 1 and attention bias matrix M 1. Input to the cross-image full-area interaction module. Specifically, Figure 4 This is the fourth flowchart illustrating the robot position recognition method provided in this application. Encoder1 and Encoder2 have the same structure; therefore, only the internal structure of Encoder1 is shown, as follows: Figure 4 As shown, the first aggregation feature A 1. Input into Reshape2, Reshape2 aggregates the first feature. A After dimensional rearrangement, the output is a fifth aggregated feature with a size of 14B×1×768. A 5. The fifth aggregation feature A 5 and attention bias matrix M 1. Input is fed into Encoder1 for processing, and the output size is 14. B The sixth aggregation feature of ×1×768 ; and the sixth aggregation feature and attention bias matrix M Input 1 is fed into Encoder2 for processing, and the output size is 14. B The seventh aggregation feature of ×1×768 .
[0070] Encoder l The processing procedure can be recorded as follows: ,in, l =1,2. Here, the fifth aggregation feature is used. A5 and attention bias matrix M The processing steps in Encoder1 will be explained using the sixth aggregated feature as an example. and attention bias matrix M The processing procedure in Encoder2 is the same as that of the fifth aggregated feature. A 5 and attention bias matrix M The same applies to Encoder1, so I won't go into details here.
[0071] Specifically, the size is 14 B The fifth aggregation feature of ×1×768 A 5 is sent to MHA1, where the fifth aggregation feature is located. A 5. Features are mapped through linear mapping layers Q_Linear1, K_Linear1, and V_Linear1, and each outputs a scale of 14. B ×1×768th one Query feature vector Q 1. No. one Key feature vector K 1 and the one Value eigenvector V 1, Will Q 1. K 1 and V 1. After being input into Reshape_Q1, Reshape_K1, and Reshape_V1 for dimensional rearrangement, the output dimensions are 1×16×14 respectively. B ×48 one Multi-head query characteristics MQF 1. Dimensions are 1×16×48×14 B First multi-headed key features MKF 1, and dimensions of 1×16×14 B The first multi-head value feature of ×48 MVF 1. Finally, MQF 1 and MKF 1. Input into MatMul_QK1, MatMul_QK1 will... MQF 1 and MKF After matrix multiplication and scaling, the output has a size of 1×16×14. B ×14 B First attention logits AL 1. AL 1 and attention bias matrix M 1. Enter into AddBias l In the torch library, a learnable parameter is initialized using the nn.Parameter function. The initial value is 2.0, and after applying the softplus operation, it is compared with the attention bias matrix. M Multiply by 1, then multiply by 1 AL Adding them together gives a size of 1×16×14. B ×14 B First attention matrix AM 1.
[0072] Will AM 1. Enter the information sequentially into Softmax. l After normalization and random deactivation using Dropout_Attn1, the resulting size is 1×16×14. B ×14 B First attention weight AW 1, AW 1 and MVF 1. Performing matrix multiplication yields a result of size 1×16×14. B ×48 First attentional features AH 1, AH 1. Using the reshape operation in the torch library, a size of 14 is obtained. B The first multi-head attention fusion feature of ×1×768 AF 1, AF 1. After passing through Out_Linear l Mapping and random deactivation using Dropout_Out1 result in a size of 14. B ×1×768 First attention output feature AO 1.
[0073] Will AO 1 and the fifth aggregation feature The input is fed into AddNorm1 for residual summation and normalization, resulting in a size of 14. B ×1×768 First attention residual features AR 1. AR 1. Input to FFN1, FFN1 corresponds to... AR After performing feature upscaling, activation, and dimensionality reduction, a size of 14 is obtained. B ×1×768th one Feedforward features FW 1. FW 1 and first attention residual features AR 1. Send to AddNorm2, AddNorm2 responds to... FW 1 and AR 1. After summing and normalizing the residuals, a size of 14 is obtained. B The sixth aggregation feature of ×1×768 .
[0074] Similarly, after the above iterations, Encoder2 finally outputs the seventh aggregated feature. .
[0075] To eliminate feature scale differences and redundant responses caused by multi-head mechanisms and multi-level iterations, this embodiment performs dimensional rearrangement and L2 norm constraint processing on the seventh aggregated feature to output an initial descriptor for each image. Then, the initial descriptor is processed to output the final global descriptor.
[0076] Specifically, the seventh aggregation feature The input is fed into Reshape3, which aggregates the seventh feature. Dimensional rearrangement is performed to output the initial descriptor for each image. The initial descriptor is then input into L2Norm2, which performs L2Norm2 normalization on the initial descriptor. Finally, the output is a set of global descriptors for each image with a size of B×10752. F g .
[0077] The robot position recognition method provided in this application extracts regional features from an image set, establishes regional associations between different images based on a cross-image attention mechanism, and guides cross-image regional feature interaction by combining an attention bias matrix, enabling regions with corresponding relationships in the same location image to undergo sufficient information fusion. After the interaction is completed, a global descriptor corresponding to each image in the image set is obtained, thereby reducing the impact of environmental occlusion and viewpoint changes on the position recognition results, improving the discrimination ability of the global descriptor and the accuracy of robot position recognition.
[0078] In some embodiments, the global feature extraction network further includes an image pair weak supervision module, which includes a fourth-dimensional rearrangement unit, a temperature scaling unit, and a label comparison unit; the training phase of the global feature extraction network includes the following steps: The sample attention bias matrix is input into the fourth dimension rearrangement unit. The sample attention bias matrix is rearranged in dimension based on the fourth dimension rearrangement unit, and the part associated with the whole image features of the sample image set is extracted to obtain the global local correlation tensor. The global local correlation tensor is input into the temperature scaling unit. The temperature scaling unit takes the maximum value of the global local correlation tensor in the region dimension and performs scaling to obtain the image pair response matrix. The location identifier of each sample image in the sample image set is input into the label comparison unit, and a weak label matrix of image pairs is generated based on the comparison of location identifiers between different sample images by the label comparison unit. The parameters of the global feature extraction network are updated based on the loss value between the image pair response matrix and the image pair weak label matrix.
[0079] Specifically, the sample image set is a pre-collected set of images with real-world location markers used to train the global feature extraction network.
[0080] The sample attention bias matrix is calculated during the training phase and is used to characterize the similarity between local regions of different images in the sample image set.
[0081] The global local correlation tensor is a multidimensional data array extracted based on the sample attention bias matrix, which specifically reflects the correlation strength between the whole image features of the sample image and the local region features of other images.
[0082] The image pair response matrix is obtained by extracting the maximum value of the region dimension and scaling the global local correlation tensor. It is used to quantitatively represent the comprehensive matching score between any two sample images.
[0083] The weak label matrix for images is constructed based on the true location identifiers of the sample images and is used to indicate the true classification answer matrix whether any two sample images belong to the same location.
[0084] During the training phase of the global feature extraction network, image-level location labels can only characterize whether two images belong to the same location, but cannot directly provide the region-level correspondence between the two images. Therefore, this embodiment introduces a weak supervision constraint mechanism at the image pair level through an image pair weak supervision module. Based on the image-level location labels, a supervision signal is constructed to constrain the overall matching relationship of image pairs. This allows the global feature extraction network to learn the effective region association relationships between images of the same location during training, thereby improving the discriminative ability of cross-image feature interaction results and global descriptors.
[0085] The image pair weak supervision module includes a fourth-dimensional rearrangement unit Reshape4, a temperature scaling unit Temp1, and a label comparison unit Compare1. Reshape4 is obtained using the reshape operation in the torch library. Temp1 includes a learnable temperature parameter, with an initial value preferably of 5.0. Compare1 compares the two input matrices element by element to see if they are equal. If they are equal, the corresponding position of the element is 1, and if they are not equal, it is 0. It is used to compare whether the location labels of any two images in the training set are the same.
[0086] In this embodiment, the image pair weak supervision module is only used during the network training phase to apply weak supervision constraints at the image pair level to the cross-image region association learning process based on image-level location labels. It does not participate in the feature extraction, descriptor generation, and similarity calculation processes during the network inference phase.
[0087] Specifically, the sample attention bias matrix is input into Reshape4. Reshape4 rearranges the dimensions of the sample attention bias matrix and extracts the slices that represent the similarity between the whole image features of one image and the features of each region of another image, and outputs the global local correlation tensor.
[0088] The global-local correlation tensor is input into Temp1. Temp1 performs max pooling on the dimension representing the local region of the tensor, that is, takes the maximum value to represent the overall response matching strength of the two images, and uses the temperature parameter to scale to control the numerical distribution, and outputs the image response matrix.
[0089] The sample image's actual location label is input into Compare1. Compare1 constructs a classification answer matrix that labels whether the locations are the same or not through matrix expansion and pairwise comparison, and outputs a weak label matrix for the image pairs.
[0090] The binary cross-entropy loss function is used to measure the difference between the image pair response matrix and the image pair weak label matrix, and the parameters are updated through backpropagation, which prompts the network to spontaneously learn to find stable regional associations between images of the same location.
[0091] For example, training batches contain B Each sample image, with a size of 14 B ×14 B The sample attention bias matrix is input into Reshape4. Reshape4 then rearranges the dimensions of the sample attention bias matrix to obtain a matrix of size [size missing]. B ×14× B The first correlation tensor of ×14 Q 1, and from the first relevant tensor Q In the second dimension of 1, the part with index 0 is selected, that is, the part associated with the features of the whole image, and the output size is... B × B ×14 global local correlation tensor Corr 1.
[0092] global and local correlation tensors Corr 1. Enter into Temp1, Temp1 is in Corr 1. Taking the maximum value of the last dimension yields the size. B × B The initial image to the response matrix P 1. And utilize learnable temperature parameters to... P After scaling, the output size is B × B Image response matrix P 2.
[0093] The size isB First location label of each sample image Label 1. Enter the value into Compare1. Compare1 will then... Label Add a new dimension to the second dimension of 1 and copy it. B The size obtained is B×B Second location label Label 2; in Label Add a new dimension to the first dimension of 1 and copy it. B The size obtained is B × B Third location tag Label 3; Label 2 and Label 3. After comparison using Compare1, the dimensions are obtained as follows: B × B Image pairs of weak label matrices Y , among which, when Y ij When =1, it indicates that the two images corresponding to the horizontal and vertical coordinates of this location correspond to the same location identifier. Y ij When =0, it indicates that the location corresponding to the horizontal and vertical coordinates of this position is different in the two images.
[0094] In this embodiment, the global feature extraction network uses the GSV-Cities dataset as training data. During training, training samples are constructed on a location-by-location basis, with four sample images randomly sampled from each location, and these four sample images are assigned the same location label. When using batch training, it is preferable that each training batch includes 72 locations, with each location including 4 sample images. B =72 4 = 288. It is a multiplication sign.
[0095] The training process of the global feature extraction network can use the Adam optimizer, with an initial learning rate of 0.0001, 10 training epochs, and a stepped learning rate decay strategy, where the learning rate decreases to 0.5 times the current learning rate every 3 generations. During training, ... B A training batch of sample images, consisting of a set of training sample images, is fed into the current global feature extraction network to obtain the global descriptor of the sample image set. F g The size is B ×10752; Image-to-response matrix P 2, dimensions are B × B Image pairs with weak label matrix Y The size is B × BFor each image, the loss function of the global feature extraction network is... L Including multi-similarity loss Image-based weak supervision loss As shown in the formula below: ;
[0096] ; in, Represents the feature similarity function. A global descriptor representing an image. A global descriptor representing a positive sample image. A global descriptor representing negative sample images. M This indicates the number of positive samples used in the loss calculation. N This indicates the number of negative samples used in the loss calculation. , and The adjustment parameters in the multi-similarity loss are preferably set to 1, 50, and 0, respectively. We employ a binary cross-entropy loss with logits. This represents the Sigmoid function. The weight coefficients, which represent the changes in weights over training epochs, are preferably set to... =0.005×(epoch+1), where epoch represents the current training epoch. This represents the index number of the image currently being processed, with values ranging from 1 to B. Indicates the index number of the positive sample image. The index number representing the negative sample image. and These represent the row and column indices in a matrix or tensor, respectively, both ranging from 1 to... B .
[0097] The robot position recognition method provided in this application generates weak labels for image pairs based on location identifiers, and applies weak supervision constraints at the image pair level during the training phase through the weak supervision module for image pairs. This guides the global image descriptor to pay more attention to the association of stable and visible regions in the same location image, reducing the impact of local occlusion or irrelevant interference regions on the position recognition results, thereby improving the robustness and accuracy of robot position recognition.
[0098] The robot position recognition device provided in the embodiments of this application is described below. The robot position recognition device described below can be referred to in correspondence with the robot position recognition method described above.
[0099] Figure 5This is a schematic diagram of the robot position recognition device provided in the embodiments of this application, as shown below. Figure 5 As shown, the device includes an acquisition module 510, a network module 520, and an identification module 530.
[0100] The acquisition module 510 is used to acquire a set of images of the robot's current environment; Network module 520 is used to input the image set into the global feature extraction network and obtain the global descriptor corresponding to each image in the image set output by the global feature extraction network; The recognition module 530 is used to determine the robot's position based on the distance between the global descriptor of the latest image in the image set and the global descriptor of the historical images in the database; The global feature extraction network includes a multi-scale feature extraction module, a similarity bias generation module, and a cross-image full-region interaction module. The multi-scale feature extraction module is used to extract the whole image features and region features of each image in the image set, and concatenates the whole image features and region features of each image to obtain the first aggregated feature of the image set; The similarity bias generation module is used to generate an attention bias matrix based on the first aggregated feature. The attention bias matrix is used to characterize the similarity between different regions of different images in the image set. The cross-image full-region interaction module is used to perform feature interaction between regions of different images in the image set based on the first aggregated feature and the attention bias matrix, and generate a global descriptor for each image in the image set.
[0101] Specifically, according to the embodiments of this application, any and multiple modules among the acquisition module 510, network module 520 and identification module 530 can be combined into one module, or any one of them can be split into multiple modules.
[0102] Alternatively, at least some of the functionality of one or more of these modules can be combined with at least some of the functionality of other modules and implemented in a single module.
[0103] According to embodiments of this application, at least one of the acquisition module 510, network module 520, and identification module 530 can be at least partially implemented as hardware circuitry, such as a Field Programmable Gate Array (FPGA), Programmable Logic Array (PLA), System-on-a-Chip, System-on-a-Substrate, System-on-Package, Application Specific Integrated Circuit (ASIC), or any other reasonable means of integrating or packaging the circuitry, or implemented in any one of the three methods of software, hardware, and firmware, or in a suitable combination of any of them.
[0104] Alternatively, at least one of the acquisition module 510, network module 520, and identification module 530 may be implemented at least partially as a computer program module, which can perform corresponding functions when the computer program module is run.
[0105] In some embodiments, the multi-scale feature extraction module includes a self-attention unit, a multi-scale generalized mean pooling unit, and a channel concatenation operation unit; it extracts the whole-image features and region features of each image in the image set, and concatenates the whole-image features and region features of each image to obtain the first aggregated feature of the image set, including: The image set is input into the self-attention unit, and features are extracted from each image in the image set based on the self-attention unit to obtain the whole image features and spatial feature map; The spatial feature map is input into a multi-scale generalized mean pooling unit. Based on the multi-scale generalized mean pooling unit, the spatial feature map is divided into regions of different scales to obtain multiple region blocks. The multiple region blocks are then pooled to obtain region features. The whole image features and region features are input into the channel stitching operation unit. Based on the channel stitching operation unit, the whole image features and multiple region features of each image are stitched together to obtain the first aggregated feature.
[0106] In some embodiments, the similarity bias generation module includes a first dimension rearrangement unit, a first Euclidean norm unit, and a first matrix unit; generating an attention bias matrix based on the first aggregated feature includes: The first aggregated feature is linearly mapped and normalized to obtain the second aggregated feature; The second aggregated feature is input into the first dimension rearrangement unit, and the second aggregated feature is rearranged in dimension based on the first dimension rearrangement unit to obtain the third aggregated feature. The third aggregated feature is input into the first Euclidean norm unit, and the third aggregated feature is normalized by the second norm based on the first Euclidean norm unit to obtain the fourth aggregated feature. The fourth aggregated feature is input into the first matrix unit, and the fourth aggregated feature is multiplied by the transpose of the fourth aggregated feature based on the first matrix unit to output the attention bias matrix.
[0107] In some embodiments, the cross-image full-region interaction module includes a first cross-image attention unit and a second cross-image attention unit; based on the first aggregated features and the attention bias matrix, it performs feature interaction on regions between different images in the image set to generate a global descriptor for each image in the image set, including: The fifth aggregate feature is obtained by rearranging the dimensions of the first aggregate feature; The fifth aggregated feature and the attention bias matrix are input into the first cross-image attention unit. Multi-head attention calculation is performed on the fifth aggregated feature based on the first cross-image attention unit. The attention bias matrix is introduced in the multi-head attention calculation process to obtain the first attention output feature. The first attention output feature and the fifth aggregated feature are fused to obtain the sixth aggregated feature. The sixth aggregated feature and the attention bias matrix are input into the second cross-image attention unit. Multi-head attention calculation is performed on the sixth aggregated feature based on the second cross-image attention unit. The attention bias matrix is introduced in the multi-head attention calculation process to obtain the second attention output feature. The second attention output feature and the sixth aggregated feature are fused to obtain the seventh aggregated feature. A global descriptor for each image in the image set is generated based on the seventh aggregated feature.
[0108] In some embodiments, the cross-image full-region interaction module further includes a third-dimensional rearrangement unit and a second Euclidean norm unit; and generates a global descriptor for each image based on the seventh aggregated feature, including; The seventh aggregated feature is input into the third dimension rearrangement unit, and the seventh aggregated feature is rearranged in dimension based on the third dimension rearrangement unit to obtain the initial descriptor of each image. The initial descriptor is input into the second Euclidean norm unit, and the initial descriptor is normalized by the second Euclidean norm unit to obtain the global descriptor for each image.
[0109] In some embodiments, the global feature extraction network further includes an image pair weak supervision module, which includes a fourth-dimensional rearrangement unit, a temperature scaling unit, and a label comparison unit; the training phase of the global feature extraction network includes the following steps: The sample attention bias matrix is input into the fourth dimension rearrangement unit. The sample attention bias matrix is rearranged in dimension based on the fourth dimension rearrangement unit, and the part associated with the whole image features of the sample image set is extracted to obtain the global local correlation tensor. The global local correlation tensor is input into the temperature scaling unit. The temperature scaling unit takes the maximum value of the global local correlation tensor in the region dimension and performs scaling to obtain the image pair response matrix. The location identifier of each sample image in the sample image set is input into the label comparison unit, and a weak label matrix of image pairs is generated based on the comparison of location identifiers between different sample images by the label comparison unit. The parameters of the global feature extraction network are updated based on the loss value between the image pair response matrix and the image pair weak label matrix.
[0110] It should be noted that the robot position recognition device provided in this application embodiment can implement all the method steps implemented in the above robot position recognition method embodiment and can achieve the same technical effect. Here, the parts that are the same as those in the method embodiment and the beneficial effects will not be described in detail.
[0111] Figure 6 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application, such as... Figure 6 As shown, the electronic device may include a processor 610, a communication interface 620, a memory 630, and a communication bus 640, wherein the processor 610, the communication interface 620, and the memory 630 communicate with each other through the communication bus 640. The processor 610 can call the computer program in the memory 630 to execute the above-described method.
[0112] Furthermore, the logical instructions in the aforementioned memory can be implemented as software functional modules and sold or used as independent products, and can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to related technologies, or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0113] On the other hand, this application also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can perform the methods provided in the above embodiments.
[0114] On the other hand, embodiments of this application also provide a non-transitory computer-readable storage medium storing a computer program for causing a processor to execute the methods provided in the above embodiments.
[0115] The non-transitory computer-readable storage medium can be any available medium or data storage device that the processor can access, including but not limited to magnetic memory (e.g., floppy disk, hard disk, magnetic tape, magneto-optical disk (MO)), optical memory (e.g., CD, DVD, BD, HVD), and semiconductor memory (e.g., ROM, EPROM, EEPROM, non-volatile memory (NAND FLASH), solid-state drive (SSD)).
[0116] The device embodiments described above are merely illustrative. The modules described as separate components may or may not be physically separate. The components shown as modules may or may not be physical modules; that is, they may be located in one place or distributed across multiple network modules. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0117] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the parts that contribute to the related technology, can be embodied in the form of software products. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0118] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.
Claims
1. A robot position recognition method, characterized in that, include: Obtain a set of images of the robot's current environment; The image set is input into a global feature extraction network to obtain a global descriptor corresponding to each image in the image set output by the global feature extraction network; The robot's position is determined based on the distance between the global descriptor of the latest image in the image set and the global descriptor of historical images in the database; The global feature extraction network includes a multi-scale feature extraction module, a similarity bias generation module, and a cross-image full-region interaction module. The multi-scale feature extraction module is used to extract the whole image features and region features of each image in the image set, and to concatenate the whole image features and region features of each image to obtain the first aggregated feature of the image set; The similarity bias generation module is used to generate an attention bias matrix based on the first aggregated feature. The attention bias matrix is used to characterize the similarity between different regions of different images in the image set. The cross-image full-region interaction module is used to perform feature interaction between regions of different images in the image set based on the first aggregated feature and the attention bias matrix, and generate a global descriptor for each image in the image set; The similarity bias generation module includes a first dimension rearrangement unit, a first Euclidean norm unit, and a first matrix unit; it generates an attention bias matrix based on the first aggregated feature, including: The first aggregated feature is linearly mapped and normalized to obtain the second aggregated feature; The second aggregated feature is input into the first dimension rearrangement unit, and the second aggregated feature is rearranged in dimension based on the first dimension rearrangement unit to obtain the third aggregated feature; The third aggregated feature is input into the first Euclidean norm unit, and the third aggregated feature is normalized by the second norm based on the first Euclidean norm unit to obtain the fourth aggregated feature. The fourth aggregated feature is input into the first matrix unit, and the attention bias matrix is output by performing matrix multiplication on the fourth aggregated feature and its transpose based on the first matrix unit.
2. The robot position recognition method according to claim 1, characterized in that, The multi-scale feature extraction module includes a self-attention unit, a multi-scale generalized mean pooling unit, and a channel concatenation operation unit; it extracts the whole-image features and region features of each image in the image set, and concatenates the whole-image features and region features of each image to obtain the first aggregated feature of the image set, including: The image set is input into the self-attention unit, and feature extraction is performed on each image in the image set based on the self-attention unit to obtain the whole image feature and spatial feature map; The spatial feature map is input into the multi-scale generalized mean pooling unit. Based on the multi-scale generalized mean pooling unit, the spatial feature map is divided into regions of different scales to obtain multiple region blocks. The multiple region blocks are then pooled to obtain the region features. The whole image features and the region features are input to the channel stitching operation unit. Based on the channel stitching operation unit, the whole image features and multiple region features of each image are stitched together to obtain the first aggregated feature.
3. The robot position recognition method according to claim 1, characterized in that, The cross-image full-region interaction module includes a first cross-image attention unit and a second cross-image attention unit; the step of performing feature interaction on regions between different images in the image set based on the first aggregated feature and the attention bias matrix to generate a global descriptor for each image in the image set includes: The fifth aggregate feature is obtained by rearranging the dimensions of the first aggregate feature; The fifth aggregated feature and the attention bias matrix are input into the first cross-image attention unit. Multi-head attention calculation is performed on the fifth aggregated feature based on the first cross-image attention unit, and the attention bias matrix is introduced in the multi-head attention calculation process to obtain the first attention output feature. The first attention output feature and the fifth aggregated feature are fused to obtain the sixth aggregated feature. The sixth aggregated feature and the attention bias matrix are input into the second cross-image attention unit. The multi-head attention calculation is performed on the sixth aggregated feature based on the second cross-image attention unit, and the attention bias matrix is introduced in the multi-head attention calculation process to obtain the second attention output feature. The second attention output feature and the sixth aggregated feature are fused to obtain the seventh aggregated feature. A global descriptor for each image in the image set is generated based on the seventh aggregation feature.
4. The robot position recognition method according to claim 3, characterized in that, The cross-image full-region interaction module further includes a third-dimensional rearrangement unit and a second Euclidean norm unit; it generates a global descriptor for each image in the image set based on the seventh aggregated feature, including: The seventh aggregated feature is input into the third dimension rearrangement unit, and the seventh aggregated feature is rearranged in dimension based on the third dimension rearrangement unit to obtain the initial descriptor of each image; The initial descriptor is input into the second Euclidean norm unit, and the initial descriptor is normalized to L2 based on the second Euclidean norm unit to obtain the global descriptor of each image.
5. The robot position recognition method according to any one of claims 1 to 4, characterized in that, The global feature extraction network further includes an image pair weakly supervised module, which comprises a fourth-dimensional rearrangement unit, a temperature scaling unit, and a label comparison unit; the training phase of the global feature extraction network includes the following steps: The sample attention bias matrix is input into the fourth dimension rearrangement unit. Based on the fourth dimension rearrangement unit, the sample attention bias matrix is rearranged in dimensions, and the part associated with the whole image features of the sample image set is extracted to obtain the global local correlation tensor. The global local correlation tensor is input to the temperature scaling unit. The temperature scaling unit takes the maximum value of the region dimension of the global local correlation tensor and performs scaling processing to obtain the image pair response matrix. The location identifier of each sample image in the sample image set is input into the label comparison unit, and a weak label matrix of image pairs is generated based on the comparison of location identifiers between different sample images by the label comparison unit. The parameters of the global feature extraction network are updated based on the loss value between the image pair response matrix and the image pair weak label matrix.
6. A robot position recognition device, characterized in that, include: The acquisition module is used to acquire a set of images of the robot's current environment; The network module is used to input the image set into the global feature extraction network to obtain the global descriptor corresponding to each image in the image set output by the global feature extraction network; The recognition module is used to determine the position of the robot based on the distance between the global descriptor of the latest image in the image set and the global descriptor of the historical images in the database; The global feature extraction network includes a multi-scale feature extraction module, a similarity bias generation module, and a cross-image full-region interaction module. The multi-scale feature extraction module is used to extract the whole image features and region features of each image in the image set, and to concatenate the whole image features and region features of each image to obtain the first aggregated feature of the image set; The similarity bias generation module is used to generate an attention bias matrix based on the first aggregated feature. The attention bias matrix is used to characterize the similarity between different regions of different images in the image set. The cross-image full-region interaction module is used to perform feature interaction between regions of different images in the image set based on the first aggregated feature and the attention bias matrix, and generate a global descriptor for each image in the image set; The similarity bias generation module includes a first dimension rearrangement unit, a first Euclidean norm unit, and a first matrix unit; it generates an attention bias matrix based on the first aggregated feature, including: The first aggregated feature is linearly mapped and normalized to obtain the second aggregated feature; The second aggregated feature is input into the first dimension rearrangement unit, and the second aggregated feature is rearranged in dimension based on the first dimension rearrangement unit to obtain the third aggregated feature; The third aggregated feature is input into the first Euclidean norm unit, and the third aggregated feature is normalized by the second norm based on the first Euclidean norm unit to obtain the fourth aggregated feature. The fourth aggregated feature is input into the first matrix unit, and the attention bias matrix is output by performing matrix multiplication on the fourth aggregated feature and its transpose based on the first matrix unit.
7. An electronic device comprising a memory and a processor, characterized in that, The memory stores a computer program, and the processor is configured to execute the robot position recognition method according to any one of claims 1 to 5 through the computer program.
8. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the robot position recognition method as described in any one of claims 1 to 5.
9. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the robot position recognition method as described in any one of claims 1 to 5.
Citation Information
Patent Citations
Underground space position identification method based on multi-head attention feature enhancement
CN117576203A
Robot position identification method and device, electronic equipment and storage medium
CN118505790A