Visual navigation method, system and medium based on fault-tolerant target positioning

Through cluster analysis and sparse self-attention mechanism, different categories of objects with similar appearance to the target type can be identified and reduced, and the target positioning fault tolerance and navigation efficiency of the agent are improved.

CN119832391BActive Publication Date: 2025-05-20SHANDONG UNIV +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510300945.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-14
Publication Date
2025-05-20
Estimated Expiration
2045-03-14

AI Technical Summary

Technical Problem

When existing visual navigation technologies face different categories of objects that look similar to the target type, it is difficult to identify their differences and avoid interference, resulting in insufficient navigation efficiency or navigation to the wrong target.

Method used

Clustering comprehensive comparison and analysis of potential target object information, and reduce the interference of confusing targets to the real target through sparse self-attention mechanism, and improve the target positioning fault tolerance of the agent.

Benefits of technology

It improves the navigation success rate and navigation efficiency of the agent, effectively reducing the interference and misleading of non-target objects on positioning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119832391B_ABST
    Figure CN119832391B_ABST
Patent Text Reader

Abstract

The present application belongs to the field of visual navigation technology, and specifically relates to a visual navigation method, system and medium based on fault-tolerant target positioning, including: using an appearance and position similarity calculation module to generate the appearance and position similarity between all stored target detection records; using a clustering module to perform cluster analysis on detection records of different similarities; based on the clustering results, using a sparse self-attention module to generate a masked attention matrix and update the detection records; using a target cross-attention module, using the embedded representation of the target object as a query vector to perform attention operations on the updated detection records; and finally using a strategy module to generate specific intelligent body movement or steering actions. The model disclosed by the present invention improves the navigation success rate and navigation efficiency of the intelligent body by reducing the interference and misleading of non-target category objects on the positioning of real targets.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of visual navigation technology, and specifically relates to a visual navigation method, system and medium based on fault-tolerant target positioning. Background Technology

[0002] Visual navigation technology is a method that uses real-time visual input from an agent or device to achieve autonomous navigation. It analyzes image data from the first-person perspective, identifies environmental features and target objects, and guides the agent to reach the predetermined target accurately and efficiently. This technology is widely used in fields such as robots, unmanned vehicles, and augmented reality. However, in actual scenarios, it is often easy to have confusing objects that are similar to the target object but of different types, which interfere with and mislead the positioning of the target. Therefore, fault-tolerant target positioning capabilities need to be taken into consideration in the design of navigation models.

[0003] There are currently two main methods for target positioning of visual navigation agents: time domain convolutional network type and attention mechanism type. The former uses multi-scale convolution kernels to perform convolution operations on the stored target detection records, which is suitable for fixed-length detection record memory buffer pools. The latter uses a multi-head attention mechanism, which can comprehensively analyze detection records of different capacity sizes and refine the directional information of the generated target. However, the existing methods do not consider how to identify the difference between objects of different categories but similar in appearance to the target type in the scene and avoid interference of such objects in target positioning. This results in insufficient navigation efficiency of the agent or even navigation to the wrong target. SUMMARY OF THE INVENTION

[0004] In response to the above technical problems, the present invention provides a visual navigation method based on fault-tolerant target positioning, which uses clustering to comprehensively compare and analyze the potential target object information detected during the navigation process, and uses a sparse self-attention mechanism to avoid the interference of confusing targets on the real target, thereby improving the target positioning fault tolerance of the intelligent body, and improving the navigation efficiency and success rate.

[0005] To achieve the above purpose, the technical solution of the present invention is as follows:

[0006] A visual navigation method based on fault-tolerant target positioning includes the following steps:

[0007] S1. Appearance and position similarity calculation: Generate the appearance and position similarity between all target detection records stored up to the current moment, generate the query vector, key vector and value vector as well as the initial attention matrix;

[0008] S2. Clustering: Use the query vector and key vector generated in step S1 to calculate the distance between the detection records and perform cluster analysis on these records;​

[0009] S3. Sparse self-attention calculation: Use the clustering results obtained in step S2 to perform a masking operation on the initial attention matrix generated in step S1 to obtain an updated matrix, and perform sparse self-attention calculation on the value vector obtained in step S1;

[0010] S4. Target cross-attention calculation: Use the embedding representation of the target object as the query vector to perform cross-attention calculation on the sparse self-attention result output in step S3, so as to reduce or eliminate the influence of similar objects that are prone to confusion on locating the real target;

[0011] S5. Action generation: Based on a recurrent neural network and a fully connected neural network, generate specific agent movement or turning actions.

[0012] Preferably, in step S1, the target object detection records collected by the agent through the target detector during navigation are used , to obtain a position query vector and a position key vector , :

[0013] ;

[0014] ;

[0015] Among them, is L2 norm normalization; is the planar coordinates of the target detection box of the i-th detected target object in the image; includes the relative spatial position coordinates of the agent and the rotation and pitch angles of the camera when the i-th target object is detected; is a learnable parameter matrix; , are the query vector and the key vector for calculating the position similarity respectively;

[0016] Use the target object detection records collected by the agent during navigation to obtain an appearance query vector and an appearance key vector , :

[0017] ;

[0018] ;

[0019] Among them, is the appearance feature of the i-th detected target object; is a learnable parameter matrix; , They are respectively a query vector and a key vector for calculating appearance similarity;

[0020] Calculate the initial attention matrix :

[0021] ;

[0022] where is the sigmoid activation function; is a learnable parameter for balancing position and appearance similarity, and T is the transpose.

[0023] Preferably, in step S1, use the target object detection records collected by the agent during navigation to obtain a value vector for comprehensively refining the target information :

[0024] ;

[0025] where is the appearance feature of the i-th detected target object; is a learnable parameter matrix; is a value vector that synthesizes the position and appearance information of the i-th detected object.

[0026] Preferably, the clustering step in step S2 is as follows:

[0027] S21. Calculate the distance between any two target object detection records ;

[0028] ;

[0029] S22. Apply the DBSCAN algorithm to cluster all detection records under the set hyperparameters of the maximum neighbor distance Epsilon and the minimum neighborhood number MinPts.

[0030] Preferably, the sparse self-attention calculation step in step S3 is as follows:

[0031] S31. Perform a masking operation on the initial attention matrix to obtain an updated attention matrix :

[0032] ;

[0033] S32. Use the masked attention matrix and the value vector of the detection record to perform sparse self-attention calculation to obtain an updated detection record :

[0034] ;

[0035] Among them is used to convert the input into a probability distribution, ensuring that each element of the output is non - negative and the sum is 1; refers to the dimensionality size of the key vector; is the updated detection record.

[0036] Preferably, for the target cross - attention calculation in step S4:

[0037] Use the target - category semantic vector obtained from the Glove dataset to query and perform cross - attention calculation on the updated detection record so as to reduce or even eliminate the influence of easily confused objects with low similarity to the true target, and obtain a directional vector that is not easily interfered by false detections :

[0038] ;

[0039] Among them is a learnable parameter mapping matrix; is the confidence vector of the detected target object; is the normalized attention weight vector; is the vector concatenation operation; is the number of heads in the multi - head attention calculation; is the directional vector that finally plays the function of fault - tolerant target positioning.

[0040] Preferably, the action generation in step S5 is as follows:

[0041] S51. Use the global image features extracted by the multi - layer convolutional network and the action feature vector executed by the agent at the previous moment , combined with the directional vector , to generate the state vector of the agent at the current moment :

[0042] ;

[0043] Among them, is the fully - connected layer; is the Layernorm normalization layer;

[0044] S52. Generate the action that the agent will execute next:

[0045] ;

[0046] ;

[0047] Among them is a long short-term memory recurrent neural network; is the hidden state and cell state vector of the agent at the previous moment; is the output state, hidden state and cell state vector at the current moment; is a multi-layer perceptron network; is the action vector to be executed in the next step.

[0048] A vision navigation system based on fault-tolerant target localization is used to implement the vision navigation method based on fault-tolerant target localization described in this application.

[0049] A storage medium, when running on a computer, executes the steps in the vision navigation method based on fault-tolerant target localization described in this application.

[0050] Compared with the prior art, the beneficial effects of this application are as follows:

[0051] The model disclosed by the present invention improves the navigation success rate and navigation efficiency of the agent by reducing the interference and misguidance of non-target category objects on the localization of real targets. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] Figure 1 is an overall schematic diagram of a vision navigation method based on fault-tolerant target localization disclosed in an embodiment of the present invention;

[0053] Figure 2 is a schematic diagram of a sparse self-attention calculation module in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0054] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0055] Embodiment 1: The present invention provides a vision navigation method based on fault-tolerant target localization, as Figure 1 shown. The model improves the fault-tolerant ability of the agent's target localization and the navigation efficiency and success rate by using clustering to comprehensively compare and analyze the potential target object information detected during the navigation process and using the sparse self-attention mechanism to avoid the interference of confused targets on real targets.

[0056] A vision navigation method based on fault-tolerant target localization includes the following steps:

[0057] (1) Appearance and position similarity calculation: Generate the appearance and position similarities between all the object detection records stored up to the current moment, and generate query vectors, key vectors, value vectors, and an initial attention matrix.

[0058] The specific method of step (1) is as follows:

[0059] (1.1) Use the object detection records collected by the agent through the object detector during navigation , to obtain the position query vector and position key vector for calculating the clustering distance in the subsequent step (2.1) , :

[0060] ;

[0061] ;

[0062] Among them, is L2 norm normalization; are the planar coordinates of the object detection box of the i-th detected object in the image; contains the relative spatial position coordinates of the agent and the rotation and pitch angles of the camera when the i-th object is detected; is a learnable parameter matrix; , are respectively used for calculating the position similarity.

[0063] (1.2) Use the object detection records collected by the agent during navigation to obtain the appearance query vector and appearance key vector for calculating the clustering distance in the subsequent step (2.1) , :

[0064] ;

[0065] ;

[0066] Among them, is the appearance feature of the i-th detected object; is a learnable parameter matrix; , are respectively the query vector and key vector for calculating the appearance similarity.

[0067] (1.3) Use the object detection records collected by the agent during navigation to obtain the value vector for comprehensively refining the object information :

[0068] ;

[0069] Among them, is the appearance feature of the i-th detected target object; is a learnable parameter matrix; is a value vector that synthesizes the position and appearance information of the i-th detected object.

[0070] (1.4) Use the query vector and key vector pairs obtained in (1.1) and (1.2) to calculate the initial attention matrix :

[0071] ;

[0072] Where is the sigmoid activation function; is a learnable parameter used to balance the position and appearance similarities.

[0073] (2) Clustering: Calculate the distances between detection records using the query vectors and key vectors generated in step (1), and perform clustering analysis on these records using the DBSCAN algorithm;

[0074] The specific method of step (2) is as follows:

[0075] (2.1) Use the calculation generated in (1) to calculate the distance between any two target object detection records :

[0076] ;

[0077] Where is the sigmoid activation function; is consistent with that in step (1.4).

[0078] (2.2) Use the distances obtained in (2.1) , and apply the DBSCAN algorithm to cluster all detection records under the settings of the maximum neighbor distance Epsilon and the minimum neighborhood number MinPts.

[0079] (3) Sparse self-attention calculation: Use the clustering results obtained in step (2) to perform a masking operation on the initial attention matrix generated in step (1), and perform sparse self-attention calculation on the value vector obtained in step (1); as Figure 2 shown, similar objects belonging to different categories can be distinguished and separated in the sparse attention matrix, thus avoiding the interference of false targets on real targets.

[0080] The specific method of step (3) is as follows:

[0081] (3.1) According to the clustering results obtained in (2.2), perform a masking operation on the initial attention matrix in (1.4) to obtain an updated attention matrix. :

[0082] .

[0083] (3.2) Use the masked attention matrix in (3.1) to perform sparse self-attention calculation on the detection record value vector obtained in (1.3) to obtain an updated detection record. :

[0084] ;

[0085] where is used to convert the input into a probability distribution, ensuring that each element of the output is non-negative and the sum is 1; refers to the dimension size of the key vector; is the updated detection record.

[0086] (4) Target cross-attention calculation: Use the embedding representation of the target object as the query vector to perform cross-attention calculation on the sparse self-attention result output in step (3), so as to reduce or eliminate the influence of similar objects that are prone to confusion on locating the real target.

[0087] The specific method of step (4) includes:

[0088] Use the target category semantic vector obtained from the Glove dataset for query, and perform cross-attention calculation on the updated detection record obtained in step (3.2) so as to reduce or even remove the influence of confusing objects with low similarity to the real target, and obtain a directional vector that is not easily interfered by false detections. :

[0089] ;

[0090] where is a learnable parameter mapping matrix; is the confidence vector of the detected target object; is the normalized attention weight vector; is the vector concatenation operation; is the number of heads in the multi-head attention calculation; is the directional vector that finally plays the function of fault-tolerant target positioning.

[0091] (5) Action generation: Based on the recurrent neural network and the fully connected neural network, generate specific agent movement or turning actions.

[0092] The specific method of step (5) is as follows:

[0093] (5.1) Use the global image features extracted by the multi-layer convolutional network and the action feature vector executed by the agent at the previous moment , combined with the directional vector obtained in step (4) , to generate the state vector of the agent at the current moment :

[0094] ;

[0095] Among them, is the fully connected layer; is the Layernorm normalization layer.

[0096] (5.2) Use the state vector obtained in (5.1) and the historical state information of the agent to generate the action that the agent will execute next:

[0097] ;

[0098] ;

[0099] Among them is the long short-term memory recurrent neural network; is the hidden state and cell state vector of the agent at the previous moment; is the output state, hidden state and cell state vector at the current moment; is the multi-layer perceptron network; is the action vector that will be executed next.

[0100] Embodiment 2: The present invention also provides an embodiment of a fault-tolerant target positioning visual navigation system, which adopts the navigation method described in the above embodiment. The navigation system includes:

[0101] (1) Appearance and position similarity calculation module: Generate the appearance and position similarity between all target detection records stored up to the current moment, and generate a query vector, a key vector, a value vector, and an initial attention matrix;

[0102] (2) Clustering module: Calculate the distance between detection records using the query vector and key vector generated in step (1), and perform clustering analysis on these records using the DBSCAN algorithm;

[0103] (3) Sparse self-attention calculation module: Use the clustering results obtained in step (2) to perform a masking operation on the initial attention matrix generated in step (1), and obtain an updated matrix to perform sparse self-attention calculation on the value vectors obtained in step (1). Figure 2 That is, it is a schematic diagram of the target feature generation module in the embodiment of the present invention;

[0104] (4) Target cross-attention calculation module: Use the embedding representation of the target object as the query vector to perform cross-attention calculation on the sparse self-attention results output in step (4), so as to reduce or eliminate the influence of similar objects that are prone to confusion on locating the true target;

[0105] (5) Action generation module: Based on a recurrent neural network and a fully connected neural network, generate specific agent movement or turning actions.

[0106] The method embodiments described above are only illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative work.

[0107] The system is constructed to run the method of this application. Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the above technical solution, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0108] To evaluate the effectiveness of the model proposed in the present invention, we designed the following experiments: The experimental environment is based on the Ai2Thor and RoboThor platforms. In Ai2Thor, there are a total of 30 different rooms, of which 20 rooms are used to train the model, 5 rooms are used to verify the model performance, and the remaining 5 rooms are used for the final test. On the RoboThor platform, there are 75 floors. We selected 60 of them as the training set, 5 floors for the verification stage, and another 10 floors for test evaluation. In this way, we can comprehensively test the actual application effect of the model.

[0109] The experimental results use SR and SPL to evaluate the performance of the model, which are the most commonly used evaluation metrics in visual navigation; to reflect the navigation ability of the agent in large scenes and long paths, we conducted experiments according to the optimal path from the starting position to the target end point. The results are shown in Tables 1 and 2.

[0110] TAMSA and NTWA in Tables 1 and 2 are relatively commonly used target localization methods in end-to-end visual navigation. From the results in the tables, it can be seen that the model of the present invention is superior to the other two methods in the ability to navigate to unseen objects, fully verifying the effectiveness of the model proposed by the present invention.

[0111] Table 1 Experimental Results of Ai2Thor

[0112] 。

[0113] Table 2 Experimental Results of RoboThor

[0114] 。

[0115] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but will be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A visual navigation method based on fault-tolerant target positioning, characterized in that: The following steps are involved: S1. Appearance and position similarity calculation: Generate the appearance and position similarity between all target detection records stored up to the current moment, generate the query vector, key vector and value vector as well as the initial attention matrix; S2. Clustering: Use the query vector and key vector generated in step S1 to calculate the distance between the detection records and perform cluster analysis on the detection records; S3. Sparse self-attention calculation: Use the clustering result obtained in step S2 to perform a mask operation on the initial attention matrix generated in step S1 to obtain an updated matrix, and perform sparse self-attention calculation on the value vector obtained in step S1; S4. Target cross attention calculation: Use the embedded representation of the target object as the query vector and perform cross attention calculation on the sparse self-attention result output in step S3; S5. Action generation: Generate specific agent movement or turning actions based on recursive neural networks and fully connected neural networks.

2. The visual navigation method based on fault-tolerant target positioning according to claim 1 is characterized in that: Step S1 uses the target object detection records collected by the agent during navigation through the target detector , get the position query vector and position key vector , : ; ; in, is L2 norm normalization; is the plane coordinate of the target detection box of the i-th target object detected in the image; Contains the relative spatial position coordinates of the agent when the i-th target object is detected, as well as the rotation and pitch angles of the camera; is a learnable parameter matrix; , are the query vector and key vector used to calculate location similarity, respectively; Use the target object detection records collected by the agent during navigation to obtain the appearance query vector and appearance key vector , : ; ; in, is the appearance feature of the i-th target object detected; is a learnable parameter matrix; , They are the query vector and key vector used to calculate the appearance similarity respectively; Calculate the initial attention matrix : ; in is the sigmoid activation function; Learnable parameters for balancing location and appearance similarity.

3. The visual navigation method based on fault-tolerant target positioning according to claim 1, characterized in that: In step S1, the target object detection records collected by the agent during navigation are used to obtain a value vector for comprehensive and refined target information : ; in, is the appearance feature of the i-th target object detected; is a learnable parameter matrix; It is a value vector that integrates the position and appearance information of the i-th detected object.

4. The visual navigation method based on fault-tolerant target positioning according to claim 1, characterized in that: The clustering steps in step S2 are as follows: S21. Calculate the distance between any two target object detection records ; ; S22. Under the set maximum neighbor distance Epsilon and minimum number of neighbors MinPts hyperparameters, the DBSCAN algorithm is applied to cluster all detection records.

5. The visual navigation method based on fault-tolerant target positioning according to claim 1, characterized in that: The sparse self-attention calculation steps in step S3 are as follows: S31. Perform mask operation on the initial attention matrix to obtain the updated attention matrix : ; S32. Use the masked attention matrix to detect the recorded value vector Perform sparse self-attention calculation to obtain updated detection records : ; in Used to convert the input into a probability distribution, ensuring that each element of the output is non-negative and the sum is 1; Refers to the dimension size of the key vector; This is the updated test record.

6. The visual navigation method based on fault-tolerant target positioning according to claim 1, characterized in that: Step S4 target cross attention calculation: Use the target category semantic vector obtained from the Glove dataset Query the updated test records Perform cross-attention calculations to reduce or even remove the influence of confusing objects with low similarity to the real target, and obtain a directional vector that is not easily disturbed by false detection. : ; ; in is the learnable parameter mapping matrix; is the confidence vector of the detected target object; is the normalized attention weight vector; is a vector concatenation operation; is the number of heads in the multi-head attention calculation; It is the directional vector that ultimately performs the fault-tolerant target positioning function.

7. The visual navigation method based on fault-tolerant target positioning according to claim 1, characterized in that: Step S5: Action generation steps are as follows: S51. Global features of images extracted using multi-layer convolutional networks and the action feature vector performed by the agent at the last moment , combined with the directivity vector , generates the state vector of the agent at the current moment : ; in, is a fully connected layer; It is the Layernorm normalization layer; S52. Generate the next action that the agent will perform: ; ; in is a long short-term memory recurrent neural network; is the hidden state and cell state vector of the agent at the previous moment; is the output state, hidden state and cell state vector at the current moment; is a multi-layer perceptron network; is the action vector to be executed next.

8. A visual navigation system based on fault-tolerant target positioning, characterized in that: Used to implement the visual navigation method based on fault-tolerant target positioning as described in any one of claims 1-7.

9. A storage medium, characterized in that: When the storage medium is run on a computer, the steps in the visual navigation method based on fault-tolerant target positioning according to any one of claims 1 to 7 are executed.

Citation Information

Patent Citations

  • Self-adaptive target navigation method and system for service robot

    CN114460943A

  • Visual tracking method based on attention adaptive selection Transform

    CN116485839A