Geospatial brain-like navigation route planning methods, devices, equipment and storage media
By extracting typical scene features from navigation task data using a brain-like scene recognition model, the problem of inaccurate route planning in visual navigation algorithms under visual interference is solved, and accurate navigation is achieved in low light or in bad weather.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BEIJING NORMAL UNIVERSITY
- Filing Date
- 2025-09-03
- Publication Date
- 2026-05-26
AI Technical Summary
Existing visual navigation algorithms suffer from inaccurate route planning due to visual interference such as insufficient lighting or inclement weather, where noisy image information leads to problems.
Typical scene features of each image in the navigation task dataset are extracted by a brain-like scene recognition model. The similarity between the feature map extracted by the brain-like neurons and the human attention map is less than a preset similarity threshold. The typical scene features are then input into the brain-like behavior decision model to obtain the target route.
This improves the model's robustness under various visual disturbances and ensures the correctness of navigation decisions.
Smart Images

Figure CN121384009B_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of artificial intelligence technology, specifically relating to a geospatial brain-like navigation route planning method, device, equipment, and storage medium. Background Technology
[0002] Currently, in order to overcome the reliance on auxiliary information such as satellite signals, visual navigation algorithms are generally used to plan corresponding routes based on the collected image data and all the information in the image data.
[0003] However, due to environmental factors such as insufficient lighting and severe weather, various visual interferences can occur, resulting in noise in some image information. Consequently, routes planned based on all the information in the image data may be inaccurate. Summary of the Invention
[0004] This application proposes a geospatial brain-like navigation route planning method, device, equipment, and storage medium, which can solve the technical problem that various visual interferences caused by insufficient lighting, bad weather, etc., result in noise in some image information, thus making the route planned based on all the information in the image data inaccurate.
[0005] The first aspect of this application proposes a geospatial brain-like navigation route planning method, including:
[0006] Typical scene features of each image data in the navigation task dataset are extracted by a brain-like scene recognition model. The feature maps extracted by the multiple brain-like neurons in the brain-like scene recognition model have a similarity of less than a preset similarity threshold with human attention maps.
[0007] The typical scene features of each image data are input into the brain-like behavior decision model to obtain the target route corresponding to the navigation task dataset.
[0008] An embodiment of the second aspect of this application provides a geospatial brain-like navigation route planning device, comprising:
[0009] The extraction module is used to extract typical scene features of each image data in the navigation task dataset through a brain-like scene recognition model. The feature maps extracted by the multiple brain-like neurons in the brain-like scene recognition model have a similarity of less than a preset similarity threshold with the human attention map.
[0010] The input module is used to input the typical scene features of each image data into the brain-like behavior decision model to obtain the target route corresponding to the navigation task dataset.
[0011] An embodiment of the third aspect of this application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the method described in the first aspect above.
[0012] An embodiment of the fourth aspect of this application provides a computer-readable storage medium having a computer program stored thereon, the program being executed by a processor to implement the method described in the first aspect above.
[0013] The technical solutions provided in this application embodiment have at least the following technical effects or advantages:
[0014] This application proposes a geospatial brain-like navigation route planning method, apparatus, device, and storage medium, the method comprising:
[0015] Typical scene features of each image in the navigation task dataset are extracted using a brain-like scene recognition model. The feature maps extracted by the multiple brain-like neurons in the brain-like scene recognition model have a similarity of less than a preset similarity threshold with human attention maps. The typical scene features of each image are then input into a brain-like behavior decision model to obtain the target route corresponding to the navigation task dataset. This embodiment of the application extracts typical scene features of each image using a brain-like scene recognition model comprising multiple brain-like neurons. Since typical scene features conform to human visual attention, the number of typical scene features is small, and they are all crucial local details for navigation decision-making. Therefore, while ensuring the correctness of navigation decisions, the robustness of the model under various visual interferences can be improved.
[0016] Additional aspects and advantages of this application will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of this application. Attached Figure Description
[0017] Various other advantages and benefits will become apparent to those skilled in the art upon reading the following detailed description of preferred embodiments. The accompanying drawings are for illustrative purposes only and are not intended to limit the scope of this application. Furthermore, the same reference numerals denote the same parts throughout the drawings. In the drawings:
[0018] Figure 1 This invention provides a flowchart of a geospatial brain-like navigation route planning method according to an embodiment of the present application.
[0019] Figure 2 This illustration shows a schematic diagram of the structure of a ResNet network provided in one embodiment of this application;
[0020] Figure 3This illustration shows a schematic diagram of the structure of a geospatial brain-like navigation route planning device according to an embodiment of this application;
[0021] Figure 4 This illustration shows a schematic diagram of the structure of an electronic device according to an embodiment of this application;
[0022] Figure 5 A schematic diagram of a storage medium provided in one embodiment of this application is shown. Detailed Implementation
[0023] Exemplary embodiments of this application will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of this application are shown in the drawings, it should be understood that this application may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided to enable a more thorough understanding of this application and to fully convey the scope of this application to those skilled in the art.
[0024] It should be noted that, unless otherwise stated, the technical or scientific terms used in this application shall have the ordinary meaning as understood by one of ordinary skill in the art to which this application pertains.
[0025] In recent years, with the rapid development of fields such as autonomous driving and smart cities, the demand for intelligent navigation technologies with strong autonomy, efficiency, and environmental adaptability has become increasingly prominent. Currently, mainstream navigation technologies mainly rely on auxiliary information such as satellite signals to achieve geolocation, direction estimation, and path planning. Without auxiliary information such as satellite signals, navigation functions cannot be achieved.
[0026] To avoid navigation failures, a pure visual navigation algorithm has been proposed. This algorithm provides rich visual information, minimizing reliance on external devices while incorporating deep learning and reinforcement learning algorithms. It can autonomously recognize environmental information such as color, depth, brightness, and semantics, achieving accurate environmental perception, location positioning, and navigation. This makes the pure visual navigation algorithm applicable to various indoor and outdoor geographical environments, especially in scenarios lacking satellite signals. To date, visual navigation algorithms have undergone several development stages. Early visual navigation models extracted manually designed visual features, primarily suitable for navigation tasks in static environments. The rise of feature extraction and image matching algorithms in the 1980s enabled models to achieve preliminary dynamic 3D environmental perception. More recently, the rise of convolutional neural networks has significantly improved the accuracy of visual feature extraction, leading to the emergence of end-to-end learning models that directly learn from visual image input to navigation behavior decisions.
[0027] Despite significant progress in visual navigation algorithms, these models suffer from slow computation speed and poor interpretability. Furthermore, the complexity of visual image scenes, variations in viewpoint, lighting, weather, and seasonal changes contribute to problems such as insufficient model adaptability and poor robustness.
[0028] To address the aforementioned technical problems, this application proposes a geospatial brain-like navigation route planning method, apparatus, device, and storage medium. The method includes: extracting typical scene features from each image data in a navigation task dataset using a brain-like scene recognition model, wherein the similarity between the feature maps extracted by the multiple brain-like neurons in the brain-like scene recognition model and human attention maps is less than a preset similarity threshold; and inputting the typical scene features of each image data into a brain-like behavior decision model to obtain the target route corresponding to the navigation task dataset. This application's embodiment extracts typical scene features from each image data using a brain-like scene recognition model comprising multiple brain-like neurons. Since typical scene features conform to human visual attention, the number of typical scene features is small, and they are all crucial local detail features for navigation decision-making. Therefore, it can improve the robustness of the model under various visual interferences while ensuring the correctness of navigation decisions.
[0029] The following description, in conjunction with the accompanying drawings, describes a geospatial brain-like navigation route planning method, apparatus, device, and storage medium according to embodiments of this application. This application uses an electronic device as the execution subject to illustrate the above-mentioned geospatial brain-like navigation route planning method.
[0030] See Figure 1 The method specifically includes the following steps:
[0031] S101. Extract typical scene features of each image data in the navigation task dataset using a brain-like scene recognition model.
[0032] The feature maps extracted by the multiple brain-like neurons in the brain-like scene recognition model have a similarity of less than a preset similarity threshold with human attention maps.
[0033] In stark contrast to the challenges of machine navigation, biological navigation (such as bird migration and human spatial cognition) demonstrates advantages such as centimeter-level accuracy, strong generalization, and extremely low energy consumption. Research shows that humans construct topological memories by extracting salient environmental features (such as landmarks) and dynamically integrating multimodal sensory inputs (visual, vestibular, etc.), thereby significantly reducing reliance on a single sensor.
[0034] Therefore, neurons related to human visual attention can be combined to obtain a brain-like scene recognition model, which can then extract local detail features that are crucial for navigation decisions.
[0035] The navigation task dataset typically includes multiple image datasets. Each image dataset is obtained by navigating to different locations. Image datasets can be collected at preset intervals or at preset distances. The specific image dataset collection method can be flexibly set based on actual conditions, and will not be elaborated further here. The preset intervals and preset distances can also be flexibly set based on actual conditions.
[0036] The typical scene features are those that humans visually pay attention to. For example, if the image data includes roads, road signs, and trees, the features that humans visually pay attention to are those corresponding to trees. Correspondingly, the feature maps extracted by multiple brain-like neurons based on this image are similar to human attention maps, both being feature maps where tree features are prominent.
[0037] The similarity threshold can be flexibly set based on the actual situation, such as 90%, 95%, etc.
[0038] In some embodiments, the human attention map can be obtained based on the location of human visual attention on the image data.
[0039] The location corresponding to human visual attention can be obtained in the following way:
[0040] Focus on eye movement metrics:
[0041] Fixation point and fixation duration: Fixation point is the location where the eyes rest, and fixation duration is the length of time the eyes remain at that location. A longer fixation duration indicates a higher level of interest and attention from the participant to that area. For example, in tourism landscape studies, participants tend to fixate longer on areas such as text, main buildings, and sculptural faces, indicating that these are the focal points of their visual attention.
[0042] First fixation time: This refers to the time it takes for the eyes to first fixate on a particular area. A shorter first fixation time may indicate that the area is more attractive to the subject and can quickly attract their attention.
[0043] Fixation count: This indicates the number of times the eyes fixate on a particular area. A higher fixation count indicates that the area attracts the subject's repeated attention.
[0044] Saccade speed and direction: Saccades are the rapid movement of the eyes from one fixation point to another. The direction and speed of saccades can reflect a subject's attentional shifting strategies for different areas.
[0045] Define the areas of interest:
[0046] Based on the research objectives and experimental materials, the visual scene was divided into different Areas of Interest (AOIs). Then, eye movement metrics of the participants in each AOI, such as average fixation duration, total fixation time, and number of fixations, were analyzed to determine the distribution of their visual attention. For example, in a study of map reading, the map was divided into different geographical feature regions, and the participants' eye movement data in these regions were analyzed to construct a visual attention model.
[0047] Analyzing eye movement trajectories:
[0048] Trajectory shape and direction: The shape (such as straight line, curve, spiral, etc.) and direction of eye movement trajectories can reflect the subject's focus of attention and visual search strategies. For example, when a subject focuses on a specific area, the eye trajectory may appear as a spiral or circular movement around that area.
[0049] Trajectory complexity: Complex eye-tracking trajectories usually indicate that the subject is multitasking or repeatedly comparing and searching multiple areas, while simple trajectories may indicate that attention is highly focused on a single area.
[0050] Visualization using heatmaps and trajectory plots
[0051] Heatmaps: Eye-tracking data is visualized as heatmaps. Warm-colored areas (such as red and orange) represent areas where subjects fixate for longer periods and have concentrated attention, while cool-colored areas (such as blue and green) represent areas where fixation is shorter and attention is less. Heatmaps provide a clear visual representation of the distribution of subjects' attention in visual scenes.
[0052] Eye movement diagrams: These diagrams show the sequence and path of the participants' eye movements. The numbers inside the circles indicate the order of fixation, and areas with more concentrated circles indicate areas that were fixated on multiple times. Eye movement diagrams can reveal the sequence of fixation on different areas and the process of attention shifting.
[0053] After determining the location of human visual attention on the image data, the corresponding human attention map can be obtained through a feature extraction model.
[0054] Traditional convolutional neural networks activate neurons by inputting a global visual image and acquiring a corresponding visual representation of the scene. This process is similar to the early visual information processing stage of human scene recognition. However, navigation scenes are composed of different local fine-grained information, and the method of inputting a global image easily obscures the effective local information representing the scene. In contrast, humans, after processing the navigation scene as a whole in a coarse manner, are more likely to capture the local information representative of the scene in the navigation image by adjusting attention and feature integration mechanisms. Compared to activating neurons by inputting a global visual image, typical scene features are less numerous and are all local details that are crucial to navigation decisions. Therefore, this method can improve the robustness of the model under various visual interferences while ensuring the correctness of navigation decisions.
[0055] S102. Input the typical scene features of each image data into the brain-like behavior decision model to obtain the target route corresponding to the navigation task dataset.
[0056] In some embodiments, the brain-like behavioral decision-making model, upon receiving typical scene features of each image data, can determine the scene representation of the typical scene features for each image data, and further, based on the scene representation of each image data, determine the final target route.
[0057] This application proposes a geospatial brain-like navigation route planning method. The method includes: extracting typical scene features from each image data in a navigation task dataset using a brain-like scene recognition model; wherein the similarity between the feature maps extracted by the multiple brain-like neurons in the brain-like scene recognition model and human attention maps is less than a preset similarity threshold; and inputting the typical scene features of each image data into a brain-like behavior decision model to obtain the target route corresponding to the navigation task dataset. This application's embodiment extracts typical scene features from each image data using a brain-like scene recognition model comprising multiple brain-like neurons. Since typical scene features conform to human visual attention, the number of typical scene features is small, and they are all crucial local detail features for navigation decision-making. Therefore, it can improve the robustness of the model under various visual interferences while ensuring the correctness of navigation decisions.
[0058] In some embodiments, the brain-like behavior decision model further includes: a feature activation network, which inputs typical scene features of each image data into the brain-like behavior decision model to obtain the target route corresponding to the navigation task dataset, including: inputting at least one typical scene feature of the first image data into the feature activation network to obtain scene adaptive feature values corresponding to each of the at least one typical scene feature, wherein the first image data is any image data in each image data; integrating the at least one scene adaptive feature value to obtain the navigation scene representation corresponding to the first image data; and obtaining the target route corresponding to the navigation task dataset based on the navigation scene representation corresponding to each image data.
[0059] In some embodiments, after obtaining the typical scene features of each image data, it is necessary to activate the typical scene features to obtain the corresponding feature representation, so as to determine the corresponding target route based on the feature representation.
[0060] It is understandable that the feature activation network can include feature representations corresponding to each typical scene feature. By inputting at least one typical scene feature of the first image data into the feature activation network, scene adaptive feature values corresponding to each typical scene feature are obtained. That is, the feature activation network can determine the scene adaptive feature values of each typical scene feature based on the feature attributes of each typical scene feature. Furthermore, the at least one scene adaptive feature value is integrated.
[0061] The integration methods could include determining the encoding weights of each typical scene feature and integrating them based on the encoding weights and scene adaptive feature values of each typical scene feature; or directly concatenating the scene adaptive feature values of each typical scene feature for integration, etc.
[0062] In some embodiments, the feature activation network includes multiple activation matrices corresponding to each feature. At least one typical scene feature of the first image data is input into the feature activation network to obtain scene-adaptive feature values corresponding to each of the at least one typical scene feature. This includes: determining the target feature category and spatial information content corresponding to each of the at least one typical scene feature, where the spatial information content represents the proportion of information corresponding to the typical scene feature in the image data and the relevance of the information corresponding to the typical scene feature to the navigation task; determining the activation weight corresponding to each of the at least one typical scene feature based on the spatial information content corresponding to each of the at least one typical scene feature; and activating the at least one typical scene feature according to the at least one activation weight using the activation matrices corresponding to each of the at least one target feature category to obtain scene-adaptive feature values corresponding to each of the at least one typical scene feature.
[0063] Generally, typical scene features can be divided into three dimensions: two-dimensional (2D), three-dimensional (3D), and semantic dimension. 2D mainly focuses on two-dimensional information in the image, including features such as edges, depth, haze, and vanishing point detection. 3D mainly focuses on three-dimensional spatial information in the image, including features such as spatial layout, key points, and curvature. Semantic dimension focuses on high-level semantic features in the image, such as object categories and semantic categories. Combining these features can help us better understand how human navigation interacts with external navigation scene information.
[0064] The specific features and meanings corresponding to the three types of features are shown in Table 1:
[0065]
[0066] Table 1
[0067] Each of the above features corresponds to a feature activation matrix, and each feature activation matrix can be activated based on an attention mechanism. The feature activation network can be constructed by integrating multiple scene features extracted from the training set.
[0068] Specifically, this application embodiment, based on a portion of the Taskonomy and Cityscapes label datasets, trains a ResNet50 neural network to extract 20 low-level and high-level visual features related to human navigation processes. This ResNet network uses convolutional layers, max-pooling layers, and residual blocks to propagate and update parameters layer by layer, thereby forming a representation of a certain visual attribute of the image. The network mainly consists of an encoder and a decoder, with the encoder composed of seven modules, such as... Figure 2 As shown: First, the street view image is input into the network and compressed into a 1×256×256×3 tensor. Then, it passes through the first module, Conv1, and the second module, Pool1. The first two modules perform convolution and max pooling operations on the input image, respectively. Next, the data enters four residual modules, each composed of convolutional kernels with different numbers, lengths, and output dimensions, and these modules are residually connected. This aims to enhance network performance, reduce the gradient vanishing problem, and improve the model's image representation capabilities. Finally, the data passes through the last max pooling layer, Pool2, and outputs the encoded features.
[0069] This study extracted neuronal activation feature maps of the Satge4 residual block of the network coding layer as visual features of the image. The feature map size of this layer is 1×16×16×2048. Previous studies have shown that this network layer has a high degree of similarity to the visual coding of brain regions such as the occipital lobe and parietal lobe of the human brain.
[0070] After the feature activation network obtains the features of a typical scene, the target feature category corresponding to the typical scene features is determined. The activation matrix corresponding to the target feature category is used to activate the typical scene features to obtain the scene adaptive feature value corresponding to the typical scene features.
[0071] It is understandable that each typical scene feature has a corresponding amount of spatial information. Spatial information is usually used to describe the richness of spatial information carried by different features in a scene, and can be used to determine the activation level of each typical scene feature.
[0072] In some embodiments, the spatial information content of typical scene features can be quantified using the following methods:
[0073] Information entropy-based methods:
[0074] Information entropy is a classic method for measuring the amount of information and can be used to quantify the spatial information content of scene features.
[0075] Calculation steps:
[0076] Statistical feature distribution: the frequency of occurrence of typical scene features at each spatial location in a statistical scenario.
[0077] Calculating probability: Normalizing frequencies to probabilities
[0078] Calculate entropy: Substitute into the entropy formula to calculate.
[0079] Complexity-based methods:
[0080] Complexity can be used to measure the amount of spatial information in scene features; the higher the complexity, the greater the amount of spatial information.
[0081] Definition: Various complexity metrics can be used, such as fractal dimension, Lempel-Ziv complexity, etc.
[0082] Fractal dimension: The fractal dimension measures the complexity of a characteristic distribution. For example, box counting is a commonly used method for calculating fractal dimension.
[0083] Lempel-Ziv complexity: a measure of the complexity of a feature sequence using a compression algorithm.
[0084] Calculation steps:
[0085] Choose a complexity measurement method: Select an appropriate complexity measurement method based on specific needs.
[0086] Computational complexity: The complexity of calculating the feature according to the selected method.
[0087] Of course, there are other ways to calculate the spatial information of typical scene features, which will not be elaborated here.
[0088] Furthermore, activation weights corresponding to at least one typical scene can be determined based on the spatial information content corresponding to each of the features of at least one typical scene. The spatial information content can be positively correlated with the activation weights.
[0089] In some embodiments, the calculation process of the scene adaptive feature value of the i-th typical scene feature is as follows:
[0090]
[0091] SAA i Let be the scene adaptive feature value of the i-th typical scene feature, act be the activation weight of a certain feature neuron, and neur be the feature value of the corresponding type activation matrix.
[0092] For example, suppose the i-th typical scene feature is green trees, and the corresponding feature types include: 2D edge, color and object category, and the corresponding activation weights are 0.5 for 2D edge, 0.8 for color and 0.2 for object category. The feature value corresponding to 2D edge is feature value 1, the feature value corresponding to color is feature value 2, and the feature value corresponding to object category is feature value 3. Then the corresponding scene adaptive feature value is 0.5*feature value 1 + 0.8*feature value 2 + 0.2*feature value 3.
[0093] In some embodiments, integrating at least one scene adaptive feature value to obtain a navigation scene representation corresponding to the first image data includes: determining the encoding weights corresponding to the typical scene features included in at least one target feature category based on the spatial information corresponding to each of the at least one typical scene features; integrating at least one scene adaptive feature value based on the encoding weights corresponding to the typical scene features included in at least one target feature category to obtain a navigation scene representation corresponding to the first image data.
[0094] In some embodiments, each typical scene feature also has a corresponding encoding weight, which can also be determined based on the amount of spatial information, that is, based on the proportion of information corresponding to the typical scene feature in the image data and the relevance of the information corresponding to the typical scene feature to the navigation task.
[0095] Correspondingly, the greater the proportion and relevance of information, the greater the encoding weight, thus allowing more attention to typical scenario features with a large proportion of information and relevance to the navigation task in the final navigation decision.
[0096] Specifically, Where w i Encoding weights for different visual features.
[0097] The process of generating the above navigation scene representation is as follows: First, the typical scene features of each image data in the navigation task dataset are extracted by a brain-like scene recognition model. At least one typical scene feature of the first image data is input into the feature activation network to obtain the scene adaptive feature value corresponding to each of the at least one typical scene feature. Based on the spatial information corresponding to each of the at least one typical scene feature, the encoding weights corresponding to the typical scene features included in at least one target feature category are determined. Based on the encoding weights corresponding to the typical scene features included in at least one target feature category, the at least one scene adaptive feature value is integrated to obtain the navigation scene representation corresponding to the first image data.
[0098] Furthermore, the brain-like behavioral decision-making model determines the target route corresponding to the navigation task dataset based on the navigation scene representation of each image data.
[0099] The brain-like behavioral decision-making model includes a multilayer perceptron and a Q-network. The training process of the Q-network is as follows:
[0100] The navigation scene representation is input into a multilayer perceptron for initial decoding, outputting a distribution of possible navigation action values (e.g., forward, backward, left, etc.). These action value distributions are then input into an online Q-network and a target Q-network. The online Q-network samples the actions and makes navigation behavior decisions, while the target Q-network evaluates these decisions. Based on the error loss from both networks, the network uses gradient descent to continuously optimize and update the model parameters.
[0101] To evaluate the effectiveness of the model's navigation decision-making, this study uses four behavioral decision-making evaluation metrics: cumulative success rate, total average reward, moving average reward, and number of decisions per round.
[0102] The Cumulative Success Rate (CSR) refers to the proportion of times an agent successfully completes a navigation task during model training or testing. A higher CSR indicates that the model completes the navigation task more times.
[0103] Total Average Reward (TAR) refers to the average of all immediate rewards obtained during the model training process (as shown in Equation 5-3). A higher total average reward indicates that the model is more efficient in completing the navigation decision task in each round.
[0104]
[0105] Where episode is the number of training rounds, r i As a reward for a certain round of training, p i This is a punishment for a particular round of training.
[0106] Moving Average Reward (MAR) refers to the average of all immediate rewards obtained by the model during training within a sliding window of a certain number of training epochs (as shown in Equation 5-4). Compared to the total average reward, moving average reward focuses on evaluating the stability of the training process.
[0107]
[0108] Where t is a training round number, w is the sliding window size (set to 20 in this study), and r i As a reward for a certain round of training, p i This is a punishment for a particular round of training.
[0109] Actions Per Episode (APE) refers to the actual number of decision steps taken to complete a navigation task during a given training round. A lower APE indicates higher efficiency in completing the navigation task and a more complete understanding of the scene by the model.
[0110] In some embodiments, the method further includes: identifying multiple brain-like neurons from multiple neurons of a preset neural network.
[0111] In some embodiments, determining multiple brain-like neurons from multiple neurons of a preset neural network includes: acquiring a masking pixel block of a preset scale; moving the masking pixel block from the upper left position of the original image to the lower right position of the original image until all pixels of the original image have been traversed; acquiring a masking image obtained for each position the masking pixel block traverses, to obtain multiple masking images; inputting the multiple masking images and the original image into multiple neurons of the preset neural network to obtain multiple masking feature maps and original feature maps corresponding to each neuron, wherein the multiple masking feature maps correspond one-to-one with the multiple masking images, and the original feature maps correspond to the original image; and determining multiple brain-like neurons among the multiple neurons based on the multiple masking feature maps and original feature maps corresponding to each neuron.
[0112] The preset neural network can be a convolutional neural network, a recurrent neural network, a long short-term memory network, etc.
[0113] Generally, a preset neural network includes multiple neurons, including brain-like neurons. In order to obtain typical scene features, it is necessary to extract brain-like neurons from multiple neurons, and then extract typical scene features of image data based on a brain-like scene recognition model composed of brain-like neurons.
[0114] The preset scale can be flexibly set based on actual conditions, and the choice of the size of the masked pixel block and the sliding step size is very important. Excessively large pixel blocks and sliding steps will result in coarser granularity in the interpretability of pixel attributes to the scene. Conversely, excessively small pixel blocks and sliding steps will lead to a sharp increase in computational power requirements. This application, taking into account both interpretation granularity and computational power requirements, sets the preset pixel block size to 16*16 pixels and the sliding step size to 16 pixels, masking the entire original image sequentially from the upper left to the lower right.
[0115] It is understandable that each time the masking pixel block slides, a masking image is obtained. If the masking pixel block masks from the top left to the bottom right, multiple masking images are obtained.
[0116] Multiple masked images are passed through each neuron to obtain multiple masked feature maps. At the same time, the original features are passed through each neuron to obtain multiple original feature maps.
[0117] Furthermore, multiple brain-like neurons are identified from the multiple occlusion feature maps and original feature maps corresponding to each neuron, as well as the human visual attention location of the original image.
[0118] In some embodiments, multiple brain-like neurons are determined from multiple neurons based on multiple masking feature maps and original feature maps corresponding to each neuron, including: calculating the feature similarity between each masking feature map and the first original feature map for multiple masking feature maps and the first original feature map of the first neuron, where the first neuron is any neuron among the multiple neurons; determining multiple target masking feature maps whose feature similarity with the first original feature map is greater than a feature similarity threshold; determining the masking pixel block position corresponding to each of the multiple target masking feature maps as the encoding position of the first neuron; obtaining the human visual attention position and the attention feature map corresponding to the human visual attention position in the original image; determining multiple target neurons whose encoding positions overlap with the human visual attention position among the multiple neurons; determining the target feature map corresponding to each of the multiple target neurons based on the encoding position corresponding to each of the multiple target neurons; and determining the target neurons whose feature map similarity threshold with the attention feature map is greater than a preset feature map similarity threshold as brain-like neurons.
[0119] Specifically, for each neuron, multiple masking feature maps and original feature maps are used to calculate the feature similarity between each masking feature map and the original feature map. The feature similarity can be determined using KL divergence (KL divergence). The smaller the KL divergence, the greater the feature similarity. Multiple target masking feature maps whose feature similarity to the first original feature map is greater than a feature similarity threshold can be identified as masking feature maps whose KL divergence value to the first original feature map is less than a preset KL divergence value. The masking pixel block positions corresponding to each of the multiple target masking feature maps are then determined as the encoding positions of the first neuron.
[0120] The formula for calculating the KL divergence is as follows:
[0121]
[0122] Where P is the original feature map distribution, Q is the masked feature map distribution, and x represents the feature value in the two distributions.
[0123] By obtaining the encoding position of each neuron in the above manner, the human visual attention position in the original image is further obtained, and multiple target neurons whose encoding positions overlap with the human visual attention positions are identified among multiple neurons.
[0124] Although the encoding locations of multiple target neurons overlap with human visual attention locations, these multiple target neurons cannot all be identified as brain-like neurons because their encoding locations also include other locations in the original image.
[0125] To minimize interference from other locations in the original image at the encoding positions of multiple target neurons, target feature maps corresponding to each target neuron can be determined based on their respective encoding positions. Target neurons whose feature map similarity threshold with the attention feature map is greater than a preset feature map similarity threshold are identified as brain-like neurons.
[0126] This application also provides a geospatial brain-like navigation route planning device, which is used to perform the above-mentioned... Figure 1 The embodiment provides a geospatial brain-like navigation route planning method. For example... Figure 3 As shown, the device includes an extraction module 301 and an input module 302.
[0127] The extraction module 301 is used to extract typical scene features of each image data in the navigation task dataset through a brain-like scene recognition model. The feature maps extracted by the multiple brain-like neurons included in the brain-like scene recognition model have a similarity of less than a preset similarity threshold with the human attention map.
[0128] The input module 302 is used to input the typical scene features of each image data into the brain-like behavior decision model to obtain the target route corresponding to the navigation task dataset.
[0129] This application proposes a geospatial brain-like navigation route planning device. The method includes: extracting typical scene features of each image data in a navigation task dataset using a brain-like scene recognition model, wherein the similarity between the feature maps extracted by the multiple brain-like neurons in the brain-like scene recognition model and human attention maps is less than a preset similarity threshold; and inputting the typical scene features of each image data into a brain-like behavior decision model to obtain the target route corresponding to the navigation task dataset. This application's embodiment extracts typical scene features of each image data using a brain-like scene recognition model comprising multiple brain-like neurons. Since typical scene features conform to human visual attention, the number of typical scene features is small, and they are all crucial local detail features for navigation decision-making. Therefore, it can improve the robustness of the model under various visual interferences while ensuring the correctness of navigation decisions.
[0130] In some embodiments, the brain-like behavioral decision-making model further includes: a feature activation network, and the input module 302 is specifically used for:
[0131] At least one typical scene feature of the first image data is input into the feature activation network to obtain scene adaptive feature values corresponding to each of the at least one typical scene feature, wherein the first image data is any image data in each image data;
[0132] Integrate at least one of the scene adaptive feature values to obtain the navigation scene representation corresponding to the first image data;
[0133] Based on the navigation scene representation corresponding to each image data, the target route corresponding to the navigation task dataset is obtained.
[0134] In some embodiments, the feature activation network includes activation matrices corresponding to multiple feature categories, and the input module 302 is further specifically used for:
[0135] The target feature category and spatial information content corresponding to each of the at least one typical scene feature are determined. The spatial information content is used to represent the proportion of information corresponding to the typical scene feature in the image data and the correlation between the information corresponding to the typical scene feature and the navigation task.
[0136] The activation weights corresponding to the at least one typical scene are determined based on the spatial information content corresponding to each of the at least one typical scene features.
[0137] By activating the at least one typical scene feature according to the activation matrix corresponding to each of the at least one target feature type and the activation weight corresponding to the at least one typical scene feature, the scene adaptive feature value corresponding to each of the at least one typical scene feature is obtained.
[0138] In some embodiments, the input module 302 is further specifically used for:
[0139] Based on the spatial information corresponding to each of the at least one typical scene feature, determine the encoding weights corresponding to each of the typical scene features included in at least one of the target feature categories;
[0140] Based on the encoding weights corresponding to the typical scene features included in at least one of the target feature categories, the adaptive feature values of the at least one scene are integrated to obtain the navigation scene representation corresponding to the first image data.
[0141] In some embodiments, the above-described apparatus includes: a determining module, used for
[0142] Multiple brain-like neurons were identified from multiple neurons in a pre-defined neural network.
[0143] In some embodiments, the determining module is specifically used for:
[0144] Obtain the masking pixel block of a preset size;
[0145] The masking pixel block is moved from the upper left position of the original image to the lower right position of the original image until all pixels of the original image have been traversed;
[0146] Each time the masked pixel block passes through a position, a masked image is obtained to obtain multiple masked images;
[0147] The multiple masked images and the original image are input into multiple neurons of a preset neural network to obtain multiple masked feature maps and original feature maps corresponding to each neuron. The multiple masked feature maps correspond one-to-one with the multiple masked images, and the original feature maps correspond to the original images.
[0148] Based on the multiple occlusion feature maps and the original feature map corresponding to each neuron, and the human visual attention position of the original image, multiple brain-like neurons are determined among the multiple neurons.
[0149] In some embodiments, the determining module is further specifically used for:
[0150] For multiple occlusion feature maps and a first original feature map of the first neuron, calculate the feature similarity between each occlusion feature map and the first original feature map, where the first neuron is any neuron among the multiple neurons;
[0151] Identify multiple target occlusion feature maps whose feature similarity to the first original feature map is greater than a feature similarity threshold;
[0152] The positions of the occlusion pixel blocks corresponding to each of the multiple target occlusion feature maps are determined as the encoding positions of the first neuron;
[0153] Obtain the human visual attention location in the original image and the attention feature map corresponding to the human visual attention location;
[0154] Among the plurality of neurons, a plurality of target neurons whose encoding positions overlap with the human visual attention positions are identified;
[0155] The target feature map corresponding to each of the multiple target neurons is determined based on the encoding position corresponding to each of the multiple target neurons;
[0156] Target neurons whose feature map similarity threshold with the attention feature map is greater than a preset feature map similarity threshold are identified as brain-like neurons.
[0157] This application also provides an electronic device for performing the above-described geospatial brain-like navigation route planning method. Please refer to... Figure 4 It illustrates a schematic diagram of an electronic device provided by some embodiments of this application. For example... Figure 4 As shown, the electronic device 7 includes: a processor 700, a memory 701, a bus 702 and a communication interface 703. The processor 700, the communication interface 703 and the memory 701 are connected through the bus 702. The memory 701 stores a computer program that can run on the processor 700. When the processor 700 runs the computer program, it executes the geospatial brain-like navigation route planning method provided in any of the foregoing embodiments of this application.
[0158] The memory 701 may include high-speed random access memory (RAM) or non-volatile memory, such as at least one disk storage device. Communication between this device network element and at least one other network element is achieved through at least one communication interface 703 (which can be wired or wireless), such as the Internet, wide area network, local area network, or metropolitan area network.
[0159] Bus 702 can be an ISA bus, PCI bus, or EISA bus, etc. Buses can be divided into address buses, data buses, control buses, etc. Memory 701 is used to store programs. After receiving execution instructions, processor 700 executes the program. The geospatial brain-like navigation route planning method disclosed in any of the aforementioned embodiments of this application can be applied to processor 700, or implemented by processor 700.
[0160] The processor 700 may be an integrated circuit chip with signal processing capabilities. In implementation, each step of the above method can be completed by the integrated logic circuitry in the hardware of the processor 700 or by instructions in software form. The processor 700 may be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it may also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), an off-the-shelf programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor may be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this application can be directly embodied in the execution of a hardware decoding processor, or executed by a combination of hardware and software modules in the decoding processor. The software modules may reside in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. The storage medium is located in memory 701. Processor 700 reads the information in memory 701 and, in conjunction with its hardware, completes the steps of the above method.
[0161] The electronic device provided in this application embodiment and the geospatial brain-like navigation route planning method provided in this application embodiment are based on the same inventive concept and have the same beneficial effects as the methods they adopt, operate or implement.
[0162] This application also provides a computer-readable storage medium corresponding to the geospatial brain-like navigation route planning method provided in the foregoing embodiments. Please refer to... Figure 5 The computer-readable storage medium shown is an optical disc 30, on which a computer program (i.e., a program product) is stored. When the computer program is run by a processor, it executes the geospatial brain-like navigation route planning method provided in any of the aforementioned embodiments.
[0163] It should be noted that examples of computer-readable storage media may also include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other optical and magnetic storage media, which will not be elaborated here.
[0164] The computer-readable storage medium provided in the above embodiments of this application and the geospatial brain-like navigation route planning method provided in the embodiments of this application are based on the same inventive concept and have the same beneficial effects as the methods adopted, run or implemented by the applications stored therein.
[0165] It should be noted that:
[0166] Numerous specific details are set forth in the specification provided herein. However, it will be understood that embodiments of this application may be practiced without these specific details. In some instances, well-known structures and techniques have not been shown in detail so as not to obscure the understanding of this specification.
[0167] Similarly, it should be understood that, for the sake of brevity and to aid in understanding one or more of the various inventive aspects, in the above description of exemplary embodiments of this application, various features of this application are sometimes grouped together in a single embodiment, figure, or description thereof. However, this disclosure should not be construed as reflecting a schematic diagram in which the claimed application requires more features than expressly recited in each claim. Rather, as reflected in the following claims, inventive aspects lie in fewer than all features of a single foregoing disclosed embodiment. Therefore, the claims following the detailed description are hereby expressly incorporated into that detailed description, wherein each claim itself is a separate embodiment of this application.
[0168] Furthermore, those skilled in the art will understand that although some embodiments herein include certain features included in other embodiments but not others, combinations of features from different embodiments are intended to be within the scope of this application and form different embodiments. For example, in the following claims, any of the claimed embodiments can be used in any combination.
[0169] The above are merely preferred embodiments of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A geospatial brain-like navigation route planning method, characterized in that, include: Typical scene features of each image data in the navigation task dataset are extracted by a brain-like scene recognition model. The feature maps extracted by the multiple brain-like neurons in the brain-like scene recognition model have a similarity of less than a preset similarity threshold with human attention maps. The typical scene features of each image data are input into the brain-like behavior decision model to obtain the target route corresponding to the navigation task dataset. The brain-like behavior decision-making model further includes a feature activation network, wherein typical scene features of each image data are input into the brain-like behavior decision-making model to obtain the target route corresponding to the navigation task dataset, including: At least one typical scene feature of the first image data is input into the feature activation network to obtain scene adaptive feature values corresponding to each of the at least one typical scene feature, wherein the first image data is any image data in each image data; Integrate at least one of the scene adaptive feature values to obtain the navigation scene representation corresponding to the first image data; Based on the navigation scene representation corresponding to each image data, the target route corresponding to the navigation task dataset is obtained; The feature activation network includes activation matrices corresponding to multiple feature categories. The step of inputting at least one typical scene feature from the first image data into the feature activation network to obtain scene-adaptive feature values corresponding to each of the at least one typical scene feature includes: The target feature category and spatial information content corresponding to each of the at least one typical scene feature are determined. The spatial information content is used to represent the proportion of information corresponding to the typical scene feature in the image data and the correlation between the information corresponding to the typical scene feature and the navigation task. The activation weights corresponding to the at least one typical scene are determined based on the spatial information content corresponding to each of the at least one typical scene features. By activating the at least one typical scene feature according to the at least one activation weight using the activation matrix corresponding to each of the at least one target feature types, the scene adaptive feature value corresponding to each of the at least one typical scene feature is obtained; The step of integrating at least one of the scene adaptive feature values to obtain the navigation scene representation corresponding to the first image data includes: Based on the spatial information corresponding to each of the at least one typical scene feature, determine the encoding weights corresponding to each of the typical scene features included in at least one of the target feature categories; Based on the encoding weights corresponding to the typical scene features included in at least one of the target feature categories, the adaptive feature values of the at least one scene are integrated to obtain the navigation scene representation corresponding to the first image data.
2. The method according to claim 1, characterized in that, The method further includes: Multiple brain-like neurons were identified from multiple neurons in a pre-defined neural network.
3. The method according to claim 2, characterized in that, The process of identifying multiple brain-like neurons from multiple neurons in a preset neural network includes: Obtain the masking pixel block of a preset size; The masking pixel block is moved from the upper left position of the original image to the lower right position of the original image until all pixels of the original image have been traversed; Each time the masked pixel block passes through a position, a masked image is obtained to obtain multiple masked images; The multiple masked images and the original image are input into multiple neurons of a preset neural network to obtain multiple masked feature maps and original feature maps corresponding to each neuron. The multiple masked feature maps correspond one-to-one with the multiple masked images, and the original feature maps correspond to the original images. Based on the multiple occlusion feature maps and the original feature map corresponding to each neuron, and the human visual attention position of the original image, multiple brain-like neurons are determined among the multiple neurons.
4. The method according to claim 3, characterized in that, The step of determining multiple brain-like neurons among the multiple neurons based on multiple occlusion feature maps and original feature maps corresponding to each neuron, and the human visual attention position of the original image, includes: For multiple occlusion feature maps and a first original feature map of the first neuron, calculate the feature similarity between each occlusion feature map and the first original feature map, where the first neuron is any neuron among the multiple neurons; Identify multiple target occlusion feature maps whose feature similarity to the first original feature map is greater than a feature similarity threshold; The positions of the occlusion pixel blocks corresponding to each of the multiple target occlusion feature maps are determined as the encoding positions of the first neuron; Obtain the human visual attention location in the original image and the attention feature map corresponding to the human visual attention location; Among the plurality of neurons, a plurality of target neurons whose encoding positions overlap with the human visual attention positions are identified; The target feature map corresponding to each of the multiple target neurons is determined based on the encoding position corresponding to each of the multiple target neurons; Target neurons whose feature map similarity threshold with the attention feature map is greater than a preset feature map similarity threshold are identified as brain-like neurons.
5. A geospatial brain-like navigation route planning device, characterized in that, include: The extraction module is used to extract typical scene features of each image data in the navigation task dataset through a brain-like scene recognition model. The feature maps extracted by the multiple brain-like neurons in the brain-like scene recognition model have a similarity of less than a preset similarity threshold with the human attention map. The input module is used to input the typical scene features of each image data into the brain-like behavior decision model to obtain the target route corresponding to the navigation task dataset; The brain-like behavioral decision-making model further includes a feature activation network, wherein the input module is specifically used for: At least one typical scene feature of the first image data is input into the feature activation network to obtain scene adaptive feature values corresponding to each of the at least one typical scene feature, wherein the first image data is any image data in each image data; Integrate at least one of the scene adaptive feature values to obtain the navigation scene representation corresponding to the first image data; Based on the navigation scene representation corresponding to each image data, the target route corresponding to the navigation task dataset is obtained; The feature activation network includes activation matrices corresponding to multiple feature categories, and the input module is further specifically used for: The target feature category and spatial information content corresponding to each of the at least one typical scene feature are determined. The spatial information content is used to represent the proportion of information corresponding to the typical scene feature in the image data and the correlation between the information corresponding to the typical scene feature and the navigation task. The activation weights corresponding to the at least one typical scene are determined based on the spatial information content corresponding to each of the at least one typical scene features. By activating the at least one typical scene feature according to the at least one activation weight using the activation matrix corresponding to each of the at least one target feature types, the scene adaptive feature value corresponding to each of the at least one typical scene feature is obtained; The input module is also used for: Based on the spatial information corresponding to each of the at least one typical scene feature, determine the encoding weights corresponding to each of the typical scene features included in at least one of the target feature categories; Based on the encoding weights corresponding to the typical scene features included in at least one of the target feature categories, the adaptive feature values of the at least one scene are integrated to obtain the navigation scene representation corresponding to the first image data.
6. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, The processor executes the computer program to implement the method as described in any one of claims 1-4.
7. A computer-readable storage medium having a computer program stored thereon, characterized in that, The program is executed by a processor to implement the method as described in any one of claims 1-4.
Citation Information
Patent Citations
Brain-computer fusion enhanced visual navigation method, electronic equipment and storage medium
CN118840515A
Navigation path generation method, electronic equipment and computer readable storage medium
CN120008605A