Unmanned cluster brain-like navigation map fusion construction and cooperative positioning method
Through bionic neural coding and heterogeneous hypergraph modeling, combined with multi-agent collaborative optimization and reinforcement learning, the high-precision map construction and positioning of unmanned clusters in dynamic environments is solved, and efficient navigation and positioning is achieved to adapt to task requirements in complex scenarios.
Patent Information
- Application Number
- CN202510747124.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-05
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-06-05
AI Technical Summary
The existing unmanned system navigation technology is difficult to achieve high-precision map construction and positioning in a dynamic and changeable real environment. Especially when multiple agents collaborate on maps, there are problems such as heterogeneous spatiotemporal encoding conflicts in local maps, accumulation of positioning noise and insufficient map timeliness, which affects the application efficiency of unmanned clusters in battlefield reconnaissance and disaster rescue scenarios.
Bionic neural coding, heterogeneous hypergraph modeling and multiagent collaborative optimization methods are used to generate advanced brain-like semantic maps through multimodal perceptual data, and aggregate spatiotemporal features using heterogeneous hypergraph convolutional networks, combining reinforcement learning and Monte Carlo particle filtering algorithms to optimize the distribution of positioning particles, eliminate map conflicts and improve positioning accuracy.
It significantly improves the navigation and positioning efficiency of unmanned clusters in dynamic and complex environments, enhances the robustness and positioning accuracy of environmental characterization, reduces computing resource consumption, ensures geometric-semantic consistency of the global map, and improves task reliability and coordination efficiency.
Smart Images

Figure CN120351920A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of brain-inspired navigation and multi-agent collaborative positioning, and particularly relates to a method for constructing and fusing a brain-inspired navigation map and collaborative positioning for an unmanned cluster. Background Art
[0002] Existing unmanned system navigation technologies mostly rely on pre-set maps or single-agent SLAM frameworks, and it is difficult to adapt to the dynamically changing real environment. The traditional methods have the following significant limitations: First, when multi-agent collaborative mapping is carried out, the heterogeneous spatio-temporal encoding of local maps easily leads to global fusion conflicts, such as geometric coordinate offsets or semantic label contradictions; Second, under the interference of dynamic obstacles, the positioning noise of a single agent will accumulate during the collaborative process, reducing the global consistency of map and position estimation; Third, existing technologies are difficult to effectively cope with the dynamic changes of the environment and cannot update the map in a timely manner, resulting in insufficient timeliness of the map and making it difficult to meet the application requirements in complex dynamic scenarios. These above problems seriously restrict the application efficiency of unmanned clusters in dynamic scenarios such as battlefield reconnaissance and disaster relief.
[0003] Based on this, there is an urgent need for a method for constructing and fusing a brain-inspired navigation map and collaborative positioning for an unmanned cluster, which is suitable for high-precision map construction and positioning of multi-agent autonomous navigation in a dynamically unknown environment. Summary of the Invention
[0004] To solve the above problems, the present invention aims to propose a method for constructing and fusing a brain-inspired navigation map and collaborative positioning for an unmanned cluster, which realizes high-precision navigation and positioning in a dynamically unknown environment through bionic neural coding, heterogeneous hypergraph modeling, and multi-agent collaborative optimization.
[0005] To achieve the above object, the technical solution of the present invention is realized as follows:[[]]
[0006] A method for constructing and fusing a brain-inspired navigation map and collaborative positioning for an unmanned cluster, comprising the following steps:[[]]
[0007] S1. Multi-modal brain-inspired neural coding mechanism: Collect multi-modal perception data, respectively obtain visual, acoustic, and pose information through a binocular camera, radar, and inertial sensor, simulate the brain neural coding mechanism, and generate a pulse sequence and feature vector in combination with spatio-temporal information;
[0008] S2. Heterogeneous hypergraph semantic modeling: Construct multi-modal coding units as hypergraph nodes, and based on the dynamic hyper-edge connection topology relationship, aggregate spatio-temporal features through a heterogeneous hypergraph convolutional network to generate a high-order brain-inspired semantic map of a single agent;
[0009] S3. Multi-agent collaborative fusion and conflict resolution: Multiple agents share local brain-like maps through distributed communication, detect geometric and semantic conflicts in overlapping areas, perform spatio-temporal alignment based on anchor benchmarks, eliminate feature contradictions using probability distribution matching, and generate a globally consistent high-confidence brain-like map.
[0010] S4. Reinforcement learning for optimized positioning: Based on the fused global brain-like map, use the Monte Carlo particle filter algorithm to generate candidate positioning point clouds, and combine with an improved reinforcement learning strategy to dynamically optimize the distribution of positioning particles and output the optimal position estimate.
[0011] Further, in step S1, the image information collected by the binocular camera simulates the collaborative mechanism of the entorhinal cortex grid cells and the visual cortex V1 / V2 regions to encode the visual feature F vis , which is specifically expressed as follows:
[0012]
[0013] where x, y, and z respectively represent the coordinate values and depth information of two-dimensional image pixels. is the grid period of the grid cells, and the hierarchical scale increases by λ (k) = 1.4 k · 30 cm, b is the baseline length, f v is the focal length, d v is the parallax, c p is the fitting parameter, φ k is the weight for obtaining different grid cell encoding patterns through training, with an initial value of 0.2, and γ vis is the initial value of the visual gradient attenuation factor, which is 0.1. is the gradient of the visual information S(x, y), and ||·|| represents the Euclidean norm.
[0014] The radar collects acoustic wave information, simulates the sound source localization and material reflection characteristic classification in the auditory center, encodes the radar echo information, and the radar encoding feature F rad , which is specifically expressed as follows:
[0015]
[0016] where α m is the material reflection coefficient, representing the reflection intensity of different materials, M is the number of different echo information, ReLU(x) = max(0, x) is the rectified linear unit function, f d is the Doppler frequency shift, θ r is the angle of the target object, v noise is the noise velocity, β r is the coefficient used to adjust the input of the ReLU function, with an initial value of 0.5, and λrad The radar time decay factor is set to 0.2, and the size is adjusted in combination with the time interval Δt;
[0017] The inertial sensor obtains the attitude information of the brain-like intelligent agent, simulates the vestibular grid cells and the cerebellar motor coordination mechanism for encoding, and the inertial encoding feature F iner , which is specifically expressed as follows:
[0018]
[0019] where a iner (τ) is the acceleration, ω iner (τ) is the angular velocity, is the rotation matrix, η iner is the vestibular-proprioceptive coupling coefficient, which is taken as η when stationary iner = 0.1, and η is taken as iner = 0.6 during high-speed movement, and LSTM(·) is the long short-term memory network;
[0020] Fuse and encode multi-modal information Simulate the hippocampus-entorhinal cortex loop to achieve spatio-temporal synchronous multi-modal encoding, which is specifically expressed as follows:
[0021]
[0022] where Sync(·) is the pulse synchronization function, θ vis , θ rad , θ iner are the visual pulse phase, radar pulse phase, and inertial pulse phase respectively, θ0 is the reference phase for comparison, and W k = diag(W vis , W rad , W iner ) is the modal diagonal weight matrix, where is the visual feature projection matrix, is the radar auditory feature projection matrix, and is the inertial feature projection matrix, d x is the dimension of the x-modal feature, d h is the dimension of the hidden layer feature, k is the index of the modality, T is a preset sequence length, and τ is a time step.
[0023] Furthermore, in step S2, on the basis of spatio-temporal synchronization, map different modal encoding units to the nodes of the hypergraph. The specific representation of the hypergraph node set V is as follows:
[0024] V = V vis ∪V rad ∪V iner
[0025] Among them, V vis ={v p |v p ∈F vis} represents the visual nodes, V rad ={r q |r q ∈F rad} represents the auditory nodes, V iner ={i s |i s ∈F iner} represents the inertial nodes;
[0026] Based on spatio-temporal correlation and modal complementarity, the hyperedge e m is dynamically generated to connect the multimodal nodes v p , r q , i s . The calculation of the spatio-temporal correlation degree comprehensively considers the feature similarity and time factor between nodes. Only when the correlation degree is greater than the dynamic correlation threshold and the timestamp difference is within the tolerance range, the hyperedge will be generated, simulating the dynamic synchronous activation of cross-modal neural clusters in the cerebral cortex. The hyperedge set E(t) at time t is specifically represented as follows:
[0027]
[0028] Among them, represents the spatio-temporal correlation degree between the node v p and the node r q , λ edge is the hyperedge time decay coefficient, taking the classical value λ edge =0.05s -1 , t represents the current time, t p and t q are the timestamps of the nodes v p and r q respectively, θ e (t) is the dynamic correlation threshold, which is adaptively adjusted, and the initial value is θ e (0)=0.6, represents the timestamp difference of the inertial node i s , and ΔT is the maximum tolerance interval;
[0029] Design convolutional kernels for modal perception to hierarchically aggregate heterogeneous features. Features of different modalities can be independently learned and projected, so as to extract and integrate multimodal information, which is specifically represented as follows:
[0030]
[0031] Among them, H (l+1) represents the hypergraph feature matrix at the (l + 1)-th layer, which is composed of the feature matrix H at the l-th layer(l) obtained by convolution, and σ(·) is the pulse trigger function of the threshold firing mechanism θ th is the threshold, and D v is the node degree matrix, and A (l) is the hyperedge-node incidence matrix, and A (l)T is the transpose matrix, and D e is the hyperedge degree matrix, and W (l) =diag(W vis ,W rad ,W iner ) is the modal diagonal weight matrix;
[0032] Through spatio-temporal attention pooling, the input hypergraph feature matrix is processed, and important feature information is screened out by means of the attention mechanism. The pooling function Pool attn (·) and the semantic category prototype matrix C sem is specifically expressed as follows:
[0033]
[0034] C sem =[c1,c2,...,c K
[0035] where |V| is the total number of nodes, h n is the nth eigenvector in the hypergraph feature matrix H, and u q is the learnable query vector with dimension d h , and C sem represents the semantic category prototype matrix, which is composed of K semantic category prototype vectors c1,c2,...,c K ; MLP(·) is a multi-layer perceptron, which is a neural network structure composed of multiple fully connected layers, represents the eigenvector related to the kth semantic category, is the feature concatenation operator;
[0036] Dynamically select important feature information according to different query requirements, and use the predefined semantic category prototype matrix for semantic binding to generate the agent's local high-order brain-like map M, which is specifically expressed as follows:
[0037]
[0038] where M is the agent's local high-order brain-like map, is the nth eigenvector in the last-layer feature matrix H of the hypergraph (L) , and ⊙ represents the semantic binding operation, that is, element-wise multiplication.
[0039] Further, in step S3, the brain-like local map established by the intelligent monomer i at time t forms a map representation that includes spatial, semantic, and uncertainty information, specifically as follows:
[0040]
[0041] Among them, is a spatio-temporal hypergraph with objects as nodes and geometric topological relations as hyperedges, is a semantic feature tensor, is a probability confidence matrix, where p is the observation distribution and q is the reference distribution, representing the semantic feature distributions of agents i and j in the overlapping region Ω ij respectively;
[0042] By calculating the geometric difference and semantic feature difference of agents i and j in the overlapping region Ω ij the conflict degree D is obtained ij Specifically, it is represented as follows:
[0043]
[0044] Among them, d geo is the geometric difference based on the hyperedge topological Hausdorff distance, k is the semantic category index, and α d = 0.6, β d = 0.4, sup represents the supremum, inf represents the infimum, and x and y are points in the spatio-temporal hypergraph respectively;
[0045] To achieve spatio-temporal alignment between agents and ensure the consistency of their map information in space and time, it is necessary to use the significant feature points in the brain-like local map as the anchor point set geometry A ij to find the minimum geometric error, and then solve the optimal SE(3) transformation matrix T ij and the time delay compensation amount Δt ij , specifically represented as follows:
[0046]
[0047] Among them, A ij is the anchor point set composed of anchor points a with high confidence, taking the confidence Ψ a of the anchor point a ≥ 0.85, T ij is the SE(3) transformation matrix from agent i to j, Δt ij is the time delay compensation amount, and ||·|| F represents the Frobenius norm;
[0048] In the case of certain characteristic contradictions, the semantic features are fused and processed by means of probability distribution, simulating the information integration mechanism of the hippocampal-neocortical loop, and the Gaussian mixture model is used to probabilistically reconstruct the semantic features of agent j. Specifically, it is expressed as follows:
[0049]
[0050] Among them, is the reconstructed result of the semantic features of agent j at time t. represents the k-th Gaussian distribution with mean μ k and covariance matrix Σ k , and ι k is the Gaussian distribution weight. Through experiments, the conflict sensitivity coefficient γ is calibrated to be 1.2.
[0051] The reference agent is selected and marked as reference agent 0. By suppressing the weights in the high-conflict areas with agent 0, the global consistency of the map is achieved, and the local brain-like maps of multiple agents are fused into a globally consistent high-confidence brain-like map. It is expressed as follows:
[0052]
[0053] Among them, is the globally consistent brain-like map. is the union operation, representing the merging operation of the local maps of agents with subscript i from i = 1 to i = N. T i0 is the transformation matrix from agent i to reference agent 0. Through prefrontal cortex experiments, λ is calibrated to be 0.8. represents the composite operation, that is The meaning is to rotate and translate the local map through the transformation matrix T i0 so that it is aligned with the map of the reference agent.
[0054] Furthermore, in step S4, particle filtering is initialized, and a particle set is randomly sampled and initialized from the spatial feasible region. Initial uniform sampling is adopted, and the specific representation is as follows:
[0055]
[0056] Among them, N is the total number of particles in the feasible region, and each particle contains three-dimensional coordinates and semantic information, expressed as The initial weights are evenly sampled and equally distributed;
[0057] The movement of the brain-like agent in the environment requires the state prediction of particles over time for the three-dimensional coordinates of the particles Make a prediction, which is specifically expressed as follows:
[0058]
[0059] Among them, represents three-dimensional position information, and f(·) is the motion model, taking the classical uniformly accelerated model equation where x is the position in a certain direction, is the instantaneous velocity in the x-axis direction, a x is the acceleration in the x-axis direction, Δt is the time step, θ is the angle of the agent, is the angular velocity of the agent, is the particle state at the previous moment, u k is the control input including linear velocity and angular velocity, ∈ k is Gaussian noise, where the covariance Q is calibrated by the sensor error;
[0060] The estimation of the three-dimensional space position needs to consider the matching degree between the sampled particle position and the predicted particle position, as well as the semantic confidence of the current position, and then the particle weight can reflect the credibility of the space position represented by the particle. The observation update and weight correction are as follows:
[0061]
[0062] Among them, is the weight of the i-th particle at the k-th moment, updated from the particle weight at the (k - 1)-th moment, κ k is the observed three-dimensional coordinate, is the predicted three-dimensional coordinate, is the noise standard deviation, η(·) is the semantic confidence weight function, is the semantic confidence of the current position, specifically expressed as:
[0063]
[0064] Among them, is the particle 's observed value on the m-th semantic feature, is the global map 's reference value at the corresponding position, is the noise standard deviation, and the experimentally calibrated gain coefficients α η = 0.5, β η = 0.7.
[0065] Furthermore, the strategy π for optimizing the particle distribution is obtained through the method of reinforcement learning * , by defining the state space s t , the action space a tGiven a state space \(s\), an action space \(a\), and a reward function \(r(\cdot)\), the optimal policy is found using reinforcement learning algorithms as follows:
[0066]
[0067] where, \(E\) is for expectation, \(\pi\) is the policy function, and \(\gamma\) t \(\in [0, 1]\) is the discount factor, with an initial value of \(\gamma\) t \(= 0.95\), the state space \(s\) t , the action space \(a\) t and the reward function \(r(\cdot)\) are as follows:
[0068]
[0069] a t \(= [\Delta\tau\) t , \(\Delta\sigma\) x , \(\Delta\sigma\) y , \(\Delta\sigma\) θ T
[0070] where, \(e\) t is the positioning error at the current moment, \(N\) eff is the number of effective particles, is the conflict degree between the current particle set and the map, \(\lambda_1, \lambda_2, \lambda_3, \lambda_4\) are the weights for measuring the positioning accuracy, particle diversity, conflict penalty, and policy stability of the reward function, with initial values of \(\lambda_1 = 0.4\), \(\lambda_2 = 0.3\), \(\lambda_3 = 0.1\), \(\lambda_4 = 0.2\), is the particle entropy, is the semantic gradient, \(e\) t-K is the historical error, \(\Delta\tau\) t is the incremental adjustment amount of the resampling threshold \(\tau\) t , \(\Delta\sigma\) x , \(\Delta\sigma\) y , \(\Delta\sigma\) θ is the adjustment amount of the semantic noise covariance \(\Sigma\) sem ;
[0071] According to the optimal policy \(\pi\) * the particles are resampled, and the particle distribution is dynamically adjusted to improve the diversity and effectiveness of the particles, achieving particle resampling and the optimal position estimation is expressed as follows:
[0072]
[0073] where, is the optimal position estimation, is the indicator function, which takes the value of 1 when the condition is satisfied and 0 otherwise, \(\tau\) th The threshold value is taken as 0.8.
[0074] Furthermore, in step S2, hierarchical compression is performed on the high-order brain-like semantic map:
[0075] Geometry layer: An octree structure is used to store the space occupancy information;
[0076] Semantic layer: The knowledge distillation technique is used to compress the semantic prototype matrix into a low-rank representation;
[0077] Compression ratio: 15%-30% of the original map size.
[0078] To achieve the above object, the present invention also provides a brain-like cluster collaborative navigation and positioning system, including the following steps:
[0079] Multi-modal brain-like neural coding module: Collect multi-modal perception data, respectively obtain visual, acoustic and pose information through binocular cameras, radars and inertial sensors, simulate the human brain neural coding mechanism, and generate pulse sequences and feature vectors by combining spatio-temporal information;
[0080] High-order brain-like map establishment module: Construct multi-modal coding units as hypergraph nodes, based on the dynamic hyper-edge connection topology relationship, aggregate spatio-temporal features through a heterogeneous hypergraph convolutional network, and generate a high-order brain-like semantic map of a single intelligent agent;
[0081] Multi-agent collaborative fusion and conflict resolution module: Multiple intelligent agents share local brain-like maps through distributed communication, detect geometric and semantic conflicts in overlapping areas, perform spatio-temporal alignment based on anchor point benchmarks, and eliminate feature contradictions using probability distribution matching to generate a globally consistent high-confidence brain-like map;
[0082] Reinforcement learning optimization positioning module: Based on the fused global brain-like map, use the Monte Carlo particle filter algorithm to generate candidate positioning point clouds, combine with an improved reinforcement learning strategy, dynamically optimize the positioning particle distribution, and output the optimal position estimate.
[0083] To achieve the above object, the present invention also provides an application of a brain-like cluster collaborative navigation and positioning system, and the brain-like cluster collaborative navigation and positioning system is applied to high-precision map construction and positioning for unmanned cluster collaborative reconnaissance and intelligent logistics scheduling or multi-agent autonomous navigation in a dynamically unknown environment.
[0084] Beneficial effects: Through the brain-inspired neural coding mechanism and multi-agent collaborative optimization strategy, the present invention significantly improves the navigation and positioning efficiency of unmanned clusters in dynamic and complex environments; based on neural pulse coding and heterogeneous hypergraph modeling technologies, it fuses multi-modal perception data, effectively solves the semantic fragmentation problem caused by sensor noise and local perception blind spots, and enhances the robustness of environmental representation; by dynamically optimizing the Monte Carlo particle distribution through reinforcement learning and combining multi-dimensional weight-driven strategies, it significantly improves the positioning accuracy and reduces the consumption of computing resources, and maintains stable output under the interference of dynamic obstacles; using distributed collaborative conflict resolution technology, it eliminates the spatio-temporal coding contradictions of multi-agent maps and ensures the geometric-semantic consistency of the global map. The present invention can be applied to collaborative reconnaissance of unmanned clusters and intelligent logistics scheduling, greatly improving the task reliability and collaborative efficiency in complex scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0085] The accompanying drawings, which form a part of this invention, are used to provide a further understanding of the invention. The schematic embodiments and descriptions thereof of the invention are used to explain the invention and do not constitute an improper limitation of the invention. In the drawings:
[0086] Figure 1 is the main flowchart of the method for constructing and collaborative positioning of brain-inspired navigation maps of unmanned clusters according to the embodiments of the present invention;
[0087] Figure 2 is a schematic diagram of an agent in the method for constructing and collaborative positioning of brain-inspired navigation maps of unmanned clusters according to the embodiments of the present invention;
[0088] Figure 3 is a schematic diagram of the motion scenario of an agent in the method for constructing and collaborative positioning of brain-inspired navigation maps of unmanned clusters according to the embodiments of the present invention;
[0089] Figure 4 is a schematic diagram of the structure of the brain-inspired cluster collaborative navigation and positioning system according to the embodiments of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0090] It should be noted that, without conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other.
[0091] The present invention will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.
[0092] Embodiment 1
[0093] See Figures 1-3 : A method for constructing and collaborative positioning of brain-inspired navigation maps of unmanned clusters, including the following steps:
[0094] S1. Multi-modal cerebral nerve coding mechanism: Collect multi-modal perception data, obtain visual, acoustic and pose information through binocular cameras, radars and inertial sensors respectively, simulate the human brain nerve coding mechanism, and generate pulse sequences and feature vectors by combining spatio-temporal information;
[0095] S2. Higher-order brain-like map establishment: Construct multi-modal coding units as hypergraph nodes, based on the dynamic hyper-edge connection topology relationship, aggregate spatio-temporal features through a heterogeneous hypergraph convolutional network, and generate a higher-order brain-like semantic map of a single intelligent agent;
[0096] S3. Multi-agent collaborative fusion and conflict resolution: Multiple intelligent agents share local brain-like maps through distributed communication, detect geometric and semantic conflicts in overlapping regions, perform spatio-temporal alignment based on anchor benchmarks, and use probability distribution matching to eliminate feature contradictions, generating a globally consistent high-confidence brain-like map;
[0097] S4. Reinforcement learning for optimized positioning: Based on the fused global brain-like map, use the Monte Carlo particle filter algorithm to generate candidate positioning point clouds, combine with an improved reinforcement learning strategy, dynamically optimize the distribution of positioning particles, and output the optimal position estimate.
[0098] Multi-modal fusion: Integrate visual, radar, and inertial sensor data, simulate the human brain nerve coding mechanism, and improve the comprehensiveness and robustness of environmental perception.
[0099] Heterogeneous hypergraph modeling: Capture complex spatio-temporal relationships through dynamic hyper-edge connections, construct a higher-order semantic map, and enhance the environmental representation ability.
[0100] Collaborative conflict resolution: Distributed multi-intelligent agents share local maps and eliminate conflicts, generating a globally consistent map to adapt to dynamic and complex environments.
[0101] Intelligent optimized positioning: Combine particle filtering and reinforcement learning, dynamically adjust the positioning strategy, and improve the navigation accuracy and stability.
[0102] In a specific example, in step S1, the image information collected by the binocular camera simulates the collaborative mechanism of the entorhinal cortex grid cells and the visual cortex V1 / V2 regions to encode the visual feature F vis , which is specifically expressed as follows:
[0103]
[0104] Among them, x, y, z respectively represent the coordinate values and depth information of two-dimensional image pixels, is the grid period of the grid cells, and the hierarchical scale increases by λ (k) = 1.4 k ·30 cm, b is the baseline length, f v is the focal length, dv is the parallax, c p is the fitting parameter, φ k is the weight for obtaining different grid cell encoding patterns through training, with an initial value of 0.2, γ vis is the initial value of the visual gradient attenuation factor, which is 0.1 is the gradient of the visual information S(x, y), and ||·|| represents the Euclidean norm;
[0105] The radar collects acoustic wave information, simulates the sound source localization and material reflection characteristic classification in the auditory center, encodes the radar echo information, and the radar coding feature F rad , which is specifically expressed as follows:
[0106]
[0107] where, α m is the material reflection coefficient, representing the reflection intensity of different materials, M is the number of different echo information, ReLU(x) = max(0, x) is the rectified linear unit function, f d is the Doppler frequency shift, θ r is the angle of the target object, v noise is the noise velocity, β r is the coefficient used to adjust the input of the ReLU function, with an initial value of 0.5, λ rad is the radar time attenuation factor, which is 0.2, and is adjusted in combination with the time interval Δt;
[0108] The inertial sensor obtains the attitude information of the brain-like intelligent agent, simulates the vestibular grid cells and the cerebellar motor coordination mechanism for encoding, and the inertial coding feature F iner , which is specifically expressed as follows:
[0109]
[0110] where, a iner (τ) is the acceleration, ω iner (τ) is the angular velocity, is the rotation matrix, η iner is the vestibular-proprioceptive coupling coefficient, which is η when stationary iner = 0.1, and η when moving at high speed iner = 0.6, LSTM(·) is the long short-term memory network;
[0111] Fuse and encode the multi-modal information Simulate the hippocampus-entorhinal cortex loop to achieve spatio-temporal synchronous multi-modal coding, which is specifically expressed as follows:
[0112]
[0113] where Sync(·) is the pulse synchronization function, and θ vis , θ rad , θ iner are the visual pulse phase, the radar pulse phase, and the inertial pulse phase respectively, θ0 is the reference phase for comparison, and W k = diag(W vis , W rad , W iner ) is the modal diagonal weight matrix, where is the visual feature projection matrix, is the radar-auditory feature projection matrix, and the inertial feature projection matrix, d x is the dimension of the x-modal feature, d h is the dimension of the hidden layer feature. k is the index of the modality, T is a preset sequence length, and τ is a time step.
[0114] Bionic mechanism optimization:
[0115] Visual coding simulates grid cells and the visual cortex to achieve depth and hierarchical feature extraction.
[0116] Radar coding combines material reflection and Doppler frequency shift to enhance the obstacle recognition accuracy.
[0117] Inertial coding simulates cerebellar motor coordination through LSTM to improve the dynamic adaptability of attitude estimation.
[0118] Pulse synchronization fusion: Using the hippocampus-entorhinal cortex loop model to achieve multi-modal spatio-temporal synchronous coding, reducing information redundancy and delay.
[0119] In a specific example, in step S2, on the basis of spatio-temporal synchronization, different modal coding units are mapped to the nodes of a hypergraph. The specific representation of the hypergraph node set V is as follows:
[0120] V = V vis ∪ V rad ∪ V iner
[0121] where V vis = {v p | v p ∈ F vis} represents the visual nodes, V rad = {r q | r q ∈ F rad} represents the auditory nodes, and V iner = {i s | i s ∈ F iner} represents the inertial nodes;
[0122] Dynamically generate hyperedge e based on spatio-temporal correlation and modal complementarity m Connect multimodal nodes v p ,r q ,i s , the calculation of spatio-temporal correlation comprehensively considers the feature similarity and time factor between nodes. Only when the correlation degree is greater than the dynamic correlation threshold and the timestamp difference is within the tolerance range, will a hyperedge be generated, simulating the dynamic synchronous activation of cross-modal neural clusters in the cerebral cortex. The hyperedge set E(t) at time t is specifically represented as follows:
[0123]
[0124] Among them, represents the spatio-temporal correlation degree between node v p and node r q , λ edge is the hyperedge time decay coefficient, taking the classical value λ edge = 0.05s -1 , t represents the current time, t p and t q are the timestamps of node v p and r q respectively, θ e (t) is the dynamic correlation threshold, which is adaptively adjusted, and the initial value is θ e (0) = 0.6, represents the timestamp difference of inertial node i s , ΔT is the maximum tolerance interval;
[0125] Design convolutional kernels for modal perception to hierarchically aggregate heterogeneous features. Features of different modalities can be independently learned and projected, so as to extract and integrate multimodal information, which is specifically represented as follows:
[0126]
[0127] Among them, H (l+1) represents the hypergraph feature matrix at the l+1 layer, which is obtained by convolving the feature matrix H (l) at the l layer. σ(·) is the pulse trigger function of the threshold firing mechanism θ th is the threshold, D v is the node degree matrix, A (l) is the hyperedge-node incidence matrix, A (l)T is the transpose matrix, D e is the hyperedge degree matrix, W (l) = diag(W vis ,W rad ,W iner ) is the modal diagonal weight matrix;
[0128] Process the input hypergraph feature matrix through spatio-temporal attention pooling, and use the attention mechanism to screen out important feature information. The pooling function is Pool attn (·) and the semantic category prototype matrix C sem Specifically, it is expressed as follows:
[0129]
[0130] C sem =[c1, c2,..., c K
[0131] where |V| is the total number of nodes, h n is the nth feature vector in the hypergraph feature matrix H, u q is a learnable query vector with a dimension of d h , C sem represents the semantic category prototype matrix, which is composed of K semantic category prototype vectors c1, c2,..., c K MLP(·) is a multi-layer perceptron, which is a neural network structure composed of multiple fully connected layers. represents the feature vector related to the kth semantic category, is the feature concatenation operator;
[0132] Dynamically select important feature information according to different query requirements, and use the predefined semantic category prototype matrix for semantic binding to generate the agent's local high-order brain-like map M. Specifically, it is expressed as follows:
[0133]
[0134] where M is the agent's local high-order brain-like map, is the nth feature vector in the last layer feature matrix H of the hypergraph (L) ⊙ represents the semantic binding operation, that is, element-wise multiplication.
[0135] Dynamic hyperedge connection: Generate hyperedges based on spatio-temporal correlation and time tolerance threshold, simulate cross-modal dynamic activation of brain regions, and adapt to environmental changes.
[0136] Modal perception convolution: Hierarchically aggregate heterogeneous features, retain multi-modal independence, and improve the efficiency of semantic feature extraction.
[0137] Attention pooling: Dynamically screen key features through query-driven and semantic prototype binding, and enhance the semantic interpretability of the map.
[0138] In a specific example, in step S3, the brain-like local map established by the intelligent monomer i at time t A map representation containing spatial, semantic, and uncertainty information is formed as follows:
[0139]
[0140] Among them, The spatio-temporal hypergraph has objects as nodes and geometric topological relationships as hyperedges. is the semantic feature tensor. is the probability confidence matrix, where p is the observation distribution and q is the reference distribution, representing the semantic feature distributions of agents i and j in the overlapping region Ω ij respectively;
[0141] By calculating the geometric difference and semantic feature difference of agents i and j in the overlapping region Ω ij the conflict degree D is obtained. ij Specifically, it is represented as follows:
[0142]
[0143] Among them, d geo is the geometric difference based on the hyperedge topological Hausdorff distance, k is the semantic category index, and α d = 0.6, β d = 0.4, sup represents the supremum, inf represents the infimum, and x and y are points in the spatio-temporal hypergraph respectively;
[0144] To achieve spatio-temporal alignment between agents and ensure the consistency of their map information in space and time, it is necessary to use the significant feature points in the brain-like local map as the anchor point set geometry A ij to minimize the geometric error, and then solve for the optimal SE(3) transformation matrix T ij and the time delay compensation amount Δt ij , specifically represented as follows:
[0145]
[0146] Among them, A ij is the anchor point set composed of anchor points a with high confidence, and the confidence Ψ a of the anchor point a is taken as Ψ ij ≥ 0.85, T ij is the SE(3) transformation matrix from agent i to j, Δt F is the time delay compensation amount, and ||·||
[0147] In the case of certain characteristic contradictions, the semantic features are fused and processed by means of probability distribution, simulating the information integration mechanism of the hippocampal-neocortical loop, and the Gaussian mixture model is used to probabilistically reconstruct the semantic features of agent j. Specifically, it is expressed as follows:
[0148]
[0149] Among them, is the reconstructed result of the semantic features of agent j at time t. represents the k-th Gaussian distribution with mean μ k and covariance matrix Σ k , and ι k is the Gaussian distribution weight. Through experiments, the conflict sensitivity coefficient γ is calibrated to be 1.2.
[0150] The reference agent is selected and marked as reference agent 0. By suppressing the weights in the high-conflict areas with agent 0, the global consistency of the map is realized, and the local brain-like maps of multiple agents are fused into a globally consistent high-confidence brain-like map. It is expressed as follows:
[0151]
[0152] Among them, is the globally consistent brain-like map. is the union operation, representing the merging operation of the local maps of agents with subscript i from i = 1 to i = N. T i0 is the transformation matrix from agent i to reference agent 0. Through experiments in the prefrontal cortex, λ is calibrated to be 0.8. represents the composite operation, that is, The meaning is to rotate and translate the local map through the transformation matrix T i0 to align it with the map of the reference agent.
[0153] Geometric-semantic joint conflict detection: Combining the Hausdorff distance and the semantic distribution difference to accurately identify the conflicts in the multi-agent maps.
[0154] Anchor point spatio-temporal alignment: Optimizing the transformation matrix based on high-confidence anchor points to achieve low-error spatio-temporal synchronization of cross-agent maps.
[0155] Probability fusion mechanism: Reconstructing semantic features through the Gaussian mixture model, suppressing the weights in the conflict areas, simulating the neural loop integration mechanism, and enhancing the global consistency.
[0156] In a specific example, in step S4, particle filtering is initialized, and a particle set is randomly sampled and initialized from the spatial feasible region. Initial uniform sampling is adopted, and the specific representation is as follows:
[0157]
[0158] Among them, N is the total number of particles in the feasible region. Each particle contains three-dimensional coordinates and semantic information, which is represented as is the equal distribution of the initial weight for uniform sampling;
[0159] The movement of the brain-inspired agent in the environment requires the prediction of the particle state over time. The three-dimensional coordinates of the particle are predicted, and the specific representation is as follows:
[0160]
[0161] Among them, represents the three-dimensional position information, and f(·) is the motion model, taking the classical uniformly accelerated model equation where x is the position in a certain direction, is the instantaneous velocity in the x-axis direction, a x is the acceleration in the x-axis direction, Δt is the time step, θ is the angle of the agent, is the angular velocity of the agent, is the particle state at the previous moment, u k is the control input, including the linear velocity and the angular velocity, ∈ k is the Gaussian noise, and the covariance Q is calibrated by the sensor error;
[0162] The estimation of the three-dimensional space position needs to consider the matching degree between the sampled particle position and the predicted particle position, as well as the semantic confidence of the current position. Furthermore, the particle weight can reflect the credibility of the space position represented by the particle. The observation update and weight correction are as follows:
[0163]
[0164] Among them, is the weight of the i-th particle at the k-th moment, updated from the particle weight at the (k - 1)-th moment, κ k is the observed three-dimensional coordinate, is the predicted three-dimensional coordinate, is the noise standard deviation, η(·) is the semantic confidence weight function, is the semantic confidence of the current position, and the specific representation is as follows:
[0165]
[0166] Among them, is the particle The observed value on the m-th semantic feature is the reference value at the corresponding position in the global map , is the noise standard deviation, and the experimentally calibrated gain coefficients α η = 0.5, β η = 0.7
[0167] Semantic enhancement weight update: Combine the coordinate matching error and semantic confidence to dynamically adjust the particle weights and improve the positioning robustness
[0168] Motion model adaptability: Adopt a uniformly accelerated model and sensor noise calibration to improve the state prediction accuracy in a dynamic environment
[0169] In a specific example, the strategy π for optimizing the particle distribution is obtained by reinforcement learning * , by defining the state space s t , the action space a t and the reward function r(·), and the optimal strategy is found using the reinforcement learning algorithm as follows
[0170]
[0171] where is to find the expectation, π is the policy function, and γ t ∈ [0, 1] is the discount factor, and the initial value is taken as γ t = 0.95. The state space s t , the action space a t and the reward function r(·) are as follows
[0172]
[0173] a t = [Δτ t , Δσ x , Δσ y , Δσ θ T
[0174] where e t is the positioning error at the current moment, N eff is the number of effective particles is the conflict degree between the current particle set and the map, and λ1, λ2, λ3, λ4 are the weights for measuring the positioning accuracy, particle diversity, conflict penalty, and policy stability of the reward function. The initial values are taken as λ1 = 0.4, λ2 = 0.3, λ3 = 0.1, λ4 = 0.2 is the particle entropy is the semantic gradient, e t-K is the historical error, Δτ t is the resampling threshold τt Incremental adjustment amount, Δσ x , Δσ y , Δσ θ is the semantic noise covariance Σ sem Adjustment amount;
[0175] According to the optimal policy π * Resample the particles, dynamically adjust the particle distribution to improve the diversity and effectiveness of the particles, and achieve particle resampling and optimal position estimation It is expressed as follows:
[0176]
[0177] where is the optimal position estimation, is the indicator function, which takes the value of 1 when the condition is satisfied and 0 otherwise, τ th The threshold is taken as 0.8.
[0178] Multi-objective reward function: Integrate multi-dimensional indicators such as positioning error, particle diversity, and conflict penalty to guide the policy to optimize in the direction of high precision and low conflict.
[0179] Dynamic resampling mechanism: Adjust the particle distribution and noise parameters according to the optimal policy, balance the particle convergence speed and diversity, and avoid local optima.
[0180] In a specific example, in step S2, hierarchical compression of the high-order brain-like semantic map is performed:
[0181] Geometry layer: Use an octree structure to store space occupancy information;
[0182] Semantic layer: Use knowledge distillation technology to compress the semantic prototype matrix into a low-rank representation;
[0183] Compression ratio: 15%-30% of the original map size.
[0184] Octree storage: The geometry layer compresses the original data volume and supports real-time updates;
[0185] Knowledge distillation: The low-rank representation of the semantic layer retains key features and greatly reduces the transmission bandwidth requirements.
[0186] Embodiment 2
[0187] To achieve the above object, see Figure 4 : This embodiment also provides a brain-like cluster collaborative navigation and positioning system, including the following steps:
[0188] Multi-modal cerebral nerve coding module: Collect multi-modal perception data, obtain visual, acoustic and pose information through binocular cameras, radars and inertial sensors respectively, simulate the human brain nerve coding mechanism, and generate pulse sequences and feature vectors by combining spatio-temporal information;
[0189] Higher-order brain-like map building module: Construct multi-modal coding units as hypergraph nodes, based on the dynamic hyper-edge connection topology relationship, aggregate spatio-temporal features through a heterogeneous hypergraph convolutional network, and generate a higher-order brain-like semantic map of a single intelligent agent;
[0190] Multi-agent collaborative fusion and conflict resolution module: Multiple intelligent agents share local brain-like maps through distributed communication, detect geometric and semantic conflicts in overlapping areas, perform spatio-temporal alignment based on anchor benchmarks, and use probability distribution matching to eliminate feature contradictions, generating a globally consistent and highly confident brain-like map;
[0191] Reinforcement learning optimization positioning module: Based on the fused global brain-like map, use the Monte Carlo particle filter algorithm to generate candidate positioning point clouds, and combine with an improved reinforcement learning strategy to dynamically optimize the distribution of positioning particles and output the optimal position estimate.
[0192] The brain-like navigation map fusion construction and collaborative positioning system of this embodiment has the same advantages as the above-mentioned brain-like navigation map fusion construction and collaborative positioning method for unmanned clusters compared with the prior art, and will not be elaborated here.
[0193] Embodiment 3
[0194] To achieve the above object, this embodiment also provides an application of a brain-like cluster collaborative navigation and positioning system, and the brain-like cluster collaborative navigation and positioning system is applied to high-precision map construction and positioning for unmanned cluster collaborative reconnaissance and intelligent logistics scheduling or multi-agent autonomous navigation in a dynamically unknown environment.
[0195] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for constructing and collaborative positioning of brain-inspired navigation maps for unmanned clusters, characterized in that Including the following steps: S1. Multi-modal cerebral nerve coding mechanism: Collect multi-modal perception data, obtain visual, acoustic and pose information through binocular cameras, radars and inertial sensors respectively, simulate the human brain nerve coding mechanism, and generate pulse sequences and feature vectors by combining spatio-temporal information; S2. Establishment of high-order brain-like maps: Construct multi-modal coding units as hypergraph nodes, and based on the dynamic hyper-edge connection topology relationship, aggregate spatio-temporal features through heterogeneous hypergraph convolutional networks to generate high-order brain-like semantic maps of single agents; S3. Multi-agent collaborative fusion and conflict resolution: Multiple agents share local brain-like maps through distributed communication, detect geometric and semantic conflicts in overlapping areas, perform spatio-temporal alignment based on anchor benchmarks, and use probability distribution matching to eliminate feature contradictions to generate a globally consistent high-confidence brain-like map; S4. Reinforcement learning optimization for positioning: Based on the fused global brain-like map, use the Monte Carlo particle filter algorithm to generate candidate positioning point clouds, and combine with an improved reinforcement learning strategy to dynamically optimize the distribution of positioning particles and output the optimal position estimate.
2. The method for constructing and collaborative positioning of the brain-inspired navigation map fusion of the unmanned cluster according to claim 1, wherein In step S1, the image information collected by the binocular camera encodes the visual feature F by simulating the cooperative mechanism between the entorhinal cortex grid cells and the visual cortex V1 / V2 regions vis , which is specifically represented as follows: Among them, x, y, and z respectively represent the coordinate values and depth information of two-dimensional image pixels. is the grid period of grid cells, and the hierarchical scale is in accordance with λ (k) = 1.4 k ·30 cm increment. b is the baseline length, f v is the focal length, d v is the parallax, c p is the fitting parameter, φ k is the weight for obtaining different grid cell coding patterns through training, with an initial value of 0.2, γ vis is the initial value of the visual gradient attenuation factor, which is 0.
1. is the gradient of the visual information S(x, y), and ||·|| represents the Euclidean norm. The radar collects acoustic wave information, simulates the sound source localization of the auditory center and classifies the material reflection characteristics, encodes the radar echo information, and the radar coding feature F rad , which is specifically expressed as follows: Among them, α m is the material reflection coefficient, representing the reflection intensity of different materials. M is the number of different echo information. ReLU(x) = max(0, x) is the rectified linear unit function, and f d is the Doppler frequency shift, θ r is the angle of the target object, v noise is the noise velocity, β r is a coefficient used to adjust the input of the ReLU function, with an initial value of 0.
5. λ rad is the radar time decay factor, taking 0.2, and is adjusted in combination with the time interval Δt; The inertial sensor obtains the attitude information of the brain-inspired intelligent agent, simulates the vestibular grid cells and the cerebellar motor coordination mechanism for encoding, and the inertial encoding feature F iner , which is specifically expressed as follows: Among them, a iner (τ) is the acceleration, ω iner (τ) is the angular velocity, is the rotation matrix, η iner is the vestibular-proprioceptive coupling coefficient, taking η iner = 0.1 at rest, and taking η iner = 0.6 during high-speed motion; LSTM(·) is the long short-term memory network; Fuse and encode multimodal information Simulate the hippocampal-entorhinal cortex loop to achieve spatiotemporal synchronous multimodal encoding, which is specifically expressed as follows: where Sync(·) is the pulse synchronization function, and θ vis , θ rad , θ iner are the visual pulse phase, radar pulse phase, and inertial pulse phase respectively, θ0 is the reference phase for comparison, and W k = diag(W vis , W rad , W iner ) is the modal diagonal weight matrix, where is the visual feature projection matrix, is the radar-auditory feature projection matrix, and is the inertial feature projection matrix, d x is the dimension of the x-modal feature, d h is the dimension of the hidden layer feature, k is the index of the modality, T is a preset sequence length, and τ is a time step.
3. The method for constructing and collaborative positioning of the brain-inspired navigation map fusion of the unmanned cluster according to claim 1, characterized in that, In step S2, on the basis of spatio-temporal synchronization, map different modal coding units into the nodes of the hypergraph. The specific representation of the hypergraph node set V is as follows: V = V vis ∪V rad ∪V iner Among them, V vis ={v p |v p ∈F vis} represents the visual nodes, V rad ={r q |r q ∈F rad} represents the auditory nodes, V iner ={i s |i s ∈F iner} represents the inertial nodes; Dynamically generate a hyperedge e based on spatio-temporal correlation and modal complementarity m Connect multi-modal nodes v p ,r q ,i s , the calculation of the spatio-temporal correlation degree comprehensively considers the feature similarity and time factors between nodes. Only when the correlation degree is greater than the dynamic correlation threshold and the timestamp difference is within the tolerance range, a hyperedge will be generated, which simulates the dynamic synchronous activation of cross-modal neural clusters in the cerebral cortex. The hyperedge set E(t) at time t is specifically represented as follows: Among them, represents the spatio-temporal correlation degree between node v p and node r q , λ edge is the time decay coefficient of the hyperedge, taking the classical value λ edge = 0.05s -1 , t represents the current moment, t p and t q are the timestamps of node v p and r q respectively, θ e (t) is the dynamic correlation threshold, which is adaptively adjusted, and the initial value is θ e (0) = 0.6, represents the timestamp difference of the inertial node i s , ΔT is the maximum tolerance interval; Design convolutional kernels for modal perception to hierarchically aggregate heterogeneous features. Features of different modalities can be independently learned and projected, so as to extract and integrate multi-modal information. The specific representation is as follows: Among them, H (l+1) represents the hypergraph feature matrix at the (l + 1)-th layer, which is obtained by convolving the feature matrix H (l) at the l-th layer. σ(·) is the pulse trigger function of the threshold-firing mechanism θ th is the threshold, D v is the node degree matrix, A (l) is the hyperedge-node incidence matrix, is the transpose matrix, D e is the hyperedge degree matrix, W (l) = diag(W vis , W rad , W iner ) is the modal diagonal weight matrix; Process the input hypergraph feature matrix through spatio-temporal attention pooling, and use the attention mechanism to screen out important feature information, with the pooling function Pool attn (·) and the semantic class prototype matrix C sem The specific representation is as follows: C sem = [c1, c2,..., c K where |V| is the total number of nodes, h n is the n-th eigenvector in the hypergraph feature matrix H, u q is a learnable query vector with dimension d h , C sem represents the semantic class prototype matrix, which consists of K semantic class prototype vectors c1, c2,..., c K ; MLP(·) is a multi-layer perceptron, which is a neural network structure composed of multiple fully connected layers, represents the feature vector related to the k-th semantic class, is the feature concatenation operator; Dynamically select important feature information according to different query requirements, and use a predefined semantic category prototype matrix for semantic binding to generate the local high-order brain-like map M of the agent. The specific representation is as follows: where M is the local high-order brain-like map of the agent, which is the feature matrix H of the last layer of the hypergraph (L) the nth eigenvector in, and ⊙ represents the semantic binding operation, i.e., element-wise multiplication.
4. The method for constructing and collaborative positioning of the brain-inspired navigation map fusion of the unmanned cluster according to claim 1, characterized in that, In step S3, the brain-inspired local map established by intelligent monomer i at time t forms a map representation that includes spatial, semantic, and uncertainty information, as follows: Among them, the spatio-temporal hypergraph includes objects as nodes and geometric topological relationships as hyperedges, is the semantic feature tensor, is the probability confidence matrix, p is the observation distribution, q is the reference distribution, respectively representing the semantic feature distributions of agents i and j in the overlapping region Ω ij among them; By calculating the geometric difference and semantic feature difference of agents i and j in the overlapping area, the conflict degree D is obtained ij Specifically, it is expressed as follows ij Specifically, it is as follows Among them, d geo is the geometric difference based on the hyperedge topological Hausdorff distance, k is the semantic category index, and the multi-modal integration proportional weight coefficient α d = 0.6, β d = 0.4, sup represents the supremum, inf represents the infimum, and x and y are points in the spatio-temporal hypergraph respectively; To achieve spatio-temporal alignment between agents and ensure the consistency of their map information in space and time, it is necessary to use the significant feature points in the brain-inspired local map as the anchor set geometry A ij with the minimum geometric error, and then solve for the optimal SE(3) transformation matrix T ij and the time delay compensation amount Δt ij , which is specifically expressed as follows: Among them, A ij is the set of anchor points a composed of anchor points a with high confidence, and the confidence Ψ of the anchor point a is taken a ≥ 0.85, T ij is the SE(3) transformation matrix from agent i to j, and Δt ij is the time delay compensation amount, and ||·|| F represents the Frobenius norm; In the case of certain characteristic contradictions, the semantic features are fused and processed by means of probability distribution, simulating the information integration mechanism of the hippocampal-neocortical loop, and the Gaussian mixture model is used to probabilistically reconstruct the semantic features of agent j. The specific representation is as follows: Among them, is the semantic feature reconstruction result of agent j at time t, represents the k-th Gaussian distribution with mean μ k and covariance matrix Σ k , and ι k is the Gaussian distribution weight. Through experiments, the conflict sensitivity coefficient γ is calibrated to be 1.2; The selected reference agent is marked as reference agent 0. By suppressing the weights in the regions with high conflict with agent 0, the global consistency of the map is achieved, and the local brain-like maps of multiple agents are fused into a globally consistent high-confidence brain-like map. It is expressed as follows: Among them, is a globally consistent brain-like map, is the union operation, representing the merging operation of the local maps of the agents with subscript i from i = 1 to i = N, T i0 is the transformation matrix from agent i to the reference agent 0, calibrated by prefrontal cortex experiments with λ = 0.8, represents the composition operation, that is, the meaning is to rotate and translate the local map i0 through the transformation matrix T to align it with the map of the reference agent.
5. The method for constructing and collaborative positioning of the brain-inspired navigation map fusion of the unmanned cluster according to claim 1, characterized in that, In step S4, particle filtering is initialized, and a particle set is initialized by randomly sampling from the spatial feasible region. Initial uniform sampling is adopted, and the specific representation is as follows: where N is the total number of particles in the feasible region, and each particle contains three-dimensional coordinates and semantic information expressed as is the uniform sampling of the initial weights and evenly distributed; The movement of the brain-inspired agent in the environment requires the prediction of the state of particles over time, and the three-dimensional coordinates of the particles are predicted, which is specifically expressed as follows: Among them, represents three-dimensional position information, and f(·) is the motion model taking the classical uniformly accelerated model equation where x is the position in a certain direction, is the instantaneous velocity in the x-axis direction, a x is the acceleration in the x-axis direction, Δt is the time step, θ is the angle of the agent, is the angular velocity of the agent, is the particle state at the previous moment, u k is the control input including linear velocity and angular velocity, ∈ k is Gaussian noise, where the covariance Q is calibrated by the sensor error; The estimation of the three-dimensional spatial position needs to consider the matching degree between the sampled particle position and the predicted particle position as well as the semantic confidence of the current position, and then the particle weight which can reflect the credibility of the spatial position represented by the particle. The observation update and weight correction are as follows: Among them, is the weight of the i-th particle at the k-th moment, which is updated from the particle weight at the (k - 1)-th moment, κ k is the observed three-dimensional coordinate, is the predicted three-dimensional coordinate, is the noise standard deviation, η(·) is the semantic confidence weight function, is the semantic confidence at the current position, specifically expressed as: wherein, is the observation value of the particle on the m-th semantic feature, is the reference value at the corresponding position in the global map , and is the noise standard deviation, and the experimentally calibrated gain coefficients α η = 0.5, β η = 0.
7.
6. The method for constructing and collaborative positioning of the brain-inspired navigation map fusion of the unmanned cluster according to claim 5, characterized in that, The policy π for optimizing the particle distribution is obtained by reinforcement learning * , by defining the state space s t , the action space a t and the reward function r(·), the optimal policy is found using the reinforcement learning algorithm as follows: Among them, To find the expectation, π is the policy function, γ t ∈ [0, 1] is the discount factor, and the initial value is taken as γ t = 0.95, the state space s t , the action space a t and the reward function r(·) are as follows: a t = [Δτ t , Δσ x , Δσ y , Δσ θ T where, e t is the positioning error at the current moment, N eff is the number of effective particles, is the conflict degree between the current particle set and the map, and λ1, λ2, λ3, λ4 are the weights for measuring the positioning accuracy, particle diversity, conflict penalty, and policy stability of the reward function. The initial values are taken as λ1 = 0.4, λ2 = 0.3, λ3 = 0.1, λ4 = 0.2, is the particle entropy, is the semantic gradient, e t-K is the historical error, Δτ t is the resampling threshold τ t is the incremental adjustment amount of Δσ x , Δσ y , Δσ θ is the semantic noise covariance Σ sem is the adjustment amount; According to the optimal policy π * Resample the particles, dynamically adjust the distribution of the particles, to improve the diversity and effectiveness of the particles, and achieve particle resampling and the optimal position estimation It is expressed as follows: wherein, is the optimal position estimate, is an indicator function that takes the value of 1 when the condition holds and 0 otherwise, and τ th is the threshold value taken as 0.
8.
7. The method for constructing and collaborative positioning of the brain-inspired navigation map fusion of the unmanned cluster according to claim 5, characterized in that, In step S2, perform hierarchical compression on the high-order brain-like semantic map: Geometry layer: Use an octree structure to store space occupancy information; Semantic layer: Use knowledge distillation technology to compress the semantic prototype matrix into a low-rank representation; Compression ratio: 15%-30% of the original map size.
Citation Information
Patent Citations
Multi-unmanned aerial vehicle cooperative task planning method based on hybrid enhanced intelligence
CN113324545A
Multi-unmanned aerial vehicle cooperative mapping and sensing method and system based on semantic consistency
CN117152249A
Digital-analog hybrid unmanned cluster brain-like crowd-sourcing collaborative navigation method
CN118643858A
Cited By
Distributed multi-agent cooperation method based on mixed implicit neural field
CN120562462A
Post practical training culture system and culture method based on AR virtual vision
CN120726869A
Risk assessment method and system for multi-source heterogeneous airspace data of unmanned aerial vehicle
CN120833691A
Multi-robot collaborative navigation method based on anonymous perception enhancement and adaptive decision
CN120993904A
Two-dimensional map annotation collaborative editing method based on geometric perception
CN121095390A