Station intelligent navigation and decision support method and system fusing audio positioning and multi-modal data

By extracting multimodal features through a sound source spatial analysis network and a structure perception network, and combining them with a graph attention network for cross-modal dynamic transmission path selective fusion, the problem of insufficient navigation accuracy and reliability in complex station environments is solved, and high-precision autonomous navigation is achieved.

CN121898428APending Publication Date: 2026-04-21JIANYUAN FUTURE CITY INVESTMENT DEV CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
JIANYUAN FUTURE CITY INVESTMENT DEV CO LTD
Filing Date
2026-01-30
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

In complex field environments, existing technologies rely on single-modal sensor navigation schemes which are susceptible to interference and lack semantic understanding capabilities. Multimodal fusion methods fail to delve into the deep spatiotemporal correlations and semantic consistency between audio and vision, resulting in insufficient navigation accuracy and reliability.

Method used

Multimodal features are extracted by sound source spatial analysis network and structure perception network, and cross-modal dynamic transmission path selective fusion is performed by graph attention network to generate scene-adaptive multimodal decision features. Semantic association is performed by event-target decoupling module to construct adaptive navigation decision.

Benefits of technology

It achieves high-precision and high-reliability autonomous navigation in complex site environments, enhances the system's robustness to noise, interference and dynamic obstacles, and is suitable for various complex site environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121898428A_ABST
    Figure CN121898428A_ABST
Patent Text Reader

Abstract

The invention provides a station intelligent navigation and decision support method and system fusing audio positioning and multi-modal data. The method comprises the following steps: preprocessing a station environment multi-mode signal to obtain an audio frequency spectrum feature and a three-dimensional visual feature; processing the audio features through a sound source space analysis network and obtaining spatially enhanced audio features; processing the three-dimensional visual features through a structure sensing network and obtaining structure-guided visual fusion representation; constructing a cross-modal dynamic conduction path based on reliability measurement and space consistency of the two types of features, realizing adaptive feature fusion, and generating a scene-adaptive multi-modal decision feature; inputting the multi-modal decision features into an event-target decoupling module, generating a sound source-target incidence matrix and a target segmentation mask through a graph attention network, and applying semantic association constraints; and generating a navigation sequence through hierarchical decision based on the sound source position information analyzed by the incidence matrix, the segmentation mask, the three-dimensional semantic map and the real-time dynamic information. According to the method, the sensing precision, the fusion efficiency and the scene self-adaptive capability of the navigation system in a complex station environment are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of intelligent navigation technology, specifically relating to a method and system for intelligent navigation and decision support at stations that integrates audio positioning and multimodal data. Background Technology

[0002] With the continuous expansion of logistics warehouses, transportation hubs, and large factories, and the ongoing improvement of automation levels, the demand for highly reliable and high-precision intelligent navigation for autonomous mobile robots or unmanned equipment is becoming increasingly urgent. Traditional navigation solutions mainly rely on single-modal sensors, such as SLAM based on LiDAR, visual SLAM, or pre-laid magnetic guidance. However, in dynamic, complex, and unstructured site environments, such solutions have significant limitations. For example, in scenarios with numerous visually similar areas, frequently moving dynamic obstacles, drastic changes in lighting, or strong reflection interference, purely visual methods are easily affected by interference, leading to positioning failures; solutions relying solely on lasers struggle to effectively perceive special structures such as glass and shelf gaps, and lack semantic understanding of the environment. Furthermore, the background noise, acoustic reverberation, and multipath reflection effects prevalent in site environments also pose serious challenges to audio signal-based positioning assistance technologies.

[0003] In recent years, multimodal fusion navigation technology has gradually become a research hotspot, aiming to improve the environmental perception integrity and overall robustness of the system by fusing complementary information from multiple sensors such as vision, hearing, and inertial measurement units. However, most existing multimodal fusion methods still employ relatively simple feature stitching or post-decision-level fusion strategies, failing to deeply explore the deep spatiotemporal correlations and semantic consistency between different modal data, especially in the dynamic collaboration and feature-level fusion of audio and visual modalities, which remains insufficient. Audio signals contain rich location and event information, but are susceptible to environmental noise and reverberation interference; visual signals can provide accurate spatial structure and rich semantic information, but are sensitive to occlusion, lighting changes, and areas lacking texture. Therefore, how to effectively enhance and purify audio location information at the feature level, and dynamically and adaptively fuse it with the geometric structure and semantic information of vision to generate reliable navigation decisions highly adapted to the scene, has become a key technical challenge that urgently needs to be overcome in the field of high-precision adaptive navigation for current sites.

[0004] In summary, there is an urgent need for an intelligent navigation method that can fully utilize the complementary advantages of multimodal information, possess strong anti-interference capabilities, deep cross-modal understanding capabilities, and efficient scene adaptive decision-making capabilities, in order to address the challenges of high-precision and high-reliability autonomous navigation in complex site environments. Summary of the Invention

[0005] To overcome the shortcomings of the prior art, the present invention provides a site navigation method and system that can deeply integrate auditory and visual information, has strong anti-interference capabilities and scene adaptability, so as to achieve robust and accurate autonomous navigation in complex environments.

[0006] In a first aspect, embodiments of this application provide a method for intelligent navigation and decision support at a facility that integrates audio positioning and multimodal data, the method comprising: The collected multimodal signals from the field environment are preprocessed to obtain the spectral characteristics of multichannel audio and the three-dimensional visual characteristics of multi-view video. The spectral features are processed by a sound source spatial analysis network, wherein the sound source spatial analysis network includes a azimuth analysis unit based on multi-channel interferometry and direction of arrival estimation, and an anti-interference filtering unit that performs sound field feature purification, and outputs spatially enhanced audio features. The three-dimensional visual features are processed by a structure-aware network to extract the three-dimensional geometric structure features and semantic object features of the scene, and a hierarchical feature aggregation mechanism is used to generate a structure-guided visual fusion representation. Based on the reliability metric of the spatially enhanced audio features and the spatial consistency information of the structure-guided visual fusion representation, a cross-modal dynamic transmission path is constructed. Through a parameterized selection unit, multimodal features are selectively fused along this path to generate scene-adaptive multimodal decision features. The multimodal decision features are input into the event-target decoupling module, which extracts the temporal distribution information of acoustic events and the spatial composition information of visual targets respectively. The graph attention network is used to associate the sound source location hypothesis with the target instance, outputting the sound source-target association matrix and the target segmentation mask. Correlation constraints are applied to ensure the semantic association between the two outputs. Based on the sound source location information, target segmentation mask, station 3D semantic map and real-time dynamic information parsed from the sound source-target correlation matrix, a hierarchical strategy is used to generate a model output navigation decision sequence.

[0007] Secondly, embodiments of this application provide a station intelligent navigation and decision support system that integrates audio positioning and multimodal data, applied to the method described in the first aspect, the system comprising: The multimodal signal preprocessing module is used to preprocess the acquired multimodal signals from the site environment to obtain the spectral characteristics of multi-channel audio and the three-dimensional visual characteristics of multi-view video. A sound source spatial analysis network module is used to process the spectral features, wherein the sound source spatial analysis network module includes a azimuth analysis unit based on multi-channel interferometry and direction of arrival estimation, and an anti-interference filtering unit that performs sound field feature purification and outputs spatially enhanced audio features. The structure-aware network module is used to process the three-dimensional visual features, extract the three-dimensional geometric structure features and semantic object features of the scene, and generate a structure-guided visual fusion representation using a hierarchical feature aggregation mechanism. The cross-modal dynamic fusion module is used to construct a cross-modal dynamic transmission path based on the reliability metric of the spatially enhanced audio features and the spatial consistency information of the structure-guided visual fusion representation. The parameterized selection unit selectively fuses multimodal features along the path to generate scene-adaptive multimodal decision features. The event-target decoupling module is used to input the multimodal decision features into the event-target decoupling module, extract the temporal distribution information of acoustic events and the spatial composition information of visual targets respectively, associate the sound source location hypothesis with the target instance through a graph attention network, output the sound source-target association matrix and the target segmentation mask, and apply correlation constraints to ensure the semantic association between the two outputs; The navigation decision generation module is used to generate a navigation decision sequence based on the sound source location information, target segmentation mask, station 3D semantic map and real-time dynamic information parsed from the sound source-target correlation matrix, and through a hierarchical strategy to generate a model outputting a navigation decision sequence.

[0008] Thirdly, embodiments of this application provide an electronic device, including: processor; Memory used to store processor-executable instructions; The processor is configured to implement the intelligent station navigation and decision support method as described in the first aspect when executing the instructions.

[0009] Fourthly, embodiments of this application provide a computer-readable storage medium storing a program that instructs a device to execute the intelligent navigation and decision support method for stations that integrates audio positioning and multimodal data as described in the first aspect.

[0010] Compared with existing technologies, the advantages of this invention are as follows: By constructing a specialized sound source spatial analysis network, it achieves refined extraction of directional information in audio signals and effective suppression of interference components, significantly improving the reliability and accuracy of auditory perception. It abandons simple feature splicing and innovatively introduces dynamic transmission paths and parameterized selection units based on reliability metrics and spatial consistency, realizing adaptive and on-demand deep fusion of audio and visual features at key nodes, enhancing the system's adaptability to complex scenes. Through an event-target decoupling module, acoustic events and visual targets are separated and extracted in a correlated manner, simultaneously generating a sound source-target association matrix and segmentation mask with strong semantic association, providing richer and more structured environmental semantic information for navigation decisions. Adopting a three-layer decision architecture (task-path-motion control) combined with a real-time feedback mechanism, it achieves coherent and adaptive planning from high-level task objectives to low-level motion actions, effectively responding to dynamic environmental changes. Through the optimization design of the entire process from feature purification, dynamic fusion, semantic decoupling to hierarchical decision-making, the system as a whole has stronger robustness to noise, interference, dynamic obstacles and scene changes, and is suitable for a variety of complex site environments. Attached Figure Description

[0011] Figure 1 This is a schematic flowchart of a site intelligent navigation and decision support method that integrates audio positioning and multimodal data, provided as an embodiment of this application.

[0012] Figure 2 The architecture diagram of the intelligent navigation and decision support system for the station that integrates audio positioning and multimodal data provided in this application.

[0013] Figure 3 A schematic diagram of an electronic device provided in an embodiment of this application. Detailed Implementation

[0014] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them.

[0015] It should be noted that in the embodiments of this application, "at least one" refers to one or more, and "more than one" refers to two or more. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used in the specification of this application is for the purpose of describing particular embodiments only and is not intended to be limiting of this application.

[0016] Based on the embodiments described in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0017] Example 1

[0018] Figure 1 This is a schematic flowchart of a site intelligent navigation and decision support method that integrates audio positioning and multimodal data, provided as an embodiment of this application. Figure 1 As shown, a site intelligent navigation and decision support method integrating audio positioning and multimodal data includes: S1. Preprocess the acquired multimodal signals from the site environment to obtain the spectral characteristics of multi-channel audio and the three-dimensional visual features of multi-view video. Preliminary processing of the multi-channel audio and multi-view video from the site environment extracts spectral features and three-dimensional visual features (H, W, D), providing standardized input data for subsequent analysis.

[0019] Specifically, in this embodiment, the hardware configuration and data are synchronized. The audio acquisition device uses a 4-microphone circular array, mounted on top of the inspection robot, with an array diameter of 30cm, a sampling rate of 48kHz, and 16-bit precision. A lightweight binocular stereo matching network (such as RAFT-Stereo) or sparse feature matching + triangulation is employed to balance real-time performance and accuracy. All sensors achieve millisecond-level synchronization via hardware trigger signals, ensuring that audio frames and video frames are strictly aligned at the same moment.

[0020] Multichannel audio preprocessing includes: a) Framing and noise reduction: The continuous 4-channel audio stream is framed in 50ms lengths with a 25ms frame shift. Preliminary noise reduction is performed on each frame using spectral subtraction to suppress steady-state background noise (such as fan noise). b) Spectral feature extraction: A 512-point STFT (Short-Time Fourier Transform) is performed on each channel's audio frame to obtain a complex spectrum. A 64-dimensional Mel spectrum (Mel scale) is extracted, and its logarithmic energy is calculated to obtain features with a shape of [4,64]. Further calculation of the MFCC (Mel frequency cepstral coefficients, taking the first 13 dimensions) for each channel enhances the expressive power of speech and event features. The final output is a multichannel spectral feature containing the Mel spectrum and MFCC, with a shape of [4,77] (64-dimensional Mel + 13-dimensional MFCC).

[0021] Multi-view video preprocessing includes: a) Image correction and alignment: Distortion correction is performed on the raw images from each camera (based on pre-calibrated intrinsic parameters). Using the extrinsic parameter calibration results, the images from the four views are projected onto the same world coordinate system to form an intermediate representation of the panoramic image stitching. 3D visual feature reconstruction: A real-time stereo matching algorithm (Semi-Global Matching) is used, combining front-to-back and left-to-right camera pairs to calculate a dense disparity map. Based on the camera geometry, the disparity map is converted into a 3D point cloud (approximately 300,000 points are generated per second), with each point containing (x, y, z) coordinates and (R, G, B) color information. The point cloud is voxelized and downsampled (voxel grid size 3 cm), preserving the color mean to generate a regular 3D voxel grid. The output 3D visual features are a floating-point tensor of shape (Dx, Dy, Dz, C), where (Dx, Dy, Dz) are the spatial grid size, and C is the number of feature channels (e.g., color channels or the feature dimension extracted later).

[0022] Spatiotemporal alignment and output are performed, including: a) Timestamp alignment: A unified timestamp (1ms precision) is applied to each audio frame and its corresponding 3D visual feature block. b) Normalization: The audio Mel-ray spectrum is normalized for mean and variance (using historical statistical values). 3D voxel color values ​​are normalized to the range [0,1].

[0023] It also includes anomaly handling mechanisms, including: a) Frame loss detection: If data from a sensor is lost, the system uses linear interpolation or nearest neighbor copying to fill in the gaps and marks the confidence level of that frame as a partial estimate. Data buffering: A 2-second sliding window buffer is set to ensure that subsequent modules can acquire a continuous and stable multimodal data stream.

[0024] Hardware synchronization ensures spatiotemporal consistency, laying the foundation for subsequent cross-modal alignment. Lightweight 3D reconstruction balances real-time performance with scene representation accuracy. Standardized output and a unified tensor format facilitate direct processing by deep learning models. The complete preprocessing workflow, from raw sensor data to standardized and aligned multimodal features, provides high-quality input for subsequent sound source analysis, structure perception, and fusion decision-making.

[0025] S2. The spectral features are processed through a sound source spatial analysis network, wherein the sound source spatial analysis network includes a azimuth analysis unit based on multi-channel interferometry and direction-of-arrival estimation, and an anti-interference filtering unit that performs sound field feature purification, outputting spatially enhanced audio features. The sound source spatial analysis network processes the audio features, separates multiple channels and models sound field interference, estimates the sound source azimuth and suppresses environmental reverberation and multipath interference, outputting enhanced spatial audio features.

[0026] Audio features are processed through a spatial source resolution network, sound source localization is performed based on a multi-channel interference model, and environmental reverberation and multipath interference are suppressed to output enhanced spatial audio features. Specifically, in this embodiment, the processing of the spectral features through the spatial source resolution network includes a directional resolution unit processing procedure, specifically including: S2.1. Perform spatiotemporal-frequency decomposition on the spectral characteristics of the multi-channel audio signal, and calculate the time delay difference between each channel using the generalized cross-correlation method to generate an initial time delay difference matrix. First, obtain the complex spectrum through short-time Fourier transform, and then calculate the time delay difference between any two microphone channels using the generalized cross-correlation-phase transform (GCC-PHAT) algorithm. Specifically: for the m-th and n-th channels, calculate their cross-power spectrum, perform a phase transform, and then obtain the generalized cross-correlation function through inverse Fourier transform. The peak position of this function is the estimated time delay τ between the two channels. mn By pairwise combining the four channels, a total of six delay estimates are generated, forming a 4×4 symmetric delay difference matrix T, where T[m,n]=τ mn The diagonal elements are all 0. This matrix reflects the relative time difference of sound waves arriving at each microphone.

[0027] S2.2. Based on the initial time delay difference matrix, the initial azimuth and elevation angles of the sound source are estimated using a direction-of-arrival (DOA) estimation algorithm. Based on the geometric layout of the microphone array (a 30cm diameter ring array is used in this embodiment), a geometric relationship model between the sound source direction and the theoretical time delay is established. The time delay difference matrix T is input into the Multiple Signal Classification (MUSIC) algorithm. By constructing a spatial spectrum function P(θ,φ), a spectrum peak search is performed within the range of azimuth θ∈[-180°, 180°] and elevation φ∈[-45°, 45°] to find the angle pair that maximizes P(θ,φ). , The initial azimuth and elevation angles of the sound source are estimated using these values. To improve real-time performance, the search step size is set to 5° for azimuth and 10° for elevation, and a local fine-tuning search is performed near the historical estimates.

[0028] S2.3. A deep neural network is used to jointly model the spectral features after the spatiotemporal decomposition and the initial time delay difference matrix to extract deep acoustic features containing spatial information. The four-channel Mel spectrum is used as the first input. The time delay difference matrix is ​​flattened into a 6-dimensional vector and normalized, then expanded to a 64-dimensional vector through a fully connected layer as the second input. A two-stream convolutional neural network is used for processing: the time-frequency branch uses two convolutional layers (kernel size 3×3) to extract spectral features; the time delay branch first performs feature transformation through two fully connected layers, and then concatenates the output of the convolutional branch with the feature vector in the feature dimension. The concatenated features are further processed by a residual module to extract a high-dimensional representation. Finally, a 256-dimensional deep acoustic feature vector is output. Specifically, the dual-stream convolutional neural network uses global average pooling at the end of the time-frequency branch to convert the feature map into a vector, which is then concatenated with the transformed vector from the time-delay branch. This vector is then passed through a fully connected layer in the residual module and finally projected uniformly to 256 dimensions. This feature simultaneously encodes audio content information and spatial cues.

[0029] S2.4. The initial azimuth and elevation angle estimates are fused with the depth acoustic features to form azimuth coding features. The initial azimuth angle... and pitch angle The angles are respectively converted into 4-dimensional vectors using sine and cosine encoding: sin(θ), cos(θ), sin(φ), cos(φ). These two 4-dimensional vectors are then concatenated to form an 8-dimensional angle-encoded vector. This angle-encoded vector is then copied and extended to be combined with the depth acoustic feature vector. Alignment is performed in the time dimension. Finally, the angle encoding vector is aligned with... The features are concatenated along the feature dimensions to form a 264-dimensional (256+8) orientation coding feature. This feature not only contains depth information of the audio, but also explicitly encodes the three-dimensional direction information of the sound source, providing accurate spatial guidance for subsequent anti-interference processing.

[0030] Specifically, in this embodiment, the process of processing the spectral features through the sound source spatial analysis network further includes an anti-interference filtering unit process, specifically including: S2.5. Based on the aforementioned azimuth coding features, reverberation modes in the spectrum are identified using a time-frequency attention mechanism. The azimuth coding features... (The image, in shape T, 2^64, where T is the number of time frames,) is reconstructed into a time-frequency representation, assuming T time frames and F frequency sub-bands. This is input to a dual-channel time-frequency attention module, which includes a time attention head and a frequency attention head. The time attention head calculates the importance weight for each time frame, focusing on periods of high direct sound activity; the frequency attention head calculates the importance weight for each frequency sub-band, focusing on clear frequency bands less affected by reverberation. The outputs of the two attention heads are fused through a learnable gating mechanism to generate a comprehensive attention map. (Shapes T, F). This attention map assigns lower weights to time-frequency regions with severe reverberation (characterized by slow energy decay and diffuse distribution) and higher weights to regions dominated by direct sound, thereby identifying reverberation modes in the spectrum.

[0031] S2.6. Dynamically generate an adaptive filter coefficient matrix based on the identified reverberation patterns. (Using a reverberation pattern attention map...) As the primary input, combined with the directional information contained in the orientation coding features (especially the angle coding part), it is fed into a lightweight multilayer perceptron (MLP). This MLP is based on the reverberation intensity of the current frame (via... The mean value assessment and spatial distribution characteristics of reverberation (combined with the source direction information to infer the possible source direction of early reflections) dynamically generate a set of two-dimensional filter coefficient matrices. The size of the filter coefficients is matched with the convolution kernel size (3×3 in this embodiment) and the number of output channels. The generation process is performed frame by frame, allowing the filter to adapt to changes in the reverberation state of the acoustic environment in real time. For example, a milder suppression coefficient is used in open areas, while a stronger suppression coefficient is used in enclosed workshops.

[0032] S2.7. The azimuth coding features are subjected to convolutional spatial filtering using the adaptive filter coefficient matrix to suppress reverberation components. The azimuth coding features are treated as a multi-channel time-frequency map, and a dynamically generated filter coefficient matrix is ​​used. As the convolution kernel, a two-dimensional convolution operation is performed on the time-frequency map. This convolution is specifically designed for the time-frequency domain, and the convolution operation attenuates the reverberation mode attention map. The energy is labeled as the high reverberation region while preserving the signal components reaching the direct sound region. After convolution, a nonlinear activation (ReLU) and layer normalization are performed to stabilize the feature distribution and enhance the sparsity of the features. The output is a spatially filtered feature with preliminary reverberation suppression. .

[0033] S2.8. The spatially filtered features are input into the multipath interference suppression module. A method combining geometric acoustics-based multipath reflection modeling and Wiener filtering is used to predict and suppress multipath reflection interference components. First, based on the geometric layout of the microphone array and the initial azimuth and elevation angles of the sound source, combined with a preset site environment geometric model (such as the approximate positions of walls and equipment), the main multipath reflection paths are calculated using the mirror source method. For each predicted reflection path, its theoretical time delay and attenuation coefficient relative to the direct sound are calculated to construct the impulse response model of multipath reflection. Then, the spatially filtered features are used... As input, a multipath suppression filter is constructed in the frequency domain using Wiener filtering theory. This filter, based on a predicted reflection model, weights the signal in the frequency domain, attenuating frequency components related to the reflection path. Specifically, by calculating the power spectral density ratio of the signal to the noise (referring to multipath reflections), the filter gain for each frequency band is dynamically adjusted to suppress discrete multipath reflection interference while preserving the direct sound. The final output further suppresses the characteristics of multipath interference. .

[0034] S2.9. Enhance and reconstruct the interference-suppressed features, outputting the final spatially enhanced audio features. Input features after multipath suppression. A lightweight U-Net-based encoder-decoder network is used for feature enhancement and reconstruction. The encoder uses three convolutional layers and downsampling to compress features into a low-dimensional bottleneck representation, extracting robust core sound source information. The decoder uses three deconvolutional layers and upsampling to progressively restore feature resolution. Skip connections are used between the encoder and decoder to transfer early azimuth-encoded features. High-frequency detail information is incorporated into the decoding process to compensate for useful signal components that may be lost during filtering and suppression. The final output of the decoder is then processed through a 1×1 convolutional layer to adjust the dimensions, resulting in the final spatially enhanced audio features. (Shape is T, 256). This feature has a higher signal-to-noise ratio, clearer spatial directivity, and retains complete sound source orientation cues, providing high-quality and reliable auditory perception input for subsequent cross-modal fusion and event localization.

[0035] It also includes a feature quality control mechanism. After the anti-interference filtering unit outputs the final spatially enhanced audio features, the system introduces a feature quality evaluation and control module to ensure the reliability of the output features. Specifically, it includes: S2.10. A feature quality evaluation module is set at the output of the anti-interference filtering unit to calculate the signal-to-noise ratio and azimuth consistency index of the features. Specifically, the spatially enhanced audio features (denoted as...) are input to the output of the anti-interference filtering unit. ), and the azimuth angle of the current frame estimated by the azimuth resolution unit in step S2.4. Quality metric calculation: Signal-to-noise ratio (SNR) metric: from The signal and noise subspaces are separated. This can be achieved by eigenvalue decomposition of the characteristic covariance matrix; larger eigenvalues ​​correspond to the signal, and smaller eigenvalues ​​correspond to the noise. Signal-to-noise ratio (SNR) is the key performance indicator. Calculated as the logarithm of the ratio of signal subspace energy to noise subspace energy. Orientation consistency index: from... Inferring an azimuth angle from the middle and back This can be achieved using a lightweight inverse regression network that takes augmenting features as input and outputs a direction estimate. (Calculation...) Compared with the original azimuth angle The absolute difference Δθ. Orientation consistency index. Defined as 1.0 / (1.0+Δθ), the smaller the difference, the higher the score (closer to 1). Output two scalar indicators. and .

[0036] S2.11. When the signal-to-noise ratio (SNR) is lower than a preset threshold, the feature re-extraction process is initiated. The system presets two thresholds: the SNR threshold and the feature re-extraction threshold. (e.g., the rating corresponding to 15dB) and orientation consistency threshold (e.g., 0.8). Scenario 1: Signal-to-noise ratio is too low, when... < When this happens, the feature re-extraction process is initiated. The specific steps are: a. The system temporarily discards the current frame's features. b. Return to the input of the anti-interference filtering unit (i.e., the original orientation-coded features). c. Enable a more aggressive but computationally more computationally intensive set of alternative filtering parameters (e.g., larger convolution kernels, stronger attenuation coefficients), and re-execute the filtering, suppression, and reconstruction steps. d. Output the reprocessed features. Then, a quality assessment is performed again. If it still fails to meet the standards, the frame features are marked as low quality and given lower weights in subsequent fusion. Scenario 2: Anomalous orientation consistency, when < When this happens, the azimuth calibration process is initiated. The specific steps are: a. The system determines the original azimuth angle of the current frame. The estimate may be unreliable. b. The most recent N frames (e.g., N=5) extracted from the historical cache were confirmed as high-quality azimuth sequences. c. Using a Kalman filter or a simple sliding window weighted average, smooth and predict the historical azimuth information to obtain the calibrated azimuth angle. d. using Alternative The system then feeds back the information to the azimuth-coded feature generation step (S2.5). The system regenerates the azimuth-coded features using the new azimuth angle and processes them again through the anti-interference filtering unit (standard parameters can be selected). e. Outputs the enhanced features regenerated after azimuth calibration.

[0037] S2.12. When the azimuth consistency index is abnormal, initiate the azimuth calibration process to correct the current azimuth estimate using historical azimuth information. Only when... >= and >= At that time, the current frame Only those features are marked as high-quality features and directly passed to the downstream cross-modal fusion module. All output features are accompanied by a comprehensive quality confidence label, which is... and The weighted sum is used to guide the feature trust assignment of subsequent modules (such as cross-modal fusion).

[0038] S3. The 3D visual features are processed through a structure-aware network to extract the 3D geometric structure features and semantic object features of the scene, and a hierarchical feature aggregation mechanism is used to generate a structure-guided visual fusion representation. Geometric structure information and semantic object information are extracted from the 3D visual features, and a structure-guided visual representation is formed through hierarchical fusion, enhancing the understanding of scene space and semantics. The input undergoes preprocessing. The input data comes from video streams synchronously acquired by a multi-view camera array (such as four wide-angle cameras) mounted on the site inspection robot. Real-time visual SLAM technology is used to fuse multi-view video frames to generate a dense point cloud, generating 3D visual features of the scene (i.e., point cloud data with color information) once per second.

[0039] Specifically, in this embodiment, the above steps include: S3.1. The three-dimensional visual features are processed into point cloud voxels and converted into a three-dimensional voxel mesh. A three-dimensional convolutional neural network is used to extract local geometric features to obtain an initial geometric feature map. Then, a geometric attention mechanism is used for spatial weighting to obtain enhanced geometric features. Specifically, the input is a dense point cloud (approximately 100,000 points) for the current second. The point cloud space is divided into a uniform voxel mesh (voxel size set to 5cm), with the average color and density of points within each voxel as initial values. A sparse convolutional network (MinkowskiNet) or a lightweight geometric feature extraction network based on Octree is used. This network includes a downsampling encoder and an upsampling decoder, ultimately outputting an initial geometric feature map, where each spatial location contains 256-dimensional geometric context features (such as surface curvature, normal, and connectivity). A geometric attention module is applied to the feature map. This module calculates a spatial importance weight map based on the spatial location of each voxel, local curvature changes, and relative distance to the robot body. This weight is multiplied element-wise with the initial geometric feature map to obtain the enhanced geometric features. Output enhanced geometric features Its spatial dimensions are (H,W,D), which is (200,200, 50) defined in S1.

[0040] S3.2. Multi-level semantic features of the 3D visual features are extracted using a multi-scale feature pyramid network. Simultaneously, an object detection module identifies semantic objects and their spatial locations in the scene to generate object feature vectors. These object feature vectors are then fused with the semantic features of their corresponding spatial locations to obtain semantically enhanced object features. Specifically, multiple original RGB images from the same time stamp are input. Each image is input into a pre-trained feature pyramid network (FPN) (with a ResNet-50 backbone) to extract three semantic feature maps at different resolutions (corresponding to low, medium, and high-level semantics). Using camera parameters, the 2D semantic features of each viewpoint and scale are back-projected into 3D space and aligned with a voxel grid. Multi-level 3D semantic features are then generated through voting and averaging. A real-time object detector (such as YOLOv7) is run in parallel on a 2D image to identify common objects within the depot (such as valves, pipes, toolboxes, and personnel), obtaining their category labels and 2D bounding boxes. Depth information (from point clouds) is used to upscale the detected object bounding boxes to 3D bounding boxes, locating their voxel regions. The corresponding 3D semantic features are extracted from these regions, concatenated with the object category embedding vectors provided by the detector, and then fused using a small MLP to generate semantically enhanced object features. (Each detected object corresponds to a 256-dimensional vector and its 3D center coordinates). Output 3D semantic features. and a set of semantically enhanced object features { }

[0041] S3.3 Input the enhanced geometric features and semantic enhancement object features into the hierarchical aggregation module, construct a feature aggregation weight map based on the spatial distribution information of the enhanced geometric features, spatially reweight the semantic enhancement object features according to the weight map, and then align the reweighted semantic features and the enhanced geometric features through channel concatenation and cross-modal attention mechanism. After multiple levels of iterative aggregation processing, the structure-guided visual fusion representation is generated.

[0042] Specifically, input augmented geometric features 3D semantic features Semantic enhancement object features { Construct an aggregated weight graph to... Based on this, a spatial weight map is generated through a 1x1x1 3D convolutional layer. The weighted map has higher values ​​in highly structured regions with significant geometric features (such as equipment edges and pipe junctions), and lower values ​​in open or simple-textured regions. Semantic feature space reweighting... With weighted graph Element-wise multiplication focuses semantic features on structurally important regions, yielding... .Will and The features are concatenated along the feature channel dimension to form a fused feature of [Dx, Dy, Dz, 512]. Then, this feature is combined with the object features { A cross-modal attention module is input together. This module uses fused features as the query and object features as the key and value, enabling 3D voxel features to absorb semantic information of relevant objects and align them to the same semantic space. Iterative aggregation processing is performed, with the above weighted-stitching-alignment process repeated at two levels: Level 1: Performed on higher resolution features (4x downsampling rate), focusing on local structural details. Level 2: Spatially downsampled from the Level 1 output, repeated on lower resolution features (16x downsampling rate), capturing the global scene layout. Final representation generation involves upsampling the output features from both levels to the original voxel grid size, then stitching them together, and finally reducing the dimensionality through a fusion convolutional layer. This forms a fused feature of [H, W, D, 512], outputting a structure-guided visual fusion representation. The dimensions are [H, W, D, 128]. This representation simultaneously encodes the fine geometric structure of the scene, multi-level semantic information, and the precise location and category semantics of key objects.

[0043] S4. Based on the reliability metric of the spatially enhanced audio features and the spatial consistency information of the structure-guided visual fusion representation, a cross-modal dynamic transmission path is constructed. A parameterized selection unit selectively fuses multimodal features along this path to generate scene-adaptive multimodal decision features. Combining the reliability metric of audio features and the spatial consistency of visual features, a dynamic transmission path is constructed to adaptively fuse multimodal information and generate joint decision features adapted to the current scene.

[0044] Specifically, in this embodiment, the multimodal decision features for scene adaptation include: audio input: spatially enhanced audio features output from the sound source spatial resolution network. (i.e., the characteristics of the S2.9 output, with a shape of (T, 256)), and its associated signal-to-noise ratio (SNR) specification. Orientation Consistency Index Visual input: Structure-guided visual fusion representation from the output of the structure-aware network. (i.e., the features output by S3.3, with shape [Dx,Dy,Dz, 128]), and spatial consistency information obtained by calculating the variance of its internal features (reflecting the smoothness and confidence of the three-dimensional features in space).

[0045] S4.1 Calculate the reliability metric based on the spatially enhanced audio features, calculate the spatial consistency information based on the structure-guided visual fusion representation, and generate a cross-modal confidence score by comparing the reliability metric with the spatial consistency information using a cross-validation module. Specifically, the audio reliability metric... The calculation formula is: α and β are learnable parameters, initially set to 0.6 and 0.4 respectively, reflecting the contribution of signal-to-noise ratio and orientation consistency to reliability. A scalar value is output, ranging from [0,1], with values ​​closer to 1 indicating more reliable audio features. Visual spatial consistency information. The calculation method is as follows: The variance of local feature blocks is calculated by sliding along the spatial dimension to obtain a consistency map. The average value of this map is then taken and normalized to obtain... Output a scalar value ranging from [0,1]. The closer to 1, the more spatially consistent and reliable the visual features are. Mapping to 3D space: Using the sound source azimuth angle estimated in the current frame, the audio features are projected onto... In the corresponding three-dimensional spatial grid, a sparse audio spatial feature is generated. In the projection area, calculate and The cosine similarity of corresponding location features is used to obtain the local consistency score. The local consistency score is then combined with... and It outputs cross-modal confidence scores by fusing data from a small neural network (two-layer MLP). Output rating The range is [0,1]. This score quantifies the degree of trustworthiness of collaboration between audio and video modalities at the current moment.

[0046] S4.2. Based on the cross-modal confidence score, construct a main transmission path dominated by audio features and an auxiliary transmission path supplemented by visual features, setting parameterized selection units at key nodes of the transmission paths. Specifically, construct dynamic transmission paths and parameterized selection units. The system presets two parallel processing paths: the main transmission path: dominated by audio features... Dominant pathway. This pathway assumes reliable sound source information and focuses on utilizing spatial cues from the audio source to guide the allocation of attention to visual features. Secondary pathway: based on visual features. To supplement this, this path enhances the autonomous perception capability of vision when audio is unreliable. Three key fusion nodes are set on each path: a feature alignment node, a context aggregation node, and a decision generation node. Each node has a parameterized selection unit (PSU).

[0047] S4.3 At each key node, the parameterized selection unit dynamically adjusts the activation threshold of the gating parameter matrix based on the cross-modal confidence score, and performs weighted fusion of audio and visual features through a soft selection mechanism, wherein the fusion weight is dynamically determined based on the real-time evaluation results of the reliability metric and spatial consistency information. Specifically, through node fusion and dynamic weight calculation, each parameterized selection unit performs the following operations: Gating parameter adjustment: The unit receives... As input. According to The value of dynamically adjusts the activation threshold of a gating parameter matrix. For example, when When the threshold is high, lower the threshold to make the gate easier to open and promote information flow; When the threshold is low, increase the threshold to constrain information transmission. Soft selection and weighted fusion: For the subset of audio features input to this node... and visual feature subset Calculate the real-time quality score for each, and combine it with the fusion weights. Output fusion features: A horizontal connection is provided between corresponding nodes in the main path and secondary paths, allowing... Control information interacts across paths, avoiding decision-making fragmentation caused by completely independent paths.

[0048] S4.4. The features fused from each node are input into the scene adaptation module. The feature fusion strategy is adjusted according to the current scene type, and adaptive normalization is performed. The resulting multimodal decision features for scene adaptation are output with unified dimensions. Specifically, the final features after fusion from each node are the converged output of two paths. The scene adaptation module first performs scene type identification, specifically based on… The dominant semantic category (e.g., dense pipe area, open passage, equipment operation area) and the robot's real-time position determine the current scene type. Then, a strategy adjustment is performed, with pre-set fusion strategies for different scenes. For example, in the equipment operation area, sound (e.g., alarms, operating sounds) is a key cue, and the strategy slightly increases the audio weight. In the open passage, vision is more reliable, and the strategy increases the visual weight. Next, adaptive normalization is performed: layer-level normalization of the converged features for scene perception is applied. The normalization parameters (gain and bias) are determined by a lightweight network based on the scene type and... Dynamic generation ensures the feature distribution is better suited to the current scenario. Finally, it outputs multimodal decision features adapted to the scenario. It is a fixed-dimensional vector (e.g., 1024-dimensional), which uniformly encapsulates multimodal perception information that has undergone quality assessment, dynamic fusion, and scene optimization.

[0049] S5. Input the multimodal decision features into the event-target decoupling module to extract the temporal distribution information of acoustic events and the spatial composition information of visual targets. Use a graph attention network to associate the hypothetical sound source location with the target instance, outputting a sound source-target association matrix and a target segmentation mask. Apply correlation constraints to ensure the semantic correlation between the two outputs. Extract the temporal distribution of acoustic events and the spatial composition of visual targets from the fused features, simultaneously generating a sound source distribution map and a target segmentation mask, and maintaining their semantic correlation through constraints. Input data comes from the scene adaptation multimodal decision features of the cross-modal dynamic fusion module. (The shape is, for example, [T, 1024], where T is the time step, corresponding to the audio frame sequence), and the structure-guided visual fusion representation aligned with it. (As a spatial context reference).

[0050] Specifically, in this embodiment, the above steps include: S5.1 Extract the temporal distribution features of acoustic events from the multimodal decision features through temporal convolutional networks and temporal attention mechanisms, and extract the spatial composition features of visual targets through spatial convolutional networks and spatial attention mechanisms. Establish a cross-branch information exchange channel between the temporal and spatial feature extraction processes to perform feature interaction alignment.

[0051] The acoustic event branching process includes: processing multimodal decision features through a learnable linear projection layer. The feature subset encoding the main acoustic event information is extracted from a shape [T, 512], where T is the number of time steps. First, a 4-layer Temporal Convolutional Network (TCN) is used, with each layer containing dilated convolutions (dilation coefficients of 1, 2, 4, and 8), weight normalization, and ReLU activation, progressively expanding the temporal receptive field to 31 time steps to extract multi-scale temporal context features. Next, the 256-dimensional features output by the TCN are input into a temporal attention module. This module calculates the query-key relevance of features at each time step to generate a temporal attention weight map, highlighting active periods of events (such as device startup sounds and alarm durations) and suppressing background noise periods. The final output is a temporal distribution feature. (Shape [T,256]) encodes the intensity distribution and dynamic evolution pattern of acoustic events.

[0052] The visual target branching process includes: input structure guiding visual fusion representation. (Shape: [Dx, Dy, Dz, 128]). First, a lightweight 3D spatial convolutional network is used, consisting of three 3D convolutional layers (kernel size 3×3×3). Each layer is followed by batch normalization and ReLU activation to progressively extract multi-scale spatial structural features. Then, the output 256-dimensional features are input into a spatial attention module. This module generates a 3D spatial attention weight map based on the spatial saliency of the features (calculated through 1×1×1 convolutions to obtain the importance score for each voxel), enhancing the feature response of structural regions such as target edges and corners. The final output is a spatial composition feature. (with shape [Dx,Dy,Dz,256]), which encodes the three-dimensional geometry, spatial location, and local detail information of visual targets in the scene.

[0053] The cross-branch information exchange process involves establishing a bidirectional feature interaction mechanism between the third layer of the TCN and the second layer of the 3D convolutional network. Temporal-to-spatial interaction: [The text abruptly ends here, likely due to an incomplete sentence or missing information.] After linear transformation, it becomes a Query vector Qt. After being flattened, these vectors are used as Key and Value vectors Ks and Vs. Cross-attention (Attention(Qt, Ks, Vs)) is then calculated, enabling spatial features to perceive the temporal dimension of acoustic event activity patterns. This spatial-temporal interaction... Global average pooling is performed to obtain the spatial context vector Cs, which is then used as the query vector Qs. As key and value vectors Kt and Vt, a cross-attention function Attention(Qs,Kt,Vt) is calculated to make the temporal features aware of the spatial layout context of the scene. The interacting features are then fused with their respective original features through residual connections to achieve preliminary alignment and semantic interaction of spatiotemporal features.

[0054] S5.2. Input the temporal distribution features and spatial composition features into a graph attention network, where the source location hypothesis decoded from the temporal distribution features is used as one set of nodes, and the target instance hypothesis decoded from the spatial composition features is used as another set of nodes. Calculate the attention weights between the nodes to generate a source-target association matrix representing the strength of the association between the source and the target. From the temporal distribution features... In the middle, a lightweight decoder (two-layer fully connected network) decodes the hypothetical set of sound source locations S={ }, each of which Includes estimated three-dimensional coordinates And the existence confidence level From the perspective of spatial composition characteristics In the process, the target instance hypothesis set O = { is decoded through the Region Proposal Network (RPN). }, each of which It includes a 3D bounding box, class probabilities, and feature vectors.

[0055] Construct a bipartite graph G=( using elements in S and O as nodes). ∪ Vs corresponds to the source node, and Vo corresponds to the target node. The initial features of each node are its corresponding feature vectors (source nodes use a combination of their coordinate encoding and confidence, while target nodes use a concatenation of their region features and category embeddings). A two-layer graph attention network (GAT) is used to model the relationships between nodes. The first layer calculates the attention within nodes of the same type (source-source, target-target) to enhance the contextual representation of node features; the second layer calculates the attention between nodes of different types (source-target), and the calculation process is as follows: For each sound source node Query its relationship with all target nodes { Attention coefficient of} , Where W is the learnable weight matrix, a is the attention vector, and || represents the concatenation operation. The original features of the source node and the target node are linearly transformed separately to obtain the transformed features. . This is the core computation of the attention mechanism. The concatenated vector is multiplied by the attention vector a to obtain a scalar score, representing the original attention level of the source node s to the target node o. This is the activation function for the leakage linear rectification.

[0056] Based on attention coefficient The target node features are aggregated to update the source node features, and the attention coefficient matrix A (dimension |S|×|O|) is output as the source-target correlation matrix. Each element A[i,j] of this matrix represents the source. With the goal The correlation strength is such that the higher the value, the greater the likelihood that the sound source was generated by the target.

[0057] S5.3. Input the spatial composition features into the target segmentation generator, and generate a target segmentation mask through an encoder-decoder structure and pixel-level classification. The spatial composition features... Input a 3D U-Net segmentation network. The encoder contains four downsampling stages, each consisting of two 3×3×3 convolutional layers and max pooling, progressively extracting high-level semantic features. The decoder contains four upsampling stages, each restoring resolution through deconvolution and fusing with features from the corresponding encoder layers via skip connections, progressively restoring spatial details. The final output layer uses 1×1×1 convolutions and softmax activation to generate a target segmentation mask (shape [Dx, Dy, Dz, K+1]), where K is the number of target categories (e.g., equipment, people, vehicles), and +1 represents the background category. At each voxel location, an (K+1)-dimensional probability distribution is output, representing the confidence level of belonging to each category.

[0058] S5.4 Establish mutual information maximization constraints and geometric consistency constraints, and jointly optimize the correlation matrix generation process and the target segmentation mask generation process through a multi-task loss function to ensure the semantic correlation between the correlation matrix and the target segmentation mask.

[0059] Mutual information maximization constraint: A mutual information estimation method based on contrastive learning is employed. The global pooling vector of the correlation matrix A is used. Global feature vectors of the target segmentation mask As a positive sample pair, The mask feature vectors of other samples within the batch are used as negative sample pairs. The mutual information of positive sample pairs is maximized using the InfoNCE loss function. : , Where sim is the cosine similarity and τ is the temperature parameter. The mask features are those of the k-th negative sample.

[0060] Geometric consistency constraint: Filtering out high-associations based on association matrix A ( Source-target pairs with a value >0.5. For each such pair... The Euclidean distance between the predicted coordinates of the sound source and the centroid coordinates of the corresponding target instance in the target segmentation mask is calculated and used as the geometric consistency loss. : , in This is a double summation symbol. It iterates through and sums all combinations of source indices i and target indices j. Let be the value of the element in the i-th row and j-th column of the correlation matrix A. This represents the hypothesis of the i-th sound source. With the j-th target hypothesis The strength of the association (confidence level) between them. Assuming the i-th sound source 3D coordinates . For the j-th target instance The centroid's three-dimensional coordinates are calculated. These coordinates are obtained by averaging the voxel positions belonging to the j-th target in the target segmentation mask.

[0061] Multi-task loss function: The total loss is a weighted sum of the segmentation loss, the association loss, and the constraint loss mentioned above. , in To balance hyperparameters, This is a combination of cross-entropy loss and Dice loss for the segmentation mask. The binary cross-entropy loss (if labeled) or self-supervised contrastive loss is used for the correlation matrix. For mutual information constraint loss, This represents the geometric consistency loss.

[0062] Through the above multi-task joint optimization, the model can accurately segment visual targets while learning to generate a sound source correlation matrix that is highly consistent with the spatial distribution of the targets. This ensures a strong semantic correlation between acoustic events and visual targets, providing a reliable multimodal environment understanding for subsequent navigation decisions.

[0063] S6. Based on the sound source location information, target segmentation mask, 3D semantic map of the site, and real-time dynamic information parsed from the sound source-target correlation matrix, a navigation decision sequence is generated by a hierarchical strategy, and the movement of the mobile robot or unmanned equipment is controlled based on this sequence. The 3D semantic map of the site is either a pre-generated static point cloud map with semantic annotations or a dynamic map constructed online and semantically annotated in real time using visual SLAM.

[0064] Specifically, in this embodiment, the step of generating the navigation decision sequence output by the model through a hierarchical strategy includes: S6.1 Based on the sound source direction information represented by the sound source-target correlation matrix, and combined with the target three-dimensional spatial information provided by the target segmentation mask, the possible regions of the sound source in the three-dimensional space are analyzed. For example, for a sound source highly correlated with the alarm target, its possible region is limited to the spatial range where the alarm device is located. The target segmentation mask and the three-dimensional semantic map of the site are processed respectively to generate instance information containing target spatial bounding boxes and semantic categories, as well as a navigation environment map with semantic annotations. Specifically, non-maximum suppression and probability threshold filtering are applied to the sound source distribution map to extract several high sound source probability regions. A weighted centroid is calculated for each region and a probability value is assigned to it to form a three-dimensional sound source location set {( )},in For location confidence, To estimate altitude or pitch angle mapping values, the system outputs a sound source location probability distribution map, serving as a crucial basis for task triggering. Target instance information extraction involves connected component analysis of the target segmentation mask, extracting for each independent target instance: 3D bounding box (minimum bounding box), semantic category (e.g., alarm, mobile robot, worker), and set of occupied grid points. A structured list `target_list` is output, with each target containing [bbox, class, voxels]. Navigation environment map construction includes fusing a static 3D semantic map with real-time target instance information: static layers (passages, restricted areas, equipment areas) remain unchanged. Dynamic layers overlay the 3D occupied areas of all targets in the current `target_list`, labeling their semantic categories. The output is a dynamically updated navigation environment map, `NavMap`, a 3D raster map with semantic labels.

[0065] S6.2 Construct a three-layer decision architecture comprising a task planning layer, a path planning layer, and a motion control layer. The task planning layer determines the navigation task objective and priority based on the sound source location information and target instance information. The path planning layer generates a set of feasible paths under the constraints of the task objective by combining the navigation environment map and real-time dynamic information. The motion control layer generates a sequence of navigation actions based on the set of feasible paths and the current state.

[0066] Specifically, for the task planning layer, the inputs are: a probability distribution map of sound source locations, a target instance list `target_list`, the current robot state, and a predefined task library (such as inspection, obstacle avoidance, and anomaly detection). The decision logic is: a. Task triggering: If the probability of a certain sound source... If the value exceeds a threshold and a semantically relevant target (such as an alarm) is nearby, an anomaly investigation task is triggered. b. Multi-objective optimization: Using the Pareto optimization algorithm, simultaneously optimizing: task completion benefits (such as covering more sound source points), path estimation time, and safety risks (such as approaching personnel areas). c. Outputting the current optimal task objective (e.g., proceeding to sound source A to inspect device X) and its priority. For the path planning layer, the inputs are: task objective, navigation environment map NavMap, and real-time dynamic obstacle trajectory prediction. Planning algorithm: An improved version of the spatiotemporal A* algorithm is adopted: a. Searching on a 3D grid map, the cost function comprehensively considers: path length, semantic passage cost (such as deceleration in the device area), and dynamic obstacle collision risk (based on the probability of trajectory prediction); b. Outputting multiple feasible paths and their risk assessments, forming a path set PathSet. For the motion control layer, the inputs are: path set PathSet and the robot's current state (pose, velocity, sensor data). Control strategy: Using a pre-trained reinforcement learning policy network: a. The network takes the current state and preferred path features as input. b. Output the sequence of low-level control commands, including: linear velocity, angular velocity, gimbal rotation, etc. c. The commands are output in real time at a fixed frequency (e.g., 10Hz) to control the robot's movement along the path.

[0067] S6.3 In the three-layer decision-making architecture, the task planning layer employs a multi-objective optimization algorithm, the path planning layer uses an improved A* algorithm combined with dynamic obstacle prediction, and the motion control layer executes the path-to-action conversion through a reinforcement learning policy network. Specifically, the implementation details of each layer are as follows: Task planning layer multi-objective optimization algorithm: adopts the NSGA-II framework, with the fitness function integrating task reward, time cost, and safety indicators. Path planning layer improved A* algorithm: introduces the concept of a time window to avoid conflict with predicted trajectories of dynamic obstacles; sets a semantic cost layer to ensure the planning results conform to the site behavior specifications. Motion control layer reinforcement learning network: trained using the PPO algorithm, with a reward function design that balances path tracking accuracy, smoothness, energy consumption, and safety.

[0068] S6.4. Establish a feedback mechanism between the levels of the three-layer decision-making architecture, so that the upper-layer decisions provide constraints for the lower layers, and the execution results of the lower layers are fed back to the upper layers, thereby dynamically optimizing and outputting the final navigation decision sequence. The inter-layer feedback mechanism includes: a) Top-down constraint transmission: The task planning layer transmits the safety distance constraint to the path planning layer. The path planning layer transmits the maximum curvature constraint to the motion control layer. b) Bottom-up state feedback: The motion control layer feeds back tracking errors and unexpected obstacle information to the path planning layer in real time. The path planning layer feeds back the path as unreachable or too risky to the task planning layer, triggering task replanning. c) Dynamic optimization loop: A layer-by-layer collaborative update is performed every 200 milliseconds, dynamically adjusting the task, path, and control commands based on the latest perception information and execution status. The final output is a rolling navigation decision sequence, containing the target, path points, and control commands for the next few seconds. The final output navigation decision sequence is organized in a timeline and includes: a task target sequence, a path point sequence, and a low-level control command sequence. The visualized decision flow can be displayed in real time on the monitoring interface, showing the robot's task intent, planned path, and perception results.

[0069] It also includes a human-machine collaborative interaction and progressive parameter optimization module. This module receives user commands through a semantic interaction interface, transforms user intentions into guidance signals for system decision-making, and adjusts the system output to match user expectations based on a multimodal collaborative optimization objective function. During navigation execution, based on environmental perception confidence, it dynamically adjusts navigation parameters using a progressive optimization algorithm, forming a human-machine collaborative autonomous navigation closed loop. Specifically, it performs the following operations: S7.1 User command reception and semantic parsing: Understanding the interaction format, specifically, the station operator can issue commands via voice commands (e.g., go to Pump Room 3 for inspection) or graphical selection (selecting an area on the map) through a handheld terminal app. The semantic parsing module uses speech recognition and an ASR model enhanced with specialized station terminology to convert speech into text. The intent understanding module, based on a predefined station task grammar framework, parses out: action type (e.g., go to, inspect, follow), target object (e.g., Pump Room 3, alarm equipment), constraints (e.g., avoid Area A, complete within ten minutes), and spatial association: associating the semantic target with entities in the station's 3D semantic map to determine its specific coordinates or area. Structured user commands are then output.

[0070] S7.2 User intent is converted into decision guidance signals. Guidance signal generation methods include: Task-level guidance: If the user instruction is task-level (e.g., checking equipment), it is directly inserted into the task planning layer queue as a high-priority task and marked as user-specified. Path-level guidance: If the instruction contains path constraints (e.g., avoiding area A), it is converted into a cost map modification signal in the path planning layer, temporarily increasing the passage cost in area A. Motion-level guidance: If the instruction is for fine-grained control (e.g., slow passage), it is converted into a parameter adjustment signal in the motion control layer, temporarily reducing the maximum speed. Multimodal collaborative optimization objective function adjustment: A user intent matching term is introduced into the system's overall optimization objective function. ,in Measure the current decision of the system With user instructions The degree of alignment between task objectives, paths, and behaviors is represented by γ, which is the user trust coefficient (default 0.3, adjustable based on user identity). This loss term is incorporated into the multi-objective optimization function of the task planning layer, enabling the system to proactively favor the direction desired by the user during autonomous decision-making.

[0071] S7.3 Environmental Perception Confidence Assessment. Input: Sound source reliability. Visual spatial consistency Cross-modal confidence The overall confidence level is calculated using the following formula: , The product approach is used to reinforce the "weakest link" effect; low confidence in any modality will lower the overall score. The output is the environmental perception confidence level at the current moment. ∈[0,1].

[0072] S7.4 Progressive dynamic adjustment of navigation parameters. The system maintains a set of adjustable navigation parameters, including: safety radius. Maximum speed Number of planned retries User intent weight γ, adjustment strategy (based on Piecewise functions): High confidence patterns ( >0.7): Adopt an active exploration strategy: slightly increase ,reduce This allows the robot to pass through obstacles more quickly and closely. The user intent weight γ remains at its default value, with the system prioritizing autonomous efficiency. Medium confidence mode (0.4 ≤ ≤0.7): A robust cruise strategy is adopted: using default parameters and increasing the safety margin during planning. The user intent weight γ is appropriately increased to enhance the role of manual guidance. Low confidence mode ( <0.4): Adopt a cautious preservation strategy: a. Significantly reduce ,expand b. Increase a. During route planning, try more alternative solutions. b. Increase the user intent weight γ to above 0.6, prioritizing explicit user instructions and reducing the weight of autonomous decision-making. c. Simultaneously, proactively initiate confirmation through the interactive interface.

[0073] S7.5 human-machine collaboration closed loop is formed. First, the decision-making, execution, and perception loop is implemented, specifically: the system, based on the current... Adjust parameters and generate decisions. Continuously collect new sensing data during execution. Update confidence levels. Then, the system proceeds to the next round of adjustments. Next, user intervention and learning are implemented, including allowing users to directly control or demonstrate operations when the system frequently requests confirmation due to low confidence. The system records user actions in these boundary scenarios for online fine-tuning of the reinforcement learning reward function, enabling the policy network to gradually learn to mimic user behavior in similar situations. Finally, a closed-loop output is generated, which includes not only the navigation action sequence but also system status reports (e.g., currently in cautious mode, executing user instructions to check pump room 3) and suggestion requests (e.g., suggesting cleaning the left-side camera).

[0074] User commands are not simply overridden by the system, but rather integrated as soft constraints into the optimization objectives, enabling flexible collaboration between human and host systems or between host and machine systems. The system dynamically adjusts the level of risk based on its perceived trustworthiness, autonomously balancing efficiency and safety. By recording user intervention behavior, the system can continuously optimize its strategies over long-term operation, gradually reducing its reliance on human intervention.

[0075] Example 2

[0076] like Figure 2 As shown, this application provides an architecture diagram of a station intelligent navigation and decision support system that integrates audio positioning and multimodal data. It is applied to the station intelligent navigation and decision support system that integrates audio positioning and multimodal data as described in Embodiment 1, including: a multimodal signal preprocessing module 210, a sound source spatial analysis network module 220, a structure perception network module 230, a cross-modal dynamic fusion module 240, an event-target decoupling module 250, and a navigation decision generation module 260.

[0077] The multimodal signal preprocessing module 210 is used to preprocess the acquired multimodal signals of the station environment to obtain the spectral characteristics of multi-channel audio and the three-dimensional visual characteristics of multi-view video.

[0078] The sound source spatial analysis network module 220 is used to process the spectral features, wherein the sound source spatial analysis network module includes a directional analysis unit based on multi-channel interferometry and direction of arrival estimation, and an anti-interference filtering unit that performs sound field feature purification and outputs spatially enhanced audio features.

[0079] The structure-aware network module 230 is used to process the three-dimensional visual features, extract the three-dimensional geometric structure features and semantic object features of the scene, and generate a structure-guided visual fusion representation using a hierarchical feature aggregation mechanism.

[0080] The cross-modal dynamic fusion module 240 is used to construct a cross-modal dynamic transmission path based on the reliability metric of the spatially enhanced audio features and the spatial consistency information of the structure-guided visual fusion representation. The multimodal features are selectively fused along the path by a parameterized selection unit to generate scene-adaptive multimodal decision features.

[0081] The event-target decoupling module 250 is used to input the multimodal decision features into the event-target decoupling module, extract the temporal distribution information of acoustic events and the spatial composition information of visual targets respectively, associate the sound source location hypothesis with the target instance through a graph attention network, output the sound source-target association matrix and the target segmentation mask, and apply correlation constraints to ensure the semantic association between the two outputs.

[0082] The navigation decision generation module 260 is used to generate a navigation decision sequence by using a hierarchical strategy to generate a model based on the sound source location information, target segmentation mask, station 3D semantic map and real-time dynamic information parsed from the sound source-target correlation matrix.

[0083] Figure 3 This is an electronic device provided in one embodiment of this application. For example... Figure 3 As shown, the electronic device includes at least the following components: processor 301 and memory 300, communication interface 303, and bus 302.

[0084] In this embodiment of the application, memory 300 is used to store executable instructions of processor 301, which is configured to implement the method as described in the first aspect when executing instructions.

[0085] In embodiments of this application, a computer-readable storage medium includes instructions that instruct a device to perform the method as described in the first aspect. For example, the instructions instruct the device to perform... Figure 1 The method is shown in the process steps.

[0086] In one embodiment of this application, the program operating in the electronic device may be a program that controls a central processing unit (CPU) or similar device to achieve the functions of the above-described embodiments of the present invention (a program that enables the computer to function). Information processed by these systems is then temporarily stored in random access memory (RAM) during processing, and subsequently stored in various ROMs such as read-only memory (FlashROM) and hard disk drives (HDDs), and read, corrected, and written by the CPU as needed.

[0087] It should be noted that a portion of the electronic device described in the above embodiments can also be implemented using a computer. In this case, the program for implementing the control function can be recorded on a computer-readable recording medium, and the program recorded on the recording medium can be read into the computer and executed.

[0088] It should be noted that the computer mentioned here refers to a computer built into an electronic device, employing hardware including an operating system and peripheral devices. Furthermore, computer-readable recording media refers to removable media such as floppy disks, magneto-optical disks, ROMs, and CD-ROMs, as well as storage systems such as hard drives built into the computer.

[0089] Furthermore, computer-readable recording media can include: media that dynamically stores programs for short periods of time, such as communication lines used when transmitting programs via networks like the Internet or communication lines like telephone lines; and media that store programs for fixed periods of time, such as volatile memory inside a computer that serves as a server or client in this case. In addition, the aforementioned program can be a program used to implement the above-mentioned functions, or it can be a program that can implement the above-mentioned functions by combining them with programs already recorded in the computer.

[0090] Furthermore, the electronic device in the above embodiments can also be implemented as an assembly (system group) composed of multiple systems. Each system constituting the system group can possess some or all of the functions or functional blocks of the electronic device in the above embodiments. As a system group, it is sufficient to have all the functions or functional blocks of the electronic device.

[0091] Those skilled in the art should recognize that the above embodiments are only used to illustrate this application and are not intended to limit this application. Any appropriate changes and variations made to the above embodiments within the essential spirit and scope of this application fall within the scope of protection claimed in this application.

Claims

1. A method for intelligent navigation and decision support at a facility that integrates audio positioning and multimodal data, characterized in that, Includes the following steps: The collected multimodal signals from the field environment are preprocessed to obtain the spectral characteristics of multichannel audio and the three-dimensional visual characteristics of multi-view video. The spectral features are processed by a sound source spatial analysis network, wherein the sound source spatial analysis network includes a azimuth analysis unit based on multi-channel interferometry and direction of arrival estimation, and an anti-interference filtering unit that performs sound field feature purification, and outputs spatially enhanced audio features. The three-dimensional visual features are processed by a structure-aware network to extract the three-dimensional geometric structure features and semantic object features of the scene, and a hierarchical feature aggregation mechanism is used to generate a structure-guided visual fusion representation. Based on the reliability metric of the spatially enhanced audio features and the spatial consistency information of the structure-guided visual fusion representation, a cross-modal dynamic transmission path is constructed. Through a parameterized selection unit, multimodal features are selectively fused along this path to generate scene-adaptive multimodal decision features. The multimodal decision features are input into the event-target decoupling module, which extracts the temporal distribution information of acoustic events and the spatial composition information of visual targets respectively. The graph attention network is used to associate the sound source location hypothesis with the target instance, and outputs the sound source-target association matrix and the target segmentation mask that represent the association strength between the sound source direction information and the visual target. Correlation constraints are applied to ensure the semantic association between the two outputs. Based on the sound source location information, target segmentation mask, 3D semantic map of the site and real-time dynamic information parsed from the sound source-target correlation matrix, a hierarchical strategy is used to generate a model output navigation decision sequence, and the movement of the mobile robot or unmanned equipment is controlled based on the sequence.

2. The method according to claim 1, characterized in that, The process of processing the spectral features through the sound source spatial analysis network includes the azimuth analysis unit processing procedure, specifically comprising: The spectral characteristics of the multi-channel audio signal are decomposed into space-time-frequency components, and the time delay difference between each channel is calculated using the generalized cross-correlation method to generate an initial time delay difference matrix. Based on the initial time delay difference matrix, the initial azimuth and elevation angles of the sound source are estimated using the direction of arrival estimation algorithm. A deep neural network is used to jointly model the spectral features and the initial time delay difference matrix after the spatiotemporal decomposition, and to extract deep acoustic features containing spatial information. The initial azimuth and elevation angle estimates are fused with the depth acoustic features to form azimuth coding features.

3. The method according to claim 2, characterized in that, The process of processing the spectral features through the sound source spatial analysis network also includes an anti-interference filtering unit process, specifically including: Based on the aforementioned azimuth coding features, reverberation patterns in the spectrum are identified through a time-frequency attention mechanism; Dynamically generate an adaptive filter coefficient matrix based on the identified reverberation pattern; The azimuth coding features are subjected to convolutional spatial filtering using the adaptive filtering coefficient matrix to suppress reverberation components; The spatially filtered features are input into the multipath interference suppression module. A method combining geometric acoustics-based multipath reflection modeling and Wiener filtering is used to predict and suppress multipath reflection interference components. The features after interference suppression are enhanced and reconstructed to output the final spatially enhanced audio features.

4. The method according to claim 3, characterized in that, It also includes a feature quality control mechanism, specifically including: A feature quality evaluation module is set at the output end of the anti-interference filtering unit to calculate the signal-to-noise ratio and azimuth consistency index of the feature. When the signal-to-noise ratio is lower than the preset threshold, the feature re-extraction process is initiated. When the azimuth consistency index is abnormal, the azimuth calibration process is initiated to correct the current azimuth estimate using historical azimuth information.

5. The method according to claim 1, characterized in that, The process of processing the 3D visual features through a structure-aware network to extract the 3D geometric structure features and semantic object features of the scene, and generating a structure-guided visual fusion representation using a hierarchical feature aggregation mechanism, specifically includes: The three-dimensional visual features are processed into point cloud voxels and converted into three-dimensional voxel meshes. Local geometric features are extracted using a three-dimensional convolutional neural network to obtain an initial geometric feature map. Then, a geometric attention mechanism is used to perform spatial weighting to obtain enhanced geometric features. Multi-level semantic features of the three-dimensional visual features are extracted by a multi-scale feature pyramid network. At the same time, the semantic objects and spatial locations in the scene are identified by the object detection module to generate object feature vectors. The object feature vectors are fused with the semantic features of the corresponding spatial locations to obtain semantically enhanced object features. The enhanced geometric features and semantically enhanced object features are input into the hierarchical aggregation module. A feature aggregation weight map is constructed based on the spatial distribution information of the enhanced geometric features. The semantically enhanced object features are spatially reweighted according to the weight map. The reweighted semantic features and the enhanced geometric features are concatenated by channels and aligned through a cross-modal attention mechanism. After multiple levels of iterative aggregation processing, the structure-guided visual fusion representation is generated.

6. The method according to claim 1, characterized in that, The multimodal decision features for scene adaptation specifically include: The reliability metric is calculated based on the spatially enhanced audio features, and the spatial consistency information is calculated based on the structure-guided visual fusion representation. The cross-validation module is used to compare the reliability metric with the spatial consistency information to generate a cross-modal confidence score. Based on the cross-modal confidence score, a main conduction path dominated by audio features and an auxiliary conduction path supplemented by visual features are constructed, and parameterized selection units are set at key nodes of the conduction path. At each critical node, the parameterized selection unit dynamically adjusts the activation threshold of the gating parameter matrix based on the cross-modal confidence score, and performs weighted fusion of audio features and visual features through a soft selection mechanism, wherein the fusion weight is dynamically determined based on the real-time evaluation results of the reliability metric and spatial consistency information. The features fused from each node are input into the scene adaptation module. The feature fusion strategy is adjusted according to the current scene type and adaptive normalization is performed. The module outputs multimodal decision features for scene adaptation with unified dimensions.

7. The method according to claim 1, characterized in that, The process involves inputting the multimodal decision features into the event-target decoupling module, extracting the temporal distribution information of the acoustic events and the spatial composition information of the visual targets, associating the sound source location hypothesis with the target instance through a graph attention network, outputting a sound source-target association matrix and a target segmentation mask, and applying correlation constraints to ensure the semantic relevance of the two outputs. Specifically, this includes: The temporal distribution features of acoustic events are extracted from the multimodal decision features through temporal convolutional networks and temporal attention mechanisms. At the same time, spatial composition features of visual targets are extracted through spatial convolutional networks and spatial attention mechanisms. A cross-branch information exchange channel is established between the temporal and spatial feature extraction processes to perform feature interaction alignment. The temporal distribution features and spatial composition features are input into a graph attention network, where the source location hypothesis decoded from the temporal distribution features is used as one set of nodes, and the target instance hypothesis decoded from the spatial composition features is used as another set of nodes. The attention weights between nodes are calculated to generate a source-target association matrix that characterizes the association strength between the source and the target. The spatial features are input into the target segmentation generator, and a target segmentation mask is generated through an encoder-decoder structure and pixel-level classification. Mutual information maximization constraint and geometric consistency constraint are established, and the correlation matrix generation process and the target segmentation mask generation process are jointly optimized through a multi-task loss function to ensure the semantic correlation between the correlation matrix and the target segmentation mask.

8. The method according to claim 1, characterized in that, The process of generating the navigation decision sequence from the model using a hierarchical strategy specifically includes: Based on the sound source direction information represented by the sound source-target association matrix, and combined with the target three-dimensional spatial information provided by the target segmentation mask, the possible regions of the sound source in the three-dimensional space are analyzed. The target segmentation mask and the station three-dimensional semantic map are processed respectively to generate instance information containing target spatial bounding boxes and semantic categories, as well as a navigation environment map with semantic annotations. A three-layer decision architecture is constructed, comprising a task planning layer, a path planning layer, and a motion control layer. The task planning layer determines the navigation task objective and priority based on the sound source location information and target instance information. The path planning layer generates a set of feasible paths under the constraints of the task objective by combining the navigation environment map and real-time dynamic information. The motion control layer generates a sequence of navigation actions based on the set of feasible paths and the current state. In the three-layer decision architecture, the task planning layer adopts a multi-objective optimization algorithm, the path planning layer uses an improved A* algorithm combined with dynamic obstacle prediction, and the motion control layer performs the transformation from path to action through a reinforcement learning policy network. A feedback mechanism is established between the levels of the three-layer decision-making architecture, so that the upper-level decisions provide constraints for the lower-level decisions, and the execution results of the lower-level decisions are fed back to the upper-level decisions, thereby dynamically optimizing and outputting the final navigation decision sequence.

9. The method according to claim 1, characterized in that, Also includes: The system receives user commands through a semantic interaction interface, transforms user intentions into guiding signals for system decision-making, and adjusts the matching degree between the system output and user expectations based on a multimodal collaborative optimization objective function. During navigation execution, navigation parameters are dynamically adjusted based on environmental perception confidence, using a progressive optimization algorithm to form a human-machine collaborative autonomous navigation closed loop.

10. A station intelligent navigation and decision support system integrating audio positioning and multimodal data, applied to the method described in any one of claims 1 to 9, characterized in that, The system includes: The multimodal signal preprocessing module is used to preprocess the acquired multimodal signals from the site environment to obtain the spectral characteristics of multi-channel audio and the three-dimensional visual characteristics of multi-view video. A sound source spatial analysis network module is used to process the spectral features, wherein the sound source spatial analysis network module includes a azimuth analysis unit based on multi-channel interferometry and direction of arrival estimation, and an anti-interference filtering unit that performs sound field feature purification and outputs spatially enhanced audio features. The structure-aware network module is used to process the three-dimensional visual features, extract the three-dimensional geometric structure features and semantic object features of the scene, and generate a structure-guided visual fusion representation using a hierarchical feature aggregation mechanism. The cross-modal dynamic fusion module is used to construct a cross-modal dynamic transmission path based on the reliability metric of the spatially enhanced audio features and the spatial consistency information of the structure-guided visual fusion representation. The parameterized selection unit selectively fuses multimodal features along the path to generate scene-adaptive multimodal decision features. The event-target decoupling module is used to input the multimodal decision features into the event-target decoupling module, extract the temporal distribution information of acoustic events and the spatial composition information of visual targets respectively, associate the sound source location hypothesis with the target instance through a graph attention network, output the sound source-target association matrix and the target segmentation mask that characterize the association strength between the sound source direction information and the visual target, and apply correlation constraints to ensure the semantic association between the two outputs; The navigation decision generation module is used to generate a navigation decision sequence based on the sound source location information, target segmentation mask, station 3D semantic map and real-time dynamic information parsed from the sound source-target correlation matrix, and through a hierarchical strategy to generate a model outputting a navigation decision sequence.

Citation Information

Patent Citations

  • Urban environment digital management system based on big data

    CN118379170A

  • Personnel positioning method and system based on 3D Gaussian splash model and video fusion

    CN120318327A

  • Robot sensing and decision-making method based on lightweight multi-modal large model

    CN120612683A