Visual language navigation method for large bridge unmanned aerial vehicle autonomous inspection

CN122813852APending Publication Date: 2026-09-25HARBIN INST OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611058680.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-16
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

一旦遇到复杂环境或障碍物,通常需要人工干预,无法自主做出飞行决策

Benefits of technology

1、支持将输入的中文巡检指令,输出严格的包含巡检任务类型、语言标识、目标列表、巡检顺序、安全约束、观测约束和飞行约束JSON对象。摆脱传统无人机巡检对人工底层摇杆遥控或预设绝对刚性坐标航迹的依赖。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122813852A_ABST
    Figure CN122813852A_ABST
Patent Text Reader

Abstract

The application provides a visual language navigation method for large bridge unmanned aerial vehicle (UAV) autonomous inspection. The method deeply integrates visual-language-action model and embodied intelligence, so that the UAV can understand the abstract task requirements in the form of natural language or high-level instructions in a complex bridge environment, and autonomously decompose the requirements into a series of specific navigation, positioning and observation actions. Through the construction of the bottom logic closed loop of instruction understanding-feature extraction-evidence alignment, and the hierarchical navigation and dynamic optimization framework of the flight path, the whole-chain UAV intelligent inspection solution from natural language instruction to autonomous perception, planning and execution is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of bridge structural health monitoring, autonomous UAV inspection, embodied intelligence, visual language multimodal large models, and UAV trajectory planning and dynamic optimization, particularly to a visual language navigation method for autonomous UAV inspection of large bridges. The method can be directly applied to fields including intelligent transportation infrastructure maintenance, UAV visual navigation in complex scenarios, intelligent detection of bridge surface defects, and high-fidelity digital twin simulation testing. Background Technology

[0002] With the continuous development of my country's infrastructure, a large number of bridge structures have undergone decades of service and face multiple impacts from environmental corrosion, material deterioration, natural disasters, and human factors, resulting in various types of structural damage, including concrete cracks, surface spalling, exposed rebar, anchor cable corrosion, and fatigue cracks in steel structures. Bridge structural health monitoring is crucial for traffic safety. Traditional bridge inspections mainly rely on manual visual inspection or contact sensor monitoring, which has inherent drawbacks such as extremely low efficiency, high risks associated with working at heights, numerous blind spots, and high costs. Unmanned aerial vehicle (UAV) inspections have been widely used in civil engineering due to their ability to significantly reduce personnel safety risks and improve inspection coverage. Using UAVs for inspections can significantly reduce the dangers to human operators. UAVs can carry sensors on or near the bridge contact surface to thoroughly inspect dangerous areas or hard-to-reach locations, and the inspection speed is generally faster. It provides a multifunctional and efficient way to collect and monitor data, enabling detailed inspections, detection of surface damage, and even assessment of dynamic parameters. The generated information can be integrated into a broader management system to improve maintenance and decision-making.

[0003] Traditional drones mostly rely on pre-set flight paths and lack intelligent decision-making capabilities when performing inspection tasks. When encountering complex environments or obstacles, they typically require human intervention and cannot make autonomous flight decisions. Traditional drone path planning is usually based on simple preset rules, such as straight lines or manual adjustments, and cannot make timely adjustments to the drone's attitude when encountering dynamic obstacles. Furthermore, during close-range bridge inspections, they are prone to positioning drift or even loss of control due to obstructed satellite signals, severely impacting the continuity and safety of inspections. They lack deep physical semantic understanding of complex bridge environments and cannot autonomously find targets and avoid obstacles according to the issued inspection task requirements. They also cannot directly translate high-level natural language inspection commands into low-level drone flight, obstacle avoidance, and observation actions to execute bridge inspection tasks.

[0004] To achieve truly autonomous and unmanned bridge inspection, it is essential to transform UAVs from simple automated surveying tools into embodied intelligent agents with a closed-loop perception-cognition-decision system. This requires a deep interdisciplinary integration of cutting-edge large-scale model high-order cognitive reasoning capabilities, dynamic trajectory planning technology coupled with semantic information, and the professional needs of bridge engineering inspection. To overcome these technical bottlenecks, this invention proposes a visual-language navigation method for autonomous bridge UAV inspection. By deeply integrating a visual-language-action model with embodied intelligence, the UAV can understand abstract task requirements issued in natural language or high-level commands in complex bridge environments and autonomously decompose them into a series of specific navigation, positioning, and observation actions. By constructing a low-level logical closed loop of command understanding, feature extraction, and evidence alignment, as well as a hierarchical navigation and dynamic trajectory optimization framework, a full-chain intelligent UAV inspection solution is achieved, from natural language commands to autonomous perception, planning, and execution. Summary of the Invention

[0005] The purpose of this invention is to solve the problems in the prior art and to propose a visual language navigation method for autonomous inspection of large bridges by unmanned aerial vehicles (UAVs).

[0006] This invention is achieved through the following technical solution: This invention proposes a visual language navigation method for autonomous inspection of large bridges using unmanned aerial vehicles (UAVs), the method comprising: Step 1: Construct a multimodal bridge inspection task representation and natural language command parsing framework; A large language model is introduced as the cognitive decision-making center of the UAV, which receives natural language input or high-level commands from the inspection personnel, and parses and decomposes the abstract inspection task into a set of objectives, sequence constraints, safety constraints and observation constraints that can be executed by the embodied intelligent system. Step 2: Construct a multi-source visual feature dual-stream extraction and segmentation network for bridge scenarios; By using UAV visual sensors to acquire bridge scene images, a dual-stream target detection architecture for bridge components and apparent damage is established. Combined with a lightweight semantic segmentation model, high-precision multi-source visual features of the bridge environment are extracted. Step 3: Establish a visual-linguistic feature alignment mechanism based on cross-modal image-text semantic matching; The abstract instruction elements output by the large language model are deeply integrated with the multi-source visual features extracted in step two to achieve a precise association between language targets and visual evidence, providing a semantically anchored environmental cognition basis for the autonomous decision-making of UAVs. Step 4: Construct a UAV spatial positioning and navigation map based on point cloud and multimodal semantics; By integrating spatial point cloud data acquired by UAV sensors with visual semantic information, high-precision spatial positioning of UAVs in complex bridge environments can be achieved, and a multimodal semantic navigation map containing structural semantics, damage location and obstacle information can be constructed in real time. Step 5: Perform UAV hierarchical navigation decision-making, candidate viewpoint generation, and dynamic trajectory optimization; Based on the multimodal semantic navigation map, combined with the constraints generated in step one, candidate viewpoints that meet the inspection and observation requirements are generated; by embedding track constraints, online track construction and dynamic optimization are realized, and high-level decisions are connected to the UAV's underlying flight control to complete the embodied closed-loop navigation. Step Six: Construct a high-fidelity virtual simulation system and closed-loop verification mechanism for bridge unmanned aerial vehicles (UAVs); A 3D bridge simulation base with scalable structural system and parameterized component damage was built in Unreal Engine, enabling the export of multi-channel heterogeneous data and full-link integrated testing of the embodied navigation system.

[0007] Furthermore, step one specifically includes: Step 11: Receive the Chinese natural language inspection command input string This is then fed into a pre-tuned large language model; Steps one and two: The large language model learns from context and uses prompts to constrain the input string. Map and refactor into a JSON structured task object; Step 13: Activate the offline rule parser as an exception fallback mechanism; Step 14: Deconstruct unstructured instructions into standardized multimodal constraints, and break down the decision content of the large model into interpretable navigation variables; Step 15: Summarize system coverage, collision count, battery level, and camera status into a dynamic feedback vector to achieve a dimensionality reduction mapping from abstract intent to computable navigation variables; Step 16: Process the JSON structured task object Perform a consistency check.

[0008] Furthermore, in steps one and two, the input string is... Map and refactor into a JSON structured task object:

[0009] In the formula, For the set of inspection targets, each inspection target Includes semantic categories Target Identifier ,action Priority and expected coverage ; The topological constraints on the sequential access time of each inspection target are defined; For safety constraints, define the minimum physical safety distance between drones and the virtual bridge structure, main tower, or cables, as well as the boundaries of no-fly zones; To constrain observation, the optimal alignment field of view and minimum hovering shooting time for the camera are defined; Constraints on the drone's own flight include maximum linear velocity, maximum angular velocity, and virtual battery threshold.

[0010] Furthermore, step two specifically includes: Step 21: Based on the YOLO V11 model, construct and initialize the parallel dual-stream target detection backbone network of the multi-source sensing system; Step 22: In the apparent damage detection flow, for the coexistence of large-scale components and minor defects and uneven ambient lighting, three residual photometric enhancement IEL-C3k2 modules are embedded at the junction of shallow and deep features of the backbone network, and a streaming spatial channel attention C2PSA module with transposed self-attention is introduced in the bottleneck layer. Steps 2 and 3: The IEL spatial residual photometric enhancement sub-network inside the IEL-C3k2 module performs photometric decoupling pixel by pixel; Step 24: The damage detection head and the component detection head perform multi-scale anchor-free decoupled prediction and output the set of two-dimensional bounding box parameters of the target of interest; Step 25: Establish the EdgeLightweight U-Net pixel-level segmentation flow for thin cracks; the EdgeLightweight U-Net introduces a boundary-assisted supervision mechanism during training; the EdgeLightweight U-Net uses a hybrid loss function that includes positive sample backward gradient flow to enhance bias and spatial geometric similarity for end-to-end optimization, and the training loss adopts a combination of positive sample weighted BCE, soft Dice and boundary BCE.

[0011] Furthermore, step three specifically includes: Step 31: Construct a high-dimensional embedding mapping for language text features; Step 32: Construct a high-dimensional spatial embedding mapping for multi-source visual perception features; Step 33: Perform cross-modal feature normalization and scaling projection within the high-dimensional latent space; Steps 3 and 4: Construct a cross-modal matching score calculation model that combines the prior weights of YOLO detection categories; Step 35: Perform abnormal visual false alarm suppression and online consolidation of visual evidence; Step 36: Perform online correction and confidence update of physical target evidence scores based on a soft fusion strategy; Step 37: Perform abnormal visual false alarm suppression, status labeling, and online solidification of visual evidence.

[0012] Furthermore, step four specifically includes: Step 41: Obtain the original 3D point cloud data of the bridge environment, and combine the spatial 3D coordinates of the points with color stitching features to form an input vector. The data is input into the PointNet point cloud semantic segmentation model. Step 42: Extract the local geometric features of each point cloud particle through a shared multilayer perceptron, capture the global topological structure features through a max pooling operator, concatenate the global features with the local features, and output point-by-point semantic labels through a fully connected classification layer. Step 43: Introduce the category-weighted cross-entropy loss function Perform network parameter optimization; Step 44: Construct a structured multimodal spatial semantic navigation map data structure; Steps four and five: Construct a dynamic alignment and filling network for multi-source heterogeneous data based on spatiotemporal perspective projection; Step 46: Construct a dynamic risk score calculation model that integrates multi-dimensional indicators; Step 47: Construct a task-layer inspection value quantification model that incorporates user command intent; Step 48: Perform dual-track parallel persistent storage for risk and value and cognitive routing to the backend planner.

[0013] Furthermore, step five specifically includes: Step 51: Three-dimensional sampling of navigation candidate viewpoints based on target position and surface normal; Step 52: Dynamic trajectory constraint mapping and embedding of multimodal perception priors; Step 53: Flexible Degradation Conflict Handling and Candidate Viewpoint Serialization Evaluation; Step 54: Temporal trajectory construction and hierarchical path planning based on a global objective function; Step 55: Dynamic environment perception-driven online trajectory rule optimization; Steps 5 and 6: Execute the closed-loop mapping of the underlying API control functions based on the command bridging module; the command bridging module is responsible for mapping the optimal 3D waypoint sequence planned by the higher layers. It is transformed into a specific machine control primitive command object that can be recognized by the underlying execution system.

[0014] Furthermore, step six specifically includes: Step 61: Import a high-precision 3D digital twin mesh model of a large bridge, constructed from point clouds and geometric entities, into Unreal Engine UE4.27. This model includes the main tower, main beam, suspension cables, tension cables, lower crossbeam, and high-fidelity micro-cracks, exposed reinforcement, and peeling surface defects reconstructed by texture mapping on the mesh surface. Complete the material rendering and add physical collision objects. Step 62: Configure and load the AirSim plugin in the Unreal Engine project, using it as a simulation peripheral interface for airborne sensor data feedback and low-level action execution; use its internal dynamic model to simulate the multi-rotor flight characteristics of the UAV, and set the flight constraint state without satellite signal as well as the physical noise of the inertial measurement unit and barometer. Step 63: Define and mount virtual sensor components in the AirSim configuration file, including a forward visible light RGB airborne camera with a specified resolution and a specified field of view and a point cloud LiDAR, and perform hardware-accelerated real-time rendering through the graphics card on a high-performance workstation to output airborne video stream images, forming a multi-source data source for the embodied intelligence system. Step 64: Call the Python API programming interface provided by AirSim to realize the full-process programmed control of the UAV in the simulation environment, covering a series of key operations from basic take-off and smooth landing to complex path planning flight and precise attitude adjustment, to verify the feasibility and stability of the navigation system.

[0015] The present invention also proposes an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of the visual language navigation method for autonomous inspection of large bridge UAVs.

[0016] The present invention also proposes a computer-readable storage medium for storing computer instructions, which, when executed by a processor, implement the steps of the visual language navigation method for autonomous inspection of large bridge UAVs.

[0017] The beneficial effects of this invention are: 1. Supports inputting Chinese inspection commands and outputting a strict JSON object containing the inspection task type, language identifier, target list, inspection order, safety constraints, observation constraints, and flight constraints. This eliminates the reliance on manual low-level joystick control or preset absolute rigid coordinate tracks for traditional UAV inspections.

[0018] 2. Supports the construction of a multi-source visual feature dual-stream extraction and segmentation network for bridge scenarios, enabling accurate identification of bridge components and bridge surface damage.

[0019] 3. Supports cross-modal feature depth alignment and evidence correction based on image-text semantic matching. Introduces a pre-trained high-order cross-modal alignment model to perform multimodal spatial mapping and cross-modal association between discrete high-level language text phrases generated by the task parsing layer and low-level visual bounding boxes and pixel-level mask feature maps extracted by the multi-source perception layer, thereby calibrating the physical target evidence score.

[0020] 4. Supports the construction of multimodal semantic navigation maps based on 3D point cloud semantic segmentation, projecting discrete 2D perception and alignment results into 3D geometric space to construct multimodal inspection semantic maps containing high-dimensional semantics.

[0021] 5. Integrating multi-dimensional constraint-based dynamic trajectory optimization and low-level action closed-loop execution. Based on the constructed multimodal 3D semantic navigation map, a multi-dimensional constraint trajectory optimization cost function is designed to control the UAV to perform online dynamic trajectory optimization in a simulation environment and output waypoint closed-loop execution.

[0022] 6. Construction of a bridge digital twin simulation inspection environment based on Unreal Engine and AirSim. The feasibility and stability of the embodied intelligent navigation system were verified on the virtual platform. Attached Figure Description

[0023] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.

[0024] Figure 1 This is a flowchart of the overall visual language navigation method for autonomous inspection of large bridges by UAVs.

[0025] Figure 2 This is a schematic diagram of the multimodal bridge inspection task representation and natural language command parsing process.

[0026] Figure 3 This is a diagram of the YOLO network architecture for bridge damage identification.

[0027] Figure 4 This is a diagram of the IEL spatial network structure used for residual photometric enhancement.

[0028] Figure 5 This is a schematic diagram of a dual-stream extraction and segmentation network for multi-source visual features in a bridge scene.

[0029] Figure 6 This is a schematic diagram of the visual-linguistic feature alignment mechanism based on cross-modal image-text semantic matching.

[0030] Figure 7This is a schematic diagram of the process of constructing a multimodal semantic navigation map based on 3D point clouds.

[0031] Figure 8 This is a schematic diagram illustrating the process of implementing hierarchical navigation decision-making, candidate viewpoint generation, and dynamic trajectory optimization for unmanned aerial vehicles (UAVs).

[0032] Figure 9 This is a schematic diagram of a high-fidelity virtual simulation system for bridge drones built using Unreal Engine and AirSim. Detailed Implementation

[0033] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0034] Specifically, in combination Figures 1-9 This invention proposes a visual language navigation method for autonomous inspection of large bridges by unmanned aerial vehicles (UAVs), the specific steps of which include: Step 1: Construct a multimodal bridge inspection task representation and natural language instruction parsing framework.

[0035] A large language model is introduced as the cognitive decision-making center of the UAV, which receives natural language input or high-level commands from the inspection personnel, and parses and decomposes the abstract inspection task into a set of executable targets, sequence constraints, safety constraints and observation constraints of the embodied intelligent system.

[0036] Step 1 involves constructing a multimodal bridge inspection task representation and natural language instruction parsing framework, specifically including the following steps: Step 11: Receive the Chinese natural language inspection command input string This is then fed into a pre-tuned large language model.

[0037] Steps one and two: The large language model learns from context and uses prompts to constrain the input string. Map and refactor into a JSON structured task object:

[0038] In the formula, For the set of inspection targets, each inspection target Includes semantic categories Target Identifier ,action Priority and expected coverage ; The topological constraints on the sequential access time of each inspection target are defined; For safety constraints, define the minimum physical safety distance between drones and the virtual bridge structure, main tower, or cables, as well as the boundaries of no-fly zones; To constrain observation, the optimal alignment field of view and minimum hovering shooting time for the camera are defined; Constraints on the drone's own flight include maximum linear velocity, maximum angular velocity, and virtual battery threshold.

[0039] Step 13: Activate the offline rule parser as an exception fallback mechanism. When the format of the large language model output does not meet the standard 5-tuple or times out, the offline rule parser extracts keywords using regular expressions and forcibly assembles them into the output of the basic inspection task object to ensure the stability of the main control system.

[0040] Step 14: Deconstruct unstructured instructions into standardized multimodal constraints; the decision-making content of the large model is broken down into interpretable navigation variables. Target List Save the type, number, action, priority, and coverage requirements for each inspection target; inspection sequence. The waypoint connection order is determined; safety constraints specify the minimum distance between bridge sections. Minimum distance of cable and return-to-home battery threshold ; Observation constraints provide the imaging distance range Viewing range, minimum visibility Whether to hover and the duration of hovering; flight constraints provide the maximum speed. Maximum acceleration and maximum yaw rate The task layer determines which targets need to be inspected and their priority. The constraint layer transforms the user's verbal expressions, such as "close range," "focus on shooting," "maintain distance from cables," and "cracks at the bottom of the beam," into numerical constraints. After task decomposition, a state vector is maintained.

[0041] In the formula, The current status of the drone. Location of the drone. Yaw angle This represents the percentage of battery capacity. For the visibility of the current target, For minimum safe distance, For cumulative collision or near miss events.

[0042] Step 15: Summarize system coverage, collision count, battery level, and camera status into a dynamic feedback vector, achieving a dimensionality-reduced mapping from abstract intent to computable navigation variables. The state feedback function maps the trajectory execution process into textual and numerical feedback:

[0043] In the formula, for The system provides real-time feedback results. For state feedback function, From the start of the mission to The trajectory that has already been executed. For semantic navigation maps.

[0044] Step 16: [Regarding...] Perform consistency checks. If the target set is empty, the parser will supplement the basic targets such as bridge towers, main beams, cables, and beam bottoms based on the default bridge inspection template; if the safety distance does not appear in the natural language, a conservative threshold is used and the feedback indicates that the constraint comes from the default safety policy; if the user simultaneously requests both "rapid inspection" and "close-range detailed inspection," the parser stores the velocity constraint and observation constraint separately, rather than forcibly compromising at the text level. The validation function can be written as:

[0045] In the formula, For task object The validity judgment results This is the minimum safe distance threshold for the bridge structure. The minimum safe distance threshold for the cable is when Instead of generating flight commands, it returns readable mission correction suggestions.

[0046] Step 2: Construct a multi-source visual feature dual-stream extraction and segmentation network for bridge scenarios.

[0047] By using UAV visual sensors to acquire bridge scene images, a dual-stream target detection architecture targeting bridge components and apparent damage is established. Combined with a lightweight semantic segmentation model, high-precision multi-source visual features of the bridge environment are extracted.

[0048] Step 2 involves constructing a multi-source visual feature dual-stream extraction and segmentation network for bridge scenarios, specifically including the following steps: Step 21: Based on the YOLO V11 model, construct and initialize the parallel dual-stream target detection backbone network of the multi-source perception system. During the UAV's flight, the system input is the two-dimensional RGB image acquired by the UAV in real time. Simultaneously, the data is input into both the bridge component detection stream and the bridge surface damage detection stream. In the backbone network, convolution, downsampling, and SPPF modules are used to extract multi-scale pyramid features. Specifically, the bridge component detection stream predicts macroscopic bounding boxes for the main tower, main girder, suspension cables, stay cables, and lower crossbeams; the bridge surface damage detection stream is specifically used to extract local bridge defects such as cracks, exposed reinforcement, and spalling.

[0049] Step 22: In the apparent damage detection pipeline, to address the challenges of coexistence of large-scale components and minor defects, and uneven ambient lighting, three IEL-C3k2 (Residual Photometric Enhancement) modules are embedded at the interface between shallow and deep features in the backbone network, and a C2PSA (Streaming Spatial Channel Attention with Transposed Self-Attention) module is introduced at the bottleneck layer. The IEL-C3k2 module is responsible for enhancing the response of low-salience defect areas during feature extraction, SPPF is used to expand the receptive field and integrate multi-scale contextual information, and C2PSA further strengthens the key channels and spatial locations in the backbone output features. The Neck part adopts a feature aggregation method combining top-down and bottom-up approaches to perform cross-scale fusion of features at different levels; the Head part outputs three detection branches, corresponding to small-scale, medium-scale, and large-scale targets, respectively, thereby adapting to the detection needs of defects of different sizes such as bridge cracks, spalling, and corrosion. In the Backbone network, the original C3k2 modules are replaced with IEL-C3k2 modules at layers 5, 7, and 9, giving the network stronger detail enhancement capabilities in both shallow and deep feature learning stages. The network architecture is as follows: Figure 3 As shown.

[0050] Steps two and three: The IEL spatial residual photometric enhancement sub-network within the IEL-C3k2 module performs photometric decoupling pixel by pixel. Let the input features be... IEL first utilizes Convolution performs channel mapping to obtain the expanded feature representation:

[0051] In the formula, This indicates a pointwise convolution operation. Subsequently, on... Applying depthwise separable convolutions yields bi-branch features:

[0052] In the formula, Represents depthwise convolution. This indicates a channel partitioning operation.

[0053] Next, the two branches undergo independent depthwise convolution and Tanh nonlinear transformation, which, together with the original branch, form residual enhancement:

[0054] Subsequently, the two enhancement branches are fused element-wise by multiplication:

[0055] In the formula, This represents the Hadamard product. This operation can model the coupling relationship between different enhanced responses, so that the effective disease area gets a higher response after fusion, while the background noise is weakened by mutual suppression.

[0056] Finally, through The convolution recovers the number of channels and forms a residual connection with the input to obtain the IEL output:

[0057] In the formula, This indicates a pointwise convolution operation. This is a feature map of two-stream nonlinear coupling fusion. This is the original two-dimensional image feature map input to the IEL spatial residual photometric enhancement subnetwork. The structure of the residual photometric enhancement IEL spatial network is as follows: Figure 4 As shown.

[0058] Step 24: The damage detection head and the component detection head perform multi-scale anchor-free decoupled prediction, and output the set of parameters for the two-dimensional bounding box of the target of interest. A single bounding box is represented as Based on the detection probability and scale, calculate the physical target evidence score of the disease corresponding to each detection frame. : (10) In the formula, To determine the classification prediction confidence of the current bounding box, the molecule Calculate the absolute pixel area of ​​the disease in the image. The total pixel resolution of the entire inspection image. The evidence score is calculated based on the prior baseline score of damage risk and inspection value for this type of apparent damage as preset in the bridge maintenance technical specifications. This serves as the underlying decision-making basis for subsequent localized encrypted photography of the flight path.

[0059] Step 25: Establish the EdgeLightweight U-Net pixel-level segmentation flow for elongated slits. The bounding box output in Step 24... semantic category tags Cracks were identified and the detection confidence level was [not specified]. When the threshold is exceeded, the system automatically captures the image of that region and sends it to the encoder of EdgeLightweight U-Net. Inside the network, an SE (Squeeze-and-Excitation) channel attention module is introduced at each skip connection. This module compresses the spatial dimension through global average pooling and automatically and explicitly models the interdependencies between channels to enhance crack feature channels and suppress mixed concrete rough background textures.

[0060] Step 26: To ensure no additional computational overhead during the inference phase and to address the bottleneck of breakage, step loss, and topological discontinuity that easily occur in pixel-level segmentation of thin, elongated cracks, the EdgeLightweight U-Net introduces a boundary-assisted supervision mechanism during training. An edge contour prediction head is constructed parallel to the main segmentation head. The Laplacian operator is used to extract the edge contours of the real mask. The boundary features output by the edge contour prediction head are supervised and forced to be fused back into the main network during inference. The model retains the backbone of the original U-Net, adding SE channel recalibration only at the bottleneck layer and a boundary-assisted head at the decoding end. Global average pooling is used to obtain the channel descriptions. The channel weight is

[0061] Boundary auxiliary heads are used to mitigate crack fracture problems. Decoding end features. Simultaneous entry into the main mask head and the boundary head:

[0062] In the formula, For the predicted master mask, The weights of the convolution kernels for the main mask head, For the predicted boundary map, The weights of the convolution kernel at the boundary header.

[0063] The true boundary is represented by a binary mask. Y The morphological gradient is obtained as follows:

[0064] In the formula, For true edge labels, To perform an inflation operation on the real labels. This is to perform an erosion operation on the real label.

[0065] Step 27: The EdgeLightweight U-Net uses a hybrid loss function that includes positive sample backward gradient flow to enhance bias and spatial geometric similarity for end-to-end optimization. The training loss uses a combination of positive sample weighted BCE, softDice and boundary BCE. The total segmentation loss is calculated as follows:

[0066] In the formula, These are the weighting coefficients for the BCE loss. For weighted binary cross-entropy loss, For positive sample weights, These are the weighting coefficients for the Dice loss. To prevent smoothing constants with denominators of zero, These are the weighting coefficients for the boundary loss.

[0067] Step 3: Establish a visual-linguistic feature alignment mechanism based on cross-modal image-text semantic matching.

[0068] By deeply fusing the abstract instruction elements output by the large language model with the multi-source visual features extracted in step two, the precise association between language targets and visual evidence is achieved, providing a semantically anchored environmental cognition foundation for the autonomous decision-making of UAVs.

[0069] A visual-linguistic feature alignment mechanism based on cross-modal image-text semantic matching is established. A pre-trained high-order cross-modal alignment model is introduced to perform multimodal spatial mapping and cross-modal association between discrete high-level linguistic text phrases generated by the task parsing layer and low-level visual bounding boxes and pixel-level mask feature maps extracted by the multi-source perception layer, thereby calibrating the physical target evidence score. The steps include: Step 31: Construct a high-dimensional embedding mapping for language text features. The set of inspection text target phrases extracted from the structured task object obtained in Step 1 is subjected to text symbolization processing, and then input into the text encoder of the pre-trained SigLIP cross-modal alignment model, outputting the normalized text feature vector of the m-th text target. :

[0070] In the formula, This is a text encoder function used to map input natural language text to a unified cross-modal semantic feature space. No. An input text description or language phrase, Perform text feature vector Normalize it so that its length is This eliminates vector scale differences and facilitates cosine similarity calculation with image features.

[0071] Step 32: Construct a high-dimensional spatial embedding map of multi-source visual perception features. This involves embedding the two-dimensional bounding box regions output from the apparent damage detection stream and the key component detection stream in Step 2. The corresponding local region of interest image and pixel-level mask feature maps of the cascaded crack segmentation stream output. The j-th candidate image region is synchronously input into the visual converter encoder of the cross-modal alignment model. Local feature extraction is performed on the image encoder of the cross-modal alignment model, and the normalized image feature vector of the j-th candidate image region is output. :

[0072] In the formula, This is an image encoder used to extract visual semantic features of candidate regions and map them to a cross-modal unified feature space. For image cropping operations, the target region is extracted from the original image based on the detection box and mask. No. A pixel-level mask corresponding to a crack or damaged area. This indicates that the image feature vector is being processed. Normalize it to a length of 1 to facilitate cosine similarity calculation with the normalized text features.

[0073] Step 33: Perform cross-modal feature normalization and scaling projection in the high-dimensional latent space. To eliminate the inconsistency in feature scale outputs from different modal encoders, a spatial scaling factor built into the cross-modal alignment model is introduced. With bias term For the standard text feature vector With image feature vectors Perform matrix dot product alignment. The cross-modal alignment model employs a pairwise logistic regression loss function for joint optimization during training to calculate the cross-modal image-text similarity score between the m-th text target and the j-th candidate visual evidence. :

[0074] In the formula, This represents the Sigmoid logistic regression activation function, which forces the similarity metric between the two values ​​to be linearly and smoothly mapped to... Within the range, the closer the value is to 1, the higher the semantic match between the currently visually detected component or defect area and the password entered by the operator.

[0075] Steps 3 and 4: Construct a cross-modal matching score calculation model that combines YOLO detection category prior weights. Establish a visual language localization evaluation system, introducing a balance parameter between image-text similarity and detection category prior, and controlling the relative weights of the two in the decision-making process by adjusting the balance parameter. The multimodal association matching score between the m-th text target and the j-th candidate visual evidence is calculated. for:

[0076] In the formula, Indicates the first m A set of category aliases for each text target; This represents the prior weight of the physical detection confidence score output by the YOLO target detection head in step three; and These represent the image-text similarity weight and the category prior weight, respectively. For an indicator function, if and only if the first... j Detection categories of visual bounding boxes Belongs to the m A set of preset aliases parsed from a text phrase The value is 1 if the condition is met, and 0 otherwise. For images with insufficient training samples for the bridge-specific category, zero-shot generalization across words is provided through SigLIP's shared latent space. For complex scenarios where image-text matching is unstable, YOLO's category prior is used as a constraint to limit the range of incorrect matches.

[0077] Step 35: Perform abnormal visual false alarm suppression and online consolidation of visual evidence. Set a global alignment confidence threshold. When the calculated corrected target evidence score satisfies Instead of performing a forced mismatch hard selection, the system outputs the state of the m-th text target in the current frame image and marks it as an unlocated state. This is to adapt to the situation where bridge towers, cables, and beam bottoms are partially visible in a single frame image due to blind spots and lack of conclusive evidence.

[0078] Step 36: Perform online correction and confidence update of physical target evidence scores based on a soft fusion strategy. The main control system establishes an evidence confidence cascading propagation mechanism, utilizing the aforementioned multimodal association matching scores. The physical target evidence score output in step three Perform online corrections, decouple and recalculate the corrected higher-order cross-modal target evidence scores. :

[0079] In the formula, This is the preset soft fusion penalty weight coefficient.

[0080] Step 37: Perform abnormal visual false alarm suppression, state labeling, and online consolidation of visual evidence. Based on the corrected higher-order cross-modal target evidence score... State-based routing is deployed when the following conditions are met: If the current detection area is determined to be a false alarm instance caused by complex virtual meshes on the bridge surface, lighting clutter, or rough concrete texture, the area is automatically removed from the current perception candidate queue, or its safety confirmation status is forcibly marked as pending confirmation. Targets in the pending confirmation state are forcibly retained in the task graph to avoid losing the user's query intent, but in the subsequent step six, trajectory planning, the system assigns them a safe and conservative viewpoint to prevent them from excessively dominating the UAV's trajectory generation; when the conditions are met... At that time, the system marks the security confirmation status as confirmed, officially solidifying the area as strong visual evidence supporting the high-level inspection intent. This mechanism ensures a deep coupling between understanding instructions and understanding the environment, and the calibrated high-score target points and visual evidence vectors... It will be directly passed through to the subsequent 3D point cloud semantic map projection module.

[0081] Step 4: Construct a UAV spatial positioning and navigation map based on point cloud and multimodal semantics.

[0082] By integrating spatial point cloud data acquired by UAV sensors with visual semantic information, high-precision spatial positioning of UAVs in complex bridge environments can be achieved, and a multimodal semantic navigation map containing structural semantics, damage location, and obstacle information can be constructed in real time.

[0083] The construction of a multimodal semantic navigation map based on 3D point clouds involves projecting discrete 2D perception and alignment results into a 3D geometric space to construct a multimodal inspection semantic map containing high-dimensional semantics. Specific steps include: Step 41: Obtain the original 3D point cloud data of the bridge environment, and combine the spatial 3D coordinates of the points with color stitching features to form an input vector. The input is fed into the PointNet point cloud semantic segmentation model.

[0084] Step 42: Extract the local geometric features of each point cloud particle through a shared multilayer perceptron, capture the global topological structure features through a max pooling operator, and then concatenate the global features with the local features to output point-by-point semantic labels through a fully connected classification layer.

[0085] Step 43: To overcome the problem of extremely unbalanced data volume between main components and minor defects in the point cloud, a category-weighted cross-entropy loss function is introduced. Optimize network parameters:

[0086] In the formula, Total points The bias coefficient is the inverse of the weight of the current true class. To predict probabilities for the model, This is the smoothing constant.

[0087] Step 44: Construct a structured multimodal spatial semantic navigation map data structure. The semantic navigation map uniformly encodes language targets, detection boxes, crack masks, and point cloud semantics into a standardized set of planarable map units.

[0088] In the formula, each map unit Persistently stored and characterized as the following standard multidimensional properties, where For location, For semantic categories, Risk score, For inspection value, Mark the obstacle.

[0089] Steps 4 and 5: Construct a dynamic alignment and filling network for multi-source heterogeneous data based on spatiotemporal perspective projection. The system synchronously reads the current 3D attitude matrix of the UAV and camera extrinsic parameters transmitted from the airborne sensors via remote procedure calls. Using a full perspective projection matrix, it aligns the 2D defect topology boundary, local damage area, and high-order cross-modal target evidence score after cross-modal alignment correction in Step 4. Dynamic inverse projection and anchoring to the 3D point cloud geometric space enables index fusion and filling of 2D image targets, 3D point cloud labels, and sensor results. For a specific map unit containing a target, its mathematical mapping mechanism strictly follows the following expression:

[0090] In the formula, This represents a multidimensional risk calculation function; This represents the function for calculating the value of inspections; This represents the safety shell offset penalty term that dynamically and adaptively shrinks or expands according to the target category; the last binary variable. This indicates whether the unit is considered an absolute obstacle.

[0091] Step 46: Construct a dynamic risk score calculation model that integrates multi-dimensional indicators. The risk score within the map unit... The physical priors derived from bridge component and damage categories are jointly determined by semantic category, damage confidence, and geometric proximity. For any given spatial point cloud cell, its risk score... The following segmented nested formula is used for calculation:

[0092] In the formula, For a given category, a prior risk score is assigned. Cables, anchorage areas, and bridge towers are extremely sensitive to flight safety distances and camera angles, and are assigned high prior risk values. Supports and beam bottoms, due to their confined spaces and susceptibility to obstruction and close-range flight risks, have medium prior risk values. Ordinary bridge structures and damaged areas receive a basic safety risk score. The second item is the damaged area. The spatial Gaussian attenuation risk radiation term of surrounding voxels after inverse projection. This is a Hausdorff spatial distance metric. The third term represents the current voxel distance to the absolute set of obstacles. The reciprocal penalty of the Euclidean distance. This is achieved through the safety shell offset. Constraints are applied to the cable target to produce the maximum safe expansion offset, the bridge tower and anchorage zone to produce the medium safe offset, and the ordinary bridge body and damaged areas to produce the basic safe offset.

[0093] Step 47: Construct a task-layer inspection value quantification model that incorporates user command intent. The inspection value weight score within each map unit. The task priority dispatched by the task parsing layer, the multimodal awareness confidence, and the alignment of user text commands are all jointly determined, and its calculation formula is strictly expressed as:

[0094] In the formula, The score represents the task priority score obtained from step two; the score within square brackets represents the multitrack fusion confidence score derived from visual language matching or detection confidence. This is a Boolean indicator function used to determine whether the current map cell directly belongs to the target explicitly specified in the user-input Chinese natural language command.

[0095] Step 48: Perform dual-track parallel persistent storage for risk and value, and cognitive routing to the backend planner. The system integrates conflicting dynamic risk scores in the semantic map. Inspection value score Simultaneously, persistent storage is performed. High-risk areas are not subject to hard avoidance filtering; instead, the presence of cracks makes these units also highly valuable for inspection. The semantic map transmits this decoupled conflict information, which is both a key inspection point and a collision risk, as a structured planning unit to the backend trajectory constraint cost function. Within the same computational framework, the path planner comprehensively considers task priority, environmental obstacle avoidance risk, and trajectory observation distance. Under the premise of ensuring the safety shell boundary, it actively approaches and dynamically selects safe but still high-fidelity multi-angle observation viewpoints, endowing the UAV with spatial autonomous closed-loop update and generalized cognitive capabilities based on engineering common sense.

[0096] Step 5: Perform UAV hierarchical navigation decision-making, candidate viewpoint generation, and dynamic trajectory optimization.

[0097] Based on the multimodal semantic navigation map and combined with the constraints generated in step one, candidate viewpoints that meet the inspection and observation requirements are generated. By embedding trajectory constraints, online trajectory construction and dynamic optimization are achieved, and high-level decisions are connected to the UAV's low-level flight control to complete embodied closed-loop navigation.

[0098] Perform hierarchical navigation decision-making, candidate viewpoint generation, and dynamic trajectory optimization for the UAV. Based on the constructed multimodal 3D semantic navigation map, design a multidimensional constrained trajectory optimization cost function, control the UAV to perform online dynamic trajectory tuning in a simulation environment, and output waypoint closed-loop execution. Step five includes: Step 51: Spatial 3D Sampling of Navigation Candidate Viewpoints Based on Target Position and Surface Normal. To address the multi-angle observation requirements of bridge components, adherence to safety distance constraints, and crack visibility requirements, the system generates distance samples along the opposite direction of the normal within a preset observation distance range, based on the target's geometric center position and surface normal. Small-range offsets are added in the tangential and height directions to obtain a set of candidate viewpoints. The specific calculation steps are as follows: Step 511: Given the location of the inspection target, the system generates a set of candidate viewpoints around its three-dimensional spatial coordinates and surface normal vector. The three-dimensional position coordinates of any single candidate viewpoint are... The forward mathematical expression formula is:

[0099] In the formula, This indicates the three-dimensional spatial coordinates of the center of the inspection target; This represents the sampled value of the observation distance from the airborne camera to the inspected target; Represents the unit normal vector of the inspected target surface; Represents the tangential unit direction vector that is approximately perpendicular to the target unit normal; The unit direction vector representing the vertical direction; This represents the lateral offset scale in the tangential direction; It represents the vertical offset dimension in the vertical direction.

[0100] Step 512: Determine the desired 3D attitude angles of the candidate viewpoints. Yaw angle of the candidate viewpoints. The pitch angle of the candidate viewpoint is uniquely determined by the vector direction pointing from the current viewpoint's three-dimensional coordinates to the center point of the inspected target; The network dynamically and adaptively adjusts its angle based on the semantic category of the inspection target. When the semantic category of the inspection target belongs to the bottom of the beam or the crack target, the network automatically adopts a more obvious downward observation pitch angle.

[0101] Step 52: Dynamic Track Constraint Mapping and Embedding of Multimodal Perception Priors. The track constraint embedding module transforms the component 2D localization boxes output by YOLO in Step 2, the crack pixel-level segmentation mask generated by EdgeLightweight U-Net, and the multimodal high-dimensional visual perception priors after SigLIP cross-modal semantic alignment in Step 4 into quantifiable hard and soft constraints for UAV navigation decisions, and writes them into the cost function. The specific embedding steps are as follows: Step 521: Construct a viewpoint cost function that integrates multiple requirements of safety, visibility, damage, and pose. Its specific mathematical formula is defined as:

[0102] In the formula, Yaw angle The pitch angle, For observation distance, For visibility, For risk, For a safe distance, To ensure a safe distance from the bridge structure, For the safety distance of the cable, Penalty for exceeding the observation distance limit, The damage enhancement score.

[0103] Step 522: Align and correct the target evidence score from Step 4. Transformed into nonlinear enhancement term injection. For damage areas identified as cracks, spalling, or corrosion with high evidence scores ( The system assigns a negative cost, injecting it into the cost function as a reward, thereby actively guiding the path to shift towards a closer observation viewpoint to improve detail resolution. Conversely, for structurally sensitive areas such as cables and supports, the system imposes hard constraints by increasing the dynamic safety distance weighting coefficient. The optimal target viewpoint selection is solved for global extremum using the following formula:

[0104] In the formula, For the first The optimal observation point for each inspection target As candidate observation viewpoints, To determine a feasible set of viewpoints that meet the requirements of safe distance, UAV kinematic constraints, and field of view coverage. Let the candidate viewpoint synthesis cost function be... This indicates finding the variable that minimizes the objective function. This is the inspection target number.

[0105] Step 53: Flexible Degradation Conflict Handling and Candidate Viewpoint Serialization Evaluation. In the candidate viewpoint serialization evaluation stage, the system uses the inspection task topology and timing constraints defined by the large model decision center and rule parser. The candidate viewpoint space for each inspection target is scored and ranked in multiple dimensions, and the optimal solution of the cost function is selected. A flexible degradation mechanism is designed to address constraint conflicts caused by narrow space or dense obstacles.

[0106] Step 531: Traverse all candidate viewpoint sets corresponding to the current inspection target. If the inspection finds that there is no solution set that can fully satisfy all hard safety distances and observation constraints, the system activates the abnormal state recovery strategy and refuses to truncate the inspection task.

[0107] Step 532: Abnormal State Recovery Strategy - When the hard constraint solution set is empty, retain the current cost function. The local optimal viewpoint is forced as the recoverable solution and the specific reasons for infeasibility are encoded as status words and sent back to the constraint scoring module, and reported to the online optimization or human expert decision-making layer, so as to ensure the continuity of the inspection task under the condition of limited local perception.

[0108] Step 54: Temporal Track Construction and Hierarchical Path Planning Based on Global Objective Function. After combining viewpoints according to the analytically obtained objective order, the system connects the selected macroscopic waypoints into a continuous trajectory and adopts a deterministic hierarchical planning strategy. The specific path construction steps are as follows: Step 541: Construct the global optimization total cost function for the entire time-series trajectory Its forward calculation formula is expressed as:

[0109] In the formula, the last term constrains the flight path to be smooth, preventing the drone from making excessively sharp turns in the narrow space of the bridge.

[0110] Step 542: Let the initial position of the UAV in the simulation environment be... The selected optimal waypoint sequence is The maximum flight speed for drones is The track length and estimated time are defined as follows:

[0111] In the formula, The length of the flight path. 0.1 represents the estimated time, and 0.1 represents the minimum speed set by the system.

[0112] The trajectory objective function, which integrates path length, risk penalty, and visibility reward, can be written as follows:

[0113] In the formula, Waypoint risk score, The visibility score is the waypoint score.

[0114] Step 55: Dynamic Environment Perception-Driven Online Track Rule Optimization. The online optimization module receives the initially generated trajectory stream and multi-source sensor feedback thresholds in real time, and performs online corrections based on real-time physical safety shell relationships during the UAV's close-range inspection process. The specific optimization steps are as follows: Set waypoint visibility threshold If waypoint visibility is lower than The system will shift waypoints along a safe direction and slightly increase their altitude, generating new viewpoint numbers; if the waypoint safe distance is less than the minimum distance to the bridge structure... The system will continue to push waypoints out of the bridge's safety shell. Online optimization events will record the event type, target number, correction reason, and severity. This module corresponds to the embodied navigation integrated framework for online learning trajectory optimization. The current version is a working rule-based optimization hook, and image quality scoring, depth maps, collision sensors, wind field disturbances, and flight control status can be integrated in the future. Let the initial waypoint be... The correction amount is The optimized waypoints are Online optimization is then represented as

[0115] In the formula, To improve the direction of the view, In the direction away from the barrier shell, and To correct the scale.

[0116] Steps five and six: Closed-loop mapping and execution of the underlying API control functions based on the command bridging module. The command bridging module is responsible for mapping the optimal 3D waypoint sequence planned by the higher layers. This is transformed into a specific machine control primitive command object that the underlying execution system can recognize. The specific mapping steps are as follows: Step 561: In the final refined track, each discrete waypoint is precisely encapsulated as a location containing a unique command number, specific command type, and the target's three-dimensional spatial position. X,Y,Z ), set yaw angle Set flight speed u The underlying command object, including hover time and associated target number.

[0117] Step 562: Based on the task object from Step 1 Extracted observation constraints Automatic execution of command type branch mapping: If the constraint requires the UAV to hover and take pictures at the location of a specific component, then the command type for that waypoint is set to goto_hover_capture, which controls the UAV to remain stationary and hover for a specified duration after arriving at the waypoint in order to complete the high-fidelity capture of the hardware-accelerated rendering map; if the observation constraint does not require hovering, then the command type is set to goto, and the UAV flies directly over the waypoint.

[0118] Step 563: After completing the underlying command encapsulation of the entire flight path, the system synchronously exports the command set into a structured CSV format file that is easy for subsequent ROS nodes to read, as well as a QGCWPL-style task configuration file that retains the local coordinate system information of waypoints and the complete time sequence.

[0119] Step 564: The AirSim flight control bridge node reads the exported configuration file and directly maps and transforms the corresponding control primitives into AirSim's natively supported low-level multirotor control interface functions via the RPC protocol. For the velocity control link, the `moveByVelocityAsync()` function is called to inject the 3D velocity vector and time step; for the absolute position control link, the `moveToPositionAsync()` function is called to inject the target point's local coordinates and velocity parameters. The low-level control commands drive the core dynamics components of the multirotor virtual drone within the Unreal Engine to make attitude responses, realizing a complete semantic embodied intelligent inspection control system that integrates high-level text commands, UE4.27 environmental multi-source perception, cross-modal semantic alignment, hierarchical online planning, and AirSim low-level closed-loop action execution.

[0120] Step Six: Construct a high-fidelity virtual simulation system and closed-loop verification mechanism for bridge unmanned aerial vehicles (UAVs). A 3D bridge simulation base with scalable structural system and parameterized component damage was built in Unreal Engine, enabling the export of multi-channel heterogeneous data and full-link integrated testing of the embodied navigation system.

[0121] A high-fidelity virtual simulation system and closed-loop verification mechanism for bridge unmanned aerial vehicles (UAVs) were constructed based on Unreal Engine and AirSim. Utilizing a high-precision 3D rendering engine and a UAV physics simulation platform, the bridge inspection environment and UAV dynamic response, closely resembling the real physical world, were reproduced within the computer. Step six includes the following steps: Step 61: Import a high-precision 3D digital twin mesh model of a large bridge, constructed from point clouds and geometric entities, into Unreal Engine UE4.27. This model includes the main tower, main beam, suspension cables, tension cables, lower crossbeam, and high-fidelity micro-cracks, exposed reinforcement, and peeling surface defects reconstructed by texture mapping on the mesh surface. Complete the material rendering and add physical collision objects.

[0122] Step 62: Configure and load the AirSim plugin in the Unreal Engine project, using it as the simulation peripheral interface for airborne sensor data feedback and low-level action execution. Utilize its internal dynamics model to simulate the multi-rotor flight characteristics of the UAV, and set flight constraints without satellite signals, as well as physical noise from the inertial measurement unit and barometer.

[0123] Step 63: Define and mount virtual sensor components in the AirSim configuration file, including a forward-facing visible light RGB airborne camera with a specified resolution and field of view, and a point cloud LiDAR. Perform hardware-accelerated real-time rendering through the graphics card on a high-performance workstation to output airborne video stream images, forming a multi-source data source for the embodied intelligence system.

[0124] Step 64: Call the fully functional Python API programming interface provided by AirSim to realize the full-process programmed control of the UAV in the simulation environment, covering a series of key operations from basic take-off and smooth landing to complex path planning flight and precise attitude adjustment, to verify the feasibility and stability of the navigation system.

[0125] Example This implementation method focuses on the visual language navigation method and virtual simulation system for autonomous inspection of large bridge drones. It is illustrated using a bridge drone inspection embodied intelligent navigation scenario based on a large visual language model. The overall process includes: constructing a multimodal bridge inspection task representation and natural language command parsing framework; constructing a multi-source visual feature dual-stream extraction and segmentation network for bridge scenarios; establishing a visual-language feature alignment mechanism based on cross-modal image-text semantic matching; constructing a drone spatial positioning and navigation map based on point cloud and multimodal semantics; performing hierarchical navigation decision-making for the drone; generating candidate viewpoints and dynamically optimizing the trajectory; and constructing a high-fidelity virtual simulation system and closed-loop verification mechanism for bridge drones.

[0126] 1. Constructing a multimodal bridge inspection task representation and natural language instruction parsing framework: DeepSeek is used as a large language model for high-level task location decision-making. Its input is the user-inputted Chinese inspection instructions, and its output is a strict JSON object. This JSON object contains the inspection task type, language identifier, target list, inspection order, safety constraints, observation constraints, and flight constraints. DeepSeek's decision content is broken down into interpretable navigation variables.

[0127] 2. Constructing a multi-source visual feature dual-stream extraction and segmentation network for bridge scenarios: During UAV flight, the network efficiently and in real-time extracts the locations of key components from continuously acquired 2D video streams and simultaneously performs precise classification. When the UAV is performing a patrol mission at a distance from the bridge, the onboard YOLO model immediately generates a detection result containing location information once it detects a blurred suspected damage area in the real-time captured footage. This crucial information is uploaded to a large backend model for processing. Upon receiving the detection signal from the front end, the large model intelligently analyzes and makes decisions based on its understanding of the current task and flight path planning, automatically and dynamically adding one or more new refined reconnaissance waypoints to the current flight path, thereby guiding the UAV to conduct a close-up detailed inspection of suspected targets. EdgeLightweight U-Net is used for crack semantic segmentation for crack damage.

[0128] 3. Establish a visual-linguistic feature alignment mechanism based on cross-modal image-text semantic matching: Use the SigLIP model to achieve deep alignment of visual-linguistic features between the natural language results processed by the large model and the YOLO results for bridge component and damage identification and crack segmentation. During the training period, the cross-modal alignment model uses a pairwise logistic regression loss function for joint optimization to calculate the cross-modal image-text similarity score between the text target and candidate visual evidence.

[0129] 4. Constructing UAV spatial localization and navigation maps based on point clouds and multimodal semantics: Acquire raw 3D point cloud data of the bridge environment, and combine the spatial 3D coordinates and color features of the points to form an input vector. The data is input into the PointNet point cloud semantic segmentation model. Point-by-point semantic labels are obtained through the point cloud network. The semantic navigation map encodes language targets, detection boxes, crack masks, and point cloud semantics into a standardized set of planarable map units. Using a full perspective projection matrix, the two-dimensional lesion topological boundary, local damage region, and higher-order cross-modal target evidence scores, corrected through cross-modal alignment in step four, are integrated. Dynamic inverse projection and anchoring to the 3D point cloud geometry enables index fusion and filling of 2D image targets, 3D point cloud labels, and sensor results. For any given spatial point cloud unit, its risk score is calculated. Inspection value weight score within map units The task priority dispatched by the task parsing layer, the multimodal awareness confidence, and the alignment of user text commands are all determined together.

[0130] 5. Perform hierarchical navigation decision-making, candidate viewpoint generation, and dynamic trajectory optimization for the UAV: ​​3D sampling of navigation candidate viewpoint space based on target position and surface normal. The trajectory constraint embedding module transforms the component 2D localization boxes output by YOLO in step two, the crack pixel-level segmentation mask generated by EdgeLightweight U-Net, and the multimodal high-dimensional visual perception priors after SigLIP cross-modal semantic alignment in step four into quantifiable hard and soft constraints for UAV navigation decision-making, and writes them into the cost function. In the candidate viewpoint serialization evaluation phase, the system uses the inspection task topology and timing constraints defined by the large model decision center and rule parser. The system performs multi-dimensional scoring and ranking of candidate viewpoints for each inspection target, and selects the optimal solution for the cost function. After combining viewpoints according to the analytically obtained target order, the system connects the selected macroscopic waypoints into a continuous trajectory. The online optimization module receives the initially generated trajectory stream and multi-source sensor feedback thresholds in real time, and performs online corrections based on real-time physical safety shell relationships during the UAV's close-range inspection. The command bridging module is responsible for converting the optimal 3D waypoint sequence planned by the upper layer into specific machine control primitive command objects that can be recognized by the lower-level execution system.

[0131] 6. Construct a high-fidelity virtual simulation system for bridge drones: Based on Unreal Engine and the AirSim physics engine, a 3D bridge simulation platform with rigorous physical constraints, realistic lighting rendering, and high-precision collision detection is built. The system is then connected to drones within Unreal Engine via APIs to perform system closed-loop testing and performance verification in a simulated bridge environment.

[0132] The present invention also proposes an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of the visual language navigation method for autonomous inspection of large bridge UAVs.

[0133] The present invention also proposes a computer-readable storage medium for storing computer instructions, which, when executed by a processor, implement the steps of the visual language navigation method for autonomous inspection of large bridge UAVs.

[0134] The memory in this application embodiment can be volatile memory or non-volatile memory, or it can include both volatile and non-volatile memory. The non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. The volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous linked dynamic random access memory (SLDRAM), and direct rambus RAM (DR RAM). It should be noted that the memory used in the methods described in this invention is intended to include, but is not limited to, these and any other suitable types of memory.

[0135] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer instructions. When the computer instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium accessible to a computer or a data storage device such as a server or data center that integrates one or more available media. The available media may be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., high-density digital video discs (DVDs)), or semiconductor media (e.g., solid-state disks (SSDs)).

[0136] In implementation, each step of the above method can be completed by integrated logic circuits in the processor's hardware or by instructions in software. The steps of the method disclosed in the embodiments of this application can be directly implemented by a hardware processor, or by a combination of hardware and software modules in the processor. The software modules can reside in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. This storage medium is located in memory, and the processor reads information from the memory and, in conjunction with its hardware, completes the steps of the above method. To avoid repetition, detailed descriptions are omitted here.

[0137] It should be noted that the processor in the embodiments of this application can be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method embodiments can be completed by the integrated logic circuitry in the processor's hardware or by instructions in software form. The processor can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this application can be directly embodied as execution by a hardware decoding processor, or as a combination of hardware and software modules in the decoding processor. The software modules can be located in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. This storage medium is located in memory, and the processor reads the information in the memory and, in conjunction with its hardware, completes the steps of the above methods.

[0138] The above provides a detailed description of the visual language navigation method for autonomous inspection of large bridges by unmanned aerial vehicles (UAVs) proposed in this invention. Specific examples have been used to illustrate the principles and implementation methods of this invention. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of this invention. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this invention. Therefore, the content of this specification should not be construed as a limitation of this invention.

Claims

1. A visual language navigation method for autonomous inspection of large bridges by unmanned aerial vehicles (UAVs), characterized in that, The method includes: Step 1: Construct a multimodal bridge inspection task representation and natural language command parsing framework; A large language model is introduced as the cognitive decision-making center of the UAV, which receives natural language input or high-level commands from the inspection personnel, and parses and decomposes the abstract inspection task into a set of objectives, sequence constraints, safety constraints and observation constraints that can be executed by the embodied intelligent system. Step 2: Construct a multi-source visual feature dual-stream extraction and segmentation network for bridge scenarios; By using UAV visual sensors to acquire bridge scene images, a dual-stream target detection architecture for bridge components and apparent damage is established. Combined with a lightweight semantic segmentation model, high-precision multi-source visual features of the bridge environment are extracted. Step 3: Establish a visual-linguistic feature alignment mechanism based on cross-modal image-text semantic matching; The abstract instruction elements output by the large language model are deeply integrated with the multi-source visual features extracted in step two to achieve a precise association between language targets and visual evidence, providing a semantically anchored environmental cognition basis for the autonomous decision-making of UAVs. Step 4: Construct a UAV spatial positioning and navigation map based on point cloud and multimodal semantics; By integrating spatial point cloud data acquired by UAV sensors with visual semantic information, high-precision spatial positioning of UAVs in complex bridge environments can be achieved, and a multimodal semantic navigation map containing structural semantics, damage location and obstacle information can be constructed in real time. Step 5: Perform UAV hierarchical navigation decision-making, candidate viewpoint generation, and dynamic trajectory optimization; Based on the multimodal semantic navigation map, combined with the constraints generated in step one, candidate viewpoints that meet the inspection and observation requirements are generated; by embedding track constraints, online track construction and dynamic optimization are realized, and high-level decisions are connected to the UAV's underlying flight control to complete the embodied closed-loop navigation. Step Six: Construct a high-fidelity virtual simulation system and closed-loop verification mechanism for bridge unmanned aerial vehicles (UAVs); A 3D bridge simulation base with scalable structural system and parameterized component damage was built in Unreal Engine, enabling the export of multi-channel heterogeneous data and full-link integrated testing of the embodied navigation system.

2. The method according to claim 1, characterized in that, Step one specifically includes: Step 11: Receive Chinese natural language inspection command input string This is then fed into a pre-tuned large language model; Steps one and two: The large language model learns from context and uses prompts to constrain the input string. Map and refactor into a JSON structured task object; Step 13: Activate the offline rule parser as an exception fallback mechanism; Step 14: Deconstruct unstructured instructions into standardized multimodal constraints, and break down the decision content of the large model into interpretable navigation variables; Step 15: Summarize system coverage, collision count, battery level, and camera status into a dynamic feedback vector to achieve a dimensionality reduction mapping from abstract intent to computable navigation variables; Step 16: Process the JSON structured task object Perform a consistency check.

3. The method according to claim 2, characterized in that, In steps one and two, the input string is... Map and refactor into a JSON structured task object: In the formula, For the set of inspection targets, each inspection target Includes semantic categories Target Identifier ,action Priority and expected coverage ; The topological constraints on the sequential access time of each inspection target are defined; For safety constraints, define the minimum physical safety distance between drones and the virtual bridge structure, main tower, or cables, as well as the boundaries of no-fly zones; To constrain observation, the optimal alignment field of view and minimum hovering shooting time for the camera are defined; Constraints on the drone's own flight include maximum linear velocity, maximum angular velocity, and virtual battery threshold.

4. The method according to claim 1, characterized in that, Step two specifically includes: Step 21: Based on the YOLO V11 model, construct and initialize the parallel dual-stream target detection backbone network of the multi-source sensing system; Step 22: In the apparent damage detection flow, for the coexistence of large-scale components and minor defects and uneven ambient lighting, three residual photometric enhancement IEL-C3k2 modules are embedded at the junction of shallow and deep features of the backbone network, and a streaming spatial channel attention C2PSA module with transposed self-attention is introduced in the bottleneck layer. Steps 2 and 3: The IEL spatial residual photometric enhancement sub-network inside the IEL-C3k2 module performs photometric decoupling pixel by pixel; Step 24: The damage detection head and the component detection head perform multi-scale anchor-free decoupled prediction and output the set of two-dimensional bounding box parameters of the target of interest; Step 25: Establish the EdgeLightweight U-Net pixel-level segmentation flow for thin cracks; the EdgeLightweight U-Net introduces a boundary-assisted supervision mechanism during training; the EdgeLightweight U-Net uses a hybrid loss function that includes positive sample backward gradient flow to enhance bias and spatial geometric similarity for end-to-end optimization, and the training loss adopts a combination of positive sample weighted BCE, soft Dice and boundary BCE.

5. The method according to claim 1, characterized in that, Step three specifically includes: Step 31: Construct a high-dimensional embedding mapping for language text features; Step 32: Construct a high-dimensional spatial embedding mapping for multi-source visual perception features; Step 33: Perform cross-modal feature normalization and scaling projection within the high-dimensional latent space; Steps 3 and 4: Construct a cross-modal matching score calculation model that combines the prior weights of YOLO detection categories; Step 35: Perform abnormal visual false alarm suppression and online consolidation of visual evidence; Step 36: Perform online correction and confidence update of physical target evidence scores based on a soft fusion strategy; Step 37: Perform abnormal visual false alarm suppression, status labeling, and online solidification of visual evidence.

6. The method according to claim 1, characterized in that, Step four specifically includes: Step 41: Obtain the original 3D point cloud data of the bridge environment, and combine the spatial 3D coordinates of the points with color stitching features to form an input vector. The data is input into the PointNet point cloud semantic segmentation model. Step 42: Extract the local geometric features of each point cloud particle through a shared multilayer perceptron, capture the global topological structure features through a max pooling operator, concatenate the global features with the local features, and output point-by-point semantic labels through a fully connected classification layer. Step 43: Introduce the category-weighted cross-entropy loss function Perform network parameter optimization; Step 44: Construct a structured multimodal spatial semantic navigation map data structure; Steps four and five: Construct a dynamic alignment and filling network for multi-source heterogeneous data based on spatiotemporal perspective projection; Step 46: Construct a dynamic risk score calculation model that integrates multi-dimensional indicators; Step 47: Construct a task-layer inspection value quantification model that incorporates user command intent; Step 48: Perform dual-track parallel persistent storage for risk and value and cognitive routing to the backend planner.

7. The method according to claim 1, characterized in that, Step five specifically includes: Step 51: Three-dimensional sampling of navigation candidate viewpoints based on target position and surface normal; Step 52: Dynamic trajectory constraint mapping and embedding of multimodal perception priors; Step 53: Flexible Degradation Conflict Handling and Candidate Viewpoint Serialization Evaluation; Step 54: Temporal trajectory construction and hierarchical path planning based on a global objective function; Step 55: Dynamic environment perception-driven online trajectory rule optimization; Steps 5 and 6: Execute the closed-loop mapping of the underlying API control functions based on the command bridging module; the command bridging module is responsible for mapping the optimal 3D waypoint sequence planned by the higher layers. It is transformed into a specific machine control primitive command object that can be recognized by the underlying execution system.

8. The method according to claim 1, characterized in that, Step six specifically includes: Step 61: Import a high-precision 3D digital twin mesh model of a large bridge, constructed from point clouds and geometric entities, into Unreal Engine UE4.

27. This model includes the main tower, main beam, suspension cables, tension cables, lower crossbeam, and high-fidelity micro-cracks, exposed reinforcement, and peeling surface defects reconstructed by texture mapping on the mesh surface. Complete the material rendering and add physical collision objects. Step 62: Configure and load the AirSim plugin in the Unreal Engine project, using it as a simulation peripheral interface for airborne sensor data feedback and low-level action execution; use its internal dynamic model to simulate the multi-rotor flight characteristics of the UAV, and set the flight constraint state without satellite signal as well as the physical noise of the inertial measurement unit and barometer. Step 63: Define and mount virtual sensor components in the AirSim configuration file, including a forward visible light RGB airborne camera with a specified resolution and a specified field of view and a point cloud LiDAR, and perform hardware-accelerated real-time rendering through the graphics card on a high-performance workstation to output airborne video stream images, forming a multi-source data source for the embodied intelligence system. Step 64: Call the Python API programming interface provided by AirSim to realize the full-process programmed control of the UAV in the simulation environment, covering a series of key operations from basic take-off and smooth landing to complex path planning flight and precise attitude adjustment, to verify the feasibility and stability of the navigation system.

9. An electronic device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1-8.

10. A computer-readable storage medium for storing computer instructions, characterized in that, When the computer instructions are executed by the processor, they implement the steps of the method according to any one of claims 1-8.