A Voice-Assisted Fine-Tuning System and Method for Weld Inspection UAV in Nuclear Power Plant Construction Based on a Large Language Model
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-07
- Publication Date
- 2026-08-14
AI Technical Summary
[0004]然而,这种高度自动化的模式在实际应用中暴露出显著的刚性缺陷:
本发明针对现有技术的不足,提供一种基于大语言模型的焊缝检测无人机语音辅助位姿微调系统及方法,该系统能够通过自然语言交互,理解操作员的复杂调整意图,并结合实时环境感知,实现对无人机检测位姿的智能、精准、高效微调;与现有技术相比,本申请的技术方案具有以下有益技术效果:
Smart Images

Figure CN122575351A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of industrial non-destructive testing and intelligent robot control technology, and more specifically, to a voice-assisted fine-tuning system and method for inspecting weld seams in nuclear power plant construction using a drone. Background Technology
[0002] Nuclear power engineering is a typical example of high-safety, large-scale reinforced concrete structure engineering, and the appearance of welds is a crucial indicator of weld quality. In welding production, inspecting the appearance quality of welds is essential. Before undergoing non-destructive testing (NDT) such as ultrasonic or radiographic testing, the weld surface and the surrounding base material surface must pass an appearance quality inspection. Failure to do so will affect the accuracy and completeness of the NDT results, leading to missed detections or difficulties in assessing the internal quality of the weld. More importantly, weld appearance defects can cause stress concentration, reduce load-bearing sections, and weaken the fatigue strength of certain structures subjected to dynamic loads, directly impacting the safety of the welded product. By inspecting the appearance quality of welds and promptly identifying, eliminating, or repairing defects, we can not only reduce the impact of appearance defects on NDT but also improve the success rate of strength and tightness tests, saving construction costs and time. Traditional weld inspection mainly relies on experienced inspectors using specialized equipment (such as ultrasonic detectors) to perform close-range work on the weld. This method is not only inefficient and labor-intensive, but also poses significant personal safety risks due to harsh working environments such as high altitudes, toxic environments, enclosed spaces, and radiation. Furthermore, the test results are easily affected by the technical level and subjective state of the personnel.
[0003] To overcome the limitations of manual inspection, unmanned aerial vehicle (UAV) technology has been widely used in the inspection field in recent years, especially demonstrating significant advantages in large, high-altitude, or inaccessible steel structure scenarios. Equipped with high-definition cameras, thermal imagers, or specific non-destructive testing sensors, UAVs can replace humans in reaching hazardous areas, initially realizing the "unmanned" and "aerial" nature of inspection tasks. Existing UAV inspection solutions mainly follow two technical routes: 1. Preset fully automatic track detection mode This is currently the most mainstream technical solution. Its working principle is as follows: Before the mission begins, the operator pre-plans a complete flight path in the ground station software, based on the CAD model of the structure under test or a 3D point cloud model generated through prior manual flight scanning. This path precisely defines the UAV's spatial coordinates, flight speed, and camera angles and positions during the inspection process. During mission execution, the UAV, relying on technologies such as the Global Positioning System (GPS), Inertial Measurement Unit (IMU), and visual odometry, attempts to strictly follow this predetermined flight path for fully autonomous flight and data acquisition.
[0004] However, this highly automated model reveals significant rigidity flaws in practical applications: Poor environmental adaptability: The actual nuclear construction site environment is complex and variable. The pre-set model may deviate from the on-site steel structure due to installation errors, deformation, or temporary scaffolding. On-site lighting conditions and wind interference can also affect the accuracy of drone hovering, causing the camera's perspective to deviate from the ideal detection position.
[0005] Weak ability to respond to emergencies: If unexpected obstacles (such as temporary pipelines or scaffolding) appear on the flight path, or if the welds themselves have severe corrosion or paint peeling that alters their visual characteristics, the drone lacks intelligent response strategies and often has no choice but to interrupt the mission or cause a collision.
[0006] The adjustment process is cumbersome and inefficient: When operators discover, through real-time drone image transmission, that the current viewpoint does not meet inspection requirements (e.g., the weld seam is not centered in the frame, the image is blurry, or there are reflective obstructions), the existing adjustment methods are extremely cumbersome. Operators must first switch the drone control mode from "automatic" to "manual" remote control. Then, they need to simultaneously and precisely manipulate multiple joysticks on the remote controller to control the drone's up / down, left / right, forward / backward movement, as well as its nose direction and gimbal angle. This process not only requires operators to possess superb flying skills but also demands frequent switching of attention between the remote controller, the screen image transmission, and the on-site environment, resulting in high levels of mental stress. The adjustment process is time-consuming and severely disrupts the continuity of the inspection operation. This is essentially an "open-loop" control, lacking intelligent evaluation and feedback of the adjustment results.
[0007] 2. Purely manual remote control detection mode In some more complex scenarios, operators rely entirely on manually controlling drones for inspection. While this method offers maximum flexibility, its drawbacks are more pronounced: it demands extremely high flying skills from the operator; it is difficult to maintain a stable flight attitude to obtain clear images; inspection coverage is difficult to guarantee, leading to missed detections; and the operator's workload is extremely heavy.
[0008] In summary, the root cause of the current technological predicament lies in the significant "semantic gap" between the high-level task objective (obtaining the optimal inspection perspective) and the low-level, multi-dimensional execution actions (UAV pose control). The operator's mind is focused on questions like "Can I see this weld more clearly?" or "Avoid that shadow," but translating these into specific control commands requires complex intermediate operations. This human-machine interaction method is neither intuitive nor efficient, failing to meet the urgent demands of modern industrial inspection for intelligence, adaptability, and efficient human-machine collaboration.
[0009] Therefore, the industry urgently needs a new type of intelligent human-machine interaction system and method that can directly and seamlessly translate the high-level intentions of operators into precise actions of drones, making drones an intelligent collaborative entity that can "understand" instructions, possess "common sense," and "adjust" according to environmental feedback, thereby truly unleashing the application potential of drones in complex industrial inspections. This is precisely the technical problem that this invention aims to solve. Summary of the Invention
[0010] The purpose of this invention is to provide a voice-assisted fine-tuning system and method for inspecting weld seams in nuclear power plant construction using a drone. This system can understand the operator's complex adjustment intentions through natural language interaction and, combined with real-time environmental perception, achieve intelligent, precise, and efficient fine-tuning of the drone's inspection posture.
[0011] On the one hand, this invention provides a voice-assisted fine-tuning system for nuclear power plant construction weld inspection drones based on a large language model, comprising the following modules: The voice interaction module is configured to: acquire voice commands, preprocess the voice commands and perform voice recognition and command structuring to obtain text commands; The large language model intelligent parsing and planning module is configured to: generate a control strategy for fine-tuning the UAV pose based on the text instructions and environmental context information using a domain expert large language model; The visual perception and weld defect recognition module is configured to: acquire environmental image information in real time, use a small model to recognize welds, and perform image quality assessment to obtain weld defect recognition information including weld defect type and weld image quality. The UAV flight control and attitude execution module is configured to execute the control strategy and drive the UAV to complete attitude adjustment. The status feedback and confirmation module is configured to receive image information, weld defect identification information, and UAV attitude fine-tuning control information, provide voice feedback on the system status to the operator, and assist the operator in generating the next motion control command based on the current UAV status.
[0012] Furthermore, the preprocessing described above includes adaptive noise suppression, acoustic echo cancellation, automatic gain control, and speech activity detection.
[0013] Furthermore, the aforementioned speech recognition and instruction structuring includes informal corpus regularization, instruction boundary detection and segmentation, and preliminary keyword extraction.
[0014] Furthermore, the aforementioned domain expert language model is pre-trained using a base model fine-tuned based on corpus data from the weld inspection domain, and is used to understand technical terms and instructions related to nondestructive testing, UAV flight, and pose control.
[0015] Furthermore, the aforementioned corpus in the field of weld inspection includes a welding and inspection professional knowledge base, a UAV flight and photography terminology database, large-scale synthetic dialogue data, and real human-computer interaction logs.
[0016] Furthermore, the aforementioned small model is a lightweight convolutional neural network.
[0017] Furthermore, the aforementioned weld seam recognition includes: acquiring environmental images using multiple visual sensors; acquiring dense or semi-dense point cloud data of the scene using parallax calculation or direct measurement; detecting weld seam regions in the image in real time using a lightweight convolutional neural network; outputting its precise bounding box; performing pixel-level segmentation on the detected weld seam regions; extracting its centerline; fitting its geometric orientation in the image coordinate system; and combining the 2D position of the weld seam in the image with the 3D information provided by the depth camera, estimating the three-dimensional coordinates and orientation of the weld seam key points relative to the UAV camera coordinate system using the PnP algorithm.
[0018] Furthermore, the aforementioned image quality assessment is used to score each pose adjustment according to an assessment system that includes multiple indicators, including composition score, sharpness score, illumination quality score, and defect visibility.
[0019] Furthermore, the above composition score is calculated as follows: the pixel distance between the center of the weld area and the center of the image, and the area occupied by the weld in the image; the sharpness score is calculated as follows: the edge intensity of the image is calculated using the Tenengrad gradient function or the Brenner gradient function; the illumination quality score is calculated as follows: the histogram of the image is analyzed to assess whether the overall brightness is appropriate and the contrast is sufficient, and to detect whether there are local highlights or shadows; the defect visibility is calculated as follows: a classifier is trained to infer the detectability probability of common welding defects at the current viewing angle and resolution based on the current image features.
[0020] Secondly, the present invention also provides a method for applying the above-mentioned voice-assisted fine-tuning system for nuclear power plant construction weld inspection UAV based on a large language model, comprising the following steps: S1: System initialization and task start-up: The UAV takes off and automatically flies to the approximate starting area of weld inspection. The vision module starts working, initially identifying and tracking the weld. S2: Voice command reception and recognition: Receiving and recognizing natural language commands issued by the operator through the microphone when the operator finds that the posture needs to be adjusted based on the real-time video transmission screen; S3: Intelligent command parsing and action planning: After the voice command is recognized as text, it is sent to the large language model module for comprehensive understanding in combination with the following context, including the text semantics of the current command, real-time weld position, image quality information, and the current pose of the drone, and outputs precise control actions. S4: Safety Verification and Command Execution: Perform safety verification on the planned actions, including collision risk and flight boundary. After passing the verification, send the command to the flight control module for execution. S5: Closed-loop feedback and iterative optimization: The image quality is evaluated using a voice-assisted fine-tuning system for nuclear power plant construction weld inspection drones based on a large language model. If the image quality is still not optimal, fine compensation adjustments are made, or the operator is prompted by voice to issue further instructions until a satisfactory inspection view is obtained. S6: Task Continue or End: After the pose adjustment is satisfactory, the drone continues to perform the automatic detection task and waits for new voice commands.
[0021] The voice-assisted fine-tuning system and method for inspecting weld seams in nuclear power plant construction provided by this invention have the following beneficial effects: This invention addresses the shortcomings of existing technologies by providing a voice-assisted pose fine-tuning system and method for weld seam inspection UAVs based on a large language model. This system can understand the operator's complex adjustment intentions through natural language interaction and, combined with real-time environmental perception, achieve intelligent, precise, and efficient fine-tuning of the UAV's inspection pose. Compared with existing technologies, the technical solution of this application has the following beneficial technical effects: (1) Simplified and intuitive operation: The complex remote control operation is transformed into a natural dialogue interaction, which greatly reduces the professional skill requirements of the operator and improves the efficiency of human-computer interaction.
[0022] (2) Intelligent and precise pose adjustment: By utilizing the deep semantic understanding and reasoning capabilities of the large language model, the vague operation intention can be transformed into precise control instructions, avoiding errors and repeated trial and error in manual operation.
[0023] (3) The system has strong adaptability: combined with real-time visual feedback, the system can understand the execution effect of the instructions and has a certain degree of autonomous optimization capability to adapt to complex field environments.
[0024] (4) Improve detection efficiency and quality: Fast and accurate pose fine-tuning helps to always obtain the best detection images, thereby improving the identification rate of weld defects and the efficiency of the entire detection task. Attached Figure Description
[0025] The present invention will be further described below with reference to the accompanying drawings and embodiments. In the accompanying drawings: Figure 1 This is a block diagram of the overall architecture of the voice-assisted fine-tuning system for nuclear power plant construction weld inspection drone based on a large language model provided by the present invention; Figure 2 This is an overall flowchart of the method provided by the present invention for a voice-assisted fine-tuning system for a nuclear power plant construction weld inspection UAV based on a large language model; Figure 3 This is a logical diagram illustrating the voice command parsing and pose control generation provided by the present invention; Figure 4 This is a schematic diagram of the training curve of the semantic-visual collaborative fine-tuning network provided by the present invention. Detailed Implementation
[0026] To provide a clearer understanding of the technical features, objectives, and effects of the present invention, specific embodiments of the present invention will now be described in detail with reference to the accompanying drawings.
[0027] Figure 1 This diagram illustrates the overall architecture of the voice-assisted fine-tuning system for a nuclear power plant construction weld inspection UAV based on a large language model, as shown in this embodiment. In this embodiment, the voice-assisted fine-tuning system for a nuclear power plant construction weld inspection UAV based on a large language model includes a voice interaction module, a large language model intelligent parsing and planning module, a visual perception and weld defect recognition module, a UAV flight control and attitude execution module, and a status feedback and confirmation module.
[0028] Specifically, the large language model intelligent parsing and planning module is connected to the voice interaction module, the UAV flight control and posture execution module is connected to the large language model intelligent parsing and planning module and the visual perception and weld defect recognition module, and the status feedback and confirmation module is connected to the visual perception and weld defect recognition module and the UAV flight control and posture execution module.
[0029] Specifically, The voice interaction module is configured to: acquire voice commands, preprocess the voice commands and perform voice recognition and command structuring to obtain text commands; The large language model intelligent parsing and planning module is configured to: generate a control strategy for fine-tuning the UAV pose based on the text instructions and environmental context information using a domain expert large language model; The visual perception and weld defect recognition module is configured to: acquire environmental image information in real time, use a small model to recognize welds, and perform image quality assessment to obtain weld defect recognition information including weld defect type and weld image quality. The UAV flight control and attitude execution module is configured to execute the control strategy and drive the UAV to complete attitude adjustment. The status feedback and confirmation module is configured to receive image information, weld defect identification information, and UAV attitude fine-tuning control information, provide voice feedback on the system status to the operator, and assist the operator in generating the next motion control command based on the current UAV status.
[0030] In one exemplary embodiment, the preprocessing described above includes adaptive noise suppression, acoustic echo cancellation, automatic gain control, and speech activity detection.
[0031] In one exemplary embodiment, the above-mentioned speech recognition and instruction structuring includes informal corpus regularization, instruction boundary detection and segmentation, and preliminary keyword extraction.
[0032] In one exemplary embodiment, the aforementioned domain expert large language model is pre-trained using a base model fine-tuned based on a corpus of weld inspection, and is used to understand technical terms and instructions related to nondestructive testing, UAV flight, and pose control.
[0033] In one exemplary embodiment, the aforementioned corpus on weld inspection includes a welding and inspection expertise base, a UAV flight and photography terminology database, large-scale synthetic dialogue data, and real human-computer interaction logs.
[0034] In one exemplary embodiment, the aforementioned small model is a lightweight convolutional neural network.
[0035] In one exemplary embodiment, the lightweight convolutional neural network described above is a YOLO or SSD model.
[0036] In one exemplary embodiment, the weld seam recognition includes: acquiring environmental images using multiple visual sensors; acquiring dense or semi-dense point cloud data of the scene using parallax calculation or direct measurement; detecting weld seam regions in the image in real time using a lightweight convolutional neural network; outputting their precise bounding boxes; performing pixel-level segmentation on the detected weld seam regions; extracting their centerlines; fitting their geometric orientation in the image coordinate system; and combining the 2D position of the weld seam in the image with the 3D information provided by the depth camera, estimating the three-dimensional coordinates and orientation of the weld seam key points relative to the UAV camera coordinate system using the PnP algorithm.
[0037] In one exemplary embodiment, the image quality assessment described above is used to score each pose adjustment according to an assessment system that includes multiple indicators, including composition score, sharpness score, illumination quality score, and defect visibility.
[0038] In an exemplary embodiment, the above composition score is calculated as follows: the offset pixel distance between the center of the weld area and the center of the image, and the area of the weld in the image are calculated; the sharpness score is calculated as follows: the edge intensity of the image is calculated using the Tenengrad gradient function or the Brenner gradient function; the illumination quality score is calculated as follows: the histogram of the image is analyzed to assess whether the overall brightness is appropriate and the contrast is sufficient, and to detect whether there are local highlights or shadows; the defect visibility is calculated as follows: a classifier is trained to infer the detectability probability of common welding defects at the current viewing angle and resolution based on the current image features.
[0039] This embodiment provides a method for applying the above-mentioned voice-assisted fine-tuning system for nuclear power plant construction weld inspection drones based on a large language model, including the following steps: S1: System initialization and task start-up: The UAV takes off and automatically flies to the approximate starting area of weld inspection. The vision module starts working, initially identifying and tracking the weld. S2: Voice command reception and recognition: Receiving and recognizing natural language commands issued by the operator through the microphone when the operator finds that the posture needs to be adjusted based on the real-time video transmission screen; S3: Intelligent command parsing and action planning: After the voice command is recognized as text, it is sent to the large language model module for comprehensive understanding in combination with the following context, including the text semantics of the current command, real-time weld position, image quality information, and the current pose of the drone, and outputs precise control actions. S4: Safety Verification and Command Execution: Perform safety verification on the planned actions, including collision risk and flight boundary. After passing the verification, send the command to the flight control module for execution. S5: Closed-loop feedback and iterative optimization: The image quality is evaluated using a voice-assisted fine-tuning system for nuclear power plant construction weld inspection drones based on a large language model. If the image quality is still not optimal, fine compensation adjustments are made, or the operator is prompted by voice to issue further instructions until a satisfactory inspection view is obtained. S6: Task Continue or End: After the pose adjustment is satisfactory, the drone continues to perform the automatic detection task and waits for new voice commands.
[0040] In some embodiments, the above-described voice-assisted fine-tuning system for nuclear power plant construction weld inspection UAV based on a large language model can also be implemented in the following ways.
[0041] In this embodiment, the voice-assisted fine-tuning system for nuclear power plant construction weld inspection UAV based on a large language model includes the following modules: Voice interaction module: During the weld inspection task performed by the UAV, it synchronously receives and recognizes the operator's voice commands, forming a voice database labeled with task commands and UAV action commands; Large Language Model Intelligent Parsing and Planning Module: Connected to the voice interaction module, it is used to understand the semantics of voice commands and, in combination with environmental context information, generate control strategies for fine-tuning the UAV's pose. Visual perception and weld defect recognition module: used to acquire environmental images in real time, form a weld image database, use a small model to identify weld defect types and analyze weld quality in the images; UAV flight control and attitude execution module: Connects the large language model intelligent parsing and planning module and the visual perception and weld defect recognition module, and is used to execute the UAV attitude fine-tuning control strategy to drive the UAV to complete attitude adjustment; Status feedback and confirmation module: Connects the visual perception and weld defect recognition module with the UAV flight control and attitude execution module. It receives image information, weld defect recognition information and UAV attitude fine-tuning control information, and provides voice feedback on the system status to the operator to assist the operator in generating the next motion control command based on the current UAV status.
[0042] The aforementioned technical solution aims to provide a revolutionary UAV pose fine-tuning system. Its core mission is to overcome the human-machine interaction barriers faced by existing technologies in industrial inspection scenarios, particularly in the field of non-destructive testing (NDT) of steel structure welds. This system constructs a highly integrated and intelligent technical framework that deeply integrates multiple advanced technology fields, including natural language human-machine interaction (HCI), semantic understanding and task planning driven by large language models (LLM), real-time environmental perception based on computer vision, and precise UAV flight control.
[0043] Based on the above technical solution, the present invention can be further improved as follows.
[0044] Furthermore, the voice interaction module is configured with: hardware architecture and environmentally adaptable front-end processing. The system uses a highly directional microphone array as the core sound pickup device. This array typically consists of 4 to 8 miniature microphones arranged in a specific geometry and integrated into the operator's smart safety helmet or handheld terminal. Its core capability is adaptive beamforming. In typical industrial environments filled with fan noise, arc noise, and metallic knocking sounds, this technology can dynamically calculate the direction of the sound source (i.e., the operator's mouth) to form an electronically controllable "sound pickup beam," like an invisible "auditory spotlight," always focused on the operator, thereby greatly suppressing interference noise and reverberation effects from other directions. In addition, the system can optionally be equipped with a bone conduction sensor as an auxiliary sound pickup solution. By collecting the vibration signal of the operator's cheekbone when speaking, it is fused with the microphone signal in extremely noisy environments to further improve the reliability of voice activation and endpoint detection.
[0045] Before being fed into the recognition engine, the raw audio stream must undergo a series of rigorous preprocessing steps: Adaptive noise suppression: Using spectral subtraction or deep learning-based methods, the spectral components of background noise are estimated and filtered out.
[0046] Acoustic echo cancellation: Prevents system feedback voice played from the drone's speakers or operator terminal from being picked up by the microphone again, causing recognition confusion.
[0047] Automatic gain control: Dynamically adjusts the input volume to ensure a stable signal level regardless of whether the operator gives a low command at close range or shouts from a distance, avoiding clipping distortion due to excessive volume or insufficient signal-to-noise ratio due to insufficient volume.
[0048] Voice activity detection: Accurately distinguishes between speech segments and non-speech segments (such as breathing sounds and silence periods), and only activates recognition when there is speech, saving computing resources and reducing false triggers.
[0049] Furthermore, the voice interaction module is configured for speech recognition and command structuring. The system employs a "cloud-based collaborative" recognition strategy. Under good network conditions, it prioritizes using a large-scale cloud-based speech recognition model trained on massive amounts of data to achieve the highest accuracy and vocabulary coverage. In situations with limited network access or requiring ultra-low latency, it switches to a locally deployed lightweight edge-side ASR model. Although smaller in size, this model has been specifically optimized using a large amount of industrial speech data, achieving extremely high accuracy in recognizing specialized terms (such as "edge," "unfused," and "heat-affected zone"). The text recognized by the system through command cleaning and intent segmentation techniques is often a continuous stream full of colloquial features. This module includes a lightweight natural language understanding submodule responsible for: Informal corpus standardization: filter out filler words such as "um", "ah", "this" and repetitive words.
[0050] Instruction boundary detection and segmentation: Long sentences are broken down into atomic operation units. For example, the sentence "First rise to see the whole picture, then fly closer to check the crack" is intelligently segmented into two ordered tasks: [Instruction Unit 1: Rise to observe the whole picture] and [Instruction Unit 2: Fly closer to check the specific crack]. This lays a solid foundation for accurate parsing by the subsequent large language model.
[0051] Preliminary keyword extraction: Quickly extract core verbs (such as "fly", "turn", "look") and key nouns (such as "weld", "reflection", "top left corner") from the instructions for preliminary task classification and priority determination.
[0052] The aforementioned technical solution establishes a voice interaction module as the entry point for direct dialogue between the system and human operators. Its primary design goal is to achieve highly robust, low-latency voice recognition and understanding in complex industrial noise environments. This module is far more than a simple speech-to-text tool; it is an intelligent front-end integrating advanced signal processing, context awareness, and instruction preprocessing.
[0053] Furthermore, the large language model intelligent parsing and planning module is the core of the system. This module receives text commands from the voice interaction module. It has a built-in or integrated large language model, which is a pre-trained language model fine-tuned with corpus data from the weld inspection field. After specific training, it can understand professional terms and commands related to non-destructive testing, UAV flight, and pose control. It can understand the inspection context: understand professional terms and operational intentions related to weld inspection (such as "focus," "get closer," "track the weld," "avoid glare," etc.); parse command semantics: parse ambiguous natural language commands into precise, executable pose adjustment parameters or action sequences. For example, parse "fly to the best position where the weld root can be seen" into specific three-dimensional coordinate offsets, yaw angles, and pitch angles; generate control strategies: based on the current UAV status and visual feedback, generate safe and feasible fine-tuning flight paths or control commands.
[0054] Furthermore, the construction of domain expert large language models in the large language model intelligent parsing and planning module includes: Base Model Selection and Fine-tuning: A large pre-trained language model with moderate parameter count and high inference efficiency was selected as the base. The most critical process is domain-adaptive pre-training and instruction fine-tuning. The dataset used for fine-tuning is carefully constructed and mainly includes: a welding and inspection knowledge base: covering international standards (such as ISO 5817, AWS D1.1), national standards (such as GB / T 3323), professional textbooks, inspection report templates, etc., enabling the model to master deep domain knowledge; a UAV flight and photography terminology base: including kinematic parameters (yaw, pitch, roll), photographic parameters (focal length, depth of field), coordinate systems (world coordinate system, body coordinate system, camera coordinate system), etc.; large-scale synthetic dialogue data: through templates and rules, hundreds of thousands of dialogue data simulating inspection scenarios are automatically generated, covering various instructions, question-and-answer, and anomaly handling situations. For example, "What should I do if the current image is blurry?" -> "Suggested instruction: 'Move forward slowly until the image is clear' or 'Adjust the focus'"; real human-computer interaction logs: real dialogue data from prototype system testing is collected for reinforcement learning to further align the model's behavior with the expectations of human experts. Through the above process, the model acquires the following core capabilities: Semantic disambiguation: It can distinguish whether "clear" refers to a sharp image or the removal of obstacles; Contextual memory and reference resolution: It can remember the dialogue history and clarify the specific objects referred to by "it", "that place", and "the defect mentioned earlier"; Fuzzy quantifier quantification: It has a built-in configurable mapping table that converts subjective descriptions such as "fine-tuning", "significant", and "slight" into specific physical quantities (such as displacement of 10cm, 30cm, and 5cm).
[0055] Furthermore, the large language model's intelligent parsing and planning module also includes multi-level semantic parsing and context fusion. When the model parses instructions, it's a complex reasoning process integrating multiple information sources. For the instruction text itself, it analyzes the syntactic structure, extracting subject, verb, object, and modifiers. It understands whether the current instruction is an independent command or a modification or continuation of a previous instruction. Simultaneously, it receives and "understands" the following data in real-time: the drone's status, including current position, attitude, battery level, and signal strength; and the visual context from the visual perception and weld defect recognition module, such as the weld's position, size, clarity score, presence of obstruction / reflection, and distance to the nearest obstacle. For example, if the visual context reports "an obstacle 0.8 meters to the right," and the operator's instruction is "move to the right," the model will intelligently generate a collision-free movement path or proactively ask, "There's an obstacle on the right; we suggest moving 0.5 meters to the left to obtain a similar viewpoint. Should we execute?" Furthermore, the large language model intelligent parsing and planning module also includes control strategy generation under safety constraints. After parsing the intent, the model enters the planning phase. The model internally encodes the physical limits of the drone, such as maximum speed, maximum tilt angle, and minimum turning radius, to ensure that the generated path is smooth and feasible. The model also integrates real-time obstacle maps from the vision module and preset electronic fence information. Collision detection is performed during each planned movement. The ultimate goal of planning is to optimize the detection perspective. The model weighs multiple objectives: the centering of the weld in the image, image clarity, lighting quality, and the stability and energy consumption of the drone itself. It may generate a multi-step optimization strategy, such as, "first adjust the yaw angle to make the weld horizontal, then descend vertically to the optimal observation distance, and finally fine-tune the gimbal angle."
[0056] The above technical solution transforms a general-purpose conversational AI into a domain expert system proficient in non-destructive testing and drone operation. Its capabilities are no longer limited to simple text generation, but have been elevated to context-aware task understanding and motion planning under safety constraints.
[0057] Furthermore, the visual perception and weld condition analysis module includes multi-sensor fusion perception. UAV platforms typically carry multiple visual sensors that work collaboratively. For example, a high-definition main camera is used to acquire high-quality inspection images, with a resolution sufficient to distinguish minute welding defects. A binocular vision camera or RGB-D camera core is used for real-time 3D reconstruction and depth perception. Dense or semi-dense point cloud data of the scene is acquired through parallax calculation or direct measurement, which is crucial for accurate distance measurement and obstacle recognition. A wide-angle auxiliary camera provides a broader peripheral field of view for large-scale obstacle perception and situational awareness.
[0058] Furthermore, the visual perception and weld condition analysis module also includes deep learning-based weld recognition and tracking. A lightweight convolutional neural network (such as an improved version of YOLO or SSD) is used to detect weld regions in images in real time, outputting their precise bounding boxes. Then, the detected weld regions are further segmented pixel-level, their centerlines are extracted, and their geometric orientation (straight line, curve) in the image coordinate system is fitted. Finally, combining the 2D position of the weld in the image and the 3D information provided by the depth camera, algorithms such as PnP are used to estimate the 3D coordinates and orientation of the weld key points relative to the UAV camera coordinate system.
[0059] Furthermore, the visual perception and weld condition analysis module also includes a real-time image quality assessment system. This is one of the module's core functions; it acts like a rigorous "quality inspector," scoring each pose adjustment. The assessment system includes multiple quantifiable indicators, including but not limited to: Composition score: Calculate the pixel distance between the center of the weld area and the center of the image, as well as the area occupied by the weld in the image.
[0060] Sharpness score: The sharpness score is calculated using a no-reference image sharpness evaluation method such as the Tenengrad gradient function or the Brenner gradient function. The higher the value, the sharper the image.
[0061] Lighting quality rating: Analyze the image histogram to assess whether the overall brightness is appropriate (avoiding overexposure or underexposure), whether the contrast is sufficient, and detect the presence of localized highlights or shadows.
[0062] Defect Visibility (Speculation): An advanced feature that allows training a classifier to infer the detectability probability of common welding defects (such as cracks and porosity) at the current viewpoint and resolution based on current image features.
[0063] Through the above technical solution, the visual perception and weld condition analysis module acquires environmental images in real time using the visual sensors mounted on the UAV, and accurately identifies the weld area in the image using computer vision algorithms, calculating its three-dimensional position and orientation relative to the UAV. Simultaneously, it analyzes the current image quality, such as sharpness, contrast, presence of occlusion or overexposure, providing feedback for pose fine-tuning.
[0064] Furthermore, the UAV flight control and attitude execution module includes command translation and interface adaptation. The module receives high-level commands generated from a large language model (e.g., "Move to point (x,y,z) in the world coordinate system while simultaneously pointing the nose towards the yaw angle ψ"). The flight control module needs to translate this command into MAVLink messages or SDK function calls that the underlying flight controller (e.g., PX4, ArduPilot) can understand. This may involve coordinate system transformation (from world frame to aircraft frame), motion mode switching (e.g., switching from hovering mode to position control mode), and trajectory interpolation to ensure smooth movement.
[0065] Furthermore, the UAV flight control and attitude execution module includes high-precision attitude control. The UAV platform itself needs to possess a high-performance flight control system, typically integrating GPS, inertial measurement unit, and visual odometry / SLAM information to achieve centimeter-level positioning accuracy and stable hovering, even indoors or in environments without GPS signals. During attitude fine-tuning, the flight controller precisely drives the rotational speed of each motor through a PID controller or more advanced control algorithms, thereby achieving closed-loop control of the six degrees of freedom attitude.
[0066] Through the above technical solution, the UAV flight control and attitude execution module receives control commands generated by the large language model module, converts them into low-level commands that the UAV flight controller can understand, drives the UAV's motors and servos, and precisely controls its position, attitude and gimbal angle to perform attitude fine-tuning actions.
[0067] Furthermore, the status feedback and confirmation module includes multimodal feedback channels. The module employs a natural and fluent TTS engine to convert system status, execution results, and warning messages into voice broadcasts. Examples include: "Command received, path planning in progress," "Moved 15 cm to the left, image clarity improved by 20%," and "Warning: Only 0.5 meters of obstacle remaining on the left, movement paused, please confirm." The module overlays key information on the operator's ground station software interface: such as using AR arrows to indicate the planned movement direction, using color bars to display image quality scores, and using warning boxes to highlight approaching obstacles. Simultaneously, LEDs on the drone's body can indicate simple statuses, such as solid blue for standby, flashing green for execution, and red for warnings, providing the operator with quick auxiliary status prompts.
[0068] Furthermore, the status feedback and confirmation module includes intelligent confirmation and interaction mechanisms. When the system assesses that a command poses a certain risk (such as being very close to an obstacle) or that the intent is ambiguous, it will proactively pause execution and request instructions from the operator via voice. For example, it might announce: "It is estimated that after moving, you will be only 30 centimeters away from the pipe. Do you want to continue? Please say 'confirm' or 'cancel'." Additionally, the system can provide more proactive suggestions. For instance, if the vision module continuously detects a blurred image, but the operator has not issued an instruction, the system can proactively prompt: "Continuous image blurring has been detected. Is automatic refocusing or position adjustment necessary?" Through the above technical solution, the five modules described in detail in the first aspect of this invention constitute a complete intelligent closed loop from perception, understanding, decision-making to execution and feedback. By combining the deep semantic understanding capabilities of a large language model with the precise environmental perception capabilities of computer vision, it creatively solves the problem of precise pose adjustment for UAVs, liberating human-computer interaction from cumbersome low-level controls and moving towards a new paradigm of efficient and natural interaction oriented towards task objectives. This greatly improves the practicality, safety, and efficiency of UAV detection in complex industrial scenarios.
[0069] The above system forms a closed-loop control loop: the visual perception and weld condition analysis module takes the environmental image set and weld image set collected in real time by the UAV, and after training the small image recognition model after building the weld defect image training set, it feeds the recognition results back to the large language model module as the basis for its optimized control strategy.
[0070] In one embodiment, the above-mentioned voice-assisted fine-tuning system for nuclear power plant construction weld inspection UAV based on a large language model includes: Weld inspection drones; The visual sensor gimbal subsystem is used to acquire environmental and weld seam image information in real time; The voice interaction subsystem is used to collect natural language voice commands issued by the operator to adjust the drone's detection pose, and to return the drone fine-tuning control commands generated by the system to the operator.
[0071] A weld image defect recognition system based on small model learning is used for real-time weld defect recognition. A large language model computing platform is used to generate precise pose control commands; The robot fine-tuning control module is used to perform closed-loop attitude fine-tuning control of the UAV.
[0072] In some embodiments, the above-described method for applying the above-described voice-assisted fine-tuning system for nuclear power plant construction weld inspection UAV based on a large language model can also be implemented in the following ways.
[0073] In this embodiment, the method applied to the above-mentioned voice-assisted fine-tuning system for nuclear power plant construction weld inspection UAV based on a large language model includes the following steps: Receive natural language voice commands from the operator to adjust the drone's detection pose; The voice commands are converted into text commands, and their semantics are parsed using a large language model to generate precise pose control commands. The control commands are verified for security based on visual perception information. Drive the drone to execute verified control commands to complete pose fine-tuning; Provide feedback to the operator regarding the results of instruction execution.
[0074] Specifically, in the step of generating precise pose control commands, the large language model comprehensively references the following information: By using the natural language voice commands issued by the operator to adjust the drone's detected pose, and combining the large language model intelligent parsing and planning module, a drone voice-assisted fine-tuning control text with clear and standard drone pose control commands is generated. The drone's real-time pose database text is generated by using its own position sensors and gyroscopes. Using the visual gimbal carried by the drone, a set of environmental images and weld images are collected in real time. After training on a small image recognition model based on a nuclear-constructed weld defect image training set, an intelligent evaluation system generates real-time weld location and image quality data text.
[0075] After pose fine-tuning, the method also includes: re-evaluating the quality of the detected image through the vision module; if it does not meet the predetermined standard, automatically performing compensatory fine-tuning or prompting the operator to issue further instructions.
[0076] Figure 2 This is an overall flowchart of a method for applying a voice-assisted fine-tuning system for a nuclear power plant construction weld inspection drone based on a large language model. In this embodiment, the method applied to the aforementioned voice-assisted fine-tuning system for a nuclear power plant construction weld inspection drone based on a large language model includes the following steps: S1: System Initialization and Task Startup The drone takes off and automatically flies to the approximate starting area for weld inspection. The vision module then begins to work, initially identifying and tracking the weld.
[0077] S2: Voice command reception and recognition When the operator detects that the position needs to be adjusted based on the real-time video feed, they can issue natural language commands through the microphone, such as: "Tilt the gimbal down a little so that the weld is in the center of the screen." S3: Intelligent Command Parsing and Action Planning After the voice command is recognized as text, it is fed into the large language model module. This module performs a comprehensive understanding based on the following context: The textual semantics of the current instruction.
[0078] The vision module provides real-time weld location and image quality information.
[0079] The current pose of the drone (GPS coordinates, altitude, attitude angle).
[0080] The model outputs one or more precise control actions, such as: "Reduce the gimbal pitch angle by 3 degrees".
[0081] S4: Security Verification and Command Execution The system performs safety checks on the planned maneuvers, including collision risk and flight boundary verification. Once the checks are passed, the commands are sent to the flight control module for execution.
[0082] Step S5: Closed-loop feedback and iterative optimization After the drone performs fine-tuning, the vision module re-evaluates the image quality. If it still does not reach the optimal state, the system can automatically make minor compensation adjustments, or issue further instructions to the operator via voice prompts, forming a closed loop of "instruction-execution-evaluation" until a satisfactory detection perspective is obtained.
[0083] S6: Continue or end the task Once the pose adjustment is satisfactory, the drone continues to perform automatic detection tasks. The operator can interrupt the automatic process at any time to intervene and make new voice adjustments.
[0084] In some embodiments, the above-described nuclear power plant construction weld inspection drone voice-assisted fine-tuning system and the method applied to the system can also be implemented in the following ways.
[0085] Figure 3 This is a logical diagram illustrating the parsing of voice commands and the generation of pose control. This embodiment uses the scenario of inspecting the appearance quality of weld seams in the stainless steel wall of a nuclear power plant reactor as an example.
[0086] 1. The drone flew along the preset path to the vicinity of the stainless steel wall weld of the nuclear power plant reactor, but due to the obstruction of scaffolding, one side of the weld was clear in the camera image, while the other side was slightly blurry; 2. After observing the ground station screen, the operator said, "Move the drone half a meter to the right and keep the lens always facing the weld surface."
[0087] 3. The voice interaction module recognizes the command. After receiving the text, the large language model module parses it as follows: "Move half a meter to the right": This can be interpreted as moving 500mm in the positive Y-axis direction of the UAV's current coordinate system. "Keep the camera always facing the weld surface": This is a dynamic constraint. The model needs to combine the weld surface normal direction calculated in real time by the vision module, and dynamically adjust the yaw and pitch angles of the UAV during translation to ensure that the camera optical axis is always parallel to the local normal of the weld. 4. The model generates a translation path command with attitude constraints and sends it to the flight control module; 5. The drone executed the command smoothly, and the weld seam remained clearly centered in the image throughout the entire movement; 6. System voice announcement: "Position adjustment complete, have reached the new position."
[0088] This embodiment takes the visual inspection of the circumferential weld of the auxiliary steel structure module of the nuclear power plant construction site as an example. When the UAV has approached the weld to be inspected along the preset flight path, but there are scaffold shadows, local reflections and wind disturbances causing hovering deviations on site, the semantic-visual collaborative fine-tuning network described in this embodiment is used to input the operator's natural language commands, the UAV's real-time pose and the weld image quality assessment results into the large language model intelligent parsing and planning module to generate fine-tuning control quantities that meet safety constraints.
[0089] This embodiment further improves the large language model intelligent parsing and planning module and the visual perception and weld defect recognition module based on the aforementioned voice interaction module, large language model intelligent parsing and planning module, visual perception and weld defect recognition module, UAV flight control and posture execution module and status feedback and confirmation module, so that they form a trainable, deployable and quantifiable network model.
[0090] I. Network Model Construction Furthermore, the semantic-visual collaborative fine-tuning network constructed above includes a text instruction encoder, a weld visual quality encoder, a pose state encoder, a cross-modal gating fusion layer, a safety constraint planning layer, and a control output layer.
[0091] 1. A text instruction encoder receives text instructions u output by the voice interaction module. These text instructions are segmented using a domain lexicon and then input into a large language model that has undergone low-rank adaptation fine-tuning. The output is a semantic vector h_t. The domain lexicon includes at least weld location terms, welding defect terms, UAV action terms, gimbal control terms, and safety confirmation terms.
[0092] 2. The weld visual quality encoder receives the current frame weld image I_t and outputs the weld center point c_t, weld orientation angle theta_t, defect candidate heatmap M_t, obstacle distance d_t, and image quality vector q_t. The image quality vector includes composition score, sharpness score, illumination score, occlusion score, and defect visibility score.
[0093] 3. The pose state encoder is used to receive the current state of the UAV p_t=[x_t,y_t,z_t,phi_t,varphi_t,psi_t,gamma_t], where phi_t is the roll angle, varphi_t is the pitch angle, psi_t is the yaw angle, and gamma_t is the gimbal pitch angle, and maps this state to the pose vector h_p.
[0094] 4. The cross-modal gating fusion layer is used to determine the weights of text, visual, and pose information based on the semantics of the current task, resulting in a fusion vector h_f, which is calculated as follows: g_t=sigma(W_g[h_t;h_v;h_p]+b_g) h_f=LN(g_t⊙h_t+(1-g_t)⊙W_v[h_v;h_p]) Where sigma represents the Sigmoid function, LN represents layer normalization, ⊙ represents element-wise multiplication, and W_g, W_v, and b_g are trainable parameters. Using this formula, when the operator issues a vague instruction such as "move slightly closer," the model increases the weights of the visual quality vector and the current distance information; when the operator issues a clear instruction such as "shift 20 centimeters to the left," the model increases the weights of the text semantic vector.
[0095] 5. The safety constraint planning layer is used to evaluate candidate paths before outputting fine-tuning actions. The objective function for optimizing candidate paths is: J=Σ_{k=0}^{K-1}(||Δp_k||_R^2-ηQ(I_{k+1})+λ / (d_obs,k-d_min+ε)^2+μ||Δp_k-Δp_{k-1}||_2^2) Where Δp_k is the pose fine-tuning amount at step k, Q(I_{k+1}) is the predicted image quality score for the next frame, d_obs,k is the distance to the nearest obstacle at step k, d_min is the safe distance threshold, ε is a constant to prevent the denominator from being zero, and η, λ, and μ are weighting coefficients. This objective function simultaneously constrains motion energy consumption, detected image quality, obstacle risk, and trajectory smoothness.
[0096] 6. The control output layer outputs a six-dimensional control vector a_t=[Δx,Δy,Δz,Δψ,Δγ,v] and a risk confidence level ρ_t, where Δx, Δy, and Δz represent three-dimensional translation, Δψ represents yaw adjustment, Δγ represents gimbal pitch adjustment, and v represents the execution speed. If ρ_t is higher than a set threshold, the control input enters the flight control and attitude execution module; if ρ_t is lower than the threshold, the status feedback and confirmation module sends a secondary confirmation voice message to the operator.
[0097] II. Detailed Steps of the Training Method Step A1: Construct training samples. Collect voice commands, ASR-transcribed text, UAV poses, weld images, expert-corrected actions, and post-execution image quality data under different lighting conditions, distances, and weld types at the nuclear power plant construction site, forming a training set D={(u_i,I_i,p_i,a_i,Q_i)}. Where u_i is the text command, I_i is the weld image, p_i is the pose state, a_i is the expert action, and Q_i is the image quality score annotated by the expert or evaluated by the algorithm.
[0098] Step A2: Perform speech-text cleaning. Colloquial instructions in the training samples are processed by removing filler words, normalizing synonyms, and annotating action granularity. For example, "Get a little closer over there, don't touch the scaffolding" is labeled as the target action "approach the weld," the constraint action "maintain a safe distance," and the vague quantifier "slightly."
[0099] Step A3: Train the weld visual quality encoder. Joint supervision is performed using weld region bounding boxes, weld centerlines, defect candidate regions, and image quality labels. The visual loss function is: L_v=αL_box+βL_seg+χ(1-Q_hat)^2+δL_defect Where L_box is the weld detection bounding box loss, L_seg is the weld centerline or region segmentation loss, Q_hat is the model-predicted image quality score, L_defect is the defect candidate region classification loss, and α, β, χ, and δ are weights.
[0100] Step A4: Fine-tune the domain-specific instructions for the large language model. Welding standards, inspection reports, UAV flight control semantics, and on-site dialogue samples are converted into instruction-response pairs. A low-rank adaptation parameter matrix is used to fine-tune the base model, enabling the model to output structured control intentions. Control intentions should include at least the direction of action, the magnitude of action, constraints, confirmation requirements, and expected visual effects.
[0101] Step A5: Perform cross-modal alignment training. Align the target state described by the text instruction with the actual weld state output by the visual encoder in the same latent space. The semantic-action consistency loss is: L_c=1-cos(E_t(u_i),E_a(a_i))+κ||Q_i-Q_hat_i||_1 Where E_t is the text instruction encoding function, E_a is the action semantic encoding function, cos represents the cosine similarity, and κ is the quality score constraint weight. This loss enables semantics such as "center the weld", "avoid reflection", and "approach the crack area" to establish stable correspondences with specific displacements, yaw angles, and gimbal angles.
[0102] Step A6: Perform expert teaching and safety enhancement fine-tuning. First, use expert actions a_i for imitation learning to make the model's output action a_hat_i closely resemble the expert action; then, add disturbance conditions such as scaffolding, pipelines, wind disturbance, and local strong glare to the simulation environment, and further optimize using the safety reward function. The total training loss is: L_total=L_sft+L_v+ω||a_i-a_hat_i||_2^2+τmax(0,d_min-d_obs)^2-ξΔQ Where L_sft is the language model instruction fine-tuning loss, ω, τ, and ξ are weights, and ΔQ represents the improvement in image quality score before and after the action is executed. This formula enables the model to approach the optimal detection viewpoint while suppressing actions that are too close, too fast, or have a high risk of collision.
[0103] Step A7: Model Deployment. Deploy the trained visual quality encoder on the UAV edge computing unit, and deploy the domain adaptation layer and safety planning layer of the large language model on the ground station or edge server; when the network is limited, use the quantized local model to output conservative control values and force the secondary confirmation mechanism to be enabled.
[0104] III. Prototype Validation Data and Charts To demonstrate the feasibility of the technical solution in this embodiment, a 1:1 steel structure weld inspection prototype environment was built in a closed test scenario, setting up four typical working conditions: scaffold obstruction, localized reflection, weld deviation from the center of the image, and random wind disturbance. Each working condition was tested 50 times, and the verification data are shown in Tables 1 and 2.
[0105] Table 1: Prototype verification results for different fine-tuning methods
[0106] Table 2: Ablation Validation Results of the Model in This Example
[0107] As shown in Table 1, under the same test conditions, the network described in this embodiment reduces the average fine-tuning time from 31.6s to 8.7s compared to manual remote fine-tuning, and improves the image quality score from 0.71 to 0.91. Compared to the scheme that only uses large language model planning, this embodiment improves the defect recall rate from 91.1% to 95.8% and reduces the number of safety interruptions to 1 by introducing a weld visual quality encoder and a safety constraint planning layer.
[0108] As shown in Table 2, after removing the visual quality encoder, the model struggles to determine whether objectives such as "clearer" and "avoiding glare" have been achieved. Removing the safety constraint planning layer, while still improving image quality, significantly increases the secondary confirmation trigger rate. Removing the neighborhood instruction fine-tuning layer leads to unstable action mapping for technical terms such as "bite," "lack of fusion," and "weld root." The complete model achieves a better balance between image detection quality, path deviation, and safety confirmation efficiency.
[0109] Figure 4 This is a schematic diagram of the training curve of the semantic-visual collaborative fine-tuning network in this embodiment. The total loss gradually decreases with each training round, while the image quality score gradually increases and tends to stabilize with each training round, indicating that a stable mapping relationship has been formed between text instructions, visual quality, and control actions.
[0110] IV. Example of Specific Execution Process 1. The drone flies to a position about 1.2m in front of the weld to be inspected. The vision module detects that the weld area is located in the lower right corner of the screen. The current image quality score is 0.68 and the nearest scaffolding distance is 0.55m.
[0111] 2. The operator issues a voice command: "Move the weld to the center of the screen, a little closer, but don't touch the scaffolding on the right." 3. The voice interaction module transcribes the instructions into text and removes filler words. The large language model intelligent parsing and planning module identifies the target as "centering the weld" and "approaching it appropriately", with the constraint as "avoiding the right-side scaffolding".
[0112] 4. The semantic-visual collaborative fine-tuning network outputs a control vector a_t=[-0.08m,0.16m,0.03m,-4.5°,-2.0°,0.18m / s] based on the current weld center offset, sharpness score, and obstacle distance, with a risk confidence level ρ_t=0.94.
[0113] 5. The flight control and attitude execution module converts the above control vectors into MAVLink position control commands and gimbal control commands, and continuously reads the visual quality score during execution. If d_obs is less than 0.35m at any moment, it immediately pauses and triggers voice confirmation.
[0114] 6. After execution, the distance between the weld center and the image center decreased from 186 pixels to 24 pixels, the sharpness score improved from 0.62 to 0.87, and the overall image quality score improved to 0.91. The status feedback and confirmation module announced: "Fine-tuning completed, weld centered, obstacle distance on the right is 42 cm, current image meets inspection requirements." Through the above embodiments, this example further illustrates how the present invention unifies natural language commands, weld visual quality evaluation, and UAV pose control into a trainable closed-loop network model, enabling the UAV to automatically generate safe, accurate, and verifiable fine-tuning actions based on the operator's high-level semantic objectives in the complex nuclear power plant construction site. The embodiments of the present invention have been described above with reference to the accompanying drawings. However, the present invention is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art, under the guidance of the present invention, can make many modifications without departing from the spirit and scope of the claims, and all such modifications are within the protection scope of the present invention.
Claims
1. A voice-assisted fine-tuning system for unmanned aerial vehicle (UAV) weld inspection in nuclear power plant construction based on a large language model, characterized in that, include: The voice interaction module is configured to: acquire voice commands, preprocess the voice commands and perform voice recognition and command structuring to obtain text commands; The large language model intelligent parsing and planning module is configured to: generate a control strategy for fine-tuning the UAV pose based on the text instructions and environmental context information using a domain expert large language model; The visual perception and weld defect recognition module is configured to: acquire environmental image information in real time, use a small model to recognize welds, and perform image quality assessment to obtain weld defect recognition information including weld defect type and weld image quality. The UAV flight control and attitude execution module is configured to execute the control strategy and drive the UAV to complete attitude adjustment. The status feedback and confirmation module is configured to receive image information, weld defect identification information, and UAV attitude fine-tuning control information, provide voice feedback on the system status to the operator, and assist the operator in generating the next motion control command based on the current UAV status.
2. The system according to claim 1, characterized in that, The preprocessing includes adaptive noise suppression, acoustic echo cancellation, automatic gain control, and speech activity detection.
3. The system according to claim 1, characterized in that, The speech recognition and instruction structuring includes informal corpus regularization, instruction boundary detection and segmentation, and preliminary keyword extraction.
4. The system according to claim 1, characterized in that, The domain expert large language model is pre-trained using a base model fine-tuned based on the corpus of weld inspection, and is used to understand professional terms and instructions related to non-destructive testing, UAV flight, and pose control.
5. The system according to claim 4, characterized in that, The corpus for weld inspection includes a professional knowledge base for welding and inspection, a terminology database for UAV flight and photography, large-scale synthetic dialogue data, and real human-computer interaction logs.
6. The system according to claim 1, characterized in that, The small model is a lightweight convolutional neural network.
7. The system according to claim 1, characterized in that, The weld seam recognition includes: acquiring environmental images using multiple visual sensors; acquiring dense or semi-dense point cloud data of the scene using parallax calculation or direct measurement; detecting weld seam regions in the image in real time using a lightweight convolutional neural network and outputting their precise bounding boxes; performing pixel-level segmentation on the detected weld seam regions; extracting their centerlines and fitting their geometric orientation in the image coordinate system; and combining the 2D position of the weld seam in the image with the 3D information provided by the depth camera, estimating the three-dimensional coordinates and orientation of the weld seam key points relative to the UAV camera coordinate system using the PnP algorithm.
8. The system according to claim 1, characterized in that, The image quality assessment is used to score each pose adjustment according to an assessment system that includes multiple indicators, including composition score, sharpness score, illumination quality score, and defect visibility.
9. The system according to claim 8, characterized in that, The composition score is calculated by: calculating the pixel distance between the center of the weld area and the center of the image, and the area occupied by the weld in the image; the sharpness score is calculated by: using the Tenengrad gradient function or the Brenner gradient function to calculate the edge intensity of the image; the illumination quality score is calculated by: analyzing the image histogram to assess whether the overall brightness is appropriate and the contrast is sufficient, and detecting whether there are local highlights or shadows; the defect visibility is calculated by: training a classifier to infer the detectability probability of common welding defects at the current viewing angle and resolution based on the current image features.
10. A method for applying a voice-assisted fine-tuning system for a nuclear power plant construction weld inspection UAV based on a large language model as described in any one of claims 1-9, characterized in that, Includes the following steps: S1: System initialization and task start-up: The UAV takes off and automatically flies to the approximate starting area of weld inspection. The vision module starts working, initially identifying and tracking the weld. S2: Voice command reception and recognition: Receiving and recognizing natural language commands issued by the operator through the microphone when the operator finds that the posture needs to be adjusted based on the real-time video transmission screen; S3: Intelligent command parsing and action planning: After the voice command is recognized as text, it is sent to the large language model module for comprehensive understanding in combination with the following context, including the text semantics of the current command, real-time weld position, image quality information, and the current pose of the drone, and outputs precise control actions. S4: Safety Verification and Command Execution: Perform safety verification on the planned actions, including collision risk and flight boundary. After passing the verification, send the command to the flight control module for execution. S5: Closed-loop feedback and iterative optimization: The image quality is evaluated using a voice-assisted fine-tuning system for nuclear power plant construction weld inspection drones based on a large language model. If the image quality is still not optimal, fine compensation adjustments are made, or the operator is prompted by voice to issue further instructions until a satisfactory inspection view is obtained. S6: Task Continue or End: After the pose adjustment is satisfactory, the drone continues to perform the automatic detection task and waits for new voice commands.