Unmanned aerial vehicle dynamic environment perception response system based on end side LLM
By deploying end-side LLM and local knowledge enhancement technology on drones, the high latency and network dependence problems of drone perception systems are solved, real-time target recognition and independent decision-making are achieved, and the system's adaptability and decision-making professionalism in complex scenarios are improved.
Patent Information
- Application Number
- CN202510554082.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-29
- Publication Date
- 2025-08-15
AI Technical Summary
The existing drone perception systems rely on cloud processing, resulting in high latency and network dependence, making it difficult to achieve real-time target recognition and autonomous decision-making, and lack effective utilization of local environmental information and security specifications, making them poorly adaptable.
The end-side large language model (LLM) is used to combine local knowledge enhancement technology, and the visual perception module, warning recognition module, LLM reasoning decision-making module and two-way network communication module are used to realize localized perception and decision-making of drones. The YOLOv5 model is used for real-time object detection, and RAG technology enhances LLM reasoning, builds a local knowledge base, and realizes multi-level situation evaluation and strategy generation.
It realizes dynamic threat perception with low latency and high robustness, enhances the system's adaptability and decision-making professionalism in unstructured scenarios, reduces network communication pressure, and ensures real-time response capabilities in complex environments.
Smart Images

Figure CN120495929A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence and unmanned aerial vehicle (UAV) application technology, and in particular to a UAV dynamic environment perception and response system based on end-side LLM. Background Art
[0002] In recent years, the widespread use of drones in complex scenarios such as emergency rescue, disaster monitoring, and military reconnaissance has placed higher demands on real-time target recognition and autonomous decision-making. Currently, existing drone perception systems primarily rely on cloud-based systems for target recognition and decision-making, which suffer from high latency and network dependency. Collected data must be uploaded to the cloud for processing, resulting in insufficient real-time performance and difficulty responding to sudden threats. Existing traditional systems are also limited by the computing power of edge devices, making it difficult to run complex deep learning models and exhibiting weak generalization capabilities in complex scenarios. Existing systems often lack effective utilization of prior knowledge, are unable to incorporate local environmental information and safety regulations into decision-making, and exhibit poor adaptability.
[0003] To address the above issues, the present invention proposes a solution based on end-to-end LLM, which achieves low-latency, high-robust dynamic threat perception through local target detection, knowledge-enhanced reasoning, and real-time environmental feedback to optimize decision-making. Summary of the Invention
[0004] The purpose of this invention is to provide a UAV dynamic environment perception and response system based on end-side LLM, aiming to achieve real-time target recognition and decision-making response for UAVs in complex scenarios, and provide key technical support for the application of UAVs in emergency rescue, disaster monitoring, military reconnaissance and other fields.
[0005] To achieve the above objectives, the present invention provides a UAV dynamic environment perception and response system based on device-side LLM, including:
[0006] The visual perception module collects image data in real time through multimodal sensors, recognizes the acquired images using the YOLOv5 model, and outputs recognition information;
[0007] The warning identification module generates structured labels based on identification information and optimizes label accuracy based on historical information;
[0008] LLM reasoning and decision module: The pre-deployed LLM model performs reasoning based on labels and scenario descriptions, and provides threat assessment and decision content;
[0009] The local knowledge enhancement module pre-builds a local knowledge base containing safety specifications and typical cases, uses RAG technology to enhance LLM reasoning, distinguish categories with different labels, and provide corresponding solutions;
[0010] The two-way network communication module builds a communication link based on the TCP / IP protocol to achieve two-way communication between the drone and the ground.
[0011] The visualization and interaction module, built on Gradio, displays image recognition information and target assessment results in real time. Users send requests to the end device to obtain a real-time screen preview.
[0012] Preferably, the visual perception module includes a sensor part and a data preprocessing part. The environmental perception capability of the drone is based on the onboard camera of its sensor part and the locally deployed YOLOv5 model, which constitutes a real-time computing process as follows:
[0013] First, raw video frames are continuously acquired through the camera interface;
[0014] Then, preprocessing operations are performed on each frame to meet the model input requirements;
[0015] Next, the preprocessed image frames are fed into the YOLOv5 model for forward inference, which outputs the bounding box coordinates, category predictions, and corresponding confidence scores of potential objects in the image. The model's raw output is then post-processed and parsed to extract human-readable labels and quantified confidence scores.
[0016] Finally, the above key information is encapsulated into a standardized JSON object, and the system periodically transmits the information to the ground station at a preset frequency; at the same time, the system uses the YOLOv5 model output and the original frame to generate a visual annotation image with superimposed detection boxes and labels, and temporarily stores it in the memory buffer.
[0017] Preferably, the specific implementation steps of the warning identification module are:
[0018] First, each detected object label is semantically classified according to a predefined hazardous materials ontology library, which maps specific labels to preset hazardous categories;
[0019] Then, a multi-level confirmation protocol is set up. When a detected event meets the key dimensions of the multi-level confirmation protocol at the same time, it is confirmed as a potential dangerous symptom that requires escalation.
[0020] Secondly, through the target status tracking mechanism, the current processing stage of each confirmed dangerous target is recorded;
[0021] Finally, when a potential dangerous incident passes the multi-level confirmation protocol and the current status allows the initiation of a new process, the system immediately sends a preliminary alert message to the ground station, including the timestamp, dangerous goods label, confidence level at the time of confirmation, and its category.
[0022] Preferably, a multi-level confirmation protocol includes three key dimensions:
[0023] Time window, which requires the target to appear continuously or repeatedly within a specified short period of time;
[0024] Frequency threshold, which requires that the number of successful detections of a target within the time window reaches a minimum value;
[0025] The confidence threshold requires that the confidence of each detection that constitutes a valid detection count is higher than a preset lower limit.
[0026] Preferably, the implementation steps of the LLM reasoning and decision module are as follows: after the initial alarm is issued, the system immediately starts the cognitive processing process and interacts with the locally deployed LLM model. The key lies in the prompt word engineering: the system dynamically constructs an input prompt based on a preset template. The prompt embeds the key elements of the current situation: the background of the incident, the specific hazardous materials identified, the quantified confidence level and the tasks to be completed by the LLM, and generates a solution, which contains specific components: risk assessment, disposal steps and notification objects; then, the system calls the LLM processing process asynchronously.
[0027] Preferably, the local knowledge enhancement module adopts the retrieval enhancement generation RAG paradigm to complete reasoning locally on the drone. When the structured prompt words constructed by the LLM reasoning decision module are received, the RAG process starts:
[0028] First, in the retrieval phase, the system uses the key information in the prompt word as a query vector and performs similarity search in a pre-built local vectorized knowledge base, aiming to find the text fragments in the knowledge base that are most relevant to the current context.
[0029] Then, it enters the enhancement phase, dynamically integrating the retrieved highly relevant knowledge fragments with the original prompt word to form an enhanced prompt word;
[0030] Finally, in the generation phase, the enhanced prompt words are input into the locally deployed LLM to generate a response.
[0031] Preferably, the specific implementation steps of the visualization and interaction module are:
[0032] After LLM generates a disposal plan, the plan and related metadata are encapsulated into a JSON message of LLM response type and transmitted to the ground station through the text / JSON channel. The ground station application receives and parses the message, stores it and associates it with the corresponding preliminary alarm, and updates the status indicator of the alarm in the user interface. The operator interacts with the system through a web-based graphical user interface, which contains several situational awareness panels and a dynamically updated alarm list, showing the processing status of all confirmed dangerous time machines. When the operator selects an alarm in the list, a dedicated area on the interface will render and display the detailed disposal suggestions generated by LLM. The interface provides a request image function. Activating this function will trigger the ground station to send a request instruction to the drone through the image dedicated channel. After receiving the instruction, the drone extracts the latest visual image frame with detection annotations from the memory buffer, performs JPRG compression, and then transmits it back to the ground station and displays it in the graphical user interface.
[0033] Preferably, the two-way network communication module relies on the communication link between the UAV and the ground station, follows the TCP / IP protocol stack, and is designed with port multiplexing and traffic isolation strategies: one logical channel is designated for transmitting low-bandwidth, low-latency control signaling and structured data, including target detection metadata, hazardous materials alert identifiers, and LLM response texts; another logical channel is reserved for on-demand transmission of high-bandwidth image data streams; for long text messages that exceed the capacity of a single TCP segment, an application layer segmentation and reassembly strategy is adopted, and by adding a message end delimiter, it is ensured that the receiving end can restore the original information, and the data interaction follows the JSON format.
[0034] Therefore, the present invention adopts the above-mentioned UAV dynamic environment perception and response system based on the end-side LLM, which has the following beneficial effects:
[0035] (1) By deploying a large language model (LLM) on the device side, this system breaks through the strong dependence of traditional expert models on specific data sets. Compared with expert systems based on fixed rules, the pre-trained knowledge system of LLM gives drones the ability to understand cross-modal semantics and can independently build a dynamic environment cognition framework. This feature significantly enhances the system's adaptability in unstructured scenarios, solves the scene generalization limitations of traditional visual models caused by training data bias, and reduces the dependence on high-quality annotated data;
[0036] (2) LLM’s causal reasoning and logical deduction capabilities enable multi-level situation assessment and strategy generation;
[0037] (3) Based on the Retrieval Enhancement Generation (RAG) technology, a domain knowledge enhancement framework is constructed. By dynamically accessing the professional knowledge base, the basic LLM is equipped with the ability to transfer vertical domain knowledge. This design enables the system to autonomously adjust the knowledge weight according to the characteristics of the task scenario, while maintaining the general semantic understanding ability and ensuring the professionalism and compliance of decision-making in specific scenarios;
[0038] (4) Deploy a lightweight LLM model on the device side to reduce the latency uncertainty caused by the cloud transmission link and achieve real-time response in the perception-decision closed loop. Localized reasoning ensures efficient utilization of computing resources in complex environments. Raw perception data is extracted and reasoned on the device side, thereby reducing the pressure on network communication transmission bandwidth.
[0039] The technical solution of the present invention is further described in detail below through the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] Figure 1 Schematic diagram of the system architecture of an embodiment of the present invention;
[0041] Figure 2 Flowchart of an embodiment of the present invention. DETAILED DESCRIPTION
[0042] The following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but rather merely represents selected embodiments of the present invention. All other embodiments derived by persons of ordinary skill in the art based on the embodiments of the present invention without inventive effort shall fall within the scope of protection of the present invention.
[0043] See also Figure 1-Figure 2 The UAV dynamic environment perception and response system based on the end-side LLM includes:
[0044] The visual perception module collects image and external data in real time through multimodal sensors. The drone's environmental perception capability comes from its onboard camera and a locally deployed deep learning object detection model (YOLOv5), which constitutes a real-time computing process:
[0045] First, raw video frames are continuously acquired through the camera interface;
[0046] Then, preprocessing operations are performed on each frame, including color space conversion (for example, from the camera's native BGR format to the RGB format required by the model) and possible resizing or normalization to meet the model input requirements;
[0047] Next, the preprocessed image frames are fed into the YOLOv5 model for forward inference. As an efficient single-stage detector, YOLOv5 can quickly output the bounding box coordinates (BoundingBoxes), class predictions (Class Labels), and corresponding confidence scores of potential targets in the image. The model's raw output is post-processed and parsed to extract human-readable labels and quantified confidence scores.
[0048] Finally, the above key information (timestamp, label, confidence) is encapsulated into a standardized JSON object, and the system periodically transmits the information to the ground station at a preset frequency (for example, 5 times per second, depending on the trade-off between processing power and network bandwidth) through the text / JSON channel established by the data transmission module; at the same time, the system uses the YOLOv5 model output and the original frame to generate a visual annotation image with superimposed detection boxes and labels, and temporarily stores it in the memory buffer to respond to on-demand image retrieval requests that may be issued by the ground station.
[0049] The alert recognition module generates structured labels based on the recognition information and optimizes label accuracy by combining historical information, including:
[0050] First, each detected object label is semantically classified according to a predefined hazardous materials ontology library, which maps specific labels (such as "knife" and "fire") to pre-set hazard categories (such as "Weapons" and "Hazardous Materials").
[0051] Then, to improve the robustness of the alarm and filter out transient noise or low-confidence detections, the system sets up a multi-level confirmation protocol, which contains three key dimensions: the temporal window, which requires the target to appear continuously or repeatedly within a specified short period of time (for example, 3 seconds); the frequency threshold, which requires that the number of successful detections of the target reach a minimum value (for example, 3 times) within the time window; and the confidence threshold, which requires that the confidence of each detection that constitutes a valid detection count must be higher than a preset lower limit (for example, 0.6). When a detection event simultaneously meets the key dimensions of the multi-level confirmation protocol, it is confirmed as a potential dangerous symptom that requires escalation. Only detection events that meet all three conditions are confirmed as potential dangerous symptom that require escalation. To prevent redundant alerts and repeated LLM processing requests for the same persistent threat source, the system maintains a target status tracking mechanism that records the current processing stage of each confirmed dangerous target (such as "alarmed pending", "LLM processing", "responded") and may include cooldown period logic;
[0052] Finally, once a potential dangerous incident passes the multi-level confirmation protocol and the current status allows the initiation of a new process, the system immediately sends an InitialAlert message to the ground station via a text / JSON channel, including a timestamp, a dangerous goods label, the confidence level at the time of confirmation, and its category.
[0053] In the LLM reasoning and decision module, a pre-deployed LLM model reasoned based on labels and scenario descriptions, providing threat assessments and decision-making. Specifically, after an initial alert is issued, the system immediately initiated a cognitive processing process, centered on interaction with a locally deployed Large Language Model (LLM). Key to this process is prompt engineering: Rather than simply querying the LLM, the system dynamically constructs a highly structured, information-rich input prompt based on a pre-set template. This prompt embeds key elements of the current context: clearly identifying the context of the incident ("In this regulated environment..."), the specific hazardous material identified (category, label), a quantified confidence level, and clearly articulating the task the LLM is tasked with completing. The LLM is required to generate a structured solution, including specific components (risk assessment, action steps, and notification recipients), by referencing a "prepared solution" (implicitly referencing the RAG knowledge base). This carefully crafted prompt ensures the LLM accurately understands the task requirements and context. The system then asynchronously invokes the LLM processing pipeline. This non-blocking call allows the UAV to continue to perform its main flight control, environment monitoring, and target detection tasks while the LLM performs complex reasoning, ensuring the multi-tasking parallelism of the system.
[0054] The local knowledge enhancement module pre-builds a local knowledge base containing safety regulations and typical cases, and uses RAG technology to enhance LLM reasoning. Using retrieval enhancement to generate the RAG paradigm, reasoning is completed locally on the drone. When receiving the structured prompt words built by the LLM reasoning submodule, the RAG process is initiated:
[0055] First, in the retrieval stage, the system uses key information in the prompt word (such as dangerous goods labels and categories) as query vectors (usually generated through a lightweight embedding model) to perform similarity searches in a local pre-built vectorized knowledge base. This knowledge base is derived from a local document (such as PDF) containing professional regulations, safety plans, and disposal guidelines. After chunking and vector embedding (using models such as herald / dmeta-embedding-zh), it is stored in a vector database (such as FAISS). The retrieval process aims to find several text fragments in the knowledge base that are most relevant to the current situation semantics.
[0056] Then, in the enhancement phase, the retrieved highly relevant knowledge fragments are dynamically integrated with the original prompt word to form an enhanced prompt word with richer information and more specific context.
[0057] Finally, in the generation phase, this enhanced prompt word is fed into the locally deployed LLM (e.g., Gemma 3:4b). When generating responses, the LLM not only leverages its internal general knowledge but also prioritizes and integrates the injected domain expertise fragments. The RAG mechanism significantly improves the factual accuracy, domain relevance, and contextual adaptability of the LLM's responses, effectively suppressing "model illusions" and ensuring that the generated resolution plan is highly professional and actionable. The entire process is completed at the edge, achieving localized decision-making and low latency.
[0058] The visualization and interaction module, built on Gradio, displays image recognition information and target assessment results in real time. Users send requests to the client device to obtain a real-time preview. The specific implementation steps are as follows:
[0059] After the LLM (via the RAG process) generates a disposal plan, the plan and related metadata (such as inference time and original alarm information) are encapsulated into a JSON message of the LLM response type and transmitted to the ground station via a text / JSON channel. The ground station application receives and parses this message, stores it and associates it with the corresponding preliminary alarm, and updates the status of the alarm in the user interface. The operator interacts with the system through a web-based graphical user interface (GUI, implemented using the Gradio framework). The interface contains several situational awareness panels and a dynamically updated alarm list, showing the processing status of all confirmed dangerous time machines. When the operator selects an alarm in the list, a dedicated area on the interface will render and display detailed disposal recommendations generated by the LLM and formatted (supporting Markdown to HTML conversion for optimized typesetting). The system operates mainly on low-bandwidth text / JSON information streams to ensure communication efficiency. However, to meet the operator's need for visual confirmation in specific situations, the interface provides a request image function. Activating this function triggers the ground station to send a request command to the drone via a dedicated image channel. After receiving the command, the drone extracts the latest visual image frame with detection annotations from the memory buffer, performs JPRG compression (to balance image quality and transmission efficiency), and then transmits it back to the ground station for display in the graphical user interface. This on-demand visualization mechanism provides operators with the ability to obtain intuitive visual evidence when necessary, effectively supplementing the intelligent decision-making process that is mainly based on text.
[0060] The bidirectional network communication module transmits target information and decision results to the ground station. The data transmission module relies on establishing a stable, bidirectional communication link between the UAV (edge computing node equipped with Jetson AGX Xavier) and the ground control station (operator interface and data aggregation point). This communication follows the TCP / IP protocol stack, leveraging its connection-oriented nature to ensure reliable data transmission. The core design employs port multiplexing and traffic isolation strategies: one logical channel (via a specific TCP port, such as 5000) is designated for transmitting low-bandwidth, low-latency control signaling and structured data (encapsulated in JSON format), including target detection metadata, dangerous goods alert identifiers, and subsequent LLM response text. Simultaneously, another independent logical channel (via another TCP port, such as 5001) is reserved for on-demand transmission of high-bandwidth image data streams. This channel separation prevents sudden, high-volume image transmissions from blocking critical, time-sensitive alert and command flows, ensuring real-time responsiveness and quality of service (QoS). To address the inherent instability of wireless communication environments, the communication protocol stack implements connection state management and automatic reconnection mechanisms at the application layer. For long text messages that may exceed the capacity of a single TCP segment (such as the detailed handling plan for LLM), an application-layer segmentation and reassembly strategy is implemented, adding explicit message end delimiters to ensure that the receiving end can accurately recover the original message. Data exchange strictly adheres to the JSON (JavaScript Object Notation) format, ensuring cross-platform and cross-language interoperability and clear data structures. The system extensively utilizes concurrent programming models (multi-threading) to parallelize tasks such as network I / O, data parsing, model inference, image serving, and user interface updates, maximizing the utilization of computing resources and maintaining the system's high throughput and low latency.
[0061] In this embodiment, the drone is equipped with a Jeston AGXXavier micro-embedded edge computing device equipped with a camera and communication hardware modules. The LLM (Gemma3:12b multimodal model) and YOLOv5 model are deployed locally on the client side. This is achieved through the following steps:
[0062] Step 1: First, establish a two-way communication channel between the drone and the ground terminal. This communication link is constructed using the TCP / IP protocol socket nesting method. Physical isolation of the text and image data stream transmission channels is achieved through port splitting. Multi-threading is also established to enable parallel processing of data streams.
[0063] Step 2: The micro-embedded computer edge computing device starts YOLOv5 model inference, collects images through the camera, and performs real-time target detection and framing. The detection results are output in a structured data format, including the detection time, detection label, and confidence triplet;
[0064] Step 3: The label is determined to be a dangerous object by category. The dangerous object categories and the corresponding labels for different categories are specified in advance. For each type of dangerous object, a text solution based on the actual situation is given in the RAG knowledge base of LLM. When the YOLOv5 model determines that the dangerous object label appears continuously (3 times) within a certain period of time (3s) and the confidence level is greater than the given value (0.6), it is determined that the visual perception has detected the presence of a dangerous object in the dynamic environment. The drone immediately communicates with the ground end and sends an alarm message first;
[0065] Step 4: After the alarm is triggered, the alarm information is first sent to the ground terminal, including the time, dangerous object label, and confidence level. Simultaneously, the airborne terminal embeds the dangerous object detection results (including the label and label category information specified in step 3) into a pre-built LLM structured prompt template (e.g., in this regulated environment, {label} of the {label category} class appears, which is a dangerous object. Please follow the preparatory plan for this type of dangerous object and provide a feasible solution based on the current situation.) and activates the local LLM model Gemma3:4b to make inference decisions on the prompt.
[0066] Step 5: When making reasoning decisions, the LLM integrates a local knowledge base built using Retrieval-Augmented Generation (RAG) technology. This specialized knowledge base pre-determines different treatment options for the three major categories of labels. It provides case analysis capabilities for specific mission scenarios, reduces the illusion of LLM reasoning results, and ensures that the decision-making response information is truly feasible and professional. It can respond to specific categories of hazardous materials and generate targeted solutions.
[0067] Step 6: After the micro-embedded computer edge computing device generates an inference response solution through the local LLM, it transmits the solution to the ground end to provide ground users with real-time response recommendations for dangerous objects in dynamic environments.
[0068] Therefore, the present invention adopts the above-mentioned end-side LLM-based UAV dynamic environment perception and response system, integrates YOLOv5 real-time visual detection and deep scene understanding, and integrates the detection results into the end-side lightweight LLM that supports local retrieval enhanced generation (RAG) knowledge base for reasoning in real time, thereby realizing hazard analysis and solution generation.
[0069] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit the same. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that they can still modify or replace the technical solutions of the present invention with equivalents, and these modifications or equivalent replacements cannot cause the modified technical solutions to deviate from the spirit and scope of the technical solutions of the present invention.
Claims
1. The UAV dynamic environment perception and response system based on the end-side LLM is characterized by: include: The visual perception module collects image data in real time through multimodal sensors, recognizes the acquired images using the YOLOv5 model, and outputs recognition information; The warning identification module generates structured labels based on identification information and optimizes label accuracy based on historical information; LLM reasoning and decision module: The pre-deployed LLM model performs reasoning based on labels and scenario descriptions, and provides threat assessment and decision content; The local knowledge enhancement module pre-builds a local knowledge base containing safety specifications and typical cases, uses RAG technology to enhance LLM reasoning, distinguish categories with different labels, and provide corresponding solutions; The two-way network communication module builds a communication link based on the TCP / IP protocol to achieve two-way communication between the drone and the ground. The visualization and interaction module, built on Gradio, displays image recognition information and target assessment results in real time. Users send requests to the end device to obtain a real-time screen preview.
2. The UAV dynamic environment perception and response system based on the end-side LLM according to claim 1 is characterized in that: The visual perception module includes a sensor part and a data preprocessing part. The drone's environmental perception capability is based on the onboard camera of its sensor part and the locally deployed YOLOv5 model, which constitutes a real-time computing process as follows: First, raw video frames are continuously acquired through the camera interface; Then, preprocessing operations are performed on each frame to meet the model input requirements; Next, the preprocessed image frames are fed into the YOLOv5 model for forward inference, which outputs the bounding box coordinates, category predictions, and corresponding confidence scores of potential objects in the image. The model's raw output is then post-processed and parsed to extract human-readable labels and quantified confidence scores. Finally, the above key information is encapsulated into a standardized JSON object, and the system periodically transmits the information to the ground station at a preset frequency; at the same time, the system uses the YOLOv5 model output and the original frame to generate a visual annotation image with superimposed detection boxes and labels, and temporarily stores it in the memory buffer.
3. The UAV dynamic environment perception and response system based on the end-side LLM according to claim 1 is characterized in that: The specific implementation steps of the warning identification module are: First, each detected object label is semantically classified according to a predefined hazardous materials ontology library, which maps specific labels to preset hazardous categories; Then, a multi-level confirmation protocol is set up. When a detected event meets the key dimensions of the multi-level confirmation protocol at the same time, it is confirmed as a potential dangerous symptom that requires escalation. Secondly, through the target status tracking mechanism, the current processing stage of each confirmed dangerous target is recorded; Finally, when a potential dangerous incident passes the multi-level confirmation protocol and the current status allows the initiation of a new process, the system immediately sends a preliminary alert message to the ground station, including the timestamp, dangerous goods label, confidence level at the time of confirmation, and its category.
4. The UAV dynamic environment perception and response system based on the end-side LLM according to claim 3 is characterized in that: The multi-level confirmation protocol consists of three key dimensions: Time window, which requires the target to appear continuously or repeatedly within a specified short period of time; Frequency threshold, which requires that the number of successful detections of a target within the time window reaches a minimum value; The confidence threshold requires that the confidence of each detection that constitutes a valid detection count is higher than a preset lower limit.
5. The UAV dynamic environment perception and response system based on end-side LLM according to claim 3 is characterized in that: The implementation steps of the LLM reasoning and decision module are as follows: After the initial alert is issued, the system immediately initiates the cognitive processing process and interacts with the locally deployed LLM model. The key lies in prompt word engineering: the system dynamically constructs an input prompt based on a preset template. The prompt embeds the key elements of the current situation: the context of the incident, the specific hazardous materials identified, the quantified confidence level, and the tasks to be completed by the LLM. The solution is generated, which includes specific components: risk assessment, treatment steps, and notification targets; The system then calls the LLM process asynchronously.
6. The UAV dynamic environment perception and response system based on the end-side LLM according to claim 5 is characterized in that: The local knowledge enhancement module uses the retrieval enhancement generation (RAG) paradigm to complete reasoning locally on the drone. Upon receiving the structured prompt words constructed by the LLM reasoning and decision module, the RAG process starts: First, in the retrieval phase, the system uses the key information in the prompt word as a query vector and performs similarity search in a pre-built local vectorized knowledge base, aiming to find the text fragments in the knowledge base that are most relevant to the current context. Then, in the enhancement phase, the retrieved highly relevant knowledge fragments are dynamically integrated with the original prompt word to form an enhanced prompt word; Finally, in the generation phase, the enhanced prompt words are input into the locally deployed LLM to generate a response.
7. The UAV dynamic environment perception and response system based on the end-side LLM according to claim 6 is characterized in that: The specific implementation steps of the visualization and interaction module are: After LLM generates a disposal plan, the plan and related metadata are encapsulated into a JSON message of LLM response type and transmitted to the ground station through the text / JSON channel. The ground station application receives and parses the message, stores it and associates it with the corresponding preliminary alarm, and updates the status indicator of the alarm in the user interface. The operator interacts with the system through a web-based graphical user interface, which contains several situational awareness panels and a dynamically updated alarm list, showing the processing status of all confirmed dangerous time machines. When the operator selects an alarm in the list, a dedicated area on the interface will render and display the detailed disposal suggestions generated by LLM. The interface provides a request image function. Activating this function will trigger the ground station to send a request instruction to the drone through the image dedicated channel. After receiving the instruction, the drone extracts the latest visual image frame with detection annotations from the memory buffer, performs JPRG compression, and then transmits it back to the ground station and displays it in the graphical user interface.
8. The UAV dynamic environment perception and response system based on end-side LLM according to claim 7 is characterized in that: The two-way network communication module relies on the communication link between the drone and the ground station, following the TCP / IP protocol stack. Its design adopts port multiplexing and traffic isolation strategies: one logical channel is designated for transmitting low-bandwidth, low-latency control signaling and structured data, including target detection metadata, dangerous goods alert identifiers, and LLM response text; another logical channel is reserved for on-demand transmission of high-bandwidth image data streams; For long text messages that exceed the capacity of a single TCP segment, an application layer segmentation and reassembly strategy is adopted. By adding a message end delimiter, it is ensured that the receiving end can restore the original information, and the data interaction follows the JSON format.
Citation Information
Cited By
Intelligent security collaborative management system based on multi-source perception and language large model
CN120832408A
Unmanned aerial vehicle dynamic environment perception feedback method and system based on large language model
CN121209562A
Unmanned aerial vehicle dynamic environment perception feedback method and system based on large language model
CN121209562B
Unmanned aerial vehicle countering method and system based on inverse non-large model
CN121563021A
Large model cooperative ANIL target identification method and system for camera icing and fogging
CN122090396A