Unmanned aerial vehicle forest search and rescue command method and system based on multi-modal large model
Patent Information
- Application Number
- CN202610848173.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-12
- Publication Date
- 2026-08-18
AI Technical Summary
[0008]本发明实施例提供一种基于多模态大模型的无人机森林搜救指挥方法及系统,旨在解决现有技术在复杂林区环境中存在的感知精度低、搜寻与灭火任务割裂以及缺乏多机智能协同的技术问题,通过将森林搜寻与无人机灭火深度结合,从而提升复杂环境下的应急救援响应速度与作业效率的技术问题
[0019] First, it breaks through the perception bottleneck in complex environments, abandoning the reliance on traditional visual or infrared image recognition methods. By utilizing the comprehensive understanding capabilities of a multimodal large model, it effectively overcomes interference from high-temperature backgrounds and dense smoke obscuring the fire, accurately locating hidden fire points and significantly reducing the probability of missed detections and misjudgments. Second, it breaks down the separation between search and firefighting tasks. Through joint reasoning of a large model, it generates a composite task chain with parallel logic for search and firefighting, completely changing the traditional serial mode and achieving seamless multi-machine collaboration, avoiding missing the best opportunity for firefighting. Third, it endows the system with global cognition and powerful dynamic adaptive capabilities. When facing emergencies, it is no longer limited by static preset routes or highly dependent on manual decision-making. It can assess the status of the global resource pool in real time and automatically replan tasks with optimal logic, greatly improving the flexibility and scientific nature of emergency response.
Smart Images

Figure CN122593404A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent control and emergency search and rescue of unmanned aerial vehicles (UAVs), and in particular to a UAV forest search and rescue command method and system based on a multimodal large model. Background Technology
[0002] Forest environments are characterized by dramatic terrain and dense vegetation, making them highly susceptible to rapid fire spread. Traditional fire response methods rely primarily on manual patrols or fixed-point monitoring, which suffer from slow response times. In recent years, leveraging their high mobility, drones have been widely deployed for aerial search and firefighting operations in forest areas. Existing drone rescue solutions typically integrate visible light gimbals and infrared thermal imagers for inspection missions, utilize target recognition algorithms to locate the fire source, and employ onboard firefighting equipment to extinguish the fire. Some advanced systems have also incorporated static path planning and multi-drone formation technologies.
[0003] However, when faced with complex forest scenarios, existing technologies still face the following significant technical bottlenecks:
[0004] First, sensing capabilities are severely limited in complex physical environments. Fire scenes are often accompanied by dense smoke and complex local airflow, which not only severely obstruct visible light vision, but the high temperature background of the fire scene also causes strong interference to infrared detection equipment. Since existing technologies mostly rely on single visual or infrared image recognition, it is difficult to accurately locate hidden fire points in complex backgrounds, leading to frequent missed detections and misjudgments.
[0005] Secondly, search and firefighting tasks are often separated. Current solutions mostly treat search and reconnaissance as independent processes, generally adopting a sequential model of "reconnaissance aircraft discovering the target before dispatching firefighting aircraft." This fragmented operational process can easily lead to missed opportunities for optimal firefighting in rapidly changing fire situations. Furthermore, existing multi-aircraft coordination is mostly based on statically preset flight paths, and the system lacks the ability to adaptively adjust its flight paths should the fire situation on-site change drastically.
[0006] Finally, there is a lack of comprehensive understanding of complex on-site information. Real search and rescue sites generate a wide variety of information, and most existing systems are unable to perform joint reasoning on cross-dimensional perceived information, making it difficult to form a holistic situational awareness. This results in command and decision-making heavily relying on human intervention and insufficient intelligence.
[0007] In summary, existing technologies lack multi-dimensional information fusion and perception capabilities, making it impossible to achieve accurate perception in complex physical environments; search and firefighting tasks are disconnected from each other, failing to achieve deep integration and parallel execution; and they do not support multi-machine intelligent collaboration and dynamic adaptive adjustment, resulting in low perception accuracy, task execution disconnect, and response lag in complex forest environments. These problems urgently need to be improved. Summary of the Invention
[0008] This invention provides a method and system for commanding forest search and rescue using unmanned aerial vehicles (UAVs) based on a multimodal large model. It aims to solve the technical problems of low perception accuracy, separation of search and firefighting tasks, and lack of multi-UAV intelligent collaboration in complex forest environments. By deeply integrating forest search with UAV firefighting, it improves the emergency rescue response speed and operational efficiency in complex environments.
[0009] A UAV forest search and rescue command method based on a multimodal large model includes:
[0010] S1. Acquire and process image perception data of forest fire, extract multiple candidate hotspot targets from the processed image perception data, input the candidate hotspot targets into a multimodal large model to determine independent ignition points and generate a set of fire point coordinates; perform spatial clustering analysis based on the fire point coordinate set to generate multiple task regions, assign corresponding regional risk levels to the task regions, and output the boundary information of each task region.
[0011] S2. Obtain the capability status information of the UAV cluster, calculate the task demand intensity of each task area and the capability matching score between a certain UAV and the corresponding task area, and input them into the multimodal large model. The multimodal large model allocates UAV nodes according to the minimum survival guarantee principle and generates a composite task chain and tactical flight instructions that include search and firefighting tasks.
[0012] S3. Upon receiving information about changes in fire conditions or new target information, obtain the status change information of the global fire extinguishing resource pool and input it into the multimodal large model. Based on the preset echelon screening rules, determine the target response drone and generate emergency replanning instructions.
[0013] A drone forest search and rescue command system based on a multimodal large model includes:
[0014] The image perception and region segmentation module is used to acquire and process image perception data of forest fires, extract multiple candidate hotspot targets from the processed image perception data, input the candidate hotspot targets into a multimodal large model to determine independent ignition points and generate a set of fire point coordinates; perform spatial clustering analysis based on the fire point coordinate set to generate multiple task regions, assign corresponding regional risk levels to the task regions, and output the boundary information of each task region.
[0015] The cluster resource scheduling and tactical generation module obtains the capability status information of the UAV cluster, calculates the task demand intensity of each task area and the capability matching score between a certain UAV and the corresponding task area, and inputs it into the multimodal large model. The multimodal large model allocates UAV nodes according to the minimum survival guarantee principle and generates a composite task chain containing search and firefighting tasks and tactical flight instructions.
[0016] The status feedback and emergency replanning module, upon receiving information about changes in fire conditions or new target information, acquires the status change information of the global fire-fighting resource pool and inputs it into a multimodal large model. Based on preset echelon screening rules, it determines the target response drone and generates emergency replanning instructions.
[0017] In this invention, a multimodal large-scale model is used as the core reasoning and decision-making engine. First, target detection and adaptive region division are performed based on the multimodal large-scale model. Image perception data of forest fires are input into the multimodal large-scale model, which uses its comprehensive understanding capabilities to identify independent ignition points and eliminate false fire sources, achieving high-precision fire point location. Simultaneously, based on the spatial distribution density of the fire point coordinate set and fire evolution prediction, the large-scale model adaptively divides the global area into multiple task areas, such as high-risk attack zones, low-risk control zones, and global monitoring zones, and outputs a task area division scheme. Second, resource coordination and tactical generation are performed based on the global cognition of the large-scale model. The capability status information of the UAV swarm and the task area division scheme are input into the multimodal large-scale model as prompts. Through joint reasoning, a global resource allocation matrix for the heterogeneous UAV swarm is generated end-to-end. The large-scale model executes capability matching and minimum survival guarantee principles to ensure reasonable regional resource allocation. At the same time, it generates a composite task chain containing parallel search and firefighting logic for each UAV node and outputs specific tactical flight instructions including target allocation. Subsequently, emergency task replanning is performed based on dynamic evaluation of the large-scale model. When receiving sudden situations such as sudden changes in fire intensity or the addition of new targets, the multimodal big model takes into account the status changes of the global fire extinguishing resource pool in real time. Based on the built-in three-tier evaluation and reasoning logic with progressively decreasing priorities, it performs multi-objective optimization reasoning, dynamically locks the unique optimal response drone, generates an emergency redirection command, and forcibly interrupts its original process to achieve adaptive task redistribution and flight path adjustment.
[0018] In this invention, the above technical solution achieves the following significant beneficial effects:
[0019] First, it breaks through the perception bottleneck in complex environments, abandoning the reliance on traditional visual or infrared image recognition methods. By utilizing the comprehensive understanding capabilities of a multimodal large model, it effectively overcomes interference from high-temperature backgrounds and dense smoke obscuring the fire, accurately locating hidden fire points and significantly reducing the probability of missed detections and misjudgments. Second, it breaks down the separation between search and firefighting tasks. Through joint reasoning of a large model, it generates a composite task chain with parallel logic for search and firefighting, completely changing the traditional serial mode and achieving seamless multi-machine collaboration, avoiding missing the best opportunity for firefighting. Third, it endows the system with global cognition and powerful dynamic adaptive capabilities. When facing emergencies, it is no longer limited by static preset routes or highly dependent on manual decision-making. It can assess the status of the global resource pool in real time and automatically replan tasks with optimal logic, greatly improving the flexibility and scientific nature of emergency response. Attached Figure Description
[0020] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments of the present invention will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0021] Figure 1 This is a flowchart of a UAV forest search and rescue command method based on a multimodal large model in one embodiment of the present invention;
[0022] Figure 2 This is a structural diagram of a UAV forest search and rescue command system based on a multimodal large model in one embodiment of the present invention;
[0023] Figure 3 This is a simulation image of the original disaster situation in a forest scene according to one embodiment of the present invention;
[0024] Figure 4 This is a fire point identification and region division map generated in step S1 in one embodiment of the present invention;
[0025] Figure 5 This is an embodiment of the present invention based on the UAV task allocation and action command planning generated in step S2;
[0026] Figure 6 This is a diagram of the main front-end visual interface of the UAV forest search and rescue command system in one embodiment of the present invention;
[0027] Figure 7 This is a diagram of the front-end visualization result display interface of the UAV forest search and rescue command system in one embodiment of the present invention. Detailed Implementation
[0028] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0029] The UAV forest search and rescue command method based on a multimodal large model provided in this invention is applied to a UAV forest search and rescue command system based on a multimodal large model.
[0030] In one embodiment, such as Figure 1 As shown, a method for commanding UAV forest search and rescue based on a multimodal large model is provided, including the following steps:
[0031] S1. Acquire and process image perception data of forest fire sites, extract multiple candidate hotspot targets from the processed image perception data, input the candidate hotspot targets into a multimodal large model to determine independent ignition points and generate a set of fire point coordinates; perform spatial clustering analysis based on the fire point coordinate set to generate multiple task regions, assign corresponding regional risk levels to the task regions, and output the boundary information of each task region.
[0032] Understandably, in step S1, image perception data of the forest fire area is first acquired, and the image perception data is input into a multimodal large model to identify fire point coordinates and generate a global disaster grid and task area division results. Unlike traditional methods that rely solely on single-frame visual threshold recognition, the multimodal large model of this invention can jointly understand visible light texture, infrared high-temperature regions, smoke morphology, and external rule cues, thereby enabling semantic-level discrimination of suspected fire sources.
[0033] In one embodiment, step S1 further includes the following sub-steps:
[0034] S101. Acquire image perception data of the target forest fire area. To improve the accuracy of subsequent inference, before inputting the multimodal large model, the image perception data is registered, corrected, and georeferenced to ensure that the same scene target under different modalities is correctly associated in a unified coordinate system. The image perception data includes visible light images and corresponding geographic coordinate information.
[0035] In one embodiment, such as Figure 3 and Figure 4 As shown, step S1011 further includes the following sub-steps:
[0036] The target forest fire area is discretized into multiple grid cells to form a global disaster grid. For each grid cell Maintain the corresponding disaster feature vector The disaster feature vector It includes at least the characteristics of fire source intensity, smoke coverage, vegetation and fuel, and the trend of spread to neighboring areas.
[0037] S1012, Obtain the... Comprehensive risk value of each grid cell Its expression is:
[0038] (1)
[0039] in, Indicates the first The thermal intensity of each grid cell; This indicates the degree of smoke obscuration for that grid cell; Indicates the first Combustible load strength of each grid cell; Indicates the first The fire spread trend within the neighborhood of each grid cell; to The weighting coefficients are preset or adjusted online based on historical fire scene samples and empirical rules.
[0040] S102. Extract multiple candidate hotspot targets from the processed image perception data, and input the local image patch, temperature features and edge contour features corresponding to each candidate hotspot target into the multimodal large model to output the confidence of each candidate hotspot target as a real fire point; determine the candidate hotspots with a value greater than the preset fire point threshold as independent ignition points, obtain the fire point coordinate set; and remove the candidate hotspot targets corresponding to false fire sources.
[0041] In one embodiment, step S102 further includes the following sub-steps:
[0042] S1021, Obtain the... Fire point confidence of candidate hotspot targets Its expression is:
[0043] (2)
[0044] in, This represents the confidence score of flame morphology obtained based on visible light images; This represents the confidence score of the heat source obtained based on infrared thermal imaging; This represents the score of the multimodal large model in judging the semantic consistency of the scene; to As weighting coefficients, in If the value exceeds the preset ignition threshold, the candidate hotspot target can be identified as an independent ignition point.
[0045] S1022. After identifying and determining the independent ignition points, based on the UAV pose information and geographic coordinate information, the fire point locations in the image space are mapped to the global geographic coordinate system to obtain the fire point coordinate set. For any fire point Its global coordinates can be represented as ,in and Used to represent planar geographic coordinates.
[0046] Based on the spatial distribution of fire point coordinates and the fire evolution prediction results, the global area is adaptively divided to generate multiple task areas, and the boundary information of each task area is output. This is specifically achieved through step S103.
[0047] S103, Set of Fire Point Coordinates Spatial clustering is performed to form one or more fire point clusters; the global area is adaptively divided by combining the local density of each fire point cluster and terrain barrier factors to generate multiple task areas, assign corresponding regional risk levels to the task areas, and output the boundary information of each task area.
[0048] Among them, the Regional risk value of individual fire clusters The expression is:
[0049] (3)
[0050] in, Indicates the spatial density of the fire point cluster; This represents the observation uncertainty comprised of smoke obstruction, field-of-view occlusion, and identification uncertainty. This represents a processing difficulty factor related to terrain accessibility and communication coverage. to Weighting coefficients, regional risk values The higher the value, the more priority the corresponding fire cluster needs to be given to firefighting resources.
[0051] When the regional risk value of a certain fire cluster When the risk value of a fire cluster exceeds the first risk threshold preset based on the actual situation, it can be classified as a high-risk attack zone; when the regional risk value of a fire cluster is within the preset low-risk range, it can be classified as a low-risk control zone; the remaining areas other than the high-risk attack zone and the low-risk control zone can be uniformly classified as the whole-area monitoring zone.
[0052] Understandably, within the comprehensive monitoring area, priority can be given to deploying drone nodes with long-endurance patrol capabilities to achieve continuous scanning and risk warning of non-key areas. In other words, high-risk attack zones emphasize rapid firefighting and high-frequency feedback, low-risk control zones emphasize perimeter suppression and resurgence prevention, while comprehensive monitoring zones emphasize large-scale patrols and the discovery of new targets, thereby achieving functional stratification among different task areas.
[0053] S2. Obtain the capability status information of the UAV cluster, calculate the task demand intensity of each task area and the capability matching score between a certain UAV and the corresponding task area, and input them into the multimodal large model. The multimodal large model allocates UAV nodes according to the minimum survival guarantee principle and generates a composite task chain containing search and firefighting tasks and tactical flight instructions.
[0054] Understandably, after completing step S1, the system proceeds to step S2, which involves acquiring the capability status information of the UAV cluster and inputting this capability status information, along with the task area division results, into a multimodal large model to generate global resource allocation results corresponding to the heterogeneous UAV cluster. This heterogeneous UAV cluster can include reconnaissance UAVs, firefighting UAVs, supply UAVs, and hybrid UAVs. Different types of UAVs may differ in sensor capabilities, payload capacity, endurance, and speed. The capability status information includes at least one or more of the following: current remaining battery power, remaining payload, payload type, maximum endurance, maximum flight speed, onboard sensor type, current task status, and current location coordinates. The system can organize this capability status information into a structured resource table and input it into the multimodal large model along with the risk level of the task area, boundary information, and the number of target fire points.
[0055] In one embodiment, such as Figure 5 As shown, step S2 further includes the following sub-steps:
[0056] S201. Obtain the capability status information of the UAV cluster, organize the capability status information into a structured resource table, record the task demand intensity of each task area and the capability matching score between a certain UAV and the corresponding task area; input the task demand intensity and capability matching score into the multimodal large model, and the multimodal large model allocates UAV nodes to each task area according to the minimum survival guarantee principle.
[0057] Calculate the task demand intensity for each task region , No. Task demand intensity in each task area The expression is:
[0058] (4)
[0059] in, Indicates the area of the task region; Indicates the number of active fire points within the mission area; This represents the regional risk value obtained from equation (3); This indicates the complexity of handling issues resulting from the combined effects of communication blind spots and supply distances within the region; to Weighting coefficient; task demand intensity The larger the value, the more reconnaissance and firefighting resources are required for that area.
[0060] Calculate the first The drone and the first Capability matching score between task areas The expression for the ability matching score is:
[0061] (5)
[0062] in, This indicates the drone's sensor compatibility. This indicates the payload adaptability of the drone; This indicates the remaining battery power of the drone. This indicates the current position of the drone up to the [number]th [number]. The distance between the centers of each task area; This represents the combined score of communication quality and flight control stability. to This represents the weighting coefficient. A higher capability matching score indicates that the drone is more suitable to be assigned to the corresponding mission area.
[0063] The minimum survival guarantee principle is as follows: at any given time, each high-risk attack zone must be equipped with at least a combination of drones that meet basic reconnaissance and firefighting capabilities, requiring allocation to the [number missing] [unit missing]. A collection of drones in each mission area The following constraints must be satisfied:
[0064] (6)
[0065] in, Indicates the first The drone and the first The ability matching score between each task area Indicates the first The intensity of task requirements in each task area.
[0066] S202. Utilize a multimodal large model to generate a composite task chain for each UAV node, including search and firefighting tasks, and combine the real-time payload and capability status of each UAV node to output corresponding tactical flight commands.
[0067] Understandably, the tactical flight commands include at least target allocation information, payload consumption strategies, and node switching mechanisms, and may also include estimated arrival time and transmission frequency. For reconnaissance UAVs, the composite mission chain may include four stages: "patrol—confirmation—marking—re-check." Specifically, in the "patrol" stage, the UAV scans its assigned mission area according to the planned flight path; in the "confirmation" stage, it performs low-altitude verification of hotspot areas; in the "marking" stage, it reports the confirmed fire point coordinates, smoke range, and danger boundaries to the system; and in the "re-check" stage, it re-inspects the already treated areas to identify signs of reignition or new fire points.
[0068] For firefighting drones, the composite mission chain can include four stages: "approach—suppression—assessment—transfer / replacement". Specifically, in the "approach" stage, the drone selects a suitable approach path based on wind direction, terrain, and safe corridors; in the "suppression" stage, firefighting payloads are released sequentially according to the target priority generated by the multimodal large model; in the "assessment" stage, onboard sensors provide real-time feedback on the firefighting effect; in the "transfer / replacement" stage, if the current fire point has been suppressed and there is still payload capacity, the drone flies to the next fire point coordinates; if the current payload is insufficient or the drone's status is close to a threshold, a replacement process is triggered.
[0069] In some embodiments, for firefighting drones, the multimodal large model generates payload consumption instructions according to a strategy prioritizing the first available payload. Specifically, when the drone is simultaneously carrying water-based fire extinguishing bombs or dry powder payloads, the system can prioritize the first available payload that is in a ready state and has the lowest release cost, based on preset payloads, the current fire type, and the risk level of the mission area. In practice, the first available payload refers to the highest priority executable payload unit provided that the current flight attitude, remaining capacity, and effective range all meet the requirements.
[0070] S3. Upon receiving information about changes in fire conditions or new target information, obtain the status change information of the global fire extinguishing resource pool and input it into the multimodal large model. Based on the preset echelon screening rules, determine the target response drone and generate emergency replanning instructions.
[0071] Understandably, after completing routine resource allocation and tactical action generation, the system can continuously monitor fire situation changes or new target information. When it receives information about sudden changes in fire situation, fire line expansion, signs of reignition, or new hotspot targets, it proceeds to step S3, which involves acquiring the status change information of the global fire extinguishing resource pool, inputting the status change information into a multimodal large model, determining the target response drone based on preset echelon screening rules, and generating emergency replanning instructions.
[0072] In one embodiment, step S3 further includes the following sub-steps:
[0073] S301. Obtain the status change information of the global fire suppression resource pool. The global fire suppression resource pool can be represented as a dynamic resource set. Each resource unit It should include at least the following status fields: current location, region, task status, remaining power, remaining payload, and estimated return time; the status of the global firefighting resource pool should be refreshed in each scheduling cycle so that the multimodal large model uses the latest resource view when performing emergency replanning.
[0074] S302. The target responding UAV is determined based on preset echelon selection rules, including T1 echelon selection, T2 echelon selection, and T3 echelon selection. Understandably, T1 echelon selection targets UAVs in idle or search states; T2 echelon selection targets UAVs performing firefighting or return-to-base missions; and T3 echelon selection targets UAVs whose payloads have been exhausted and are in the process of returning to base for reloading. Through this three-level selection logic with decreasing priority, the system can ensure response speed while also controlling disturbances to existing tasks.
[0075] S3021. In the T1 echelon screening phase, among the drones in idle or search states, drones with both remaining payload and remaining battery power greater than a first preset state threshold are selected, and the drone with the highest response score is determined as a candidate response drone; the response score in the T1 phase. The expression is:
[0076] (7)
[0077] in, Indicates the first The distance from the drone to the sudden target; Indicates the remaining battery level score; Indicates the remaining load score; Indicates the communication quality score; to These are the weighting coefficients.
[0078] S3022. When the T1 echelon screening result is empty, proceed to the T2 echelon screening stage. Among the drones currently performing firefighting or return missions, screen drones whose remaining payload and remaining battery power are both greater than the second preset state threshold, and calculate the mission switching time cost for each drone. Determine the drone with the lowest time cost as the candidate response drone; the mission switching time cost of the u-th drone... The expression is:
[0079] (8)
[0080] in, Indicates the time required for the ship to veer off course and turn toward the unexpected target; This indicates the time required to re-enter a stable operating posture; Indicates the time lost due to interrupting the current task; to These are the weighting coefficients.
[0081] S3023. If the T2 echelon selection results are empty, proceed to the T3 echelon selection phase. Among the drones whose payloads have been exhausted and are in the process of returning to reload, select the drone closest to the supply base as the candidate response drone. The reason for using the criterion of closest to the supply base in the T3 phase is that although this type of drone does not yet have direct firefighting capabilities, it is most likely to be the first to re-enter an effective operational state after completing rapid resupply, thus undertaking subsequent emergency support tasks in extreme resource-constrained scenarios.
[0082] S303. The selected candidate response UAVs are identified as target response UAVs, and an emergency replanning instruction containing the coordinates of the sudden fire point is generated by the multimodal large model. The emergency replanning instruction is sent to the target response UAV to interrupt its current mission process and perform mission reassignment and flight path adjustment.
[0083] Understandably, in some embodiments, emergency replanning not only applies to the target response drone itself, but also coordinates adjustments to other drones with which it has a collaborative relationship. Specifically, when a firefighting drone is reassigned to perform an emergency response mission, the resource gap in its original mission area can be calculated simultaneously, and a replacement drone can be selected from backup nodes, or the fire priority of the remaining drones in the same area can be rearranged. This prevents the local defense line from becoming unbalanced due to the reassignment of a single node.
[0084] To improve feasibility, in some embodiments, the output of the multimodal large model can adopt a structured command template. For example, the model output is encapsulated as a JSON message and parsed by the ground control station into flight control parameters that can be directly issued. This structured command template may include at least fields such as area number, node number, mission type, fire sequence, and waypoint set to facilitate execution by the backend control program.
[0085] In one specific embodiment, assuming multiple fires occur in a forest area, the system deploys 10 heterogeneous UAVs, including 3 reconnaissance UAVs, 5 firefighting UAVs, and 2 hybrid UAVs. Each UAV establishes a communication connection with the multimodal large model inference service through a ground station and periodically reports its pose, remaining battery power, payload status, and mission execution status.
[0086] In this example scenario, the reconnaissance drone first conducts a joint visible light and infrared reconnaissance of the target forest area. Based on the collected image perception data, the system identifies nine independent ignition points and eliminates two false hotspots caused by solar reflection. Subsequently, based on the spatial distribution of these nine ignition points and local terrain information, the system generates a global disaster grid, dividing the fire area into two high-risk attack zones, one low-risk control zone, and one comprehensive monitoring zone.
[0087] In step S2, the system inputs the capability status information of the 10 UAVs along with the aforementioned task area division results into the multimodal large model. Based on the task demand intensity of the high-risk attack areas, the multimodal large model prioritizes allocating 2 firefighting UAVs and 1 reconnaissance UAV to each of the two high-risk attack areas, and allocates 1 firefighting UAV and 1 reconnaissance UAV to the low-risk control area. The remaining UAV nodes are configured as global monitoring and backup standby forces.
[0088] When a new hotspot appears outside one of the high-risk attack zones, the system detects the change in fire intensity and updates the global firefighting resource pool status. After screening by the T1 echelon, the system identifies a composite UAV currently patrolling the monitoring area that meets the first preset threshold and has the highest overall score, selecting it as the target response UAV. Subsequently, the multimodal large model generates an emergency replanning command, instructing the UAV to immediately change course and fly to the new hotspot, while simultaneously transferring its original monitoring and patrol mission to another standby reconnaissance UAV. Thus, this embodiment of the invention can achieve rapid response to sudden targets without interrupting overall operations.
[0089] In another example, if no idle or search-state UAVs meet the conditions in phase T1, the system proceeds to phase T2 for selection. For instance, a firefighting UAV performing a low-risk control zone edge suppression mission may be in mission execution mode, but its remaining battery power and remaining payload are both higher than the second preset state threshold, and the mission switching time cost calculated according to equation (8) is the lowest. Therefore, the system can identify it as the target response UAV. At the same time, the system automatically transfers its incomplete edge fire point list to another firefighting UAV or subsequent replacement node in the same area to continue execution, so as to reduce the impact of the original mission interruption.
[0090] As can be seen from the above implementation process, this embodiment is not a simple mechanical superposition of fire point identification, area division, resource allocation and emergency response, but rather a unified modeling and joint reasoning of image perception data, rule constraints and area status through a multimodal large model, giving the system stronger global cognitive ability and dynamic adaptive ability.
[0091] It should be noted that the risk threshold, state threshold, and weighting coefficients involved in the foregoing embodiments can all be preset or dynamically adjusted according to different forest environments, aircraft configurations, and task requirements. In other words, this invention does not limit the specific values of the parameters, and those skilled in the art can adaptively set the above parameters without departing from the technical concept of this invention.
[0092] Furthermore, in some optional embodiments, the multimodal large model can also be used in conjunction with an external knowledge base or rule base to enhance its understanding of historical fire experience, terrain no-fly zones, and payload applicability. The external knowledge base or rule base can be input into the multimodal large model as prompts, or the scheduling module can perform rule verification before and after model inference, thereby further improving the stability and security of command and decision-making results.
[0093] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0094] In one embodiment, such as Figure 2 , Figure 6 and Figure 7 As shown, a UAV forest search and rescue command system based on a multimodal large model is provided, including:
[0095] The image perception and region segmentation module 100 is used to acquire and process image perception data of forest fires, extract multiple candidate hotspot targets from the processed image perception data, input the candidate hotspot targets into a multimodal large model to determine independent ignition points and generate a set of fire point coordinates; perform spatial clustering analysis based on the fire point coordinate set to generate multiple task regions, assign corresponding regional risk levels to the task regions, and output the boundary information of each task region.
[0096] The cluster resource scheduling and tactical generation module 200 acquires the capability status information of the UAV cluster, calculates the task demand intensity of each task area and the capability matching score between a certain UAV and the corresponding task area, and inputs it into the multimodal large model. The multimodal large model allocates UAV nodes according to the minimum survival guarantee principle and generates a composite task chain containing search and firefighting tasks and tactical flight instructions.
[0097] The Status Feedback and Emergency Replanning Module 300, upon receiving information about changes in fire conditions or new target information, acquires the status change information of the global fire-fighting resource pool and inputs it into the multimodal large model. Based on preset echelon screening rules, it determines the target response drone and generates emergency replanning instructions.
[0098] Understandably, the system uses a multimodal large model as its core reasoning and decision engine, and works in conjunction with image sensing terminals, UAV swarm control terminals, task scheduling terminals, and communication links to complete tasks such as fire point identification, task area division, heterogeneous UAV resource allocation, and emergency replanning in the event of a sudden situation in a forest fire environment.
[0099] Specific limitations regarding the UAV forest search and rescue command system based on a multimodal large model can be found in the limitations of the UAV forest search and rescue command method based on a multimodal large model mentioned above, and will not be repeated here. Each module in the aforementioned UAV forest search and rescue command system based on a multimodal large model can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device in hardware form, or stored in the memory of a computer device in software form, so that the processor can call and execute the corresponding operations of each module.
[0100] In one embodiment, a computer device is provided, which may be a server. The computer device includes a processor, memory, a network interface, and a database connected via a system bus. The processor provides computing and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system, computer programs, and a database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The database is used for UAV forest search and rescue command based on a multimodal large model. The network interface is used for communication with external terminals via a network connection. When the computer program is executed by the processor, it implements a UAV forest search and rescue command method based on a multimodal large model.
[0101] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the UAV forest search and rescue command method based on a multimodal large model as described in the above embodiment; to avoid repetition, this will not be repeated here. Alternatively, when the processor executes the computer program, it implements the functions of each module / unit in the UAV forest search and rescue command system based on a multimodal large model embodiment, for example... Figure 2 The functions of the UAV forest search and rescue command based on the multimodal large model shown are not described again here to avoid repetition.
[0102] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When executed by a processor, this computer program implements the UAV forest search and rescue command method based on a multimodal large model as described in the above embodiment. To avoid repetition, further details are omitted here. Alternatively, when executed by a processor, the computer program implements the functions of each module / unit in the above embodiment of the UAV forest search and rescue command system based on a multimodal large model, for example... Figure 2 The functions of the UAV forest search and rescue command system based on a multimodal large model, as shown, will not be repeated here to avoid duplication.
[0103] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. This computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0104] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the system can be divided into different functional units or modules to complete all or part of the functions described above.
[0105] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.
Claims
1. A method for commanding unmanned aerial vehicle (UAV) forest search and rescue based on a multimodal large model, characterized in that, include: S1. Acquire and process image perception data of forest fire, extract multiple candidate hotspot targets from the processed image perception data, input the candidate hotspot targets into a multimodal large model to determine independent ignition points and generate a set of fire point coordinates; perform spatial clustering analysis based on the fire point coordinate set to generate multiple task regions, assign corresponding regional risk levels to the task regions, and output the boundary information of each task region. S2. Obtain the capability status information of the UAV cluster, calculate the task demand intensity of each task area and the capability matching score between a certain UAV and the corresponding task area, and input them into the multimodal large model. The multimodal large model allocates UAV nodes according to the minimum survival guarantee principle and generates a composite task chain and tactical flight instructions that include search and firefighting tasks. S3. Upon receiving information about changes in fire conditions or new target information, obtain the status change information of the global fire extinguishing resource pool and input it into the multimodal large model. Based on the preset echelon screening rules, determine the target response drone and generate emergency replanning instructions.
2. The UAV forest search and rescue command method based on a multimodal large model according to claim 1, characterized in that, Step S1 also includes the following sub-steps: S101. Acquire image perception data of the target forest fire area, and perform registration, correction and georeference mapping processing on the image perception data so that the same scene target under different modalities can be correctly associated under a unified coordinate system. S102. Extract multiple candidate hotspot targets from the processed image perception data, and input the local image patch, temperature features and edge contour features corresponding to each candidate hotspot target into the multimodal large model to output the confidence of each candidate hotspot target as a real fire point; determine the candidate hotspots with a value greater than the preset fire point threshold as independent ignition points, obtain the fire point coordinate set; and remove the candidate hotspot targets corresponding to false fire sources. S103. Perform spatial clustering on the set of fire point coordinates to form one or more fire point clusters; combine the local density of each fire point cluster with terrain barrier factors to adaptively divide the global area, generate multiple task areas, assign corresponding regional risk levels to the task areas, and output the boundary information of each task area.
3. The UAV forest search and rescue command method based on a multimodal large model according to claim 2, characterized in that, Step S2 also includes the following sub-steps: S201. Obtain the capability status information of the UAV cluster, organize the capability status information into a structured resource table, and record the task demand intensity of each task area and the capability matching score between a certain UAV and the corresponding task area. The task requirement intensity and capability matching score are input into the multimodal large model, and the multimodal large model allocates UAV nodes to each task area according to the minimum survival guarantee principle. S202. Utilize a multimodal large model to generate a composite task chain for each UAV node, including search and firefighting tasks, and combine the real-time payload and capability status of each UAV node to output corresponding tactical flight commands.
4. The UAV forest search and rescue command method based on a multimodal large model according to claim 1, characterized in that, Step S3 also includes the following sub-steps: S301. Obtain the status change information of the global fire suppression resource pool. The global fire suppression resource pool can be represented as a dynamic resource set. Each resource unit It should include at least the following status fields: current location, region, task status, remaining power, remaining payload, and estimated return time; the status of the global firefighting resource pool should be refreshed in each scheduling cycle so that the multimodal large model uses the latest resource view when performing emergency replanning; S302. Determine the target response UAV based on the preset echelon selection rules, which include T1 echelon selection, T2 echelon selection and T3 echelon selection; S303. The selected candidate response UAVs are identified as target response UAVs, and an emergency replanning instruction containing the coordinates of the sudden fire point is generated by the multimodal large model. The emergency replanning instruction is sent to the target response UAV to interrupt its current mission process and perform mission reassignment and flight path adjustment.
5. The UAV forest search and rescue command method based on a multimodal large model according to claim 4, characterized in that, Step S302 also includes: S3021. In the T1 echelon screening stage, among the drones in the idle or search state, drones with both remaining payload and remaining power greater than the first preset state threshold are selected, and the drone with the highest response score is determined as the candidate response drone. S3022. When the T1 echelon screening result is empty, enter the T2 echelon screening stage. Among the drones that are performing firefighting or return missions, select drones whose remaining payload and remaining power are both greater than the second preset state threshold, and calculate the mission switching time cost for each drone. Determine the drone with the lowest time cost as the candidate response drone. S3023. When the T2 echelon screening results are empty, proceed to the T3 echelon screening stage. Among the drones whose payloads have been exhausted and are in the process of returning to reload, identify the drone closest to the supply base as the candidate response drone.
6. A UAV forest search and rescue command system based on a multimodal large model, characterized in that, The method for commanding unmanned aerial vehicle (UAV) forest search and rescue based on a multimodal large model, as described in any one of claims 1-5, includes: The image perception and region segmentation module is used to acquire and process image perception data of forest fires, extract multiple candidate hotspot targets from the processed image perception data, input the candidate hotspot targets into a multimodal large model to determine independent ignition points and generate a set of fire point coordinates; perform spatial clustering analysis based on the fire point coordinate set to generate multiple task regions, assign corresponding regional risk levels to the task regions, and output the boundary information of each task region. The cluster resource scheduling and tactical generation module obtains the capability status information of the UAV cluster, calculates the task demand intensity of each task area and the capability matching score between a certain UAV and the corresponding task area, and inputs it into the multimodal large model. The multimodal large model allocates UAV nodes according to the minimum survival guarantee principle and generates a composite task chain containing search and firefighting tasks and tactical flight instructions. The status feedback and emergency replanning module, upon receiving information about changes in fire conditions or new target information, acquires the status change information of the global fire-fighting resource pool and inputs it into a multimodal large model. Based on preset echelon screening rules, it determines the target response drone and generates emergency replanning instructions.