Unmanned aerial vehicle expressway intelligent inspection method and system based on patrol requirements
By using large language models and dynamic behavior strategy synthesis technology, an intelligent framework for the drone inspection system is constructed, which solves the problem that the existing system cannot flexibly respond to patrol needs, realizes autonomous generation and intelligent execution, and improves the automation and emergency response capabilities of inspections.
Patent Information
- Application Number
- CN202511180468.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-22
- Publication Date
- 2025-09-23
- Estimated Expiration
- 2045-08-22
AI Technical Summary
Existing drone inspection systems are unable to flexibly respond to real-time, diverse, and even ambiguous patrol needs. They lack an intelligent integrated framework for mission planning, event detection, and emergency response, and are unable to dynamically generate autonomous inspection strategies.
It adopts multi-level reasoning and dynamic behavior strategy synthesis technology based on a large language model, and realizes the autonomous generation and intelligent execution of complex and fuzzy inspection tasks through a complete cognitive and action framework from semantic intent understanding, behavioral decision-making to closed-loop response. This includes obtaining natural language instructions, semantic parsing and intent decomposition, real-time data fusion, adaptive strategy generation, multimodal data analysis and closed-loop response.
It has significantly improved the automation level, response speed and decision-making intelligence of highway inspections, and can independently generate and execute adaptive inspection strategies, improving the ability to detect abnormal events and the accuracy of emergency response.
Smart Images

Figure CN120690202A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of drone technology, intelligent transportation systems and artificial intelligence, and in particular to a drone highway intelligent inspection method and system based on patrol needs. Background Art
[0002] As the main artery of national transportation, the safety and smooth operation of highways are of paramount importance. Drone inspections, an emerging alternative to traditional manual inspections, are increasingly being used for remote, dynamic monitoring of road conditions, traffic incidents, and ancillary facilities. This is crucial for ensuring road safety and improving traffic management efficiency.
[0003] However, existing drone inspection systems often rely on pre-set fixed routes or simple triggering rules, making them inflexible in responding to real-time, diverse, and even ambiguous patrol needs. Furthermore, mission planning, event detection, and emergency response are often fragmented, lacking an intelligent framework that can dynamically and integratedly generate and schedule inspection strategies, detection algorithms, and response actions based on inspection intent. Summary of the Invention
[0004] To solve the above problems, the present invention provides a method and system for intelligent highway inspection using drones based on patrol needs. It adopts a technical path that combines multi-level reasoning based on a large language model with dynamic behavior strategy synthesis. By constructing a complete cognitive and action framework from semantic intent understanding, behavioral decision-making to closed-loop response, it can realize the autonomous generation and intelligent execution of complex and fuzzy inspection tasks, significantly improving the automation level, response speed and decision-making intelligence of highway inspections.
[0005] The above objectives can be achieved through the following solutions: A method for intelligent highway inspection using drones based on patrol needs includes obtaining natural language instructions for patrol needs, performing semantic parsing and intent decomposition on the natural language instructions, and generating a structured task vector; constructing and generating an adaptive inspection strategy based on the structured task vector and integrating real-time traffic and environmental perception data for collaborative planning; distributing the adaptive inspection strategy to drones for autonomous inspection, performing real-time analysis on the multimodal data obtained by the drones, identifying and generating classified event alarms; and responding to the classified event alarms, matching and triggering corresponding disposal methods from a preset multi-level response plan library, and generating a closed-loop inspection task report.
[0006] Optionally, generating a structured task vector includes: using a large language model to perform multi-path reasoning on the natural language instructions to generate candidate task hypotheses; for the candidate task hypotheses, obtaining objective data for verifying the candidate task hypotheses from an external real-time data source to obtain external verification evidence; based on the external verification evidence, performing posterior probability calculation and weighted fusion on the candidate task hypotheses to generate a structured task vector.
[0007] Optionally, the construction and generation of an adaptive patrol strategy includes: constructing and generating a probabilistic task graph based on the patrol targets and probability distribution in the structured task vector; performing real-time strategy solving based on the probabilistic task graph, the real-time traffic and the environmental perception data to generate a timing behavior strategy; and performing path instantiation on the probabilistic task graph based on the timing behavior strategy to construct and generate an adaptive patrol strategy.
[0008] Optionally, the method further includes: performing path planning deduction on the probabilistic task graph under the candidate task hypothesis to generate an expected task completion index; performing statistical analysis on the expected task completion index to calculate and generate a graph confidence index of the probabilistic task graph; if the graph confidence index is lower than a preset robustness threshold, using the candidate task hypothesis as a new constraint to iteratively optimize the node transfer probability of the probabilistic task graph.
[0009] Optionally, the identification and generation of classified event alarms include: obtaining multimodal data under historical normal inspection conditions and training a data generator model; performing reverse mapping and reconstructing on the multimodal data obtained in real time through the data generator model, and calculating and generating optimal reconstructed data; performing quantitative analysis on the reconstruction error between the multimodal data and the optimal reconstructed data, and classifying events in areas where the reconstruction error exceeds a preset dynamic threshold, and identifying and generating classified event alarms.
[0010] Optionally, the calculation and generation of optimal reconstruction data includes: defining a reconstruction loss function based on the latent space of the data generator model and utilizing the difference between multimodal data and synthetic data output by the data generator model; iteratively exploring the latent space based on the gradient of the reconstruction loss function and coupling a random noise term to generate latent space state samples; performing statistical moment estimation on the latent space state samples, calculating and determining a first latent space vector, and inputting the first latent space vector into the data generator model for forward propagation to generate optimal reconstruction data.
[0011] Optionally, generating a closed-loop inspection task report includes: fusing the classified event alarm with the structured task vector to construct a comprehensive decision context; inputting the comprehensive decision context into a large language model for multi-objective reasoning to generate a dynamic response strategy; executing a disposal method according to the dynamic response strategy, and recording the disposal process to generate a closed-loop inspection task report.
[0012] Optionally, the method also includes: projecting the temporal behavior strategy into the structured task vector, calculating and generating a strategy-intention coverage matrix; performing information entropy analysis on the strategy-intention coverage matrix, quantifying and generating a cognitive coordination index; if the cognitive coordination index is lower than a preset coordination threshold, using the candidate task hypothesis as a penalty item to generate a cognitive bias feedback signal.
[0013] Optionally, the method further includes: using the cognitive bias feedback signal to dynamically reshape the reward function of the strategy generation network model to generate an updated reward function; using the updated reward function to perform online training on the strategy generation network model to generate an updated strategy generation network model.
[0014] Based on the same inventive concept, the present invention also provides a drone intelligent highway inspection system based on patrol needs, the system including: a task semantic parsing module, used to obtain natural language instructions of patrol needs, perform semantic parsing and intent decomposition on the natural language instructions, and generate a structured task vector; an adaptive strategy generation module, used to build and generate an adaptive inspection strategy based on the structured task vector and the integration of real-time traffic and environmental perception data for collaborative planning; an autonomous inspection and event recognition module, used to distribute the adaptive inspection strategy to drones for autonomous inspections, and perform real-time analysis of the multimodal data obtained by the drones, identify and generate classified event alarms; a closed-loop response and report generation module, used to respond to the classified event alarms, match and trigger corresponding disposal methods from a preset multi-level response plan library, and generate a closed-loop inspection task report.
[0015] Compared with the prior art, the present invention has the following advantages: 1. This invention uses a large language model to deeply understand the semantics of natural language instructions and dynamically generate structured task vectors containing multiple possibilities. It elevates inspection tasks from executing a preset path to autonomously synthesizing an adaptive inspection strategy that includes routes, algorithms, and behavioral logic based on probabilistic mission intent, fundamentally improving the intelligence and flexibility of handling ambiguous and complex tasks. 2. When responding to alerts, this invention doesn't simply match responses from a static response library. Instead, it fuses real-time classified event alerts with the initial structured task vector to construct a comprehensive decision context. This framework then uses a large language model for further reasoning to generate the optimal response strategy. This mechanism eliminates isolated emergency response and instead deeply couples it with the top-level intent of the task, significantly improving the accuracy and adaptability of response decisions.
[0016] Other features and advantages of the present invention will be described in the following description, and in part will become apparent from the description, or will be understood by practicing the present invention. The purpose and other advantages of the present invention can be realized and obtained by the structures pointed out in the description, claims and drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0018] Figure 1 2. It is a schematic diagram of a method for intelligent highway inspection by drones based on patrol requirements according to an embodiment of the present invention.
[0019] Figure 2 3 is a schematic diagram of updating the Bayesian posterior probability of the candidate task hypothesis according to an embodiment of the present invention.
[0020] Figure 3 It is a schematic diagram of the probabilistic task map and the optimal behavior strategy instantiation of an embodiment of the present invention.
[0021] Figure 4 4 is a schematic diagram of a decision boundary for abnormal event classification based on reconstruction error according to an embodiment of the present invention.
[0022] Figure 5 This is a schematic diagram of the structure of a drone highway intelligent inspection system based on patrol requirements in an embodiment of the present invention. DETAILED DESCRIPTION
[0023] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.
[0024] Reference Figure 1 One embodiment of the present invention proposes a method and system for intelligent highway inspection using drones based on patrol needs. It adopts a technical path that combines multi-level reasoning based on a large language model with dynamic behavior strategy synthesis. By constructing a complete cognitive and action framework from semantic intent understanding, behavioral decision-making to closed-loop response, it can achieve autonomous generation and intelligent execution of complex and fuzzy inspection tasks, significantly improving the automation level, response speed and decision-making intelligence of highway inspections.
[0025] The method of this embodiment specifically includes: Obtaining natural language instructions for patrol requirements, performing semantic parsing and intent decomposition on the natural language instructions, and generating a structured task vector; Based on the structured task vector and integrating real-time traffic and environmental perception data for collaborative planning, an adaptive inspection strategy is constructed and generated; Distributing the adaptive inspection strategy to drones for autonomous inspection, and performing real-time analysis on multimodal data acquired by the drones to identify and generate classified event alerts; In response to the classified event alarm, the corresponding disposal method is matched and triggered from the preset multi-level response plan library to generate a closed-loop inspection task report.
[0026] By adopting a technical approach that combines multi-level reasoning based on a large language model with dynamic behavior strategy synthesis, and by building a complete cognitive and action framework from semantic intent understanding, behavioral decision-making to closed-loop response, it is possible to achieve autonomous generation and intelligent execution of complex and fuzzy inspection tasks, significantly improving the automation level, response speed and decision-making intelligence of highway inspections.
[0027] Optionally, generating a structured task vector includes: Perform multi-path reasoning on the natural language instructions using a large language model to generate candidate task hypotheses; Specifically, this step aims to convert the user's potentially ambiguous or ambiguous natural language instructions into a set of clear, machine-processable alternatives. A large language model (LLM) receives the natural language instructions as input. Through a special prompt engineering technique, the model is guided not to directly output a single answer, but to conduct multi-path, divergent reasoning to generate a set of logically mutually exclusive candidate task hypotheses. Each candidate task hypothesis is a possible, complete, and structured interpretation of the user's true intent, encompassing key entities such as patrol routes, monitoring targets, and task priorities.
[0028] For the candidate task hypothesis, obtaining objective data for verifying the candidate task hypothesis from an external real-time data source to obtain external verification evidence; Specifically, this step aims to seek objective data support for the multiple subjective hypotheses generated in the previous step. A data verification process automatically calls one or more external application programming interfaces (APIs) for each candidate task hypothesis. For example, if a candidate task hypothesis is "congestion on the Suzhou section of the Beijing-Shanghai Expressway," the process will automatically call the real-time map service API to obtain the current average vehicle speed and traffic flow data for that section; if the hypothesis is "bad weather," the weather service API will be called to obtain rainfall and wind speed data. These objective data obtained from the outside in real time and capable of confirming or disproving the corresponding hypotheses together constitute external verification evidence.
[0029] Based on the external verification evidence, the candidate task hypotheses are subjected to posterior probability calculation and weighted fusion to generate a structured task vector.
[0030] Specifically, this step is the core of probabilistic, intelligent inference of user intent. A Bayesian inference process uses the initial probability of each candidate task hypothesis as the prior probability. Then, using the external verification evidence obtained in the previous step, a Bayesian update rule is used to calculate the posterior probability of each candidate task hypothesis. The core idea of a Bayesian update can be expressed as follows: , in, It is Candidate task hypothesis The posterior probability after observing the external verification evidence E; is its prior probability; It is assumed When true, evidence is observed Finally, by weighted fusion of each candidate task hypothesis and its corresponding posterior probability, a final structured task vector with richer information dimensions is generated, which contains not only the task entity but also the confidence assessment of various possibilities, such as Figure 2 As shown in the form of a radar chart, it shows how the system dynamically adjusts its confidence in the user's multiple possible intentions based on external real-time evidence, that is, the change in the posterior probability compared to the prior probability.
[0031] Optionally, the constructing and generating an adaptive inspection strategy includes: Constructing and generating a probabilistic mission map based on patrol targets and probability distributions in the structured mission vector; Specifically, this step aims to transform the generated structured task vector, which contains probabilistic interpretations of various possible user intentions, into a graph structure for behavioral decision-making. A graph construction process defines the patrol target in each candidate task hypothesis as a node in the graph. Edges between nodes represent the possibility of switching between different hypotheses or transferring between geographically adjacent patrol areas. This process uses the posterior probabilities contained in the structured task vector as the initial weights for nodes and edges. Through this step, a weighted, directed probabilistic task graph is constructed that comprehensively represents the task possibility space.
[0032] Based on the probabilistic task map, the real-time traffic and the environmental perception data, a real-time strategy solution is performed to generate a temporal behavior strategy; Specifically, the purpose of this step is to generate an optimal behavior strategy after comprehensively considering the uncertainty of the task and the dynamic changes in the real world. A policy solution process uses a deep reinforcement learning (DRL) model, such as a policy network. The network takes the topology and weights of the probabilistic task graph, as well as real-time traffic and environmental perception data as input. Its goal is to learn a policy function that can maximize long-term cumulative rewards. This reward function integrates multiple optimization objectives such as task completion, flight safety, and energy consumption. Ultimately, the temporal behavior strategy output by the policy network defines the probability of which "atomic behavior" the drone should choose when facing any possible state at each future time step.
[0033] According to the temporal behavior strategy, the probabilistic task graph is path instantiated to construct and generate an adaptive inspection strategy.
[0034] Specifically, this step is to convert the abstract behavior strategy generated in the previous step into a specific, executable inspection task plan. A path instantiation process will traverse the probabilistic task map based on the optimal action sequence given by the temporal behavior strategy to determine an optimal inspection path. At the same time, the process will configure specific sensor parameters and the intelligent detection algorithm that needs to be mounted and activated for each waypoint on the path based on the requirements for sensor actions in the temporal behavior strategy. A complete plan that includes this instantiated optimal path and the sensor and algorithm configuration bound to it constitutes the final adaptive inspection strategy, such as Figure 3 As shown, in the form of a network diagram, it is demonstrated how the present invention determines an optimal behavior strategy path through intelligent decision-making from a task map containing multiple possibilities.
[0035] Optionally, the method further includes: Performing path planning deduction on the probabilistic task graph under the candidate task hypothesis to generate an expected task completion indicator; Specifically, this step aims to conduct a comprehensive "stress test" on the generated preliminary action framework. A simulation process iterates through each generated high-probability candidate task hypothesis. For each candidate task hypothesis, the process simulates a complete path planning on the probabilistic task map and calculates a quantitative expected task completion metric based on the objectives defined by the hypothesis. This metric can be a normalized score that combines factors such as expected completion time, target coverage, and resource consumption.
[0036] Performing statistical analysis on the expected task completion index, calculating and generating a graph confidence index of the probabilistic task graph; Specifically, this step aims to quantify the robustness of the preliminary action framework generated in the previous step. A statistical analysis process calculates the variance or standard deviation of the expected task completion indicator generated in the previous step under all candidate task hypotheses. A low variance means that the current probabilistic task graph has good adaptability to a variety of possible user intentions, that is, it is robust. Conversely, a high variance means that the graph is very "fragile" and can only serve the most mainstream hypothesis well, but performs poorly for other possibilities. The final generated graph confidence index is designed to be inversely proportional to the variance.
[0037] If the graph confidence index is lower than a preset robustness threshold, the candidate task hypothesis is used as a new constraint to iteratively optimize the node transition probability of the probabilistic task graph.
[0038] Specifically, this step is a closed-loop link for enhancing the robustness of the graph. First, the graph confidence index generated in the previous step is compared with a robustness threshold. If the index is lower than the threshold, it indicates that the current probabilistic task graph is not robust enough. At this point, an iterative optimization process will be triggered. This process will identify candidate task hypotheses that lead to lower expected task completion indicators and convert them into a new set of hard constraints that must be met. Subsequently, an optimizer will re-solve and adjust the inter-node transition probabilities in the probabilistic task graph with the goal of meeting this new set of constraints, thereby generating a new version of the probabilistic task graph that is more adaptable to all high-probability intentions of the user and has enhanced robustness.
[0039] Optionally, identifying and generating a classified event alarm includes: Obtain multimodal data under historical normal inspection conditions and train a data generator model; Specifically, this step aims to build a generative model that can deeply learn and reproduce the inherent distribution patterns of "normal" inspection data. In this step, a Generative Adversarial Network (GAN) framework is used for model training. This framework consists of a generator and a discriminator. The training process is a dynamic, adversarial game process: the generator continuously attempts to generate synthetic data that is indistinguishable from real normal data, while the discriminator continuously learns to more accurately distinguish between real data and synthetic data. The goal of this training process is to minimize an adversarial loss function. An exemplary loss function can be expressed by the formula: , in, is the overall value function; is a generator; is the discriminator; Is from a real normal data distribution samples; is derived from a prior noise distribution samples; The discriminator determines The probability of being the true data; The generator is based on the noise Generated synthetic data. After training, the resulting generator is a data generator model that has the ability to generate highly realistic, normal multimodal data based on a low-dimensional latent space vector.
[0040] The multimodal data acquired in real time is reversely mapped and reconstructed through the data generator model to calculate and generate optimal reconstructed data; Specifically, the goal of this step is to find the best projection on the "normal" data manifold, that is, the "most similar normal data", for any new piece of real-time data. An inverse reconstruction process transforms this task into an optimization problem. First, a reconstruction loss function is defined, which is composed of the difference between the real-time multimodal data and the synthetic data output by the data generator model. Then, based on the gradient of the reconstruction loss function and coupled with a controlled random noise term, the latent space of the data generator model is iteratively explored through a Langevin dynamic sampling process to find the optimal latent space vector that minimizes the reconstruction loss function. Finally, the optimal latent space vector is input into the data generator model for a forward propagation, and its output is the optimal reconstructed data.
[0041] The reconstruction error between the multimodal data and the optimal reconstruction data is quantitatively analyzed, and event classification is performed on the area where the reconstruction error exceeds a preset dynamic threshold, and classified event alarms are identified and generated.
[0042] Specifically, this step is the final abnormality determination link, and an error analysis process calculates the difference between real-time multimodal data and its corresponding optimal reconstructed data, that is, the reconstruction error. When a piece of real-time data contains an abnormal pattern, since the data generator model has never learned such a pattern in training, it cannot generate a "normal" data similar to it, resulting in a significant increase in the reconstruction error. By comparing the size of the reconstruction error with a dynamically set threshold, the existence of the abnormality can be determined. Furthermore, by analyzing the distribution characteristics of the reconstruction error in different data dimensions, for example, whether the texture part of the image data has a large error, or the spectrum part of the sound data has a large error. The identified abnormal events can be preliminarily classified, and finally a classified event alarm containing the event type, confidence level, and spatiotemporal location is identified and generated, such as Figure 4 As shown in the figure, it shows how the system can intelligently distinguish different types of events such as "normal state", "traffic accident", "illegal parking" through nonlinear decision boundaries in a two-dimensional feature space composed of "image reconstruction error" and "sound reconstruction error".
[0043] Optionally, the calculating and generating optimal reconstruction data includes: defining a reconstruction loss function based on a latent space of the data generator model and utilizing the difference between the multimodal data and the synthesized data output by the data generator model; Specifically, this step aims to construct a mathematical metric that can accurately quantify the difference between any two pieces of multimodal data. The loss function definition process not only considers the direct pixel or numerical differences between the multimodal data acquired in real time and the synthetic data output by the data generator model, but also creatively introduces a perceptual loss term. This perceptual loss term extracts high-dimensional features from the two pieces of data through a pre-trained deep neural network and calculates the distance between these high-dimensional features. By weightedly summing the direct difference loss and the perceptual loss, a more comprehensive reconstruction loss function can be constructed that focuses on both low-level details and high-level semantic similarity.
[0044] Based on the gradient of the reconstruction loss function and coupled with a random noise term, the latent space is iteratively explored to generate a latent space state sample; Specifically, this step aims to find the optimal latent space vector that minimizes the reconstruction loss function through an efficient global optimization process. An iterative exploration process adopts a Langevin dynamic sampling method derived from statistical physics. In each iteration, the update of a candidate vector in the latent space not only moves a small step in the opposite direction of the gradient of the reconstruction loss function, but also is additionally imposed with a random noise term of controlled intensity sampled from a Gaussian distribution. The introduction of this random noise term enables the search process to escape the "trap" of the local optimal solution, thereby conducting a more extensive exploration in the entire latent space, and finally converges to the neighborhood of the global optimal solution, and generates a set of latent space state samples surrounding the optimal solution.
[0045] Statistical moment estimation is performed on the latent space state samples, a first latent space vector is calculated and determined, and the first latent space vector is input into the data generator model for forward propagation to generate optimal reconstructed data.
[0046] Specifically, this step is the final step in determining the optimal latent vector and completing the reconstruction. A statistical estimation process processes a set of stable latent space state samples generated after the previous iterative exploration step. By calculating the mean of this set of samples, an optimal latent space vector is obtained. This optimal latent space vector, determined after thorough exploration and optimization, is used as the final input and a forward propagation calculation is performed through the data generator model. The output is the best "normal state" reconstruction that is closest to the original real-time multimodal data, namely the optimal reconstructed data.
[0047] Optionally, generating a closed-loop inspection task report includes: fusing the classified event alert with the structured task vector to construct a comprehensive decision context; Specifically, this step aims to provide a comprehensive input for subsequent intelligent decision-making, encompassing both mission intent and current conditions. A context-building process integrates the generated classified event alerts, representing the specific current abnormal event, with the generated structured task vector, representing the user's original patrol intent, at the data level. The resulting comprehensive decision context is a machine-readable data object containing complex information such as "(task priority: high, focus target: congestion) + (real-time event: multi-vehicle rear-end collision at K105, confidence level: 98%)."
[0048] Inputting the comprehensive decision context into a large language model for multi-objective reasoning to generate a dynamic response strategy; Specifically, this step is the core of the present invention for achieving advanced intelligent decision-making. A large language model receives the comprehensive decision context constructed in the previous step as part of its prompt. The model is trained or fine-tuned to understand the expertise in the field of traffic inspection and emergency response. Through its powerful contextual understanding and multi-objective reasoning capabilities, the model is able to go beyond simple rule matching and comprehensively evaluate the urgency of the current situation, the importance of the initial task, and the available disposal resources to generate an optimal, personalized dynamic response strategy. This strategy is not a single action, but a structured behavior plan that includes multiple specific disposal actions and their recommended execution sequences.
[0049] The disposal method is executed according to the dynamic response strategy, and the disposal process is recorded to generate a closed-loop inspection task report.
[0050] Specifically, this step completes the final execution and recording of the entire "perception-decision-action" closed loop. A task execution module parses the dynamic response strategy generated in the previous step and converts each action into a call instruction for a specific hardware or software system. After all actions are executed, a report generation process compiles and structures all the information for this task, including the initial natural language instructions, the generated structured task vector, the identified classified event alerts, the generated dynamic response strategy, and the execution records and results of each action. This ultimately generates a complete and fully traceable closed-loop inspection task report.
[0051] Optionally, the method further includes: Projecting the temporal behavior strategy onto the structured task vector, calculating and generating a strategy-intention coverage matrix; Specifically, this step aims to quantify the extent to which a generated, specific behavioral strategy can satisfy the multiple possibilities contained in the user's original intention. A coverage calculation process first projects the generated temporal behavioral strategy into the multivariate probability hypothesis space defined by the generated structured task vector. For each candidate task hypothesis, the process evaluates the coverage of the temporal behavioral strategy for its core objectives. For example, if a candidate task hypothesis is "patrol a 5-kilometer-long highway", and the path length of the generated temporal behavioral strategy only covers 3 kilometers of it, its coverage is 60%. By repeating this evaluation process for all high-probability candidate task hypotheses, a strategy-intent coverage matrix can be constructed. The calculation process of a matrix element can be expressed by the formula: , in, is an element in the strategy-intention coverage matrix, representing the The core target area in the i-th candidate task hypothesis is assumed by a temporal behavior strategy coverage rate; Is the behavioral strategy in The spatial position of the drone at the moment; is an indicator function that is 1 when the drone is within the target area and 0 otherwise; The target area The total size of .
[0052] Performing information entropy analysis on the strategy-intention coverage matrix to quantify and generate a cognitive coordination index; Specifically, this step aims to evaluate whether the "focus" and "comprehensiveness" of the current decision match the "uncertainty" of the user's intention from the perspective of information theory. An information entropy analysis process normalizes each row of the strategy-intention coverage matrix generated in the previous step to obtain a probability distribution that describes the "attention allocation" of a single behavioral strategy to all possible intentions. Subsequently, the Shannon entropy of the probability distribution is calculated and weighted with the posterior probability of each candidate task hypothesis in the structured task vector to obtain the final cognitive coordination index. An exemplary calculation function is shown in the formula: , in, is the final cognitive coordination index, the higher the value, the better the coordination between strategy and intention; is the total number of candidate task hypotheses; The calculated The posterior probability of the candidate task hypothesis; is the element in the normalized strategy-intention coverage matrix, representing the strategy Hypothesis coverage probability.
[0053] If the cognitive coordination index is lower than a preset coordination threshold, the candidate task hypothesis is used as a penalty item to generate a cognitive bias feedback signal.
[0054] Specifically, to generate a cognitive bias feedback signal based on the cognitive coordination index, this step completes the decision-making link of the "meta-learning" feedback loop. First, the cognitive coordination index generated in the previous step is compared with a coordination threshold. If the index is lower than the threshold, it indicates that there is an "imbalance" between the current decision-making behavior and its cognition of the user's intention. For example, excessive attention may be paid to the task hypothesis with the highest probability, while other possibilities are completely ignored. At this point, a feedback signal generation process will be activated, which will package those "ignored" candidate task hypotheses with high posterior probabilities, along with their corresponding low coverage, into a structured cognitive bias feedback signal. This signal is used as a special "penalty term" to guide and optimize the behavioral preferences of future decision-making models, so that they can make more comprehensive and coordinated decisions in subsequent tasks.
[0055] Optionally, the method further includes: Dynamically reshape the reward function of the strategy generation network model using the cognitive bias feedback signal to generate an updated reward function; Specifically, this step aims to convert the quantified degree of "cognitive dissonance" into a direct "punishment" or "incentive" for subsequent decision-making behaviors. A reward function reshaping process receives a cognitive bias feedback signal. This signal indicates which candidate task hypotheses were "ignored" in the previous strategy. This process generates an updated reward function by adding a dynamic penalty term related to the feedback signal to the original reward function of the deep reinforcement learning model. The penalty term is designed so that if the behavior strategy generated by the policy generation network again ignores these important candidate task hypotheses in subsequent training, it will receive a significant negative reward, thereby forcing it to adjust its decision-making preferences.
[0056] The updated reward function is used to perform online training on the policy generation network model to generate an updated policy generation network model.
[0057] Specifically, this step completes the entire meta-learning loop and implements the execution link for evolving the model's "decision-making intelligence." An online reinforcement learning training process is initiated, using the updated reward function as its core optimization objective. During training, the policy generation network model, through continuous "trial and exploration," learns how to generate a new temporal behavior strategy. This new strategy must not only meet the original task objectives but also strive to maximize this dynamically reshaped reward function that incorporates rewards and penalties for "cognitive coordination." Through this online training process, the original policy generation network model is replaced by an updated one with more comprehensive decision preferences and more robust behavior in the face of ambiguous intentions.
[0058] Example 1: To demonstrate the feasibility of this invention, we applied it to the intelligent inspection and management of a newly opened 50-kilometer section of a smart highway. This highway section traverses hilly terrain and urban-rural fringe areas, and includes multiple long tunnels, viaducts, and complex ramp hubs. Traffic flow is highly dynamic and uncertain. Traditional inspection methods, which rely on manually preset fixed routes and simple rule-based triggering, are unable to cope with the complex patrol needs and emergencies of this section.
[0059] To validate the benefits of this invention, a three-month test period was selected, conducting 24 / 7 intelligent, automated inspections on this road section. The control group employed a traditional drone inspection system, in which operators manually planned routes based on instructions from a monitoring center and employed rule-based event alerting logic. The experimental group fully deployed the method and system described in this invention, conducting 24 / 7 intelligent inspections using an autonomous drone airport and a swarm of drones equipped with multimodal sensors.
[0060] In this example, when an operator at the control center inputs an ambiguous natural language instruction via voice, such as "G50 Expressway K1520 westbound, evening rush hour traffic seems a bit congested. Pay attention and see if there's an accident," a large language model immediately performs multi-path reasoning on the instruction, generating multiple candidate task hypotheses, such as {Hypothesis A: Regular evening rush hour congestion, initial probability: 0.7}, {Hypothesis B: Minor traffic accident, initial probability: 0.2}, and {Hypothesis C: Road construction ahead, initial probability: 0.1}. The system then calls the real-time traffic API of an external mapping service and finds that the average speed at K1520 is significantly lower than historical data for the same period. Using Bayesian inference, the system updates the posterior probabilities of each hypothesis, for example, increasing the probability of {Hypothesis B: Minor traffic accident} to 0.65. Ultimately, a structured task vector is generated, focusing on "accident investigation" and "traffic diversion."
[0061] Subsequently, based on this structured task vector, a probabilistic task graph was constructed, which not only included a focused inspection of the K1520 section, but also assigned a lower observation probability to the adjacent ramp exits based on the uncertainty in the task vector. At the same time, a strategy generation network model integrated real-time wind speed data and traffic flow to generate an initial temporal behavior strategy. Before execution, a robustness verification module was activated to deduce the strategy under the candidate task hypothesis of "the main road of the K1520 section is completely blocked", and found that the expected task completion rate was extremely low. Therefore, the system automatically optimized the node transfer probability of the probabilistic task graph, adding an alternative path for the drone to approach the core area from above the adjacent ramp, thereby generating a more robust adaptive inspection strategy.
[0062] During the inspection process, multimodal data acquired by the drone along its route was fed into a GAN-trained data generator model in real time. At K1522, the reconstruction error of a video frame significantly exceeded the dynamic threshold, and the system identified it as an anomaly. Analysis of the error revealed that it was primarily caused by a stationary hotspot target in the image that did not conform to the normal kinematic characteristics of a vehicle. A classified event alert was generated, designated "abnormal parking in the emergency lane." Throughout the testing period, the method also successfully identified previously unseen anomalies, such as small amounts of fallen rocks caused by landslides.
[0063] Ultimately, the "abnormal parking in the emergency lane" alert, along with the initial structured task vector with an increased "accident probability," was fed into the large language model for multi-objective reasoning. The model determined that this abnormal parking was highly likely related to a potential traffic accident. Therefore, rather than generating the standard "illegal parking evidence collection" command, it instead generated a higher-priority dynamic response strategy. This strategy included: 1. Instructing the drone to immediately adjust its attitude and use its zoom lens to capture multi-angle, high-definition video of the target vehicle and surrounding road conditions; 2. Activating the onboard loudspeaker to call the target vehicle to confirm any injuries; 3. Linking with AutoNavi Maps to push a "Multiple accidents ahead, please slow down" message to oncoming vehicles. All processes were recorded, generating a complete closed-loop inspection task report.
[0064] After three months of comparative testing, the present invention demonstrated significant technical advantages in terms of accuracy in task understanding, ability to detect abnormal events, and intelligent level of emergency response. For specific data, please refer to Tables 1, 2, and 3.
[0065] Table 1 Comparison of intelligent task understanding and planning efficiency
[0066] Table 2. Abnormal event detection performance comparison
[0067] Table 3 Comparison of emergency response effects
[0068] Based on the same inventive concept, the present invention also provides a drone highway intelligent inspection system based on patrol needs, such as Figure 5 As shown, the system includes: A task semantic parsing module is used to obtain natural language instructions for patrol requirements, perform semantic parsing and intent decomposition on the natural language instructions, and generate a structured task vector; An adaptive strategy generation module is used to construct and generate an adaptive inspection strategy based on the structured task vector and integrating real-time traffic and environmental perception data for collaborative planning; An autonomous inspection and event recognition module, which is used to distribute the adaptive inspection strategy to drones for autonomous inspection, and to perform real-time analysis on the multimodal data acquired by the drones, identify and generate classified event alarms; The closed-loop response and report generation module is used to respond to the classified event alarm, match and trigger the corresponding disposal method from the preset multi-level response plan library, and generate a closed-loop inspection task report.
[0069] It should be noted that the functional division and information interaction between the aforementioned modules are logical. Physically, they can be integrated into the same software platform or deployed in a distributed manner. The connections between them represent data and control flows, designed to collaboratively achieve the dynamic optimization of building energy consumption of the present invention. The foregoing description is merely an exemplary embodiment of the present invention and is not intended to limit its scope.
Claims
1. The intelligent inspection method of highways by drones based on patrol requirements is characterized by: The method comprises: Obtaining natural language instructions for patrol requirements, performing semantic parsing and intent decomposition on the natural language instructions, and generating a structured task vector; Based on the structured task vector and integrating real-time traffic and environmental perception data for collaborative planning, an adaptive inspection strategy is constructed and generated; Distributing the adaptive inspection strategy to drones for autonomous inspection, and performing real-time analysis on multimodal data acquired by the drones to identify and generate classified event alerts; In response to the classified event alarm, the corresponding disposal method is matched and triggered from the preset multi-level response plan library to generate a closed-loop inspection task report.
2. The UAV highway intelligent inspection method based on patrol demand according to claim 1 is characterized in that: Generating a structured task vector includes: Perform multi-path reasoning on the natural language instructions using a large language model to generate candidate task hypotheses; For the candidate task hypothesis, obtaining objective data for verifying the candidate task hypothesis from an external real-time data source to obtain external verification evidence; Based on the external verification evidence, the candidate task hypotheses are subjected to posterior probability calculation and weighted fusion to generate a structured task vector.
3. The method for intelligent highway inspection using drones based on patrol requirements according to claim 2 is characterized in that: The construction and generation of the adaptive inspection strategy includes: Constructing and generating a probabilistic mission map based on patrol targets and probability distributions in the structured mission vector; Based on the probabilistic task map, the real-time traffic and the environmental perception data, a real-time strategy solution is performed to generate a temporal behavior strategy; According to the temporal behavior strategy, the probabilistic task graph is path instantiated to construct and generate an adaptive inspection strategy.
4. The method for intelligent highway inspection using drones based on patrol requirements according to claim 3 is characterized in that: The method further comprises: Performing path planning deduction on the probabilistic task graph under the candidate task hypothesis to generate an expected task completion indicator; Performing statistical analysis on the expected task completion index, calculating and generating a graph confidence index of the probabilistic task graph; If the graph confidence index is lower than a preset robustness threshold, the candidate task hypothesis is used as a new constraint to iteratively optimize the node transition probability of the probabilistic task graph.
5. The method for intelligent highway inspection using drones based on patrol requirements according to claim 1 is characterized in that: The identifying and generating of classified event alarms includes: Obtain multimodal data under historical normal inspection conditions and train a data generator model; The multimodal data acquired in real time is reversely mapped and reconstructed through the data generator model to calculate and generate optimal reconstructed data; The reconstruction error between the multimodal data and the optimal reconstruction data is quantitatively analyzed, and event classification is performed on the area where the reconstruction error exceeds a preset dynamic threshold, and classified event alarms are identified and generated.
6. The method for intelligent highway inspection using drones based on patrol requirements according to claim 5 is characterized in that: The calculating and generating of the optimal reconstruction data comprises: defining a reconstruction loss function based on a latent space of the data generator model and utilizing the difference between the multimodal data and the synthesized data output by the data generator model; Based on the gradient of the reconstruction loss function and coupled with a random noise term, the latent space is iteratively explored to generate a latent space state sample; Statistical moment estimation is performed on the latent space state samples, a first latent space vector is calculated and determined, and the first latent space vector is input into the data generator model for forward propagation to generate optimal reconstructed data.
7. The method for intelligent highway inspection using drones based on patrol requirements according to claim 1 is characterized in that: Generating a closed-loop inspection task report includes: fusing the classified event alert with the structured task vector to construct a comprehensive decision context; Inputting the comprehensive decision context into a large language model for multi-objective reasoning to generate a dynamic response strategy; The disposal method is executed according to the dynamic response strategy, and the disposal process is recorded to generate a closed-loop inspection task report.
8. The method for intelligent highway inspection using drones based on patrol requirements according to claim 4 is characterized in that: The method further comprises: Projecting the temporal behavior strategy onto the structured task vector, calculating and generating a strategy-intention coverage matrix; Performing information entropy analysis on the strategy-intention coverage matrix to quantify and generate a cognitive coordination index; If the cognitive coordination index is lower than a preset coordination threshold, the candidate task hypothesis is used as a penalty item to generate a cognitive bias feedback signal.
9. The method for intelligent highway inspection using drones based on patrol requirements according to claim 8 is characterized in that: The method further comprises: Dynamically reshape the reward function of the strategy generation network model using the cognitive bias feedback signal to generate an updated reward function; The updated reward function is used to perform online training on the policy generation network model to generate an updated policy generation network model.
10. A drone highway intelligent inspection system based on patrol needs, applied to a drone highway intelligent inspection method based on patrol needs as claimed in any one of claims 1 to 9, characterized in that: The system comprises: A task semantic parsing module is used to obtain natural language instructions for patrol requirements, perform semantic parsing and intent decomposition on the natural language instructions, and generate a structured task vector; An adaptive strategy generation module is used to construct and generate an adaptive inspection strategy based on the structured task vector and integrating real-time traffic and environmental perception data for collaborative planning; An autonomous inspection and event recognition module, which is used to distribute the adaptive inspection strategy to drones for autonomous inspection, and to perform real-time analysis on the multimodal data acquired by the drones, identify and generate classified event alarms; The closed-loop response and report generation module is used to respond to the classified event alarm, match and trigger the corresponding disposal method from the preset multi-level response plan library, and generate a closed-loop inspection task report.
Citation Information
Patent Citations
Power plant online inspection system based on fusion of audio and video and Internet of Things
CN118430088A
Oil depot tank field inspection robot task allocation method and system based on Internet of Things
CN119417192A
Unmanned aerial vehicle autonomous target searching system and method based on large language model
CN119759083A
Unattended system based on AI model and video analysis
CN120166195A
Data center inspection robot intelligent inspection method and system based on large model
CN120181360A
Cited By
Highway ramp traffic control method based on UAV-LLM-FCS
CN120894915A
Robot interaction method and system based on spatial physical information
CN121279353A
Live broadcast goods carrying real-time detection method and system based on multi-modal fusion
CN121330409A
New energy station safety inspection management method and system
CN121582749A
New energy station safety inspection management method and system
CN121582749B